Driver bad emotion recognition method and system and medium

By using the HDF5-VGG expression recognition model in driver bad emotions recognition, combined with image preprocessing and optimized convolutional neural network structure, the problems of slow training speed, high memory consumption and insufficient accuracy in the existing technology are solved, and more efficient and accurate recognition effects are achieved.

CN120236267APending Publication Date: 2025-07-01SHANGHAI UNIVERSITY OF ELECTRIC POWER
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510345931.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art faces the problems of slow training speed, high memory consumption and insufficient accuracy in driver bad emotions recognition.

Method used

The HDF5-VGG expression recognition model is adopted, and the structure and training process of the model are optimized through image preprocessing, format conversion and convolutional neural network recognition, combined with ELU activation function, batch normalization and Adam optimizer.

Benefits of technology

The training speed of the model is accelerated, memory consumption is reduced, and the accuracy rate is significantly improved to 82.00%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236267A_ABST
    Figure CN120236267A_ABST
Patent Text Reader

Abstract

The invention relates to a driver bad emotion recognition method and system and a medium, and the method comprises the following steps: obtaining a facial expression image of a driver, and carrying out the image preprocessing, and obtaining a CSV file of the facial expression image; converting the CSV file of the facial expression image into an HDF5 format; a facial expression image in an HDF5 format is used as input of an HDF5-VGG expression recognition model to obtain an expression recognition result, the HDF5-VGG expression recognition model comprises an input layer, a first convolution block, a second convolution block, a third convolution block, a full connection layer and a Softmax classifier which are connected in sequence, the first convolution block comprises a convolution layer and a maximum pooling layer which are connected in sequence, and the second convolution block comprises a second convolution block, a third convolution block, a full connection layer and a Softmax classifier which are connected in sequence. The second convolution block comprises a convolution layer and an average pooling layer which are connected in sequence, and the third convolution block comprises a convolution layer and a maximum pooling layer which are connected in sequence. Compared with the prior art, the training speed is increased, the memory consumption is reduced, and the accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of transportation safety monitoring, and in particular, to a method, system and medium for identifying drivers' bad emotions. Background Art

[0002] With the rapid development of the transportation industry, road traffic safety issues have attracted increasing attention. The emotional state of drivers has a crucial impact on driving behavior and traffic safety. Bad emotions such as anger, anxiety, fatigue, etc. may cause drivers to be inattentive and slow to react, thus increasing the risk of traffic accidents. Therefore, accurately and timely identifying drivers' bad emotions and taking corresponding measures is of great significance for ensuring road traffic safety.

[0003] As a powerful deep learning algorithm, Convolutional Neural Network (CNN) has achieved remarkable results in the field of image recognition. It can automatically learn the feature representation of images, and has strong generalization ability and recognition accuracy. For example, patent application CN112966625A discloses an expression recognition method and device based on an improved VGG-16 network model, which acquires a facial image; uses the pre-trained improved VGG-16 network model to recognize the facial image, and obtains the expression type reflected by the facial image. However, in the practical application of bad emotion recognition, the processing of large-scale data sets and model training still face challenges such as slow training speed, large memory consumption, and insufficient accuracy. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art, and provide a method, system and medium for identifying drivers' bad emotions, which speeds up the training speed, reduces the memory consumption, and improves the accuracy.

[0005] The purpose of the present invention can be achieved by the following technical solutions:

[0006] A method for identifying drivers' bad emotions includes the following steps:

[0007] Acquire the facial expression image of the driver, and perform image preprocessing to obtain the CSV file of the facial expression image;

[0008] Convert the CSV file of the facial expression image into the HDF5 format;

[0009] Taking the facial expression images in HDF5 format as the input of the HDF5-VGG facial expression recognition model to obtain the facial expression recognition result. The HDF5-VGG facial expression recognition model includes an input layer, a first convolutional block, a second convolutional block, a third convolutional block, a fully connected layer, and a Softmax classifier connected in sequence. The first convolutional block includes a convolutional layer and a max pooling layer connected in sequence. The second convolutional block includes a convolutional layer and an average pooling layer connected in sequence. The third convolutional block includes a convolutional layer and a max pooling layer connected in sequence.

[0010] Further, the image preprocessing includes grayscale processing, image detection, and image cropping.

[0011] Further, the input layer includes a 1×1 convolutional layer with a stride of 1. The input data is preliminarily feature-extracted through the 1×1 convolutional layer with a stride of 1.

[0012] Further, the first convolutional block includes a first convolutional layer, a second convolutional layer, and a third convolutional layer. The first convolutional layer uses 32 3×3 convolutional kernels. The second convolutional layer and the third convolutional layer respectively use 64 5×5 convolutional kernels. The convolutional layer of the second convolutional block uses 64 3×3 convolutional kernels. The convolutional layer of the third convolutional block uses 128 3×3 convolutional kernels. The 5×5 convolutional kernels are used to enhance the local feature capture ability of the HDF5-VGG facial expression recognition model, and the 3×3 convolutional kernels are used to control the number of parameters of the HDF5-VGG facial expression recognition model.

[0013] Further, features are extracted through the fully connected layer, and feature combination and classification are performed to map the recognition result of the HDF5-VGG facial expression recognition model to emotion categories. The fully connected layer includes a first fully connected layer, a second fully connected layer, and a third fully connected layer. The first fully connected layer and the second fully connected layer respectively contain 64 hidden nodes, and the third fully connected layer contains 6 hidden nodes.

[0014] Further, the ELU activation function and batch normalization are applied during the training process of the HDF5-VGG facial expression recognition model. The He initialization method is used to initialize the parameters of the HDF5-VGG facial expression recognition model, and the Adam optimizer is used to automatically adjust the learning rate of the HDF5-VGG facial expression recognition model to accelerate model convergence.

[0015] Further, the emotion categories include happy emotion, surprised emotion, neutral emotion, angry emotion, afraid emotion, and sad emotion.

[0016] Further, the HDF5-VGG facial expression recognition model is trained using the FER2013 facial expression dataset. The pictures in the FER2013 facial expression dataset are grayscale images covering a variety of expressions and having a fixed size of 48×48.

[0017] According to another aspect of the present invention, there is provided a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the driver bad mood recognition method.

[0018] According to another aspect of the present invention, there is provided a driver bad mood recognition system, characterized by comprising:

[0019] An image preprocessing module for acquiring a facial expression image of a driver and performing image preprocessing to obtain a CSV file of the facial expression image;

[0020] A format conversion module for converting the CSV file of the facial expression image into the HDF5 format;

[0021] An expression recognition module for using the facial expression image in the HDF5 format as the input of the HDF5-VGG facial expression recognition model to obtain an expression recognition result. The HDF5-VGG facial expression recognition model includes an input layer, a first convolutional block, a second convolutional block, a third convolutional block, a fully connected layer, and a Softmax classifier connected in sequence. The first convolutional block includes a convolutional layer and a max pooling layer connected in sequence. The second convolutional block includes a convolutional layer and an average pooling layer connected in sequence. The third convolutional block includes a convolutional layer and a max pooling layer connected in sequence.

[0022] Compared with the prior art, the present invention has the following beneficial effects:

[0023] 1. By constructing the HDF5-VGG facial expression recognition model, which includes an input layer, a first convolutional block, a second convolutional block, a third convolutional block, a fully connected layer, and a Softmax classifier connected in sequence. The first convolutional block includes a convolutional layer and a max pooling layer connected in sequence. The second convolutional block includes a convolutional layer and an average pooling layer connected in sequence. The third convolutional block includes a convolutional layer and a max pooling layer connected in sequence. Through this hierarchical network structure, the model can extract and combine features at multiple levels and use the Softmax classifier to calculate the probability distribution of each category, thereby improving the classification accuracy of the model.

[0024] 2. The present invention obtains the facial expression image of the driver and performs image preprocessing to obtain a CSV file of the facial expression image, and converts the CSV file of the facial expression image into the HDF5 format, avoiding problems such as the text parsing overhead of the CSV file, reading partial data as needed, and loading the entire dataset into memory at one time, optimizing the efficiency of data storage and processing, accelerating the training speed, and reducing the memory consumption of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a schematic flowchart of a method for identifying the bad mood of a driver proposed by the present invention;

[0026] Figure 2 is a schematic structural diagram of the HDF5-VGG expression recognition model proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives the detailed implementation manner and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0028] English abbreviations involved:

[0029] Comma-Separated Values: CSV

[0030] Hierarchical Data Format version 5: HDF5

[0031] Exponential Linear Unit: ELU

[0032] Rectified Linear Unit: ReLU

[0033] Embodiment 1

[0034] This embodiment provides a method for identifying the bad mood of a driver, as Figure 1 shown, including the following steps:

[0035] S1. Obtain the facial expression image of the driver and perform image preprocessing to obtain a CSV file of the facial expression image.

[0036] Use an in-vehicle camera or other image acquisition devices to obtain the facial image of the driver in real time. Ensure that the resolution and frame rate of the camera can meet the requirements of subsequent processing, and at the same time adjust the position and angle of the camera to clearly capture the facial features of the driver.

[0037] Image preprocessing includes grayscale processing, image detection, and image cropping.

[0038] The color image is converted to a grayscale image through grayscale processing. Grayscale processing can reduce the dimension of image data, lower the computational complexity, and at the same time retain the main features of the image.

[0039] Use a face detection algorithm (such as Haar feature classifier, HOG feature combined with SVM classifier, or deep learning-based detection model) to locate the face region in the image. The purpose of face detection is to accurately extract the driver's facial region from the complex background, providing a basis for subsequent image cropping. The detection algorithm will output the position information of the face, including parameters such as the upper left corner coordinates, width, and height.

[0040] According to the results of face detection, the region containing the driver's face is cropped from the grayscale image. The cropped image will be the main object for subsequent analysis, removing background interference and improving the processing efficiency and accuracy. The cropping range can be adjusted according to actual needs. For example, a certain edge region can be retained to avoid information loss caused by excessive cropping.

[0041] Convert the cropped facial image data into CSV format. The CSV file contains information such as image pixel values, timestamps, and image numbers for subsequent data analysis and processing.

[0042] S2. Convert the CSV file of facial expression images into HDF5 format.

[0043] HDF5 is a data storage format, especially suitable for storing a large amount of data on disk. An HDF5 file can be regarded as a group, which contains different datasets, and these datasets can be images, tables, etc. The structure of the HDF5 group is similar to the directory hierarchy of the file system. The root directory can also contain other directories, and the corresponding datasets are stored in the node directories.

[0044] The purpose of converting the CSV file into HDF5 format is to more conveniently train a convolutional neural network on it. When using deep learning to learn image files, if the number of picture files is very large, such as thousands or even tens of thousands, importing them one by one into memory will seriously slow down the running speed of the deep learning algorithm. Just like copying a folder containing tens of thousands of files is much slower than directly copying a packed folder. To avoid I / O operations becoming a bottleneck in the improvement of deep learning training speed, a method is adopted, that is, packing a large number of files into HDF5 file format, and then directly importing data from HDF5 when performing deep learning algorithm training.

[0045] S3. Use the facial expression images in HDF5 format as the input of the HDF5-VGG expression recognition model to obtain the expression recognition result.

[0046] The structure of the HDF5-VGG expression recognition model is as Figure 2 shown, including an input layer, a first convolutional block, a second convolutional block, a third convolutional block, a fully connected layer, and a Softmax classifier connected in sequence. The first convolutional block includes a convolutional layer and a max pooling layer connected in sequence. The second convolutional block includes a convolutional layer and an average pooling layer connected in sequence. The third convolutional block includes a convolutional layer and a max pooling layer connected in sequence.

[0047] The specific structure of each layer of the HDF5-VGG expression recognition model is shown in Table 1. The input layer includes a 1×1 convolutional layer with a stride of 1. The input data is preliminarily feature-extracted through the 1×1 convolutional layer with a stride of 1.

[0048] The first convolutional block includes a first convolutional layer (Conv1), a second convolutional layer (Conv2), and a third convolutional layer (Conv3). The first convolutional layer (Conv1) uses 32 3×3 convolutional kernels. The second convolutional layer (Conv2) and the third convolutional layer (Conv3) respectively use 64 5×5 convolutional kernels. The two convolutional layers of the second convolutional block use 64 3×3 convolutional kernels. The two convolutional layers of the third convolutional block use 128 3×3 convolutional kernels. The 5×5 convolutional kernels are used to enhance the local feature capture ability of the HDF5-VGG expression recognition model, and the 3×3 convolutional kernels are used to control the number of parameters of the HDF5-VGG expression recognition model.

[0049] Feature dimensionality reduction is achieved through the max pooling layer.

[0050] The fully connected layer extracts features, combines and classifies the features, so as to map the recognition result of the HDF5-VGG expression recognition model to emotion categories. The fully connected layer includes a first fully connected layer (Fc1), a second fully connected layer (Fc2), and a third fully connected layer (Fc3). The first fully connected layer (Fc1) and the second fully connected layer (Fc2) respectively contain 64 hidden nodes. The third fully connected layer (Fc3) contains 6 hidden nodes.

[0051] The fully connected layer maps the recognition result of the HDF5-VGG expression recognition model to 6 emotion categories, including happy emotion, surprised emotion, neutral emotion, angry emotion, afraid emotion, and sad emotion.

[0052] The Softmax classifier is used to calculate the probability distribution for each class. This design enables the model to effectively perform classification tasks and output the confidence for each class. Through this hierarchical network structure, the HDF5-VGG facial expression recognition model can extract and combine features at multiple levels, thereby achieving a high-precision classification effect.

[0053] Table 1 Structure of the HDF5-VGG Facial Expression Recognition Model

[0054] Layer Name Feature Map Size Kernel Size Input 48×48×1 Conv1 48×48×32 3×3 Conv2 48×48×64 5×5 Conv3 48×48×64 5×5 MaxPooling2D 24×24×64 2×2 Conv 24×24×64 3×3 Conv 24×24×64 3×3 AveragePooling2D 12×12×64 2×2 Conv 12×12×128 3×3 Conv 12×12×128 3×3 MaxPooling2D 6×6×128 2×2 Fc1 64 Fc2 64 Fc3 6 Softmax 6

[0055] The FER2013 facial expression dataset is used to train the HDF5-VGG facial expression recognition model, which contains 35,886 facial expression images. Among them, there are 28,708 training images, and 3,589 public validation images and 3,589 private validation images respectively. Each picture consists of a grayscale image with a fixed size of 48×48, covering 7 expressions. The corresponding digital labels are 0-6, and the labels corresponding to specific expressions and their Chinese and English are as follows: 0-anger, 1-disgust, 2-fear, 3-happy, 4-sad, 5-surprised, 6-normal. Since the two emotions of "disgust" and "anger" have high visual similarity (for example, both show frowning and tense facial muscles), these two categories are combined into a new category. This adjustment not only alleviates the problem of unbalanced sample numbers but also simplifies the classification task of the model.

[0056] During the training process of the HDF5-VGG facial expression recognition model, the traditional Xavier initialization is based on the assumption of linear activation functions and is difficult to adapt to the non-linear characteristics of deep ReLU networks. The He initialization method can effectively alleviate the problem of abnormal gradient propagation and is especially suitable for deep non-linear networks. Therefore, the He initialization method is used to initialize the parameters of the HDF5-VGG facial expression recognition model.

[0057] Aiming at the defects of the SGD optimizer, such as slow convergence speed and easy to fall into local optima, the Adam optimizer is introduced. By integrating the momentum term and the adaptive learning rate, Adam can achieve a dynamic balance of second-order moment estimation in the parameter space. The learning rate of the HDF5-VGG facial expression recognition model is automatically adjusted by the Adam optimizer to accelerate the model convergence.

[0058] To solve the problem of neuron death in ReLU, the ELU activation function is adopted. Different from ReLU, ELU maintains a small slope in the negative value range, enabling neurons to respond more smoothly when facing negative values. This improvement not only helps the network better capture features but also reduces the risk of overfitting.

[0059] During the training process of the HDF5-VGG facial expression recognition model, batch normalization is also applied to ensure the stability and efficiency of the model during training.

[0060] Due to its complex network architecture and large number of parameters, the VGG model usually incurs high computational costs. However, by introducing HDF5 technology, not only the efficiency of data storage and processing is optimized, but also the accuracy of the model is significantly improved from 65.00% to 82.00%, as shown in Table 2. This improvement fully demonstrates the broad application potential of HDF5 technology in deep learning models, especially in scenarios that require frequent data reading and writing.

[0061] Table 2

[0062]

[0063] Taking traditional machine learning as the baseline, the accuracy of multiple deep learning methods and the HDF5-VGG model proposed in this invention are analyzed, and the analysis results are shown in Table 3.

[0064] Table 3

[0065]

[0066] Traditional machine learning baseline: Support Vector Machine (SVM), as a representative of traditional machine learning, has an accuracy of 82.00%.

[0067] Deep learning method comparison set:

[0068] ResNet-50: A classic residual network with an accuracy of 90.00%.

[0069] ResNet-LSTM hybrid architecture: Combining temporal modeling capabilities, with an accuracy of 72.00%.

[0070] CNN-LSTM two-stream network: Based on multi-modal fusion, with an accuracy of 89.00%.

[0071] HDF5-VGG model: Through a hierarchical parameter optimization strategy, the accuracy is further improved to 92.00%.

[0072] In another preferred embodiment, it further includes the steps of:

[0073] S4. According to the facial expression recognition result, provide professional driving advice for the driver.

[0074] Deeply explore the impact of different emotional states on driving behavior and provide professional driving advice for the driver.

[0075] Neutral Emotion: When the driver has a blank expression or is in a neutral emotional state, their cognitive ability and reaction speed are at a normal level. This state does not have a significant impact on driving and is the benchmark state for the driver to maintain safe driving. It is recommended that the driver perform deep breathing or simple relaxation exercises before departure to bring their emotions back to a neutral state.

[0076] Positive Emotion: Positive emotions such as pleasure and relaxation can enhance the driver's concentration and judgment. Moderate positive emotions can increase the driver's reaction speed by 15 - 20%, and at the same time enhance the ability to anticipate road conditions. Therefore, it is recommended that the driver maintain moderate positive emotions.

[0077] Negative Emotion: A potential threat to driving safety

[0078] (1) Surprised State: When the driver is in a surprised state, it usually means that they have encountered a sudden situation. This emotion will significantly reduce the driver's judgment. It is recommended to immediately park the vehicle in a safe area and continue driving after the emotion has subsided.

[0079] (2) Angry Emotion: Anger will seriously affect the driver's judgment and control ability. Data shows that the probability of a traffic accident occurring in an angry state is 3 times that in a normal state. It is recommended that the driver take a short break when feeling angry.

[0080] (3) Fear Emotion: Fear will cause the driver to be overly nervous and affect the control of the vehicle. It is recommended that when the driver feels fear, reduce the speed, choose a safe route, and seek assistance from others if necessary.

[0081] By analyzing the driver's facial features, the emotional state is monitored in real-time. When it is detected that the driver is in a negative emotion, a warning is issued and corresponding measures are recommended.

[0082] Embodiment 2

[0083] This embodiment provides a computer-readable storage medium with a computer program stored thereon. When the computer program is executed by a processor, it can implement the method for identifying the driver's bad emotions.

[0084] The rest is the same as Embodiment 1.

[0085] Embodiment 3

[0086] This embodiment provides a driver bad emotion identification system, including:

[0087] An image preprocessing module, used to obtain the driver's facial expression image and perform image preprocessing to obtain a CSV file of the facial expression image;

[0088] A format conversion module, used to convert the CSV file of the facial expression image into the HDF5 format;

[0089] An expression recognition module, configured to use a facial expression image in HDF5 format as the input of an HDF5-VGG expression recognition model to obtain an expression recognition result. The HDF5-VGG expression recognition model includes an input layer, a first convolutional block, a second convolutional block, a third convolutional block, a fully connected layer, and a Softmax classifier that are connected in sequence. The first convolutional block includes a convolutional layer and a max pooling layer that are connected in sequence. The second convolutional block includes a convolutional layer and an average pooling layer that are connected in sequence. The third convolutional block includes a convolutional layer and a max pooling layer that are connected in sequence.

[0090] The rest is the same as in Embodiment 1.

[0091] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art shall fall within the protection scope determined by the claims.

Claims

1. A method for identifying a driver's negative emotions, characterized in that: The following steps are involved: Acquire the driver's facial expression image, and perform image preprocessing to obtain a CSV file of the facial expression image; Convert the CSV file of the facial expression image into HDF5 format; A facial expression image in HDF5 format is used as input of an HDF5-VGG expression recognition model to obtain an expression recognition result. The HDF5-VGG expression recognition model includes an input layer, a first convolution block, a second convolution block, a third convolution block, a fully connected layer and a Softmax classifier connected in sequence. The first convolution block includes a convolution layer and a maximum pooling layer connected in sequence. The second convolution block includes a convolution layer and an average pooling layer connected in sequence. The third convolution block includes a convolution layer and a maximum pooling layer connected in sequence.

2. The method for identifying a driver's negative emotions according to claim 1, characterized in that: The image preprocessing includes grayscale processing, image detection and image cropping.

3. The method for identifying a driver's negative emotions according to claim 1, characterized in that: The input layer includes a 1×1 convolution layer with a step size of 1, and preliminary feature extraction is performed on the input data through the 1×1 convolution layer with a step size of 1.

4. The method for identifying a driver's negative emotions according to claim 1, characterized in that: The first convolution block includes a first convolution layer, a second convolution layer and a third convolution layer. The first convolution layer uses 32 3×3 convolution kernels, the second convolution layer and the third convolution layer use 64 5×5 convolution kernels respectively, the convolution layer of the second convolution block uses 64 3×3 convolution kernels, and the convolution layer of the third convolution block uses 128 3×3 convolution kernels. The 5×5 convolution kernel is used to enhance the local feature capture capability of the HDF5-VGG expression recognition model, and the 3×3 convolution kernel is used to control the parameter amount of the HDF5-VGG expression recognition model.

5. The method for identifying a driver's negative emotions according to claim 1, characterized in that: Features are extracted through the fully connected layer, and the features are combined and classified to map the recognition results of the HDF5-VGG expression recognition model to emotion categories. The fully connected layer includes a first fully connected layer, a second fully connected layer and a third fully connected layer. The first fully connected layer and the second fully connected layer respectively contain 64 hidden nodes, and the third fully connected layer contains 6 hidden nodes.

6. The method for identifying a driver's negative emotions according to claim 1, characterized in that: The ELU activation function and batch normalization are applied in the training process of the HDF5-VGG expression recognition model. The He initialization method is used to initialize the parameters of the HDF5-VGG expression recognition model. The learning rate of the HDF5-VGG expression recognition model is automatically adjusted through the Adam optimizer to accelerate the convergence of the model.

7. The method for identifying a driver's negative emotions according to claim 4, characterized in that: The emotion categories include happy emotion, surprised emotion, neutral emotion, angry emotion, afraid emotion and sad emotion.

8. The method for identifying a driver's negative emotions according to claim 1, characterized in that: The HDF5-VGG expression recognition model is trained using the FER2013 facial expression dataset, where the images in the FER2013 facial expression dataset are grayscale images covering a variety of expressions and having a fixed size of 48×48.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for identifying a driver's negative emotions as described in any one of claims 1 to 8 can be implemented.

10. A driver's negative emotion recognition system, characterized in that: include: An image preprocessing module is used to obtain the driver's facial expression image and perform image preprocessing to obtain a CSV file of the facial expression image; A format conversion module, used for converting the CSV file of the facial expression image into HDF5 format; The expression recognition module is used to use the facial expression image in HDF5 format as the input of the HDF5-VGG expression recognition model to obtain the expression recognition result. The HDF5-VGG expression recognition model includes an input layer, a first convolution block, a second convolution block, a third convolution block, a fully connected layer and a Softmax classifier connected in sequence. The first convolution block includes a convolution layer and a maximum pooling layer connected in sequence. The second convolution block includes a convolution layer and an average pooling layer connected in sequence. The third convolution block includes a convolution layer and a maximum pooling layer connected in sequence.

Citation Information

Patent Citations

  • Expression recognition method and device based on improved VGG-16 network model

    CN112966625A