Mine worker emotion recognition method based on expression, electroencephalogram and speech multi-modal fusion

By improving the backbone feature extraction network, Transformer feature enhancement, and lightweight deep separable convolutional residual neural network, and combining multimodal information fusion, the problems of facial feature differentiation, complex EEG feature acquisition, and insufficient speech samples in miner emotion recognition were solved, achieving high-precision miner emotion recognition and improving coal mine safety.

CN117195148BActive Publication Date: 2025-12-09XIAN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311152044.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-07
Publication Date
2025-12-09
Estimated Expiration
2043-09-07

AI Technical Summary

Technical Problem

Existing technologies for miner emotion recognition suffer from several problems, including variations in facial image feature information, overfitting due to the depth of convolutional neural networks, high computational resource consumption, complex acquisition of EEG time-frequency features, insufficient depth of speech sample information, and limitations and low accuracy of single-modal information. These issues result in low accuracy in miner emotion recognition and make it difficult to effectively prevent coal mine accidents.

Method used

By employing an improved backbone feature extraction network, an EEG enhancement recognition model based on Transformer feature enhancement and attention mechanisms, and a lightweight deep separable convolutional residual neural network, combined with a multimodal information fusion method, we achieve complementary and efficient recognition of multimodal information through multi-scale feature extraction, feature enhancement, and attention fusion.

Benefits of technology

It improves the accuracy and robustness of miner emotion recognition, enabling real-time identification of miners' mental state, preventing coal mine accidents, and enhancing coal mine safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117195148B_ABST
    Figure CN117195148B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, and discloses a miner emotion recognition method based on expression, electroencephalogram and voice multi-modal fusion, which comprises a multi-scale face emotion recognition network model under an improved backbone feature extraction network, an electroencephalogram enhanced emotion recognition network based on a feature enhancement and attention mechanism of a Transformer, a voice emotion recognition model based on a lightweight deep separable convolution residual neural network and a multi-modal information fusion method. The face emotion recognition method of the application extracts features based on miner facial expression features, and then uses a recognition model to achieve the purpose of face emotion recognition. The multi-modal information fusion can supplement the mental state recognition, and improve the accuracy of the mental state recognition.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to a miner emotion recognition method based on expression, electroencephalogram and voice multi-modal fusion. BACKGROUND

[0002] In recent years, about 70% of underground accidents are caused by poor emotional state of miners. The low emotional state and unstable emotions of miners can lead to major operational errors and cause coal mine accidents. It is of practical significance to timely judge the emotional state of coal mine workers and prevent accidents.

[0003] In recent years, the research, discussion and practice of brain science have developed rapidly. From the scientific point of view, electroencephalogram signals can objectively reflect the range and mode of brain activity. Therefore, more researchers pay attention to emotion recognition research through electroencephalogram signals. Voice emotion recognition is a precise grasp of emotional reflection characteristics and classification and extraction of different emotional types. The main purpose of this technology is to improve recognition accuracy. Related professional researchers are constantly exploring and innovating, and gradually achieving a driving result. Facial emotion recognition is a key physiological measure for judgment research. Through the study of emotional state, not only can the risks of daily life and work be avoided, but also the production efficiency of workers and the industrial chain can be improved. This work has gradually become the research focus of scholars in recent years.

[0004] However, the existing technology has the following problems:

[0005] 1. The differentiation of facial image feature information leads to various problems in facial information extraction: for example, it is difficult to choose the size of the convolution kernel during convolution operation; the deeper the network structure of the convolutional neural network, the more likely it is to cause overfitting; and simple stacking of convolution layers consumes a large amount of computing resources.

[0006] 2. How to obtain high-quality electroencephalogram time-frequency features and apply them to the field of electroencephalogram emotion recognition? Obtaining high-quality electroencephalogram time-frequency features is a complex and challenging task that requires the combination of appropriate signal processing and analysis methods, while considering factors such as signal complexity, noise interference and individual differences. Only accurate, stable and reliable time-frequency features can be obtained to provide information for electroencephalogram emotion recognition.

[0007] 3. How to obtain deeper voice sample information and lightweight voice emotion recognition model? The difficulty in obtaining deep voice sample feature information is mainly due to the combined effects of insufficient data, high-dimensional data, complexity of voice signals, uncertainty, data preprocessing, difficulty of label annotation, and model design and optimization. This will make the use of voice samples insufficient and will affect the prediction accuracy of the model. The lightweight of the model will reduce the operation cost and time consumption.

[0008] Four, the problem of single-modal information limitation and insufficient accuracy of single-modal algorithm. The limitations of single-modal information mainly include incomplete information, limited features, and unbalanced data. Since the data of a single modality cannot cover all the information required for the task, the model may not accurately understand the task or handle complex situations. In addition, single-modal information is easily affected by environmental noise and interference, reducing the robustness of the model and affecting the accuracy of the task. In some tasks, the data of a single modality may have the problem of data imbalance, making the model perform poorly when dealing with minority classes or samples.

[0009] To solve the above problems, Shang Yucheng et al. (Electronic World, 2021) used the EM-Xception neural network structure to realize facial emotion recognition. Xception and Inception-ResNet are both improved from the Inception v3 network structure, and EM-Xception is improved by reducing the number of residual modules in Xception and replacing the activation function RELU with ELU. From the results of this study, it can indeed achieve a certain accuracy in emotion recognition, but the acquisition of multi-scale features of human faces still needs to be further optimized, and the recognition accuracy also needs to be further improved. Linlin Gong et al. (2023) used the CNN-Transformer network structure to realize emotion recognition, which can effectively integrate the key spatial, spectral, and temporal information of electroencephalogram signals and complete emotion recognition with high precision. However, this study cannot suppress the brain electrical channels unrelated to the task to obtain higher quality electroencephalogram time-frequency features and improve recognition accuracy. Z. Li et al. (2019) used a residual network structure with SVM for speech emotion analysis, which can achieve higher accuracy. However, this study did not extract deeper speech features and the algorithm model is not lightweight. Huang Ying et al. (Computer Application, 2022) used a variable weight decision fusion algorithm for multi-modal fusion recognition. After the three channels pass through the full connection layer, the posterior probability is obtained through SoftMax, and W f , W s , W g are assigned weights W f +W s +W g =1, and then the weighted probability of fusion is used for discriminant classification. Here, W f , W s , W g are no longer fixed values, but take a variable weight strategy, satisfying W f +W s +W gThe optimal weight is automatically found under the condition that =1, the fusion of the three channels is realized, and emotion recognition is realized. However, the skeleton mode used in the research has low recognition accuracy for emotion, which affects the accuracy of the algorithm after fusion. SUMMARY

[0010] In view of the above problems in the prior art, the purpose of the present application is to provide a miner emotion recognition method based on expression, electroencephalogram and speech multi-modal fusion.

[0011] To achieve the above purpose, the present application adopts the following technical solutions:

[0012] A miner emotion recognition method based on expression, electroencephalogram and speech multi-modal fusion, comprising:

[0013] Improved multi-scale facial emotion recognition network model under improved backbone feature extraction network: the improved multi-scale emotion recognition model under the improved backbone feature extraction network comprises an improved backbone feature extraction network, four convolutional layers, two maximum pooling layers and two multi-scale feature extraction layers. After image input, the image is sequentially subjected to the improved backbone feature extraction network, the multi-scale feature extraction layer, the convolutional layer, the maximum pooling layer, the multi-scale feature extraction layer, the convolutional layer, the maximum pooling layer and the convolutional layer. Finally, a global average pooling layer is used to output features and a SoftMax function is used as a classifier to obtain facial emotion information.

[0014] EEG enhanced emotion recognition network based on feature enhancement and attention mechanism of Transformer: the EEG enhanced emotion recognition network based on feature enhancement and attention mechanism of Transformer comprises an automatic class time-frequency feature extraction module, a Transformer feature enhancement module, a deep feature transformation convolutional module and a feature fusion and classification module of attention layer. In the automatic class time-frequency feature extraction module, each EEG channel is independently assigned a scaling convolutional layer to extract the class time-frequency features of the channel. The class time-frequency feature maps of all channels are stacked in the EEG channel dimension to obtain a class time-frequency feature tensor. Then, the feature is strengthened through the Transformer feature enhancement module. Then, the EEG feature enhancement re-calibration feature is obtained by weighting and multiplying the class time-frequency features. Then, the deep feature transformation convolutional module is used to extract the deep information of the EEG signal. Finally, the feature fusion layer is performed by the attention layer. Finally, a fully connected layer and a Softmax activation function are connected to perform EEG emotion classification, and EEG emotion information is obtained.

[0015] The speech emotion recognition model based on the lightweight deep separable convolution residual neural network comprises a parallel convolution structure, a residual structure and a serial convolution structure; the parallel convolution structure comprises three parallel DSC convolution layers, and outputs of the three parallel DSC convolution layers are combined together and sent to the residual structure of the model; the residual structure comprises two DSC convolution layers in the main part; the serial convolution structure comprises four consecutive DSC convolution layers; finally, the network is set as a speech emotion classification task model by using a discrete emotion model, and speech emotion information is obtained through a Softmax layer output.

[0016] The multi-modal information fusion method adopts a multi-modal information weight adaptive decision layer information fusion algorithm to realize fusion and complementation of multi-modal information of electroencephalogram emotion information, speech emotion information and face emotion information.

[0017] Further, in the speech emotion recognition model based on the lightweight deep separable convolution residual neural network, a batch normalization layer (BN), a linear rectifier function ReLU activation layer and a pooling layer are connected after all the DSC convolution layers except the serial convolution structure part. For the selection of a specific pooling method, all the pooling methods adopt average pooling except that global average pooling is adopted at the end of the serial convolution structure of the model.

[0018] Further, the multi-scale feature extraction layer is divided into two parts, one is self-downward feature extraction from the bottom layer to the top layer, and the other is self-upward feature extraction from the top layer to the bottom layer; the self-downward feature extraction is performed through traditional convolution and pooling; the image features input through the improved backbone feature extraction network; when the top layer feature is reached, the self-upward part of the second channel is entered, the size of the feature map is expanded through the deconvolution operation, and then adjacent feature maps are fused; 1*1 convolution connection is used between each level feature layer, interpolation method is used for up-sampling operation to realize multi-scale feature extraction network to extract semantic information of high-level features and position information of bottom layer; at the same time, the high and bottom layer features are completely fused by using lateral connection; finally, the fused features are sent to the next stage of the network model through the merging layer.

[0019] Further, the cross-entropy loss function is introduced in the improved backbone feature extraction network multi-scale emotion recognition model algorithm, and the expression is as follows:

[0020]

[0021]

[0022] In the formula, S jis the j value of SoftMax output vector S, which represents the probability that the data is j class occurrence, range [1, T]; y j is the real label, which represents the probability that the sample belongs to each class, a j is the j element in the input vector a, a k is the k-th vector in the input vector a;

[0023] The cross-entropy loss function combined with the triplet loss function is used as the total function of miner face emotion recognition, and the following formula is obtained:

[0024] L = L loss + L c

[0025] Among them, the triplet loss function L loss It is suitable for expanding the distance between different categories of miner face feature vectors in Euclidean space and reducing the distance between face feature vectors of the same category in Euclidean space. Therefore, through the multi-scale feature extraction network, the multi-scale features of the target can be well extracted for learning, and the overall recognition accuracy of the model is improved.

[0026] Preferably, the automatic class time-frequency feature extraction module is composed of 32 independent scaling convolution layers.

[0027] Preferably, the Transformer feature enhancement module is composed of multi-head attention, feedforward neural network in four groups of Transformer model, and an additional average pooling layer and a fully connected layer.

[0028] Preferably, the deep feature transformation convolution module is composed of three two-dimensional convolution neural network layers.

[0029] Preferably, in the parallel convolution structure, the number of convolution kernels of the three parallel DSC convolution layers is set to 16, and the difference lies in that the sizes of their convolution kernels are 3x3, 13x1 and 1x10 respectively; In the residual structure, two DSC convolution layers, each convolution layer has 64 convolution kernels with a size of 3x3; In the serial convolution structure, the kernel size of the four consecutive DSC convolution layers is 3x3, and the kernel number is 128, 160, 256 and 300 in turn.

[0030] Further, the multi-modal information weight adaptive algorithm has the following specific strategies:

[0031] Step 1: Extract the features of different modal information, find the most suitable feature extractor for the current information feature, and the classifier obtains n recognition result probability matrices [[w 11 ,w 12 ,...,w 1j ],[w 21 ,w22 ..., w 2j ],..., [w n1 ,w n2 ..., w nj ]], wherein w nj represents the probability of the nth sample belonging to the jth class;

[0032] Step 2: Establish an initial target weight matrix w = [w1, w2,..., w n ] and an action state selection matrix A = [-Δw, Δw]. Wherein the weight w1 corresponds to the result probability matrix of modal information 1 [w 11 ,w 12 ..., w 1j ]; the weight w2 corresponds to the result probability matrix of modal information 2 [w 21 ,w 22 ..., w 2j ], and so on; and Δw is the action change amplitude value of the agent;

[0033] Step 3: Establish a Q table, and at the same time, establish a loss function loss and a reward and punishment function R

[0034]

[0035] y'(t) = w1y'1 + w2y'2 +... + w n y' n

[0036] Wherein, y(t) is the true value; y'(t) is the multi-modal information fusion determination value; N is the number of input data points; and the loss function R is as shown in the formula:

[0037]

[0038] Wherein, loss m is the loss value of the mth sample;

[0039] Step 4: According to the Q table of the current state, update the action selection based on the ε-greedy mechanism; wherein the action selection method is as shown in the formula:

[0040]

[0041] Wherein, is the selected action when the maximum Q value in the reward value Q table is taken, a random is a random action selection value, a random ∈(0,1);

[0042] Step 5: Update the Q table by using the time difference method, and the calculation formula of the value function is as shown in the formula:

[0043] V(s)←V(s)+α(R t+1 +γV(s')-V(s))

[0044] Among them, R t+1 +γV(s') is called the TD target, R t+1 +γV(s')-V(s) represents the TD deviation;

[0045] The update method for table Q is shown in the following formula:

[0046] Q(s,a)←Q(s,a)+α[γ+λMax a' Q(s',a')]

[0047] Where α is the learning rate, λ is the reward decay coefficient, and the maximum value of the next state is λMax. a' Q(s',a') is the Q reality, and Q(s,a) in the past Q table is the Q estimate;

[0048] Step 6: Repeat the above steps iteratively until the optimal reward Q value is obtained, and obtain the corresponding weight matrix w = [w1, w2, ..., w n The optimal weights for multimodal information are adaptively determined, and the final multimodal information fusion formula is as follows:

[0049] y = y1w1 + y2w2 + ... + y n w n

[0050] Among them, y n Representing different modes, weight w n The probability matrices representing each modality.

[0051] Preferably, in step 2, Δw = 0.001.

[0052] Compared with the prior art, this application has the following beneficial effects:

[0053] (1) The facial emotion recognition method summarized in this application is based on the extraction of facial expression features from miners, and then the facial expression feature data is used to achieve the purpose of facial emotion recognition through a recognition model. There is a direct connection between mental state information and different modalities of the human body. Judging mental state by integrating different modalities has better authenticity. Supplementing mental state recognition by multimodal information fusion can improve the accuracy of mental state recognition. For example, before miners go down the mine, the mental state of coal miners can be identified and judged in real time using the method of this application, thereby ensuring the working mental state of coal miners, preventing accidents in a timely manner, and taking precautions before they happen. From the perspective of related work, such as safe mining in coal mines, this has certain practical significance.

[0054] (2) The Inception-ResNet multi-level deep convolutional neural network is used as the backbone feature extraction network in the application, which can better utilize the resources inside the network. The model allows the depth and breadth of the network to be increased while keeping the computational complexity unchanged. In addition, the cross-entropy loss function is introduced in the algorithm, which is used for classifying facial expression categories and assisting the convergence of the triplet loss function, solving the problem of model convergence difficulty. Through the multi-scale feature extraction network, the multi-scale features of the target can be well extracted for learning, and the overall recognition accuracy of the model is improved.

[0055] (3) The application constructs an electroencephalogram enhanced emotion recognition model based on Transformer and attention mechanism to enhance the feature learning ability of the classic deep learning model. The electroencephalogram channels related to the electroencephalogram emotion recognition task are enhanced, while the electroencephalogram channels unrelated to the task are inhibited, obtaining higher quality electroencephalogram time-frequency features, thereby improving the accuracy of multi-channel electroencephalogram emotion recognition.

[0056] (4) The application obtains deeper features by performing log-mel spectrogram feature extraction on speech data, and proposes a lightweight depth separable convolution residual neural network model in the aspect of speech emotion recognition algorithm. The algorithm uses the DSC algorithm with fewer parameters to improve the residual network, making the algorithm more lightweight and improving the performance.

[0057] (5) The application selects decision layer information fusion when fusing multi-modal information. Its advantage lies in the independence between different modal information, and the classification model after fusion comes from the classifier information of different modal information, avoiding the accumulation of error information of different modal information classifiers, and using three modalities that have reached high accuracy in single-modal emotion recognition, so that the recognition accuracy of the fusion algorithm is further improved. BRIEF DESCRIPTION OF DRAWINGS

[0058] Other features, objects and advantages of the application will become more apparent after reading the following detailed description of non-limiting embodiments, made with reference to the following drawings:

[0059] Figure 1 is a method flowchart of the application;

[0060] Figure 2 is a pyramid feature extraction network model;

[0061] Figure 3 is an improved multi-scale emotion recognition model of the backbone feature extraction network;

[0062] Figure 4 is a backbone feature extraction network;

[0063] Figure 5 Improvement diagram of each module of the backbone feature extraction network;

[0064] Figure 6 Flow chart of EEG emotion recognition combined with Transformer and attention mechanism feature enhancement;

[0065] Figure 7 EEG emotion recognition network structure based on Transformer feature enhancement and attention mechanism;

[0066] Figure 8 Lightweight depth separable convolution residual neural network model;

[0067] Figure 9 Speech modal emotion recognition flow chart;

[0068] Figure 10 Multi-modal information weight adaptive algorithm;

[0069] Figure 11 Face emotion state recognition accuracy under different models;

[0070] Figure 12 Recognition accuracy of different models;

[0071] Figure 13 Training loss curve and accuracy curve under different models;

[0072] Figure 14 Speech emotion state recognition accuracy under different models

[0073] Figure 15 Emotion recognition results under different modal information. DETAILED DESCRIPTION

[0074] The present application will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made. These all belong to the protection scope of the present application.

[0075] Example 1

[0076] As shown in Figure 1 A miner emotion recognition method based on expression, EEG and speech multi-modal fusion, comprising:

[0077] Improved multi-scale face emotion recognition network model under the backbone feature extraction network:

[0078] When extracting features through the backbone feature extraction network, the depth of the extracted features increases with the depth of the network. However, with the increase of the depth, some deep features contain essential physical semantic feature information but lose the position information of the features; while the shallow features contain rich position feature information but lack deep semantic feature information. Therefore, the application proposes a multi-scale feature extraction network, which is divided into two parts, one is feature extraction from bottom to top, and the other is feature extraction from top to bottom. It first performs traditional convolution, pooling and bottom-up feature convolution. The image features input through the improved backbone feature extraction network. When the top feature is reached, the second channel of the top-down part is entered, which uses the deconvolution operation to expand the size of the feature map, and then fuses the adjacent feature maps. Each level of feature layer uses 1*1 convolution connection, and uses interpolation method for up-sampling operation to realize multi-scale feature extraction network to extract semantic information of high-level features and position information of bottom-level features. At the same time, the high and bottom features are completely fused by using lateral connection, so that the up-sampling can ensure that the feature resolution of the current layer and the last layer is consistent, so as to realize the superposition and fusion of different features. Finally, the fused features are sent to the next stage of the network model through the merging layer. According to the above theoretical knowledge, the application proposes a cross-entropy loss function, which can improve the convergence rate of the network under training, and the loss function expression is as follows:

[0079]

[0080]

[0081] In the formula, S j is the j value of the SoftMax output vector S, which represents the probability of the data occurring in the j class, the range is [1, T], y j is the true label, which represents the probability of the sample belonging to each class, a j is the j elements in the input vector a, a k is the k-th vector in the input vector a.

[0082] The application uses the cross-entropy loss function combined with the triplet loss function as the total function of the miner face emotion recognition, and the following formula is obtained:

[0083] L = L loss + L c

[0084] Wherein, the triplet loss function L lossThe method is suitable for expanding the distance between the feature vectors of different categories of miners in the Euclidean space and reducing the distance between the feature vectors of the same category in the Euclidean space. Therefore, the multi-scale feature extraction network can well extract the multi-scale features of the target for learning, and improve the overall recognition accuracy of the model. The pyramid feature extraction network model is as shown in Figure 2 .

[0085] The improved multi-scale face emotion recognition network model based on the improved backbone feature extraction network includes an improved backbone feature extraction network, four convolutional layers, two maximum pooling layers and two multi-scale feature extraction modules. The structure of the improved multi-scale emotion recognition model based on the improved backbone feature extraction network is as shown in Figure 3 .

[0086] As shown in Figure 4 , the backbone feature extraction network refers to an Inception-ResNet multi-cascading deep convolutional neural network, which is mainly improved as follows:

[0087] Step 1: The Inception-ResNet-A module is improved and changed, and the convolution kernel size at the top of the figure is changed from 1x1x256 to 1x1x384, and the remaining part of the model remains unchanged. Then, the Inception-ResNet-B block is improved, the organization of the model is maintained unchanged, and the convolution kernel size on the left side of the figure is changed from 1x1x128 to 1x1x192, and the three convolution kernels from bottom to top are changed to 1x1x128, 1x7x160 and 7x1x192. Finally, the convolution kernel size at the top of the figure after the fusion of the previous two steps is changed to 1x1x1154.

[0088] Step 2: Reduction-A module improvement, in this module, the 3x3 convolution kernel, from right to left, from top to bottom, the convolution dimension is changed from 256, 256 and 256 to 320, 288 and 288.

[0089] Inception-ResNet-C module improvement, the convolution kernel size on the right side of the figure is changed from bottom to top to 1x1x192, 1x3x224 and 3x1x256, and then the size of the uppermost convolution kernel is changed to 1x1x2048. The network structure after improvement of different modules of the network is shown in Figure 5 .

[0090] In order to extract the feature of the face emotion image information, the model of the application first extracts the face modal feature information through the improved main feature extraction network. The main feature of this module is that it can better utilize the resources within the network, increase the depth and breadth of the network, and keep the model calculation unchanged. In order to extract the feature information of the face modal information at the multi-scale deep level, two multi-scale feature extraction layers are introduced, which are improvements of the traditional pyramid feature extraction network model. Therefore, the extraction of feature information at different sizes can take into account the position information at the bottom layer and the physical semantic information at the high layer by fusing the bottom layer feature information and the top layer feature information through the pyramid feature extraction network model to realize the multi-scale feature extraction of the target.

[0091] Through the fusion learning of the high and low layer features, the recognition effect is greatly improved. A max pooling layer is introduced after each convolution layer to reduce the data dimension and compress the features, thereby reducing the complexity of network model learning and training. Finally, a global average pooling layer is used to output the features, and a SoftMax function is used as a classifier to obtain the classification recognition result of the emotion. The accuracy of the miner emotion recognition is effectively improved.

[0092] EEG emotion recognition model based on Transformer and attention mechanism:

[0093] As shown in Figure 6 Fig. (a) is a detailed network structure diagram, and Fig. (b) is a wrapped network structure diagram. The improved EEG emotion recognition model based on Transformer feature enhancement and attention mechanism has an input of multi-channel original EEG emotion signals and an output of positive emotion, neutral emotion or negative emotion. The main components of the network include an automatic time-frequency feature extraction module, a Transformer feature enhancement module, a deep feature transformation module, and a feature fusion and classification module based on attention mechanism. Figure 7 As shown in Fig. (a), the EEG emotion recognition network structure based on Transformer feature enhancement and attention mechanism is an end-to-end emotion recognition network. The automatic time-frequency feature extraction module is composed of 32 independent scaling convolution layers. Each EEG channel is independently assigned a scaling convolution layer to extract the time-frequency feature of the channel. The time-frequency feature maps of all channels are stacked in the EEG channel dimension to obtain a time-frequency feature tensor. Then, the feature is strengthened through the Transformer feature enhancement module. The Transformer feature enhancement module is composed of multiple attention in four groups of Transformer models, a feedforward neural network, an additional average pooling layer and a fully connected layer. Then, the EEG feature enhancement re-calibration feature is obtained by weighting and multiplying the time-frequency feature.

[0094] After the deep feature transformation module extracts the deep information of the electroencephalogram signal, the deep feature transformation module is composed of three two-dimensional convolutional neural network layers. Finally, the attention layer is used for feature fusion layer, and finally a fully connected layer and a Softmax activation function are connected for electroencephalogram emotion classification.

[0095] The electroencephalogram emotion recognition model based on the Transformer and attention mechanism fuses the time and frequency information of the multi-channel original electroencephalogram signal, automatically extracts the time-frequency features of the electroencephalogram emotion signal by using the scaling convolution layer, enhances the electroencephalogram time-frequency features by using the Transformer, and simultaneously suppresses the electroencephalogram channels irrelevant to the electroencephalogram emotion recognition task, thereby effectively improving the accuracy of electroencephalogram emotion recognition.

[0096] The speech emotion recognition model based on the lightweight deep separable convolution residual neural network:

[0097] The lightweight deep separable convolution residual neural network model built in the application mainly consists of a parallel convolution part, a residual structure part and a serial convolution part. In the research on the convolution layer, it is found that the parameter amount of DSC is less than that of the traditional convolution, and the success of Xception proves the superiority of deep separable convolution compared with traditional convolution. Therefore, the application will use it to design the network model proposed by us. As shown in Figure 8 The lightweight deep separable convolution residual neural network model for speech emotion recognition is shown. The first part of the model is a parallel convolution structure, which contains three parallel DSC convolution layers, and the number of convolution kernels of the three convolution layers is set to 16, and the difference lies in that the sizes of their convolution kernels are 3x3, 13x1 and 1x10 respectively. The outputs will be merged together and sent to the second part of the model. The second part of the model adopts the residual structure idea, and the main stem contains two DSC convolution layers, each of which has 64 convolution kernels with a size of 3x3. The third part of the model is four consecutive DSC convolution layers, and the kernel size is 3x3, and the kernel number is 128, 160, 256 and 300 in turn. It should be noted that, except for the third part, a batch normalization layer (batch normalization, BN), a linear rectifier function ReLU activation layer and a pooling layer are connected after all DSC convolution layers. For the selection of specific pooling method, in addition to the global average pooling (Global Average-Pooling, GAPool) used at the end of the third part of the model, all pooling methods use average pooling (Average Pooling, AvgPool). The last part of the model, according to the label type of the training sample, the application uses a discrete model to set the network as a speech emotion classification task model, and the probability of each speech emotion is obtained by outputting the Softmax layer.

[0098] As Figure 9 shown is a speech modal emotion recognition flow chart, after the audio file is preprocessed, log-mel feature extraction, speech emotion recognition model of lightweight depth separable convolution residual neural network, the speech modal emotion recognition is obtained.

[0099] The multi-modal information fusion method comprises the following steps:

[0100] In order to have more accurate recognition results on the emotional state of the miner, the EEG emotional information, the speech emotional information and the facial emotional information are fused and complemented. The application designs a multi-modal information weight adaptive decision layer information fusion algorithm. The algorithm realizes the weighted fusion of the EEG information decision result, the speech information decision result and the facial emotional decision result in the decision layer to realize the fusion judgment of the multi-modal information. The structure of the adaptive weight optimization algorithm is to find the best fusion weight to realize the weighted fusion of the multi-modal information. The main idea of the algorithm comes from the interaction between the agent and the environment in reinforcement learning to learn the optimal strategy and obtain the optimal solution. Reinforcement learning learns problems in the interaction process of the agent and the environment, and constantly tries and errors to obtain rewards and punishments in the next action process. In the process of obtaining the maximum reward value, the selection of the action is optimized and improved, so that the final target requirement is effectively realized. At the same time, this process does not need to establish a complex environment model. The multi-modal information weight adaptive decision layer information fusion algorithm seeks the optimal weight proportion between the multi-modal information by the learning method in reinforcement learning. The algorithm iteration update is as shown in Figure 10 .

[0101] The multi-modal information weight adaptive algorithm mainly uses the value set jointly established by the target weight matrix and the action state matrix, and updates the table optimization through the reward function in reinforcement learning and the interaction with the external environment, so as to obtain the best weight matrix to realize the fusion strategy of the multi-modal information under the optimal weight. The specific strategy of the algorithm is as follows:

[0102] Step 1: extract the features of different modal information, seek the most adaptive feature extractor of the current information features, and obtain n kinds of recognition result probability matrices [[w 11 ,w 12 ,...,w 1j ],[w 21 ,w 22 ,...,w 2j ],...,[w n1 ,w n2 ,...,w nj ]] by the classifier, wherein w nj represents the probability that the nth sample belongs to the jth category;

[0103] Step 2: Establish the initial target weight matrix w = [w1, w2, ..., w n The action state selection matrix A = [-Δw, Δw]. Where the weight w1 corresponds to the modal information 1 result probability matrix [w 11 ,w 12 ,...,w 1j The probability matrix of the result of modality information 2 corresponding to weight w2 is [w 21 ,w 22 ,...,w 2j And so on. Δw is the magnitude of the change in the agent's action; in this application, Δw = 0.001 is chosen.

[0104] Step 3: Create the Q table, and simultaneously create the loss function loss and the reward / penalty function R.

[0105]

[0106] y'(t)=w1y'1+w2y'2+...+w n y' n

[0107] Where y(t) is the true value; y'(t) is the multimodal information fusion judgment value; and N is the number of input data points. The loss function R is shown in the formula:

[0108]

[0109] Where, loss m Let be the loss value for the m-th sample;

[0110] Step 4: Based on the Q-table of the current state, select the update action using an ε-greedy mechanism. The action selection method is shown in the formula:

[0111]

[0112] in, The action to choose when taking the maximum Q value from the reward value Q table, a random A value is selected for the random action, a random ∈(0,1).

[0113] Step 5: Update the Q-table using the time difference method. The formula for calculating the value function is shown in the formula below:

[0114] V(s)←V(s)+α(R t+1 +γV(s')-V(s))

[0115] Among them, R t+1 +γV(s') is called the TD target, R t+1+ γV(s') - V(s) is the TD error. The updating of Q table is shown in the formula:

[0116] Q(s, a)←Q(s, a) + α[γ+λMax a' Q(s',a')

[0117] Wherein, α is the learning rate, λ is the reward decay coefficient. The maximum of next state λMax a' Q(s',a') as Q reality, past Q table Q(s,a) as Q estimate.

[0118] Step 6: repeat the above steps until the optimal reward Q value is taken, and the corresponding weight matrix w = [w1, w2,..., wn] is obtained. n ] is the adaptive optimal weight of multi-modal information, and the final multi-modal information fusion formula is:

[0119] y = y1w1 + y2w2 +... + y n w n

[0120] Wherein, y n represents different kinds of modalities, and the weight w n represents the probability matrix of each modality

[0121] In the present application, preprocessing, feature extraction and emotion state recognition are carried out under three modalities of electroencephalogram, face and voice. The three modalities of information are classified by the classifier to realize the classification and recognition of positive, neutral and negative emotions. Finally, the weight values of different modalities of information are optimized, and the optimal weight value is used to realize the decision layer information fusion among the multi-modal information, so as to effectively reduce the limitation of single modality information.

[0122] Embodiment 2

[0123] During the process of engaging in underground operation, the coal mine accidents caused by human factors are difficult to predict, and the working emotional state of the miners will directly affect the quality of their work, and even cause safety accidents due to misoperation. In view of the problem that the three single modalities cannot accurately recognize the emotional state of the miners, the present embodiment starts from the three emotional state aspects of the electroencephalogram emotional state under the physiological state of the miners and the face and voice emotions under the non-physiological state, and studies the problem of emotional state evaluation of the miners.

[0124] Improved multi-scale face emotion recognition network model under the backbone feature extraction network:

[0125] In order to verify the performance of the improved deep learning model proposed in this application, the improved face recognition algorithm is constructed using the TensorFlow2.1 framework deep learning space, using Window10 as the operating system, and using Python3.7 as the programming tool according to the actual needs.

[0126] In order to monitor and identify the emotional state of miners, an improved backbone feature extraction network multi-scale face emotion recognition model is built. At the same time, in order to adapt to the special environment of coal mine, a face emotion dataset suitable for monitoring the emotions of miners is constructed. The model can recognize 7 emotions: happy, angry, disgusted, neutral, sad, surprised and fearful. In order to evaluate the recognition accuracy and performance of the model, the improved backbone feature extraction network multi-scale recognition network structure and the traditional VGG-Net, Inception and other models are compared on the dataset constructed in this application. The test accuracy of the improved backbone feature extraction network multi-scale emotion recognition model on the 7 categories of happy, angry, disgusted, neutral, sad, surprised and fearful is as shown in Figure 11

[0127] Through the analysis of the accuracy of the model, it can be seen that the recognition accuracy of the model for emotions such as surprise, happiness and sadness is relatively high, which can reach 90%, 86% and 81% respectively. The recognition accuracy of the model for emotions such as disgust and fear is relatively low. This problem is mainly because the facial expression features of happy, surprised and other emotions are more obvious, so the feature extraction of the model for these emotions is also more accurate, and conversely, the recognition accuracy of disgust and fear is relatively low.

[0128] EEG emotion recognition model based on Transformer and attention mechanism:

[0129] In order to identify the EEG emotions of miners, the DEAP public dataset is used to evaluate the classification performance of the EEG emotion recognition network based on Transformer feature enhancement and attention mechanism. The DEAP dataset can be used to study human EEG emotions. The DEAP dataset only has 3 channels for recording EEG signals, which contains the ratings of subjects on emotional videos according to valence and arousal, and the corresponding emotion labels are marked on the recorded EEG signals according to these ratings.

[0130] ​The experiment selects the 60s of the EEG signal of the subject after removing the baseline as the experimental data, and divides it into 20 groups of EEG data, so that each subject can obtain 800 groups of EEG data samples, and 32 subjects can obtain 25600 groups of EEG data. Ten-fold cross-validation is used to train and test the model. The experiment is configured as: 1080Ti GPU, Intel i7-8700K CPU, TensorFlow framework, and Adam optimizer is used to optimize the edge loss.

[0131] The model proposed in the present application can classify EEG signals into positive, neutral and negative emotions. In order to evaluate the recognition accuracy and performance of the model, the improved EEG emotion recognition model based on Transformer feature enhancement and attention mechanism is compared with Convolutional Neural Network (CNN), Deep Separable Convolution DSC and Graph Convolutional Neural Network (GCNN) as a contrast model for experiment. The recognition accuracy of the improved model and the other four recognition models for the emotional state is as shown in Figure 12

[0132] The emotion recognition model proposed in the present application basically completes the three emotion recognition tasks. Among them, the recognition accuracy of positive and negative emotions is high, and the average recognition accuracy is 89.73% and 88.68% respectively, and the recognition accuracy of neutral emotion is relatively low, and the average recognition accuracy is 87.43%.

[0133] As shown in Figure 13 It can be seen that a variety of recognition models can eventually tend to be stable after a certain number of training and learning steps, which can indicate that the Transformer-based feature enhancement and attention mechanism network built for the miner EEG emotion recognition is better than other deep learning models, with higher accuracy and better performance.

[0134] Speech emotion recognition model based on lightweight deep separable convolution residual neural network:

[0135] In order to better recognize the emotional state of miners, a lightweight deep separable convolution residual neural network model for speech emotion recognition is built. The model can recognize 6 kinds of speech emotions including neutral, anger, fear, happiness, sadness and surprise. In order to evaluate the recognition accuracy and performance of the model, the model and the traditional deep learning model are compared in the speech data set in the research. The test accuracy of the model in different categories is compared with the recognition accuracy of the speech emotion state of the other three recognition models as shown in Figure 14

[0136] ​​By comparing the model accuracy under different emotional states, it can be seen that the model proposed in the application has high recognition accuracy for happiness and anger. The accuracy can reach 90.56% and 88.61% respectively. By comparing the recognition accuracy of other recognition network models, the lightweight deep separable convolution residual neural network model built in the application achieves an identification accuracy of 87.74%, which has a certain improvement in performance. The model has only a small number of parameters but can learn emotional features well, and has good effect in lightweight and achieves high accuracy.

[0137] Multi-modal adaptive fusion algorithm:

[0138] The application proposes a multi-modal adaptive weight optimization algorithm to realize the decision layer information fusion of electroencephalogram information, speech information and face modal information. Three different modal information gets separate decision information in separate network classifier, and the decision results of three modal information are weighted combined to realize information fusion, so as to improve the discrimination accuracy. The multi-modal decision adaptive fusion emotion recognition accuracy comparison result of the electroencephalogram information, speech information and face emotion information proposed in the application is shown in Figure 15 .

[0139] It can be seen from Figure 15 that under single modal information, whether it is electroencephalogram modal information in physiological state or face image and speech modal information in non-physiological state, the recognition accuracy of emotional state is not very ideal. After the multi-modal information fusion method fuses the electroencephalogram modal, speech modal and face modal information, the emotional state is recognized, and the emotional recognition accuracy is higher than that of single modal information, which also shows that the multi-modal information fusion method is feasible.

[0140] The specific embodiments of the application are described above. It should be understood that the application is not limited to the above specific embodiments, and those skilled in the art can make various modifications or changes within the scope of the claims, which does not affect the essential content of the application.

Claims

1. A method for miner emotion recognition based on multimodal fusion of facial expression, EEG, and speech, characterized in that, include: An improved backbone feature extraction network multi-scale emotion recognition model is proposed, consisting of an improved backbone feature extraction network, four convolutional layers, two max pooling layers, and two multi-scale feature extraction layers. After image input, the image passes through the improved backbone feature extraction network, multi-scale feature extraction layers, convolutional layers, max pooling layers, multi-scale feature extraction layers, convolutional layers, max pooling layers, and convolutional layers in sequence. Finally, a global average pooling layer is used to output features, and the SoftMax function is used as a classifier to obtain facial emotion information. A Transformer-based network for enhancing EEG emotion recognition and employing attention mechanisms comprises an automatic time-frequency feature extraction module, a Transformer feature enhancement module, a deep feature transformation convolution module, and an attention-based feature fusion and classification module. In the automatic time-frequency feature extraction module, each EEG channel is independently assigned a scaling convolution layer to extract the time-frequency features of that channel. The time-frequency feature maps of all channels are stacked along the EEG channel dimension to obtain a time-frequency feature tensor. Then, the Transformer feature enhancement module performs feature enhancement. Next, a weighted multiplication with the time-frequency features yields the EEG feature enhancement recalibrated features. The deep feature transformation convolution module then extracts deeper information from the EEG signals. Finally, the attention layer performs feature fusion, and a fully connected layer and a Softmax activation function are connected to classify EEG emotion information. A speech emotion recognition model based on a lightweight deep separable convolutional residual neural network includes a parallel convolutional structure, a residual structure, and a sequential convolutional structure. The parallel convolutional structure contains three parallel DSC convolutional layers, whose outputs are merged and fed into the model's residual structure. The residual structure has two DSC convolutional layers in its backbone. The sequential convolutional structure consists of four consecutive DSC convolutional layers. Finally, a discrete emotion model is used to configure the network as a speech emotion classification task model, and the speech emotion information is obtained through the output of a Softmax layer. Multimodal information fusion method: A decision-level information fusion algorithm with adaptive multimodal information weights is adopted to achieve the fusion and complementarity of multimodal information such as EEG emotion information, voice emotion information and facial emotion information.

2. The miner emotion recognition method based on multimodal fusion of facial expression, EEG, and speech as described in claim 1, characterized in that, In the speech emotion recognition model based on a lightweight deep separable convolutional residual neural network, except for the serial convolutional structure, all DSC convolutional layers are followed by batch normalization layers, linear rectified function ReLU activation layers, and pooling layers. Regarding the selection of specific pooling methods, except for the global average pooling used at the end of the serial convolutional structure of the model, all pooling methods use average pooling.

3. The miner emotion recognition method based on multimodal fusion of facial expression, EEG, and speech as described in claim 1, characterized in that, The multi-scale feature extraction layer has a structure divided into two parts: one is bottom-up feature extraction from the bottom layer to the top layer, and the other is top-down feature extraction from the top layer to the bottom layer. First, traditional bottom-up convolution with convolution and pooling is performed. Image features from the improved backbone feature extraction network are then input. Upon reaching the top-level features, the second channel's top-down portion is entered, where deconvolution is used to enlarge the feature map size, and then adjacent feature maps are fused. Each feature layer is connected by a 1*1 convolution, and upsampling is performed using interpolation to extract semantic information from high-level features and positional information from lower-level features. Simultaneously, lateral connections are used to fully fuse high- and low-level features. Finally, the fused features are fed into the next stage of the network model through a merging layer.

4. The miner emotion recognition method based on multimodal fusion of facial expression, EEG, and speech as described in claim 1, characterized in that, In the improved backbone feature extraction network multi-scale emotion recognition model algorithm, a cross-entropy loss function is introduced, with the following expression: In the formula, S j y is the j-value of the SoftMax output vector S, which represents the probability that the data belongs to class j, ranging from [1,T]; j These are the true labels, representing the probability that a sample belongs to each category. j These are the j elements in the input vector a, a k It is the k-th vector in the input vector a; Using the cross-entropy loss function combined with the triplet loss function as the overall function for miner face emotion recognition, we obtain the following formula: L=L loss +L c Among them, the triplet loss function L loss It is applicable to the distance between facial feature vectors of different categories of miners in extended Euclidean space and the distance between facial feature vectors of the same category in reduced Euclidean space.

5. The miner emotion recognition method based on multimodal fusion of facial expression, EEG, and speech as described in claim 1, characterized in that, The automatic time-frequency feature extraction module consists of 32 independent scaling convolutional layers.

6. The miner emotion recognition method based on multimodal fusion of facial expression, EEG, and speech as described in claim 1, characterized in that, The Transformer feature enhancement module consists of multi-head attention, feedforward neural networks, and an additional average pooling layer and a fully connected layer from four Transformer models.

7. The miner emotion recognition method based on multimodal fusion of facial expression, EEG, and speech as described in claim 1, characterized in that, The depth feature transformation convolution module consists of three layers of two-dimensional convolutional neural network.

8. The miner emotion recognition method based on multimodal fusion of facial expression, EEG, and speech as described in claim 1, characterized in that, In the parallel convolutional structure, the number of convolutional kernels in the three parallel DSC convolutional layers is set to 16, but their kernel sizes are 3×3, 13×1, and 1×10, respectively. In the residual structure, there are two DSC convolutional layers, each with 64 kernels of size 3×3. In the serial convolutional structure, the kernel size of the four consecutive DSC convolutional layers is 3×3, and the number of kernels is 128, 160, 256, and 300, respectively.

9. The method for miner emotion recognition based on multimodal fusion of facial expression, EEG, and speech as described in claim 1, characterized in that, The specific strategy of the multimodal information weight adaptive algorithm is as follows: Step 1: Extract features from different modalities, find the feature extractor best suited to the current information features, and the classifier obtains the probability matrix of n recognition results [[w 11 ,w 12 ,...,w 1j ],[w 21 ,w 22 ,...,w 2j ],...,[w n1 ,w n2 ,...,w nj ]], where w nj This represents the probability that the nth sample belongs to the jth category; Step 2: Establish the initial target weight matrix w = [w1, w2, ..., w n The action state selection matrix A = [-Δw, Δw]; where the weight w1 corresponds to the modal information 1 result probability matrix [w 11 ,w 12 ,...,w 1j The probability matrix of the result of modality information 2 corresponding to weight w2 is [w 21 ,w 22 ,...,w 2j ], and so on; Δw is the magnitude of the change in the agent's action; Step 3: Create the Q-table, and simultaneously create the loss function (loss) and reward / penalty function (R). y'(t)=w1y′1+w2y′2+...+w n y′ n Where y(t) is the true value; y'(t) is the multimodal information fusion judgment value; N is the number of input data points; and the loss function R is shown in the formula: loss m Let be the loss value for the m-th sample; Step 4: Based on the Q-table of the current state, select the update action using an ε-greedy mechanism; the action selection method is shown in the following formula: in, The action to choose when taking the maximum Q value from the reward value Q table, a random A value is selected for the random action, a random ∈(0,1); Step 5: Update the Q-table using the time difference method. The formula for calculating the value function is shown in the formula below: V(s)←V(s)+α(R t+1 +γV(s')-V(s)) Among them, R t+1 +γV(s') is called the TD target, R t+1 +γV(s')-V(s) represents the TD deviation; The update method for table Q is shown in the following formula: Q(s,a)←Q(s,a)+α[γ+λMax a' Q(s',a')] Where α is the learning rate, λ is the reward decay coefficient, and the maximum value of the next state is λMax. a' Q(s',a') is the Q reality, and Q(s,a) in the past Q table is the Q estimate; Step 6: Repeat the above steps iteratively until the optimal reward Q value is obtained, and obtain the corresponding weight matrix w = [w1, w2, ..., w n The optimal weights for multimodal information are adaptively determined, and the final multimodal information fusion formula is as follows: y=y1w1+y2w2+...+y n w n Among them, y n Representing different modes, weight w n The probability matrices representing each modality.

10. The miner emotion recognition method based on multimodal fusion of facial expression, EEG, and speech as described in claim 9, characterized in that, In step 2, Δw = 0.001.

Citation Information

Patent Citations

  • Multi-modal emotion recognition method and device, electronic equipment and storage medium

    CN115359576A

  • Method and apparatus for interactive monitoring of emotion during teletherapy

    US20210118323A1