Emotion recognition method and device

By collecting multiple modal data and fusion of features, multimodal fusion features are generated, the problem of insufficient single-modal data in the prior art is solved, and high-accurate emotion recognition is achieved in environments with poor light.

CN120123966APending Publication Date: 2025-06-10NANJING FORESTRY UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510137800.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing emotion recognition methods only use single-modal image data, ignore other modal data such as speech data and physiological signals. Image data acquisition is usually in an environment with sufficient light and cannot be applied to actual environments with poor light.

Method used

An emotion recognition method is adopted to collect face video data and environmental data, image data, speech data, text data, physiological signal characteristics and environmental features are extracted, and these features are spliced ​​and fused using attention mechanisms to generate multimodal fusion features, and input a pre-trained emotion detection model for emotion recognition.

Benefits of technology

Through the use of multimodal fusion features, the accuracy and robustness of emotion recognition are improved, and emotions can be effectively identified in environments with poor light, solving the problem of insufficient single-modal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123966A_ABST
    Figure CN120123966A_ABST
Patent Text Reader

Abstract

The invention provides an emotion recognition method and device. The method comprises the following steps: acquiring face video data and environment data; extracting image data, voice data and text data according to the collected face video data; extracting image features and physiological signal features from the extracted image data; extracting audio features from the extracted voice data; extracting text features from the extracted text data; environment features are extracted from the collected environment data; splicing and fusing the extracted image features, physiological signal features, audio features, text features and environment features by using an attention mechanism to generate multi-modal fusion features; and inputting the generated multi-modal fusion features into a pre-trained emotion detection model for emotion recognition, and outputting emotion contents to which the multi-modal fusion features belong. The emotion recognition method can solve the problem that an existing emotion recognition method only uses a single-mode picture and ignores other mode data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an emotion recognition method and device, belonging to the technical field of image processing. Background Art

[0002] With the progress of technology and the rapid development of the field of artificial intelligence, micro-expression recognition technology is also constantly evolving, and many research methods in the field of emotions have been successively proposed.

[0003] Emotions, as a series of complex subjective experiences, are not only affected by external factors but also play a core role in human daily life. They have a profound impact on aspects such as an individual's cognitive function, decision-making process, social interaction, and mental health. Emotion recognition technology has been widely applied in multiple fields, deeply influencing our daily life.

[0004] In the medical field, emotion recognition technology provides a new perspective for the diagnosis and treatment of mental diseases. For example, in the diagnosis of disorders of consciousness, computer-aided emotion assessment can help doctors diagnose and formulate treatment plans more accurately. In the field of distance education, emotion recognition technology identifies students' emotional responses in distance teaching and generates corresponding emotion-assisted reminders, enabling teachers to adjust teaching strategies and progress in a timely manner. In the transportation field, for occupations that require high concentration, such as astronauts, long-distance bus drivers, and pilots, negative emotions such as anger, anxiety, and sadness may affect their work performance and increase the risk of accidents. Therefore, timely monitoring of the emotional states of these personnel is of great significance for accident prevention.

[0005] In summary, the effectiveness of emotion recognition is crucial for ensuring the safety and efficiency of individuals and society.

[0006] However, in the existing emotion recognition methods, artificial feature extractors designed by experts based on experience are often used, with poor generalization and being easily affected by background images and noise; in the emotion recognition methods based on deep learning, there are often problems that the expression features of shallow convolutional neural networks are not fully extracted; in the aspect of image data selection methods in the existing technology, often only single-modal image data is used, ignoring other modal data such as voice data and physiological signals; and image data collection often focuses on environments with sufficient light and high recognition, and cannot be applied to actual environments with poor light. Summary of the Invention

[0007] The purpose of the present invention is to overcome the deficiencies in the existing technology and provide an emotion recognition method and device to solve the problem that the existing emotion recognition methods only use single-modal images and ignore other modal data.

[0008] To solve the above technical problems, the present invention is implemented by adopting the following technical solutions:

[0009] In a first aspect, the present invention provides an emotion recognition method, including:

[0010] S1: Collect face video data and environmental data;

[0011] S2: Extract image data, voice data, and text data according to the face video data collected in step S1;

[0012] S3: Extract image features from the image data extracted in step S2;

[0013] Extract physiological signal features from the image data extracted in step S2;

[0014] Extract audio features from the voice data extracted in step S2;

[0015] Extract text features from the text data extracted in step S2;

[0016] Extract environmental features from the environmental data collected in step S1;

[0017] S4: Use the attention mechanism to splice and fuse the image features, physiological signal features, audio features, text features, and environmental features extracted in step S3 to generate multi-modal fusion features;

[0018] S5: Input the multi-modal fusion features generated in step S4 into a pre-trained emotion detection model for emotion recognition, and output the emotion content to which the multi-modal fusion features belong.

[0019] For the aforementioned emotion recognition method, in step S2, extracting image data according to the face video data collected in step S1 includes:

[0020] Perform frame-by-frame processing on the face video data collected in step S1 to extract image data, extract the image data corresponding to each frame, and obtain an image frame sequence; the face video data refers to video data of a captured face.

[0021] In step S3, extracting image features from the image data extracted in step S2 includes:

[0022] T1: Perform image preprocessing on the image data extracted in step S2;

[0023] T2: Extract image features from the preprocessed image data;

[0024] Step T1 includes:

[0025] T11: Perform image enhancement processing on each image frame in the image frame sequence of step S2.

[0026] For the aforementioned emotion recognition method, the image enhancement processing for each image frame in step T11 includes:

[0027] T111: Respectively pass the image frames in the image frame sequence of step S2 through traditional processing, deep learning model processing, and large vision model processing; the traditional processing includes sequentially passing the image frames through filtering processing, block adaptive histogram equalization processing, linear function processing, gamma correction processing, and normalization processing;

[0028] T112: Merge the three images after traditional processing, deep learning model processing, and large vision model processing into one image according to a preset weight, and use it as the output image of the image enhancement processing.

[0029] For the aforementioned emotion recognition method, step T1 further includes:

[0030] T12: Calculate the structural similarity index of adjacent frames and the cosine distance of adjacent frame feature vectors after the image enhancement processing in step T11;

[0031] T13: Compare the structural similarity index calculated in step T12 with a preset structural similarity index threshold, and compare the cosine distance calculated in step T12 with a preset cosine distance threshold;

[0032] T14: Retain the adjacent frames that meet the requirements of the preset structural similarity index threshold and the preset cosine distance threshold in step T13 as the preprocessed image data.

[0033] For the aforementioned emotion recognition method, step T2 includes:

[0034] Input the preprocessed image data in step T1 into the Vision Transformer model to extract the global image features and output the global image feature vector;

[0035] Input the preprocessed image data in step T1 into the improved VGG16 model to extract the local image features and output the local image feature vector; the improved VGG16 model includes 5 sequentially connected convolutional blocks and 1 spatial-channel attention module SCA.

[0036] For the aforementioned emotion recognition method, the calculation of the spatial-channel attention module SCA includes:

[0037] Input the output feature vector of the 5th convolutional block into the spatial attention branch to obtain the spatial attention weight and into the channel attention branch to obtain the channel attention weight; the output feature vector size of the 5th convolutional block is H k *W k *C k ;

[0038] Among them,

[0039] The calculation of the spatial attention weight includes:

[0040] Perform max pooling (Maxpool) and average pooling (Avgpool) on the output feature vector of the 5th convolutional block respectively in the channel dimension, and keep the spatial dimensions H k *W k , and compress the channel dimension C k to 1; concatenate the features after max pooling (Maxpool) and average pooling (Avgpool) processing; input the concatenated features into a fully connected layer (FC) to extract features; activate the features extracted by the fully connected layer (FC) through Sigmoid to obtain the spatial attention weight;

[0041] The spatial attention weight has the following calculation formula:

[0042] ,

[0043] where is the output feature vector of the 5th convolutional block;

[0044] The calculation of the channel attention weight includes:

[0045] Perform max pooling (Maxpool) and average pooling (Avgpool) on the output feature vector of the 5th convolutional block respectively, compress the spatial dimensions H k *W k to 1, and keep the channel dimension C k ; send the features after max pooling (Maxpool) and average pooling (Avgpool) processing into a fully connected layer (FC) respectively to extract features and then add them; activate the added features through Sigmoid to obtain the channel attention weight;

[0046] The channel attention weight has the following calculation formula:

[0047] ;

[0048] According to the obtained spatial attention weight and channel attention weight, fuse the channel attention branch, the spatial attention branch, and the branch combining spatial attention and channel attention to obtain the output feature of the spatial-channel attention module (SCA);

[0049] The output feature of the spatial-channel attention module (SCA) has the following calculation formula:

[0050] .

[0051] In the foregoing emotion recognition method, in step S3, the physiological signal features include heart rate variability features;

[0052] In step S3, extracting physiological signal features from the image data extracted in step S2 includes:

[0053] Selecting the face region from the image data extracted in step S2, obtaining a three-channel time series diagram of the face region, and then using imaging photoplethysmography to analyze the color change of the skin to extract heart rate variability features;

[0054] In step S3, extracting environmental features from the environmental data collected in step S1 includes: generating environmental features by one-hot encoding the environmental data.

[0055] In the foregoing emotion recognition method, step S4 includes:

[0056] S41: Fusing the global image feature vector extracted by Vision Transformer and the local image feature vector extracted by the improved VGG16 in step S3 to obtain visual features;

[0057] S42: Fusing audio features and text features to obtain non-visual features;

[0058] S43: Concatenating and fusing visual features, non-visual features, physiological signal features, and environmental features, and after transformation by the non-linear activation RELU function, generating multi-modal fusion features.

[0059] In the foregoing emotion recognition method, in step S5, the training of the emotion detection model includes:

[0060] Training the emotion detection model with the optimization goal of minimizing the difference between the predicted emotion content obtained by emotion recognition of the sample through the emotion detection model and the actual emotion content corresponding to the sample;

[0061] Among them, the first loss function The calculation formula is:

[0062] ,

[0063] In the formula, W is the weight vector of the emotion detection model; R is the multi-modal fusion feature of the sample; bs is the batch number, and m is the number of emotion content classifications;

[0064] The second loss function The calculation formula is:

[0065] ,

[0066] In the formula, Is the feature vector of the emotion content predicted by the vision large model; is the feature vector of the actual emotional content of the sample; is a function representing the distance between two vectors. When and are the same, the function value is 0. When and are different, the function value is 1; are the model parameters of the emotion detection model;

[0067] Overall loss function The calculation formula of is:

[0068] ,

[0069] In the formula, is the loss coefficient preset by the visual large model.

[0070] In a second aspect, the present invention provides an emotion recognition device, including: an acquisition module, a data extraction module, a feature extraction module, a feature fusion module, and an emotion recognition module;

[0071] The acquisition module is used to: acquire face video data and environmental data;

[0072] The data extraction module is used to: extract image data, voice data, and text data according to the face video data acquired by the acquisition module;

[0073] The feature extraction module is used to:

[0074] Extract image features from the image data extracted by the data extraction module;

[0075] Extract physiological signal features from the image data extracted by the data extraction module;

[0076] Extract audio features from the voice data extracted by the data extraction module;

[0077] Extract text features from the text data extracted by the data extraction module;

[0078] Extract environmental features from the environmental data acquired by the acquisition module;

[0079] The feature fusion module is used to: use the attention mechanism to splice and fuse the image features, physiological signal features, audio features, text features, and environmental features extracted by the feature extraction module to generate multi-modal fusion features;

[0080] The emotion recognition module is used to: input the multi-modal fusion features generated by the feature fusion module into a pre-trained emotion detection model for emotion recognition, and output the emotion content to which the multi-modal fusion features belong.

[0081] Beneficial effects:

[0082] The emotion recognition method of the present invention splices and fuses the extracted image features, physiological signal features, audio features, text features, and environmental features to generate multi-modal fusion features, and uses the generated multi-modal fusion features for emotion recognition, solving the problem of existing emotion recognition methods that only use single-modal pictures and ignore other modal data.

[0083] The present invention designs an image preprocessing method to improve the image quality and solve the problem of poor emotion recognition effect caused by unclear images in actual application scenarios. The present invention respectively adopts three image enhancement methods for an image to improve the image quality, and fuses the images output by the three enhancement methods to generate an image, making up for the deficiencies of a single image enhancement method and effectively improving the image quality after enhancement processing; similarity screening is performed on the images after image enhancement processing, images with large similarities are deleted, and image frames with small similarities are retained, reducing the number of input image frames while retaining sufficient image frame information and improving the model operation speed.

[0084] The present invention uses the ViT model to capture global dependencies and can understand the connections between different regions in the image; an improved VGG16 model is used to capture local information such as edges and textures in the image, and the problem of insufficient extraction of expression features is reasonably solved by fully considering multi-scale features;

[0085] The present invention designs an improved VGG16 model, adds 1 spatial-channel attention module SCA, deletes the fully connected layer, and focuses on the extraction of local important features. The output features of the spatial-channel attention module SCA are used as the local feature vectors of the image output by the VGG16 model, making the improved VGG16 model focus on the extraction of local image features.

[0086] By utilizing and fusing different modal information, data features are fully extracted, complementary feature extraction is effectively carried out, the robustness to perturbations such as image deformation and illumination changes is improved, and at the same time, the robustness and generalization ability of the recognition ability are further enhanced through the supplement of the vision large model and background feature vectors. Brief description of the drawings

[0087] Figure 1 is the flowchart of the emotion recognition method in Embodiment 1 of the present invention;

[0088] Figure 2 is the flowchart of calculating the output feature vector of the spatial-channel attention module in the improved VGG model in Embodiment 1 of the present invention;

[0089] Figure 3 is the flowchart of calculating visual features by the dual-modal feature fusion module FFM in Embodiment 1 of the present invention. Detailed implementation manners

[0090] It should be noted that: The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific features in the embodiments of the present application and the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.

[0091] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0092] Embodiment 1

[0093] Figure 1 is a flowchart of the emotion recognition method in Embodiment 1.

[0094] As Figure 1 shown, this embodiment provides an emotion recognition method, including the following steps:

[0095] S1: Collect face video data and environmental data;

[0096] S2: Extract image data, voice data, and text data according to the face video data collected in step S1;

[0097] S3: Extract image features from the image data extracted in step S2;

[0098] Extract physiological signal features from the image data extracted in step S2;

[0099] Extract audio features from the voice data extracted in step S2;

[0100] Extract text features from the text data extracted in step S2;

[0101] Extract environmental features from the environmental data collected in step S1;

[0102] S4: Use the attention mechanism to splice and fuse the image features, physiological signal features, audio features, text features, and environmental features extracted in step S3 to generate multi-modal fusion features;

[0103] S5: Input the multi-modal fusion features generated in step S4 into a pre-trained emotion detection model for emotion recognition, and output the emotion content to which the multi-modal fusion features belong.

[0104] In this embodiment, the emotional content includes anger, fear, sadness, disgust, surprise, happiness, and neutrality.

[0105] The emotional recognition method of the present invention generates multi-modal fusion features by fusing visual features, non-visual features, physiological signal features, and environmental features, improving the comprehensiveness and accuracy of feature extraction; and performs emotional recognition based on the multi-modal fusion features, thereby improving the accuracy of emotional recognition.

[0106] In step S1, the environmental data includes the data information of the environment where the data is collected; when the present invention is used for driver emotional recognition, the environmental data includes weather data and traffic data; the weather conditions and traffic data are obtained through the Internet; the weather data includes the weather conditions at the time of data collection, which can be obtained through networking, such as sunny, cloudy, rainy, snowy, etc.; the traffic data includes traffic flow data, traffic congestion index, and accident conditions around the driving vehicle, which can be obtained through networking.

[0107] Step S1 includes: using a CCD camera to collect face videos. A CCD camera is a device that uses semiconductor devices to capture optical signals and convert them into digital images. The CCD camera has the characteristics of high resolution and high sensitivity. In this embodiment, a CCD camera with a frame rate (FPS) of 100 frames per second is used.

[0108] Step S2 includes:

[0109] Performing frame processing on the face video data collected in step S1 to extract image data, extracting the image data corresponding to each frame to obtain an image frame sequence;

[0110] Separating voice data and text data from the face video data collected in step S1.

[0111] Step S2 extracts image data, voice data, and text data of the same time period and duration from the same face video data.

[0112] Step S3 extracts various features from the data extracted in step S2. The process of extracting various features is executed in parallel, and there is no obvious order. The step markers in the process of extracting each feature are only used for the order division of the extraction process of each feature itself, and should not be regarded as the extraction order division between features.

[0113] In step S3, extracting image features from the image data extracted in step S2 includes:

[0114] T1: Performing image preprocessing on the image data extracted in step S2;

[0115] T2: Extracting image features from the preprocessed image data;

[0116] Step T1 includes:

[0117] T11: Separately perform image enhancement processing on each image frame in the image frame sequence of step S2;

[0118] T12: Calculate the similarity between adjacent frames after image enhancement processing;

[0119] T13: Judge the change situation of the image by comparing the similarity between adjacent frames;

[0120] T14: Retain the image frames whose change situation meets the preset requirements as the preprocessed image data;

[0121] Adjacent frames refer to two adjacent image frames in the image frame sequence of face video data. If the frame rate of face video data is 30 frames per second (fps), then there will be 30 image frames in the face video data per second, and each image frame and the previous / next frame of this image frame are adjacent frames. In step T14, the image frames that meet the preset requirements refer to the image frames with large face change amplitudes. Retaining the images with large change amplitudes for subsequent processing reduces the number of image frames for subsequent processing and is conducive to improving the efficiency of subsequent processing.

[0122] The image enhancement processing of each image frame in step T11 includes:

[0123] T111: Respectively pass the image frames in the image frame sequence of step S2 through traditional processing, deep learning model processing, and large vision model processing; the traditional processing includes sequentially passing the image frames through filtering processing, block adaptive histogram equalization processing, linear function processing, gamma correction processing, and normalization processing;

[0124] T112: Merge the three images after traditional processing, deep learning model processing, and large vision model processing into one image according to preset weights as the output image of the image enhancement processing.

[0125] In this embodiment, the traditional processing in step T111 includes:

[0126] 1) Filtering processing:

[0127] Analyze the noise type of the image frames in the image frame sequence of step S21, and select the corresponding filtering method according to the noise type to remove the noise;

[0128] Specifically, it includes: classifying the noise types of image frames into salt-and-pepper noise, Gaussian noise, and other noises; the filtering process is as follows: select a 3x3 working window, slide the working window pixel by pixel on the image frame, select the corresponding filtering method according to the noise type, calculate a certain calculated value of the corresponding values of the pixel points within the 3x3 working window, and replace the value corresponding to the central pixel point of the 3x3 window with this calculated value; perform the above filtering process on each pixel in the image frame to generate a filtered image.

[0129] In practical applications, for boundary pixels, special processing (such as mirroring, padding, etc.) needs to be adopted to expand the boundaries of the image frame so that the boundary pixels can obtain the same 3x3 working window as the internal pixels of the image frame. In this embodiment, 0 is used to fill the missing values.

[0130] Salt-and-pepper noise corresponds to the median filtering method. Median filtering is to replace the value corresponding to the central pixel point of the 3x3 window with the median of the corresponding values of the 9 pixel points in the 3x3 working window; Gaussian noise corresponds to the Gaussian filtering method. Gaussian filtering means calculating the Gaussian kernel of 3x3 through the Gaussian function, applying the calculated Gaussian kernel to the 3x3 working window, and calculating the new value of the central pixel point of the 3x3 working window through weighted averaging; other noises correspond to the mean filtering method. Mean filtering means replacing the value corresponding to the central pixel point of the 3x3 window with the mean of the 9 pixel points in the 3x3 working window.

[0131] 2) Block adaptive histogram equalization processing:

[0132] Use the block adaptive histogram equalization method on the denoised image to improve the contrast of the image. The specific process is as follows: divide the image into small blocks, determine the size of the small blocks as 24*24, calculate the number of blocks in the horizontal and vertical directions, and traverse the image. Then calculate the histogram of the small block image. The histogram is a statistical chart representing the number of pixels at each gray level in the image. For an 8-bit gray image, the gray level range is 0 to 255. Initialize an array with a length of 256 to store the number of pixels at each gray level. Traverse each pixel in the image and count the number of pixels at each gray level. Then calculate the cumulative distribution function. The cumulative distribution function is an accumulative function representing the proportion of the number of pixels less than or equal to a certain gray level in the total number of pixels. Initialize an array with a length of 256 for the cumulative distribution function to store the values of the cumulative distribution function at each gray level. Finally, calculate the values of the cumulative distribution function at each gray level, normalize the values of the cumulative distribution function to the range of 0 to 255, and adjust the value of each pixel according to the cumulative distribution function to generate a new image.

[0133] 3) Linear function processing:

[0134] The image frame after block adaptive histogram equalization processing is adjusted for brightness and contrast through a linear function; the formula of the linear function is: , where is the corresponding value of the pixel point of the image frame after block adaptive histogram equalization processing, is the corresponding value of the pixel point of the image frame after linear function processing, α is a preset contrast adjustment factor, and β is a preset brightness adjustment factor.

[0135] 4) Gamma correction processing:

[0136] The image frame after linear function processing is adjusted for brightness through gamma correction of non-linear transformation; the formula of gamma correction is: , where is the corresponding value of the pixel point of the image frame after linear function processing, is the corresponding value of the pixel point of the image frame after gamma correction processing, and γ is a preset gamma value.

[0137] 5) Normalization processing:

[0138] The image frame after gamma correction processing is improved in contrast through image normalization processing; the normalization formula is: , where is the corresponding value of the pixel point of the image frame after gamma correction processing, is the corresponding value of the pixel point of the image frame after normalization processing, is the maximum value among the corresponding values of the pixel points of the image frame after gamma correction processing, is the minimum value among the corresponding values of the pixel points of the image frame after gamma correction processing; the normalization formula is applied to traverse each pixel point in the image frame after gamma correction processing to generate the image after traditional processing.

[0139] In this embodiment, the deep learning model processing in step T111 includes:

[0140] The image frames in the image frame sequence in step S2 are input into a pre-constructed deep learning model to calculate the brightness adjustment parameters and output the image; the pre-constructed deep learning model includes a convolutional layer, a multi-scale residual block, and an attention module.

[0141] 1) Use the convolutional layer to extract the initial feature vector from the input image to generate the initial feature map.

[0142] The convolutional layer performs local feature extraction on the image in the way of a sliding window (convolution kernel) to generate the initial feature map.

[0143] 2) Input the initial feature map output by the convolutional layer into multiple multi-scale residual blocks to generate the multi-scale feature map. In this embodiment, 4 multi-scale residual blocks are adopted.

[0144] The multi-scale residual block can capture features at different scales and solve the problem of gradient disappearance in deep networks through residual connections.

[0145] 3) Input the multi-scale feature maps output by the multi-scale residual block into the attention module, calculate the brightness adjustment parameters, and output the image.

[0146] The attention module can be implemented in different ways, such as channel attention or spatial attention.

[0147] In this embodiment, the specific processing of the deep learning model in step T111 includes:

[0148] 1) Use a convolutional layer to extract the initial feature vectors from the input image. The convolutional layer consists of 64 3x3 convolutional kernels. Each convolutional kernel slides over the input image and calculates the convolution value by element-wise multiplication and summation until the convolutional kernel slides through the entire input image, generating a feature map. Each convolutional kernel corresponds to a feature map, and these feature maps retain the local feature information in the input image.

[0149] 2) Use the multi-scale residual block to extract image features from the initial feature vectors output by the convolutional layer. A total of 4 multi-scale residual blocks are constructed. Each multi-scale residual block is composed of multiple convolutional kernels of different sizes. The convolutional kernels inside the multi-scale residual block are connected in parallel, and the 4 multi-scale residual blocks are connected in series.

[0150] The multi-scale residual block includes multiple convolutional kernels of different sizes, and these different-sized convolutional kernels include: Conv-3x3, Conv-3x1, Conv-1x3, Conv-1x1, Sobelx, Sobely, and Laplacian.

[0151] Fuse the output vectors corresponding to the above convolutional kernels through an addition operation to output a multi-scale feature vector. The calculation formula for the multi-scale feature vector is as follows:

[0152] ,

[0153] Where, is the initial feature vector extracted by the convolutional layer; is the multi-scale feature vector output by the multi-scale residual block; Conv 3x3 () is a 3x3 convolutional kernel for extracting image features; Conv 3x1 () is a 3x1 convolutional kernel for extracting features in the vertical direction; Conv 1x3 () is a 1x3 convolutional kernel for extracting features in the horizontal direction; Conv 1x1() is a 1x1 convolutional kernel used to adjust the number of channels or perform feature fusion; Sobelx() is a 3x3 convolutional kernel used to detect horizontal edges; Sobely() is a 3x3 convolutional kernel used to detect vertical edges; Laplacian() uses a 1x1 convolutional kernel followed by a 3x3 convolutional kernel to simulate the Laplacian operator, which is used to detect edges in any direction without considering the directionality of the edges.

[0154] 3) Output the multi-scale feature vectors extracted by the multi-scale residual block to the attention module layer to calculate the brightness adjustment parameters. The attention module is a neural network component used to dynamically assign different weights to different parts of the input data so that the model can focus more on important information. These weights are calculated based on the input data itself, so the attention module can adaptively adjust the focus of attention.

[0155] In this embodiment, the visual large model processing in step T111 includes:

[0156] Input the image frames in the image frame sequence of step S2 into the visual large model VLM to output the brightness-enhanced image.

[0157] The large model is set with the following prompts:

[0158] Task:

[0159] Input an image;

[0160] Output a high-resolution brightness-enhanced image;

[0161] Constraints:

[0162] Ensure that the details and textures of the input image are completely retained.

[0163] In this embodiment, step T112: Merge the three images after traditional processing, deep learning model processing, and visual large model processing into one image according to preset weights as the output image of the image enhancement processing, including:

[0164] Use an image processing library to merge the three images generated by different enhancement methods into the same image according to the preset weights. The specific process includes:

[0165] Set the weight of the image enhanced by the traditional processing method to 0.3, set the weight of the image enhanced by the deep learning model processing method to 0.6, and set the weight of the image enhanced by the visual large model method to 0.1;

[0166] Use the image processing library OpenCV or Pillow (PIL) to merge the three enhanced images into one image as the output image of the image enhancement processing.

[0167] The image enhancement process in step T11 improves the image quality, facilitating the subsequent steps T12 to T13 to judge the change amplitude of the face using the enhanced-quality image, reducing the number of images, and improving the accuracy of subsequent face change amplitude judgment and image number reduction.

[0168] In this embodiment, step T12 includes:

[0169] Calculating the structural similarity index between adjacent frames and the cosine distance of adjacent frame feature vectors after the image enhancement process in step T11;

[0170] 1) The structural similarity index (SSIM) is an index for measuring the similarity between two images. It obtains a value between -1 and 1 by comparing the brightness, contrast, and structural information of the images. The closer this value is to 1, the more similar the two images are, and the closer it is to -1, the less similar the two images are. The calculation formula of the structural similarity index SSIM is as follows:

[0171] ,

[0172] where, represents a pair of adjacent frames after the image enhancement process in step T11, and x and y are the two images to be compared; is the average value of the corresponding pixel values of image frame x, is the average value of the corresponding pixel values of image frame y, and are respectively used to represent the brightness means of image frame x and image frame y; is the variance of the corresponding pixel values of image frame x, is the variance of the corresponding pixel values of image frame y, and are respectively used to measure the contrast of image frame x and image frame y; is the covariance of the corresponding pixel values of image frame x and image frame y, which is used to measure the structural similarity between image frame x and image frame y; and are preset constants, which are determined according to the range of the corresponding pixel values of the images to be compared, and are used to avoid the case where the denominator is zero.

[0173] 2) The cosine distance of adjacent frame feature vectors is the cosine distance between two consecutive frames of images in the feature space. This distance measurement method converts the cosine similarity into a distance form to more intuitively represent the difference between frames. The range of the cosine distance is from 0 to 2, where 0 means the adjacent frames are exactly the same, and 2 means the adjacent frames are completely different. The calculation formula of the cosine distance is as follows:

[0174] ,

[0175] where, is the extracted feature vector of image frame x, is the extracted feature vector of image frame y, and are n-dimensional column vectors; is the i-th element of the feature vector and is the i-th element of the feature vector where i < n.

[0176] In this embodiment, step T13 includes:

[0177] By comparing the structural similarity index calculated in step T12 with a preset structural similarity index threshold and comparing the calculated cosine distance with a preset cosine distance threshold, the change situation of the image is judged;

[0178] The preset structural similarity index threshold is 0.7, and the preset cosine distance threshold is 0.3.

[0179] In this embodiment, step T14 includes:

[0180] Keep the adjacent frames whose structural similarity index calculated in step T12 is less than 0.7 and cosine distance is greater than 0.3 as the preprocessed image data.

[0181] In this embodiment, the adjacent frames with a structural similarity index less than 0.7 and a cosine distance greater than 0.3 are image frames with a large change in the human face. Keeping the image frames with a large change for subsequent processing can retain sufficient image information while reducing the number of image frames for subsequent processing.

[0182] Step T2 includes:

[0183] Input the preprocessed image data in step T1 into the Vision Transformer model to extract the global image features and output the global image feature vector;

[0184] Input the preprocessed image data in step T1 into the improved VGG16 model to extract the local image features and output the local image feature vector;

[0185] In this embodiment, extracting the global image features through the Vision Transformer model includes:

[0186] 1) Image chunking

[0187] The size of each frame of preprocessed image data is H*W*C. Where H is the image height, W is the image width, and C is the number of channels. For a color image, C is 3, and for a grayscale image, C is 1. Cut each preprocessed image into non-overlapping image chunks, and the size of each image chunk is P*P*C. In total, it can be cut into S ( ) image patches; flatten each segmented image patch to output a one-dimensional vector, and the length of the one-dimensional vector corresponding to each image patch is P*P*C; the one-dimensional vectors of each image patch form an image one-dimensional vector sequence, and there are S image patches in total.

[0188] 2) Image patch embedding and position encoding:

[0189] Output the flattened P*P*C vector corresponding to each image patch into a D-dimensional vector through the embedding layer, and add the position encoding information to the D-dimensional vector corresponding to the image sequence. The position encoding information is defined by the Vision Transformer model.

[0190] 3) Output the added D-dimensional vector to L Transformer Encoder layers to extract feature vectors. The Transformer Encoder consists of a normalization layer (BN), a multi-head attention mechanism layer (MA), a normalization layer (BN), and a multi-layer perceptron (MLP) connected in series in sequence.

[0191] The normalization layer refers to performing normalization independently on each hidden layer of each sample.

[0192] The calculation process of the multi-head attention mechanism layer is as follows: the input vector generates query Q, key K, and value V vectors through a linear transformation W, and then is split into h vectors, where h is a multiple of 8, and then the weight vector is calculated , where refers to the j-th sub-vector obtained by splitting the query Q vector into h parts, and is the same. Multiply and to get the output of each head, and concatenate the outputs of each head to get the output result .

[0193] The multi-layer perceptron consists of a fully connected layer with an input dimension of D and an output dimension of 2048, a non-linear activation function GELU, and a fully connected layer with an input dimension of 2048 and an output dimension of D. The multi-layer perceptron is used to increase the non-linear expression ability of the model.

[0194] By capturing global dependencies, the Vision Transformer model can understand the connections between different regions in the image. In the image classification task, the model needs to understand the content of the entire image to determine its category. Global dependencies help the model identify key features in the image and ignore irrelevant background information.

[0195] In this embodiment, extracting local image features through the improved VGG16 includes:

[0196] Extract image features using the improved VGG16, which includes 5 sequentially connected convolutional blocks and 1 spatial-channel attention module SCA.

[0197] The VGG16 model is usually used to extract local features of images and output the classification results of the local features of the entire image through the fully connected layer. In this embodiment, by designing an improved VGG16 model, 1 spatial-channel attention module SCA is added, the fully connected layer is deleted, and the extraction of local important features is focused on. The output features of the spatial-channel attention module SCA are used as the local feature vectors of the image output by the VGG16 model, so that the improved VGG16 model focuses on the extraction of local image features.

[0198] Among them, the convolutional block contains multiple convolutional layers and one max pooling layer. Each convolutional layer uses a 3x3 convolutional kernel, with a stride of 1 and a padding of 1. Each max pooling layer uses a 2x2 pooling kernel, with a stride of 2. The channels of the convolutional blocks are 64, 128, 256, 512, and 512 respectively.

[0199] As Figure 2 shown, the calculation of the spatial-channel attention module SCA includes:

[0200] Input the output feature vector of the 5th convolutional block into the spatial attention branch to obtain the spatial attention weight and the channel attention branch to obtain the channel attention weight respectively; the size of the output feature vector of the 5th convolutional block is H k *W k *C k ;

[0201] Among them,

[0202] The calculation of the spatial attention weight includes:

[0203] Perform max pooling Maxpool and average pooling Avgpool processing on the output feature vector of the 5th convolutional block in the channel dimension respectively, retaining the spatial dimension H k *W k , compress the channel dimension C k to 1; splice the features after max pooling Maxpool and average pooling Avgpool processing; input the spliced features into the fully connected layer FC to extract features; activate the features extracted by the fully connected layer FC through Sigmoid to obtain the spatial attention weight;

[0204] The spatial attention weight is calculated by the formula:

[0205] ,

[0206] In the formula, is the output feature vector of the 5th convolutional block;

[0207] The calculation of the channel attention weight includes:

[0208] Perform max pooling (Maxpool) and average pooling (Avgpool) on the output feature vector of the 5th convolutional block respectively, and compress the spatial dimensions H k *W k to 1, and retain the channel dimension C k ; Feed the features after max pooling (Maxpool) and average pooling (Avgpool) into a fully connected layer (FC) respectively to extract features and then add them; Pass the added features through Sigmoid activation to obtain the channel attention weight;

[0209] , where is the output feature vector of the 5th convolutional block.

[0210] According to the obtained spatial attention weight and channel attention weight, fuse the channel attention branch, the spatial attention branch, and the branch combining spatial attention and channel attention to obtain the output feature of the spatial-channel attention module (SCA);

[0211] The output feature of the spatial-channel attention module (SCA) has the following calculation formula:

[0212] , where is the output feature vector of the 5th convolutional block.

[0213] Combining the channel attention branch and the spatial attention branch in parallel can form a hybrid attention mechanism. This mechanism can capture both channel features and spatial features simultaneously, thus understanding the input data more comprehensively. In the parallel structure, the channel attention branch and the spatial attention branch process the input feature map respectively and generate their respective attention feature maps. Then, the two attention feature maps are fused to obtain the final enhanced feature map. This fusion method enables the model to focus on important channels and key spatial positions simultaneously, thereby improving the performance of the model.

[0214] In step S3, extract physiological signal features from the image data extracted in step S2, specifically including:

[0215] In this embodiment, the physiological signal feature refers to the heart rate variability feature. It includes:

[0216] After selecting the face area from the image data extracted in step S2 to obtain a three-channel time series diagram, use imaging photoplethysmography to analyze the minute color changes of the skin and extract the heart rate variability feature.

[0217] A three-channel timing diagram is a graphical representation used to display and analyze the variation of data over time for three independent channels. In video processing, a three-channel timing diagram is used to show the variation of the RGB three color channels of each frame of an image over time, which helps analyze the color information and dynamic characteristics in the video.

[0218] Imaging photoplethysmography (iPPG) is a non-contact physiological parameter detection technology used to remotely measure physiological parameters such as heart rate and heart rate variability. Imaging photoplethysmography (iPPG) collects videos of the human body surface (such as the face or palm) through an imaging device (such as an RGB camera) and analyzes the changes in light intensity in the video to extract physiological signals. Heart rate variability (HRV) refers to the minute changes in the heart beat intervals (RR intervals). These changes reflect the regulatory effect of the autonomic nervous system on the heart and are important indicators for evaluating heart health and autonomic nerve activity. HRV features are widely used in research on heart diseases, mental diseases, emotion recognition, etc.

[0219] The acquisition of a three-channel timing diagram includes:

[0220] 1) Color space conversion: Convert the cropped face region from the RGB color space to the YUV or other suitable color space. However, in this scenario, we still retain the RGB three channels to analyze color changes.

[0221] 2) Timing diagram generation: For each frame in the video, calculate the mean or median values of the R, G, and B channels of the face region respectively to form three time series (i.e., the three-channel timing diagram).

[0222] The analysis of imaging photoplethysmography (rPPG) includes:

[0223] 1) Signal preprocessing: Filter the R, G, and B three timing diagrams to reduce the influence of noise and light changes. Commonly used filtering methods include band-pass filtering, median filtering, etc.

[0224] 2) Signal quality assessment: Evaluate the quality of the preprocessed signals to ensure they are clear enough for subsequent analysis. This can be done by calculating indicators such as the signal-to-noise ratio and power spectral density of the signals.

[0225] 3) Heart rate extraction: Apply signal processing algorithms (such as Fourier transform, wavelet transform, adaptive filtering, etc.) to extract heart rate information from the preprocessed timing diagrams. This usually involves identifying the frequency components related to the heart beat.

[0226] 4) Channel selection:

[0227] For RGB images, calculate the grayscale value of each pixel using the weighted average method. The weighting formula is as follows:

[0228] Gray = 0.11×R + 0.59×G + 0.30×B

[0229] Among them, R, G, and B are the values of the red, green, and blue color channels respectively.

[0230] For the timing diagrams of the above three colors, they respectively represent the signals of the red, green, and blue channels changing over time. The weighted value of the green channel is larger because the change of green light is more obvious when absorbed by blood and can accurately reflect the change of the pulse. Extract the alternating current signal in the timing diagram of the green light. First, extract the alternating current signal from the signal in the timing diagram. The alternating current signal is the fluctuating part of the signal changing over time. This alternating current signal is the pulse wave signal.

[0231] Heart rate variability (HRV) feature extraction includes:

[0232] 1) R-R interval calculation: Calculate the consecutive R-R intervals from the extracted heart rate information, that is, the time interval between two adjacent heartbeats.

[0233] 2) HRV feature calculation: Calculate the common features of HRV according to the R-R interval sequence, such as time domain features (SDNN, RMSSD, etc.) and frequency domain features (LF / HF ratio, etc.). These features reflect the activity of the cardiac autonomic nervous system.

[0234] The heart rate variability features in this embodiment include SDNN. SDNN refers to the standard deviation of normal R-R intervals.

[0235] The calculation formula for the standard deviation SDNN of normal R-R intervals is:

[0236] ,

[0237] In the formula, is the i-th cycle, is the average value of all RR intervals, and N is the total number of sampling points, which is also the average value of all RR intervals.

[0238] In step S3, extract audio features from the voice data extracted in step S2, including:

[0239] Use Wav2Vec to extract the audio feature vector in the video. The specific process includes:

[0240] 1) Read the audio signal

[0241] 2) Segment the audio signal into frames of fixed length. Usually, each frame contains several hundred milliseconds of audio data. There can be an overlapping part between frames to capture the continuity of the audio.

[0242] 3) Convert the audio signal of each frame into a Mel spectrogram. A Mel spectrogram is a spectral representation that maps frequencies to the Mel scale and can better simulate the human ear's perception of frequencies.

[0243] 4) Perform mean normalization on the Mel spectrogram to eliminate the statistical differences between different frames.

[0244] 5) Use a Transformer Encoder to encode the normalized Mel spectrogram to generate the context representation of each frame.

[0245] 6) Input the context representation generated by the Transformer Encoder into a quantizer. The quantizer finds the codeword closest to each context representation through nearest neighbor search. The quantizer maps the continuous context representation into a predefined discrete codebook. Each vector in the codebook is called a codeword.

[0246] 7) Concatenate the Mel spectrogram and the quantized context representation to generate the final audio features. Wav2Vec is a self-supervised learning-based audio feature extraction model. It can effectively extract audio feature vectors from clean audio and noisy recorded videos by learning the representation of audio from a large amount of unlabeled audio data, meeting the requirements of the audio scenarios in the complex emotion video dataset during driving.

[0247] In step S3, extract text features from the text data extracted in step S2. The specific process includes:

[0248] 1) Speech-to-text conversion. Use speech recognition technology such as the ready-made speech recognition library Baidu Speech Recognition to convert the speech in the collected video into text.

[0249] 2) Text cleaning. Remove special characters, punctuation marks, numbers, etc. in the text. Special characters refer to escape characters, operators, and dollar signs, etc. in the text.

[0250] 3) Word segmentation. Then split the text into words or phrases.

[0251] 4) Stop word removal. Stop words refer to auxiliary words, prepositions, conjunctions, pronouns, etc.

[0252] 5) Word embedding. Convert the text into a vector format. Use Word2Vec to convert the text into a numerical vector. First, build a vocabulary, count the occurrence frequency of all words, assign a unique index to each word, and use a pre-trained model to convert the words into a vector format.

[0253] In step S3, extracting environmental features from the environmental data collected in step S1 includes:

[0254] When the present invention is used for driver emotion recognition, the environmental data includes weather data and traffic data;

[0255] The data is used to generate environmental features through one-hot encoding, where one-hot encoding means that for a data with u states, one-hot encoding creates a binary vector of length u.

[0256] Each position in the vector corresponds to one state. If the current state matches the position in the vector, the position is 1; otherwise, it is 0.

[0257] The process of generating environmental feature vectors from weather data and traffic data through one-hot encoding (One-Hot Encoding) is to convert this type of data into numerical features that can be processed by a machine learning model. The specific process includes:

[0258] I. Data Preparation

[0259] 1) Collect weather data and traffic data:

[0260] The weather data may include weather conditions (such as sunny, rainy, snowy, etc.), temperature, humidity, wind speed, etc.

[0261] The traffic data may include traffic flow, traffic congestion index, number of traffic accidents, etc.

[0262] 2) Data cleaning:

[0263] Check and handle missing values, outliers, etc.

[0264] Ensure the accuracy and consistency of the data.

[0265] II. One-Hot Encoding Principle

[0266] One-hot encoding is a commonly used method for encoding categorical data. It converts each categorical value into a binary vector with only one element being 1 and the rest being 0. For example, for weather conditions (sunny, rainy, snowy), the one-hot encoded vectors may be [1, 0, 0] (sunny), [0, 1, 0] (rainy), [0, 0, 1] (snowy).

[0267] III. One-Hot Encoding Implementation

[0268] 1) Determine the categorical values:

[0269] For each categorical feature in the weather data and traffic data, list all its possible categorical values.

[0270] 2) Create a one-hot encoding matrix:

[0271] For each categorical feature, create a one-hot encoding matrix with a size equal to the number of categories of that feature.

[0272] Each row of the matrix corresponds to a data sample, and each column corresponds to a class value.

[0273] 3) Fill the encoding matrix:

[0274] According to the class values in the data samples, fill 1 and 0 in the one-hot encoding matrix.

[0275] For example, if the weather condition of a certain data sample is "sunny", then fill 1 in the "sunny" column of the corresponding one-hot encoding matrix for weather conditions, and fill 0 in the remaining columns.

[0276] 4) Combine the feature vectors:

[0277] Concatenate the one-hot encoding matrices of all categorical features by column to form an environmental feature vector containing all features.

[0278] If the dataset also contains numerical features (such as temperature, humidity, etc.), these features can also be added to the feature vector, but usually they need to be normalized.

[0279] In the case of poor or even no network signal, set the environmental feature to a zero vector.

[0280] Step S4: Use the attention mechanism to splice and fuse the image features, physiological signal features, audio features, text features, and environmental features extracted in step S3 to generate multi-modal fusion features, including:

[0281] S41: Pass the image feature vector extracted by the Vision Transformer model and the image feature vector extracted by the improved VGG16 model in step S3 through a bimodal feature fusion module (feature fusion module, FFM), calculate the weights required for the image, and perform a feature fusion operation on the feature maps to obtain visual features;

[0282] As Figure 3 shown, the process of the bimodal feature fusion module calculating the visual features includes:

[0283] Calculate the weights of the Vision Transformer model and the weights of the improved VGG16 model ;

[0284] Among them, the calculation process of the weights of the Vision Transformer model is as follows:

[0285] 1) The global image feature vector output by the Vision Transformer model , two feature maps are obtained through max pooling and average pooling respectively;

[0286] 2) Concatenate the two obtained feature maps;

[0287] 3) Add the maximum value and the average value of the vector elements after concatenation, and activate the added feature through the Sigmoid function to finally obtain the weights of the Vision Transformer model ;

[0288] The weights of the improved VGG16 model The calculation process is the same as that of the weights of the Vision Transformer model ;

[0289] Multiply the calculated weights of the Vision Transformer model by the global image feature vector output by the Vision Transformer model ;

[0290] Multiply the calculated weights of the improved VGG model by the local image feature vector output by the improved VGG16 model ;

[0291] Add the above two multiplication results to obtain the visual feature.

[0292] Visual feature The calculation formula is:

[0293] ;

[0294] S42: Fuse the audio feature and the text feature to obtain the non-visual feature;

[0295] Fuse the audio feature and the text feature to obtain the non-visual feature. The specific process is as follows: First, concatenate the audio feature and the text feature, then calculate the weights of the audio feature and the text feature through the attention module, perform weighted calculation according to the weights of the audio feature and the text feature, and add and calculate the weighted audio feature and text feature to obtain the non-visual feature.

[0296] S43: Concatenate and fuse the visual feature, the non-visual feature, the physiological signal feature, and the environmental feature, and generate the multi-modal fusion feature after transformation through the non-linear activation RELU function.

[0297] In step S5, the training of the emotion detection model includes:

[0298] The emotion detection model is trained with the optimization goal of minimizing the difference between the predicted emotion content obtained by the emotion recognition of the sample through the emotion detection model and the actual emotion content corresponding to the sample.

[0299] In this embodiment, the number of classifications of the emotion content recognized by the emotion detection model is 7, specifically including anger, fear, sadness, disgust, surprise, happiness, and neutral;

[0300] Among them, the first loss function The calculation formula is:

[0301] ,

[0302] In the formula, W is the weight vector of the emotion detection model; R is the multi-modal fusion feature of the sample; bs is the batch number, m is the number of classifications of the emotion content, and in this embodiment, m = 7;

[0303] The second loss function The calculation formula is:

[0304] ,

[0305] In the formula, is the feature vector of the emotion content predicted by the visual large model; is the feature vector of the actual emotion content of the sample; is a function representing the distance between two vectors. When is the same as , the function value is 0. When is different from , the function value is 1; is the model parameter of the emotion detection model;

[0306] The overall loss function The calculation formula is:

[0307] ,

[0308] In the formula, is the loss coefficient preset by the visual large model.

[0309] 1) Use the formula to calculate the loss value, where W is the model weight, b is the bias, bs is the batch size, and m is the number of seven expression categories.

[0310] 2) Calculate the loss predicted by the large model according to the label predicted by the visual large model and the actual label, where is the label predicted by the visual large model, is the actual label, It is to calculate the distance between two vectors. When is equal to , the value is 0; otherwise, the value is 1. is calculated by the model through learning.

[0311] 3) Finally, jointly define the loss function as , is the loss coefficient of the vision large model, which is used to control the proportion

[0312] The emotion recognition method of the present invention further includes step S6:

[0313] Support corresponding emotion regulation measures according to the results of emotion recognition.

[0314] The emotion regulation measures include playing music and / or voice prompts.

[0315] For example: when the emotion recognition method of the present invention is used for emotion recognition of a driver,

[0316] when the emotion content recognized is anger, play soothing music and play a voice to remind the driver to drive calmly; when the continuous duration of the recognized anger reaches a preset duration, play a voice reminder and call an emergency contact number;

[0317] when the emotion content recognized is fear, play light music to relieve negative emotions and play a voice to remind the driver to drive attentively; when the continuous duration of the recognized fear reaches a preset duration, play a voice reminder and call an emergency contact number;

[0318] when the emotion content recognized is sadness, play exciting and cheerful music and play a voice to comfort the driver's emotion; when the continuous duration of the recognized sadness reaches a preset duration, play a voice reminder and call an emergency contact number.

[0319] Embodiment 2

[0320] This embodiment provides an emotion recognition device, including a collection module, a data extraction module, a feature extraction module, a feature fusion module and an emotion recognition module;

[0321] The collection module is used for: collecting face video data and environmental data;

[0322] The data extraction module is used for: extracting image data, voice data and text data according to the face video data collected by the collection module;

[0323] The feature extraction module is used for:

[0324] extracting image features from the image data extracted by the data extraction module;

[0325] Extract physiological signal features from the image data extracted by the data extraction module;

[0326] Extract audio features from the voice data extracted by the data extraction module;

[0327] Extract text features from the text data extracted by the data extraction module;

[0328] Extract environmental features from the environmental data collected by the collection module;

[0329] The feature fusion module is used to: use the attention mechanism to splice and fuse the image features, physiological signal features, audio features, text features and environmental features extracted by the feature extraction module to generate multi-modal fusion features;

[0330] The emotion recognition module is used to: input the multi-modal fusion features generated by the feature fusion module into a pre-trained emotion detection model for emotion recognition, and output the emotion content to which the multi-modal fusion features belong.

[0331] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0332] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks the device with the functions specified.

[0333] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements in the process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks the functions specified.

[0334] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to generate a computer-implemented process, thereby providing instructions for implementing the steps of the function specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 in one block or a plurality of blocks.

[0335] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. An emotion recognition method, characterized in that: include: S1: Collect face video data and environment data; S2: extracting image data, voice data and text data according to the face video data collected in step S1; S3: extracting image features from the image data extracted in step S2; Extracting physiological signal features from the image data extracted in step S2; Extracting audio features from the speech data extracted in step S2; Extracting text features from the text data extracted in step S2; Extracting environmental features from the environmental data collected in step S1; S4: Use the attention mechanism to concatenate and fuse the image features, physiological signal features, audio features, text features, and environmental features extracted in step S3 to generate multimodal fusion features; S5: Input the multimodal fusion features generated in step S4 into the pre-trained emotion detection model for emotion recognition, and output the emotion content to which the multimodal fusion features belong.

2. The emotion recognition method according to claim 1, characterized in that: In step S2, extracting image data according to the face video data collected in step S1 includes: Perform frame processing on the face video data collected in step S1 to extract image data, extract image data corresponding to each frame, and obtain an image frame sequence; the face video data refers to video data of a face; In step S3, extracting image features from the image data extracted in step S2 includes: T1: performing image preprocessing on the image data extracted in step S2; T2: Extract image features from preprocessed image data; Step T1 includes: T11: Perform image enhancement processing on each image frame in the image frame sequence of step S2.

3. The emotion recognition method according to claim 2, characterized in that: The image enhancement processing of each image frame in step T11 includes: T111: subjecting the image frames in the image frame sequence of step S2 to traditional processing, deep learning model processing, and visual large model processing, respectively; the traditional processing includes subjecting the image frames to filtering processing, block adaptive histogram equalization processing, linear function processing, gamma correction processing, and normalization processing in sequence; T112: The three images processed by traditional methods, deep learning models, and large visual models are merged into one image according to preset weights as the output image of image enhancement processing.

4. The emotion recognition method according to claim 2, characterized in that: Step T1 also includes: T12: Calculate the structural similarity index of adjacent frames and the cosine distance of feature vectors of adjacent frames after the image enhancement processing in step T11; T13: comparing the structural similarity index calculated in step T12 with a preset structural similarity index threshold, and comparing the cosine distance calculated in step T12 with a preset cosine distance threshold; T14: retaining the adjacent frames that meet the preset structural similarity index threshold requirement and the preset cosine distance threshold requirement in step T13 as preprocessed image data.

5. The emotion recognition method according to claim 2, characterized in that: Step T2 includes: Input the image data preprocessed in step T1 into the Vision Transformer model to extract the global features of the image and output the global feature vector of the image; The image data preprocessed in step T1 is input into the improved VGG16 model to extract local features of the image, and output a local feature vector of the image; the improved VGG16 model includes 5 convolution blocks connected in sequence and 1 space-channel attention module SCA.

6. The emotion recognition method according to claim 5, characterized in that: The calculation of the spatial-channel attention module SCA includes: The output feature vector of the fifth convolutional block is input into the spatial attention branch to obtain the spatial attention weight and the channel attention branch to obtain the channel attention weight; the output feature vector size of the fifth convolutional block is H k *W k *C k ; in, The calculation of the spatial attention weight includes: The output feature vector of the fifth convolutional block is processed by maximum pooling Maxpool and average pooling Avgpool in the channel dimension, retaining the spatial dimension H k *W k , the channel dimension C k Compress to 1; concatenate the features processed by the maximum pooling Maxpool and the average pooling Avgpool; input the concatenated features into the fully connected layer FC to extract features; activate the features extracted by the fully connected layer FC with Sigmoid to obtain the spatial attention weight; The spatial attention weights The calculation formula is: , In the formula, is the output feature vector of the 5th convolutional block; The channel attention weight calculation includes: The output feature vector of the fifth convolutional block is processed by maximum pooling Maxpool and average pooling Avgpool respectively, and the spatial dimension H k *W k Compress to 1, retaining the channel dimension C k ; The features processed by the maximum pooling Maxpool and the average pooling Avgpool are sent to the fully connected layer FC to extract features and then added; the added features are activated by Sigmoid to obtain the channel attention weight; The channel attention weight The calculation formula is: ; According to the obtained spatial attention weight and channel attention weight, the channel attention branch, the spatial attention branch and the branch combining spatial attention and channel attention are fused to obtain the output features of the spatial-channel attention module SCA; The output features of the spatial-channel attention module SCA The calculation formula is: 。 7. The emotion recognition method according to claim 1, characterized in that: In step S3, the physiological signal characteristics include heart rate variability characteristics; In step S3, extracting physiological signal features from the image data extracted in step S2 includes: Selecting a face region from the image data extracted in step S2, obtaining a three-channel timing diagram of the face region, and analyzing skin color changes using imaging photoplethysmography to extract heart rate variability characteristics; In step S3, extracting environmental features from the environmental data collected in step S1 includes: generating environmental features by one-hot encoding the environmental data.

8. The emotion recognition method according to claim 6, characterized in that: Step S4 includes: S41: Fusing the global feature vector of the image extracted by the Vision Transformer in step S3 with the local feature vector of the image extracted by the improved VGG16 to obtain visual features; S42: Fusing audio features and text features to obtain non-visual features; S43: Visual features, non-visual features, physiological signal features and environmental features are concatenated and fused, and transformed through a nonlinear activation RELU function to generate multimodal fusion features.

9. The emotion recognition method according to claim 1, characterized in that: In step S5, the training of the emotion detection model includes: The emotion detection model is trained with the optimization goal of minimizing the difference between the predicted emotion content of the sample obtained through emotion recognition by the emotion detection model and the actual emotion content corresponding to the sample; Among them, the first loss function The calculation formula is: , Where W is the weight vector of the emotion detection model; R is the multimodal fusion feature of the sample; bs is the batch size, and m is the number of categories of emotional content; The second loss function The calculation formula is: , In the formula, It is the feature vector of the big visual model predicting the emotional content; is the feature vector of the actual emotional content of the sample; is a function that represents the distance between two vectors. and When the function is equal, the value is 0. and When different, the function takes the value of 1; are the model parameters of the emotion detection model; Overall loss function The calculation formula is: , In the formula, It is the loss coefficient preset by the visual large model.

10. An emotion recognition device, comprising: Acquisition module, data extraction module, feature extraction module, feature fusion module and emotion recognition module; The acquisition module is used to: collect face video data and environmental data; The data extraction module is used to extract image data, voice data and text data according to the face video data collected by the collection module; The feature extraction module is used for: extracting image features from the image data extracted by the data extraction module; extracting physiological signal features from the image data extracted by the data extraction module; extracting audio features from the speech data extracted by the data extraction module; extracting text features from the text data extracted by the data extraction module; Extracting environmental features from environmental data collected by the acquisition module; The feature fusion module is used to: use the attention mechanism to splice and fuse the image features, physiological signal features, audio features, text features and environmental features extracted by the feature extraction module to generate multimodal fusion features; The emotion recognition module is used to: input the multimodal fusion features generated by the feature fusion module into the pre-trained emotion detection model for emotion recognition, and output the emotion content to which the multimodal fusion features belong.

Citation Information

Cited By

  • User interaction method, storage medium and user interaction system

    CN120707955A

  • Music spectrum dynamic mapping method and device and intelligent cabin

    CN121075372A