Micro-expression Recognition Method and System Based on a Three-branch Network

Through the three-branch network combining deep and shallow convolutional neural network and attention module, the problems of insufficient samples and overfitting in micro-expression recognition are solved, achieving higher recognition accuracy and effective processing of difficult samples.

CN119625805BActive Publication Date: 2025-07-22SHANDONG UNIV OF FINANCE & ECONOMICS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411705395.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-07-22
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

The existing micro-expression recognition technology has problems such as insufficient samples, difficulty in learning multiple types of features, low recognition accuracy, and the deep learning model structure is complex and easy to overfit.

Method used

Using a three-branch network method, combining deep convolutional neural network with shallow 3D convolutional neural network, multiple attention modules are introduced, and the model is optimized through IM Loss and Focal Loss functions, the face, optical flow and optical strain characteristics are extracted and fused.

Benefits of technology

It significantly improves the accuracy of micro-expression recognition, solves the problems of insufficient samples and overfitting, and improves the model's ability to identify difficult-to-classify samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625805B_ABST
    Figure CN119625805B_ABST
Patent Text Reader

Abstract

The present invention proposes a micro-expression recognition method and system based on a three-branch network, which relates to the technical field of micro-expression recognition. The problem addressed is that in existing micro-expression recognition technologies, there is a lack of micro-expression samples, making it difficult to learn multiple types of features and resulting in low model recognition accuracy. This method obtains a video sequence, preprocesses the video sequence to obtain an optical flow map, further processes the optical flow map to obtain an optical strain map, inputs the video sequence, the optical flow map, and the optical strain map into a three-branch network respectively to extract face features, optical flow features, and optical strain features, fuses the face features, optical flow features, and optical strain features, and performs micro-expression classification on the fused features to obtain the classification result of micro-expressions. The present invention overcomes problems such as insufficient micro-expression samples, can utilize implicit different information, reduce the loss of subtle face features during the recognition process, enable the model to more effectively process samples that are difficult to classify, and significantly improve the accuracy of micro-expression recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of micro-expression recognition, and particularly relates to a micro-expression recognition method and system based on a three-branch network. Background Technique

[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.

[0003] Facial expressions are a way for a person to express their emotional state, and people often hide their true emotions. Psychology believes that true emotions sometimes express themselves in a very short and rapid facial expression, usually lasting from 1 / 25 s to 1 / 3 s, and appearing in specific areas of the face. This kind of expression is called a micro-expression. Compared with macro-expressions, micro-expressions are a kind of facial emotion triggered by people inadvertently, spontaneous and uncontrollable, and can reflect people's true feelings. Therefore, the study of micro-expressions has an important role and has a wide range of applications in fields such as public security, judicial criminal investigation, and people's daily lives.

[0004] Micro-expression recognition has three significant characteristics: (1) short duration, usually lasting from 1 / 25 s to 1 / 3 s, (2) small movement amplitude and low change intensity, and (3) local movement only appears in fixed facial areas. Currently, the general method for micro-expression recognition is to perform feature extraction after data preprocessing, and then classify micro-expressions based on the extracted features. In terms of feature extraction, traditional methods mostly extract features from the original image through feature descriptors, such as LBP-TOP, LBP-SIP, etc. There are also some methods that use optical flow methods to improve the accuracy of micro-expression feature extraction, such as Bi-WOOF, MDMO, etc. However, most of the features extracted by these methods are some basic and intuitive features and information describing facial expressions, and do not involve in-depth understanding of facial expressions.

[0005] To obtain richer features, most current methods adopt feature extraction methods based on deep learning. For example, Recurrent Convolutional Networks (RCN) are used to extract features from the original image sequence and optical flow respectively, and Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) are combined to extract features for micro-expression recognition. However, these network models have complex structures and require a large amount of computing resources, and are prone to overfitting problems during training. There is also a method of using a shallow convolutional network with multi-stream input information to extract features. This method is simpler and easier to understand in micro-expression recognition, but it cannot capture complex features. 3D Convolutional Neural Networks (3DCNN) can extract temporal and spatial information simultaneously and learn features from low to high levels, which helps to construct and express complex spatio-temporal features layer by layer. This network provides a good tool for micro-expression recognition. However, when applied to micro-expression datasets with a small sample size, 3DCNN is prone to the risk of overfitting. Due to the unclear boundaries of micro-expression classification and insufficient dataset samples, various current deep learning-based methods, including 3DCNN, still have limitations in terms of feature learning ability, model generalization ability, fitting ability, etc., and the recognition rate is low.

[0006] In summary, there are certain deficiencies in the existing micro-expression recognition technologies:

[0007] (1) Most models extract some basic and intuitive features and information describing facial expressions, without involving in-depth understanding of facial expressions and unable to capture complex features;

[0008] (2) Some models can extract complex micro-expression features, but the model structure is complex, requires a large amount of computing, and is prone to overfitting problems during training;

[0009] (3) Some models overcome the problems of complex structure and large amount of computing, but for micro-expression datasets with a small sample size, there are still overfitting problems, insufficient micro-expression samples, and it is difficult to learn multiple types of features, resulting in low recognition accuracy. Summary of the Invention

[0010] To overcome the deficiencies of the above-mentioned existing technologies, the present invention provides a micro-expression recognition method and system based on a three-branch network. By combining a deep convolutional neural network with a shallow 3D convolutional neural network and integrating multiple attention modules, a micro-expression recognition model based on a three-branch network is constructed. The three-branch network is used to extract and process face features, optical flow features, and optical strain features. At the same time, the IM Loss function is introduced into the model and combined with the Focal Loss function, enabling the model to more effectively process difficult-to-classify samples and solving the problems of insufficient micro-expression samples, difficulty in learning multiple types of features, and low recognition accuracy.

[0011] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:

[0012] The first aspect of the present invention discloses a micro-expression recognition method based on a three-branch network, including:

[0013] Obtain a video sequence, preprocess the video sequence to obtain an optical flow map, and further process the optical flow map to obtain an optical strain map. The video sequence includes video frames;

[0014] Input the video frames into the first branch network, and the first branch network extracts face features based on deep learning and depth-enhanced channel attention;

[0015] Input the optical flow map into the second branch network, and the second branch network extracts the features of the optical flow map using the second shallow convolutional neural network and the second facial feature enhancement attention;

[0016] Input the optical strain map into the third branch network, and the third branch network extracts the features of the optical strain map using the third shallow convolutional neural network and the third facial feature enhancement attention;

[0017] Fuse the face features, the features of the optical flow map, and the features of the optical strain map;

[0018] Perform micro-expression classification on the fused features to obtain the classification result of the micro-expression.

[0019] As a further technical solution, the video frame is a vertex frame, which refers to the frame with the most obvious intensity or manifestation of the micro-expression, or an intermediate frame sequence composed of the vertex frame and its front and rear frames.

[0020] As a further technical solution, preprocessing the video sequence includes locating the vertex frame to determine the position of the vertex frame in the micro-expression sequence and extract the time information before and after the frame;

[0021] And processing the vertex frame and the starting frame to obtain an optical flow map representing the dynamic change of the micro-expression, and processing the optical flow map to obtain an optical strain map.

[0022] As a further technical solution, the second branch network and the third branch network have the same structure, both including: three parallel streams, a channel layer, a facial feature enhancement attention module, a second fusion module, and a fully connected layer;

[0023] In the second branch network, each parallel stream uses convolutional kernels with different parameters to extract features of different sizes. After the parallel streams, the features in the three parallel streams are concatenated at the channel level to obtain a first fusion feature;

[0024] The first facial feature enhancement attention module combines and performs weighted averaging on the first fusion feature;

[0025] The feature obtained after passing through the first facial feature enhancement attention module is further fused with the first fusion feature to obtain a second fusion feature;

[0026] The second fusion feature is processed by the fully connected layer and the feature of the optical flow map is output.

[0027] As a further technical solution, the facial feature enhancement attention module includes a convolutional attention module, a dual pooling layer, and a fully connected layer. The convolutional attention module includes a channel attention module and a spatial channel attention module;

[0028] Perform max pooling on the first fusion feature to obtain the output feature of the max pooling layer;

[0029] Perform average pooling on the first fusion feature to obtain the output feature of the average pooling layer;

[0030] The output features of the max pooling layer and the average pooling layer are combined and input into the fully connected layer to obtain the output feature of the fully connected layer;

[0031] The output feature of the fully connected layer is input into the channel attention module to obtain the channel attention feature;

[0032] The channel attention feature is weighted by channel attention with the output feature of the fully connected layer to obtain the channel attention weighted feature;

[0033] The channel attention weighted feature is input into the spatial attention module to obtain the spatial attention feature;

[0034] The spatial attention feature is weighted by spatial attention with the channel attention weighted feature to obtain the spatial attention weighted feature;

[0035] The spatial attention feature and the first fusion feature are fused to obtain the second fusion feature.

[0036] As a further technical solution, the second fusion feature is processed by the fully connected layer. The fully connected layer includes three layers. The processing includes:

[0037] The first two fully-connected layers respectively map the fused second fused feature to a higher-level abstract representation and gradually reduce the feature dimension;

[0038] The third fully-connected layer then converts the features output by the first two fully-connected layers into a dimension suitable for the output of the second branch network.

[0039] As a further technical solution, before the optical flow map is input into the second branch network and the optical strain map is input into the third branch network, resampling is respectively performed.

[0040] As a further technical solution, the first branch network includes a deep convolutional neural network and a channel attention module; the channel attention module includes a global average pooling module and a max pooling module;

[0041] After the video frame is input into the deep convolutional neural network, the original face features are extracted;

[0042] The original face features are input into adaptive average pooling, average pooling is performed on each channel, and the average pooling channel attention weight is obtained through convolution and function activation processing;

[0043] The average pooling channel attention weight is multiplied by the original face features to obtain the global average pooling output features;

[0044] The original face features are input into adaptive max pooling for max pooling, and the max pooling channel attention weight is obtained through convolution and function activation processing;

[0045] The max pooling channel attention weight is multiplied by the original face features to obtain the max pooling output features;

[0046] The global average pooling output features and the max pooling output features are fused to obtain the face features.

[0047] As a further technical solution, after obtaining the micro-expression classification result, the total loss function is constructed by two loss functions, namely IM Loss and Focal Loss, to optimize the model.

[0048] In a second aspect, a micro-expression recognition system based on a three-branch network is disclosed, including:

[0049] A preprocessing module, configured to obtain a video sequence, preprocess the video sequence to obtain an optical flow map, and further process the optical flow map to obtain an optical strain map, where the video sequence includes video frames;

[0050] A feature extraction and fusion module for inputting video frames into a first branch network, which extracts face features based on deep learning and depth-enhanced channel attention; inputting optical flow maps into a second branch network, which extracts features of the optical flow maps using a second shallow convolutional neural network and second facial feature enhancement attention; inputting optical strain maps into a third branch network, which extracts features of the optical strain maps using a third shallow convolutional neural network and third facial feature enhancement attention; and fusing the face features, features of the optical flow maps, and features of the optical strain maps.

[0051] An identification and classification module for performing micro-expression classification on the fused features to obtain a classification result of the micro-expression.

[0052] The above one or more technical solutions have the following beneficial effects:

[0053] In this embodiment, by constructing a micro-expression recognition model based on a three-branch network, a three-branch network composed of a deep convolutional neural network and a shallow 3D convolutional neural network and incorporating multiple attention modules, the problems of insufficient micro-expression samples, difficulty in learning multiple types of features, low recognition accuracy, and complex model are solved.

[0054] In this embodiment, a depth-enhanced channel attention module ECANet is introduced into the deep convolutional neural network, and a new facial feature enhancement attention module is proposed in the shallow 3D convolutional neural network. By separately extracting and processing face features, optical flow features, and optical strain features through three branches and fusing them layer by layer, different information hidden in each layer of features can be fully utilized, reducing the loss of subtle face features during the recognition process and significantly improving the accuracy of micro-expression recognition.

[0055] In this embodiment, the IM Loss function is introduced into the micro-expression recognition model and combined with the Focal Loss function, enabling the model to more effectively process difficult-to-classify samples and effectively improving the model's recognition ability for difficult samples.

[0056] Experiments on the CASME II, SMIC, SAMM micro-expression datasets and the MEGC2019 composite micro-expression dataset in this embodiment show that the proposed method is superior to traditional methods and existing deep learning methods in micro-expression recognition, and the method of this embodiment can be effectively applied to multi-class micro-expression recognition.

[0057] Advantages of additional aspects of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. Description of the Drawings

[0058] The accompanying drawings of the specification, which form a part of the present invention, are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0059] Figure 1 It is a flowchart of a micro-expression recognition method based on a three-branch network for the first embodiment.

[0060] Figure 2 It is an intermediate frame sequence constructed from samples in the CASMEII dataset for the first embodiment.

[0061] Figure 3 It is a schematic diagram of the 3DCNN-FEAM network for the first embodiment.

[0062] Figure 4 It is a schematic diagram of the Facial Feature Enhancement Attention Module FEAM for the first embodiment.

[0063] Figure 5 It is a schematic diagram of the ECANet module for the first embodiment.

[0064] Figure 6 It is a confusion matrix of the MEGC2019 dataset for the first embodiment on different methods. Detailed implementation manners

[0065] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0066] It should be noted that the terms used herein are only for describing the specific implementation manners and are not intended to limit the exemplary implementation manners according to the present invention.

[0067] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0068] The first embodiment

[0069] The first embodiment discloses a micro-expression recognition method based on a three-branch network.

[0070] For a clearer elaboration of this embodiment, the micro-expression recognition process based on the three-branch network can be specifically described as follows:

[0071] The first embodiment provides a micro-expression recognition method based on a three-branch network, including:

[0072] S1, obtaining a video sequence, preprocessing the video sequence to obtain an optical flow map, and further processing the optical flow map to obtain an optical strain map, where the video sequence includes video frames;

[0073] S2. Input the video frame into the first branch network, and the first branch network extracts facial features based on deep learning and depth-enhanced channel attention.

[0074] S3. Input the optical flow map into the second branch network, and the second branch network extracts the features of the optical flow map by using the second shallow convolutional neural network and the second facial feature enhancement attention.

[0075] S4. Input the optical strain map into the third branch network, and the third branch network extracts the features of the optical strain map by using the third shallow convolutional neural network and the third facial feature enhancement attention.

[0076] S5. Fuse the facial features, the features of the optical flow map, and the features of the optical strain map.

[0077] S6. Perform micro-expression classification on the fused features to obtain the classification result of the micro-expression.

[0078] As Figure 1 、 Figure 2 shown, in step S1, obtain the video sequence, preprocess the video sequence to obtain the optical flow map, and further process the optical flow map to obtain the optical strain map. The video sequence includes video frames.

[0079] In the micro-expression video sequence, the starting frame represents the moment when the expression starts.

[0080] The video frame is the vertex frame. The vertex frame refers to the frame in which the intensity or manifestation of the micro-expression is the most obvious, or the intermediate frame sequence composed of the vertex frame and its front and back frames.

[0081] S101. Preprocess the video sequence, including locating the vertex frame to determine the position of the vertex frame in the micro-expression sequence for extracting the time information before and after this frame.

[0082] S1011. Obtain the dataset containing the video sequence.

[0083] In this embodiment, four datasets, namely CASME II, SMIC, SAMM, and MEGC2019, are adopted.

[0084] Specifically, the micro-expressions in the SMIC dataset are divided into three categories: positive, negative, and surprised.

[0085] The CASMEII and SAMM datasets are re-partitioned to unify the sample labels, and the original labels are remapped into the new label space to better avoid the ambiguity of emotion categories caused by different stimuli and environmental settings.

[0086] The specific operation is as follows: Samples of "disgust", "anger", "depression", "contempt", "sadness" and "fear" are classified as "negative" samples; Samples of "happiness" are classified as "positive" samples; Samples of "surprise" remain unchanged; Samples of "other" cannot be classified and thus are not used.

[0087] MEGC2019 is a composite dataset. 24 subjects, 16 subjects and 28 subjects are extracted from CASME II, SMIC and SAMM respectively, including 442 samples of 68 subjects from different backgrounds (environment and gender).

[0088] Each time, all the facial micro-expression data of one subject in the dataset are reserved for testing, and the rest are used for training.

[0089] S1012: Obtain the starting frame and the vertex frame of the micro-expression sequence from the dataset.

[0090] In the micro-expression sequence, the vertex frame refers to the frame where the intensity or manifestation of the micro-expression is the most obvious, and the vertex frame localization refers to determining the position of the vertex frame in the micro-expression sequence. The information contained in the vertex frame can largely represent the characteristics of the entire micro-expression sample. After determining the position of the vertex frame, the time information before and after this frame can be extracted to further analyze the temporal characteristics of the micro-expression.

[0091] Obtain the starting frame and the vertex frame from the dataset.

[0092] Among them, the starting frames and vertex frames of the CASMEII, SAMM datasets and MEGC2019 have been marked. The marked datasets show that the vertex frames of micro-expressions usually appear in the middle part of the video. For the frame sequences before and after the vertex frame of the micro-expression, the expression changes are relatively small. For the unmarked SMIC dataset, a frame in the middle of the sample video is also selected as the vertex frame.

[0093] S102: Enhance the middle frames, and replace the vertex frame with a middle frame sequence composed of the frames before and after it.

[0094] Micro-expressions have instantaneous characteristics. The vertex frame of a micro-expression is the frame with the maximum expression intensity. The expression differences between the vertex frame of the micro-expression and the frames before and after it are very small, and the changes in expression intensity are very subtle.

[0095] In this embodiment, in order to enrich the sample data, a middle frame sequence is composed of the vertex frame and the three frames before and after it to achieve middle frame enhancement. As Figure 2 shown, the middle frame sequence constructed for the samples in the CASMEII dataset.

[0096] In this embodiment, on multiple datasets, the vertex frame and the 0th, 3rd, 5th, 7th, 10th, and 15th frames before and after it are respectively selected to conduct micro-expression recognition experiments under the same conditions, and the unweighted average recall rate (UAR) is calculated as the evaluation criterion. Table 1 shows the experimental results on two datasets, CASME II and SAMM.

[0097] Table 1 Comparison of experimental results of selecting different intermediate frame sequences in CASME II and SAMM datasets

[0098]

[0099] It can be seen that compared with selecting the 0th frame before and after the vertex frame (i.e., only selecting the vertex frame), selecting the vertex frame and the frames before and after it to form an intermediate frame sequence as the input of the model can improve the UAR result of the micro-expression recognition experiment, indicating that including the frames before and after the vertex frame is effective in improving the recognition accuracy. And on the CASME II dataset, the UAR result obtained by selecting the vertex frame and the 3rd frames before and after it is the best and the time efficiency is the highest. On the SAMM dataset, although the UAR result obtained by selecting the vertex frame and the 5th frames before and after it is the best, the difference from the UAR result obtained by selecting the vertex frame and the 3rd frames before and after it is very small, and the time efficiency is much lower than that of selecting the vertex frame and the 3rd frames before and after it. Therefore, considering the evaluation results and time efficiency comprehensively, the effect of selecting the vertex frame and the 3rd frames before and after it is better.

[0100] The intermediate frame enhancement technology can select reasonable frame samples from the given micro-expression sequence set with relatively high efficiency to enrich the input of the network model, help the network model better learn the characteristics of micro-expressions, and thus obtain better detection effects. Therefore, the intermediate frame enhancement technology can, to a certain extent, improve the problem of low model prediction accuracy caused by insufficient micro-expression sample data.

[0101] S103. Process the vertex frame and the starting frame to obtain an optical flow map representing the dynamic changes of micro-expressions, and process the optical flow map to obtain an optical strain map.

[0102] S1031. Process the vertex frame and the starting frame to obtain an optical flow map representing the dynamic changes of micro-expressions.

[0103] Optical flow method is a method for motion estimation of objects in videos. It is an approximation of the image pattern based on local derivatives between two frames of images, and can display obvious motion changes between frames. The optical flow method reflects the motion of objects in the video in two consecutive frames through the optical flow field. The optical flow field is a two-dimensional motion field, representing the magnitude and direction of the motion of each pixel on the image. The optical flow field reflects more clearly than the original image the area where micro-expression movements occur on the face, enabling the model to learn more complex and abstract features related to the tiny movements and deformations of micro-expressions. For small-scale facial pixel movements, the optical flow field helps to express the degree of deformation of facial muscle tissues.

[0104] In a micro-expression video sequence, the starting frame represents the moment when the expression starts, and the apex frame is the moment when the expression reaches its maximum intensity. The frame difference between the two defines the time span of the expression movement. The selection of the starting frame and the apex frame mainly determines the time range of optical flow calculation, so both directly affect the generation quality of the optical flow map.

[0105] In this embodiment, for a micro-expression video sequence, I(x, y, t) represents the luminance value of the image at pixel point (x, y) at time t. Based on the optical flow constancy hypothesis (luminance consistency hypothesis), on the starting frame, the position of each pixel is (x, y). After reaching the apex frame, this pixel may have moved (x + dx, y + dy). In the optical flow field, u and v respectively represent the moving speeds of the pixel in the horizontal direction (dx) and the vertical direction (dy), u = dt / dx, v = dt / dy, where dt is the time difference between the starting frame and the apex frame. The optical flow field calculated from these two frames has the formula:

[0106] O i ={(u(x, y)), v(x, y))|x = 1, 2,.., X, y = 1,.., Y} (1)

[0107] where X and Y respectively represent the width and height of the frame framework f i,j and u(x, y) and v(x, y) are the optical flow components in the horizontal and vertical directions respectively.

[0108] Visualize the optical flow vector field O to generate an optical flow map, where color and intensity represent the motion direction and speed.

[0109] The optical flow field is a vector field that describes the motion direction and speed of each pixel in the image and is the core result in the calculation process.

[0110] The optical flow map is a visual representation of the optical flow field. Usually, color coding is used to display the motion direction and intensity. Color represents the motion direction, and intensity (luminance) represents the motion speed. The optical flow map is the result of visualizing the optical flow field.

[0111] After the above steps, the optical flow map can help us visually understand the direction and amplitude of facial movements.

[0112] S1032. Process the optical flow map to obtain the optical strain map.

[0113] Since the strain pattern is only related to the deformation of the face and is not easily affected by factors such as lighting conditions and facial occlusion, and has good performance in micro-expression recognition tasks, given the optical flow vector, the optical strain can be derived to describe the facial movement pattern. For sufficiently small facial pixel movements, the optical strain can represent the deformation magnitude of facial muscle tissue, and the optical strain is expressed as follows:

[0114]

[0115] where \(u = [u, v]\) T is the displacement vector, representing the projection of the displacement caused by the deformation of the facial surface in three-dimensional space on the two-dimensional image; represents the derivative of \(u\).

[0116] Expand formula (2) into matrix form:

[0117]

[0118] where the diagonal strain components (\(\varepsilon\) xx , \(\varepsilon\) yy ) are the normal strain components, and (\(\varepsilon\) xy , \(\varepsilon\) yx ) are the shear strain components.

[0119] Specifically, the normal strain measures the change in length perpendicular to a specific direction (such as the x or y direction), and the shear strain measures the change in the angular offset due to shear forces in the plane. Since the muscle movements during micro-expression may involve multiple directions, the magnitude of the optical strain for each pixel can be calculated by the sum of the squares of the normal strain component and the shear strain component, and the formula is:

[0120]

[0121] The optical strain calculation describes the deformation information of local facial movements and quantifies the degree of pixel displacement. The optical strain map converts this information into an intuitive representation through visualization, that is, the optical strain map is obtained.

[0122] After the above steps, the optical strain is visualized to obtain the optical strain map, which is usually used to analyze the deformation and intensity of the micro-expression area of the face. The optical flow map mainly reflects the movement direction and speed, while the optical strain map further quantifies the degree of deformation brought by the movement. The two together provide important feature information for micro-expression recognition.

[0123] Such as Figure 1 、Figure 5 As shown, in step S2, the video frame is input into the first branch network, and the first branch network extracts face features based on deep learning and depth-enhanced channel attention.

[0124] In this embodiment, the three-branch network includes a first branch network, a second branch network, and a third branch network.

[0125] The first branch network includes a deep convolutional neural network and a channel attention module, and the channel attention module includes a global average pooling module and a max pooling module.

[0126] In this embodiment, the first branch network introduces ResNet-18 as a transfer model and adds a depth-enhanced channel attention module ECANet for extracting face features.

[0127] ResNet-18 is a classic deep convolutional neural network suitable for image feature extraction. It has multiple convolutional and pooling layers and can automatically learn and capture features at different levels, including low-level textures and high-level semantic information. However, microexpressions often contain small but important facial dynamic changes that ResNet-18 cannot capture. The ECANet attention mechanism focuses on enhancing the network's attention to specific channels, which is particularly important for microexpression recognition because different channels may contribute differently to different facial features. ECANet can highlight the channels containing microexpression information and reduce the attention to irrelevant information by adaptively adjusting the weights of the channels. This attention mechanism helps improve the network's perception and expressiveness. Therefore, ResNet-18 is combined with the ECANet attention mechanism.

[0128] The ECANet module includes an adaptive average pooling (GAP) module and an adaptive max pooling (AMP) module for capturing the key features of the input tensor. Channel attention is calculated through adaptive average pooling (Global Average Pooling) and adaptive max pooling (Adaptive Max Pooling), thereby enhancing the representation ability of the input features. The ECANet module adopts a non-dimension-reducing local cross-channel interaction strategy, which can avoid the adverse effects of dimension reduction operations on channel attention prediction, thus maintaining good performance while effectively reducing the model complexity.

[0129] The first branch network extracts face features, and the specific process is as follows:

[0130] S201. After the video frame is input into the deep convolutional neural network, the original face features are extracted.

[0131] Such as Figure 1 、 Figure 5As shown, in this embodiment, the video frame is input into the ResNet-18 network to obtain the original face feature X, and the output feature X is respectively input into the adaptive average pooling (GAP) module and the adaptive max pooling (AMP) module of the ECANet module.

[0132] S202. Input the original face feature into the adaptive average pooling, perform average pooling on each channel, obtain the average pooling channel attention weight through convolution and function activation processing, and multiply the average pooling channel attention weight by the original face feature to obtain the global average pooling output feature.

[0133] In this embodiment, the original face feature X is input into the adaptive average pooling (GAP) module. The feature vector X with dimensions of H×W×C is averaged on each channel to generate a feature vector with a spatial dimension of 1×1×C, where C is the number of channels. Then, the channel attention weight λ1 is obtained through convolution and Sigmoid activation processing, and the weight λ1 is multiplied by the original input X to obtain the global average pooling output feature.

[0134] S203. Input the original face feature into the adaptive max pooling for max pooling, obtain the max pooling channel attention weight through convolution and function activation processing, and multiply the max pooling channel attention weight by the original face feature to obtain the max pooling output feature.

[0135] In this embodiment, the original face feature X is input into the adaptive max pooling (AMP) module. Adaptive max pooling is applied to the input feature X, and the channel attention weight λ2 is obtained through convolution and Sigmoid activation processing. The weight is multiplied by the original input X to obtain the output feature of the adaptive max pooling.

[0136] S204. Perform feature fusion on the global average pooling output feature and the max pooling output feature to obtain the face feature.

[0137] In this embodiment, the output feature of the adaptive max pooling module and the output feature of the adaptive max pooling are subjected to feature fusion to obtain the face feature.

[0138] After the above steps, deeper feature extraction and representation can be achieved in the micro-expression recognition task. ResNet-18 provides the basic feature extraction ability, while the ECANet attention mechanism further optimizes and strengthens these features. This synergy can effectively extract and capture the key information in the micro-expression images, improving the performance and robustness of the model.

[0139] Such asFigure 1 , Figure 3 and Figure 4 As shown in and

[0140] , in step S3, the optical flow map is input into the second branch network, and the second branch network uses the second shallow convolutional neural network and the second face feature enhanced attention to extract the features of the optical flow map.

[0140] As Figure 3 , Figure 4 shown, in this embodiment, the second branch network adopts a trained shallow 3DCNN model and combines a new face feature enhanced attention module FEAM (Face-Enhanced Attention Module) to extract the features of the optical flow map.

[0141] In the optical flow and optical strain features, important information about motion, shape change, and time series is included. Obviously, using a 3DCNN network helps to extract more information about these spatio-temporal characteristics. Therefore, the second and third branches of the model in this paper both introduce a 3DCNN network to extract optical flow features and optical strain features respectively.

[0142] The convolutional layers in a 3D convolutional neural network (3DCNN) are all three-dimensional. Different spatio-temporal features in different time domains are extracted by using 3D convolutional kernels with different parameters, and the features in both the time and space domains are extracted simultaneously. Each three-dimensional convolutional layer of the 3DCNN network has the ability to learn different levels of features, from low-level features such as edges and textures to high-level features, enabling the network to construct and express complex feature hierarchies layer by layer. In addition, the 3DCNN network has fewer parameters and computational amounts, so it can extract both discriminative high-level features and the details of micro-expressions through lightweight calculations.

[0143] Since micro-expression movements have the characteristic of locality and only appear in some regions of the human face, there are no effective classification features in most facial regions, and only a few regions with micro-expression movements can provide information helpful for micro-expression classification. And the attention mechanism helps to focus on specific facial regions, learn and obtain important features. Therefore, on the basis of using 3DCNN to learn optical flow features and optical strain features, a new FEAM face feature enhanced attention module is added to construct a new branch network 3DCNN-FEAM.

[0144] The second branch network and the third branch network have the same structure, both including: three parallel streams, a channel layer, a face feature enhanced attention module, a second fusion module, and a fully connected layer.

[0145] The process of extracting the features of the optical flow map is as follows:

[0146] S301. In the second branch network, each parallel stream uses convolutional kernels with different parameters to extract features of different sizes. After the parallel streams, the features in the three parallel streams are concatenated at the channel level to obtain the first fused feature.

[0147] As Figure 3 shown, in this embodiment, the optical flow image is resampled and then input into the second branch network 3DCNN-FEAM.

[0148] In the second branch network, the optical flow map passes through 3 parallel streams, and each parallel stream consists of a convolutional layer and a max pooling layer. Each parallel stream uses convolutional kernels with different parameters to extract features of different sizes, and supplements the small-scale input with different numbers of 5*5 convolutional kernels to avoid the problem of data underfitting. Then, the feature maps are stacked by channel and enter the max pooling layer. The max pooling operation can highlight the main features while eliminating redundancy. After the parallel streams, the features in the three streams are concatenated at the channel level to achieve the first fusion and obtain the first fused feature F.

[0149] S302. The first facial feature enhanced attention module combines and performs weighted averaging on the first fused feature.

[0150] The first fused feature is input into the FEAM facial feature enhanced attention module, and the features are combined and weighted averaged to enhance the model's ability to capture key information, obtaining the features of the FEAM facial feature enhanced attention module.

[0151] The facial feature enhanced attention module includes a convolutional attention module, a double pooling layer, and a fully connected layer. The convolutional attention module includes a channel attention module and a spatial channel attention module.

[0152] As Figure 4 shown, in this embodiment, on the basis of the convolutional attention module structure, the facial feature enhanced attention module FEAM introduces a double pooling layer and a fully connected layer to further improve the expression ability of the feature map and make the neural network more focused on capturing the subtle changes of the face. This module makes the neural network emphasize the extraction of facial features more during the learning process, helping to capture key information such as facial expressions and morphologies more accurately.

[0153] Among them, the double pooling layer includes average pooling and max pooling. Average pooling captures the overall information of the features, while max pooling focuses on the most significant part of the features. The double pooling operation enables the network to consider information at different scales simultaneously when processing features, thus understanding the structure of the face more comprehensively. The processing of the fully connected layer helps to further integrate the feature information, enabling the network to better understand the relationship between different features during the learning process and improving the perception ability of facial details.

[0154] S3021. Perform max - pooling on the first - fused feature to obtain the output feature of the max - pooling layer, perform average - pooling on the first - fused feature to obtain the output feature of the average - pooling layer, and combine the output feature of the max - pooling layer and the output feature of the average - pooling layer and input them into the fully - connected layer to obtain the output feature of the fully - connected layer.

[0155] In this embodiment, input the first - fused feature F into the max - pooling layer and the average - pooling layer of the FEAM module respectively to obtain the output feature F of the max - pooling layer max and the output feature F of the average - pooling layer avg . Combine the output feature F of the max - pooling layer avg and the output feature F of the average - pooling layer avg and input them into the fully - connected layer to obtain the output feature F of the fully - connected layer fully . The formula is:

[0156] F fully = ReLU(W1·(F max ⊕F avg )) (5)

[0157] where F fully represents the output feature of the fully - connected layer, F max represents the output feature of the max - pooling layer, F avg represents the output feature of the average - pooling layer, W1 represents the weight parameter in the fully - connected layer, which is a trainable matrix used for linearly transforming the feature inputs F max and F avg .

[0158] S3022. Input the output feature of the fully - connected layer into the channel attention module to obtain the channel attention feature.

[0159] Input the output feature F of the fully - connected layer fully into the channel attention module to obtain the channel attention feature M chan . The formula is:

[0160] M chan = σ(W2·F fully ) (6)

[0161] where W2 represents the weight parameter of the channel attention module, which is another trainable matrix used for adjusting the importance of each channel in F fully .

[0162] S3023. Perform channel - attention weighting on the channel attention feature and the output feature of the fully - connected layer to obtain the channel - attention weighted feature.

[0163] Input the channel attention feature M chanWith the output features F of the fully connected layer fully Channel attention weighted features F are obtained by performing channel attention weighting chan , and the formula is:

[0164]

[0165] S3024. Input the channel attention weighted features into the spatial attention module to obtain spatial attention features.

[0166] Input the channel attention weighted features F chan into the spatial attention module to obtain spatial attention features M spat , and the formula is:

[0167] M spat = σ(f conv (F chan ))(8)

[0168] where f conv is a convolution operation.

[0169] S3025. Perform spatial attention weighting on the spatial attention features and the channel attention weighted features to obtain spatially attention weighted features.

[0170] Perform spatial attention weighting on the spatial attention features M spat and the channel attention weighted features F chan to obtain spatially attention weighted features F spat , and the formula is:

[0171]

[0172] S303. Further fuse the features obtained after passing through the first facial feature enhancement attention module with the first fusion feature to obtain a second fusion feature.

[0173] Fuse the spatial attention features F spat and the first fusion feature F to obtain the final output feature F out , and the formula is:

[0174] F out = concat(F,F spat )(10)

[0175] Implement the second fusion to obtain the second fusion feature.

[0176] S304. Process the second fusion feature through a fully connected layer and output the features of the optical flow map.

[0177] Process the second fusion feature through a fully connected layer, which includes three layers. During the process: the first two fully connected layers respectively map the fused second fusion feature to a higher-level abstract representation and gradually reduce the feature dimension;

[0178] The third fully connected layer then converts the features output by the first two fully connected layers into a dimension suitable for the output of this second branch network.

[0179] Specifically, pass the second fusion feature through three fully connected layers. Among them, the first two fully connected layers, also known as two dense layers with different dimensions, namely dense layer 1 and dense layer 2, respectively map the second fusion feature to a higher-level abstract representation and gradually reduce the feature dimension. The third fully connected layer then converts the features into a dimension suitable for the output of this network to obtain the optical flow feature. Realize the extraction of the optical flow feature.

[0180] As Figure 1 、 Figure 3 and Figure 4 shown, in step S4, input the optical strain map into the third branch network, and the third branch network uses the third shallow convolutional neural network and the third face feature enhanced attention to extract the features of the optical strain map.

[0181] The third branch network adopts a trained shallow 3DCNN model and combines a new face feature enhanced attention module FEAM (Face-Enhanced Attention Module) to extract the features of the optical strain map.

[0182] The third branch network has the same structure as the second branch network, including: three parallel streams, a channel layer, a second face feature enhanced attention module, a second fusion module, and a fully connected layer;

[0183] As Figure 3 shown, in this embodiment, after resampling the optical strain image, input it into the third branch network 3DCNN-FEAM to extract the optical strain feature.

[0184] The specific extraction process is the same as that in step S301 to obtain the features of the optical strain map. Realize the extraction of the features of the optical strain map.

[0185] As Figure 1 shown, in step S5, fuse the face features, the features of the optical flow map, and the features of the optical strain map.

[0186] S501. Fuse the face features, the optical flow features, and the optical strain features to obtain an aggregated feature.

[0187] S5011. Combine the features of the optical flow map and the features of the optical strain map output by the two 3DCNN-FEAM branch networks to form a new optical flow feature block.

[0188] In this embodiment, the feature F of the optical flow map optical and the feature F of the optical strain map strain are combined through feature splicing, and the fused feature F fused is used as the input of the subsequent feature processing module. The formula is:

[0189] F fused = F optical ⊕ F strain (11).

[0190] S5012. Perform L2 regularization processing on the optical flow feature block and the face feature block output by the ResNet-18-ECA network respectively.

[0191] In this embodiment, perform L2 regularization processing on the optical flow feature block F fused and the face feature block F output by the ResNet-18-ECA network face respectively. The formula is:

[0192]

[0193] Among them, ||F||2 represents the L2 norm of the feature.

[0194] Through the L2 regularization operation, the feature values are normalized to control the model complexity and reduce the risk of overfitting, and by constraining the weight parameters during the learning process, unnecessary noise and overfitting problems of the model are prevented.

[0195] S5013. Perform feature fusion on the regularized optical flow feature block and the face feature block to obtain an aggregated feature.

[0196] In this embodiment, perform feature fusion on the regularized optical flow feature block and the face feature block to obtain an aggregated feature F final , and the formula is:

[0197]

[0198] After the above steps, effective recognition features are obtained, providing data support for subsequent feature classification.

[0199] As Figure 1 shown, in step S6, perform micro-expression classification on the fused features to obtain the classification results of micro-expressions.

[0200] S601. Probability prediction.

[0201] Use the Softmax function to predict the aggregated features. The prediction formula is:

[0202]

[0203] According to formula (14), the probability values belonging to the three micro-expression categories of positive, negative, and surprised are obtained respectively. The category with the highest probability value is the classification result of the micro-expression. The formula is:

[0204]

[0205] where y predicted represents the final predicted category (positive, negative, or surprised) of the input micro-expression video.

[0206] S602. Constrain the prediction probability through the loss function.

[0207] For each sample (i.e., a micro-expression sequence) in the micro-expression dataset, after being processed by the above network model, a corresponding probability distribution will be obtained. This distribution is composed of the probabilities of each category output by the Softmax layer of the last layer of the network. The lack of micro-expression sample data is one of the important reasons for the imbalance of the sample category distribution. To effectively solve the problem of the imbalance of the micro-expression data sample category distribution, a new loss function is proposed to constrain the prediction probability distribution of the network model to improve the accuracy of the model for micro-expression recognition. Introduce the IM Loss (Information Maximizing Loss) function into micro-expression recognition and combine it with the Focal Loss function to jointly constrain the model, which can better improve the imbalance of the sample distribution caused by the lack of sample data, and then solve the problem of low model recognition accuracy.

[0208] Among them, the IM Loss function constrains the probability distribution corresponding to the sample. Its goal is to maximize the degree of association between the predicted probability distribution and the true label distribution, that is, the amount of information shared between the two distributions, so as to prompt the probability distribution generated by the model to be as close as possible to the true distribution, and then enhance the recognition ability of the model for all categories. The Focal Loss function constrains the prediction probability of a specific category, that is, it mainly focuses on a single probability value, especially the probability values of those samples that are difficult to predict or misclassified. By adjusting their loss weights, the recognition ability of the model for minority categories can be improved. Since the probability distribution depends on the probability value, the Focal Loss function can essentially also be regarded as a constraint on the probability distribution corresponding to the sample.

[0209] The IM Loss function is introduced into micro-expression recognition for the first time. The IM Loss function consists of two parts, namely the entropy minimization part and the diversity maximization part.

[0210] The entropy minimization part is mainly to ensure that the predicted label distribution output by the model is as deterministic as possible for a given input sample, that is, to reduce the uncertainty of the model. Mathematically, entropy is a measure of the uncertainty of a random variable. For a probability distribution, the smaller its entropy, the more it tends to a certain class and the smaller the uncertainty. In a classification task, entropy minimization usually means that the model is very confident in its prediction results.

[0211] The diversity maximization part aims to ensure that the model gives different class prediction results for different input samples, that is, to increase the diversity of the model output. This is especially important for tasks in scenarios where samples are in different domains or the sample classes are imbalanced, because the model may overfit to samples of several classes and ignore minority classes. By maximizing diversity, it can be ensured that each class receives appropriate attention, thus making the overall prediction more balanced.

[0212] The purpose of the IM Loss function is to make the classification labels as diverse as possible, that is, each class is relatively average. To achieve this goal, only need to calculate the average value of the outputs of the entire target domain to make its entropy the largest, keep the samples that are not easy to classify away from the classification boundary, and increase the confidence.

[0213] The IM Loss function is defined as:

[0214]

[0215] where, L ent represents the entropy minimization part of the IM Loss function, which is used to calculate the cross-entropy loss of the predicted probability of each sample and take the average over the entire dataset; R represents all micro-expression datasets in this embodiment, R = {x1, x2,..., x V}, where the sample x t , t = 1, 2,... V represents a micro-expression video sequence in the micro-expression dataset R, and V is the total number of samples in the dataset R; represents taking the average of the sum of the losses corresponding to each sample x t in the micro-expression dataset R. The loss corresponding to each sample x t is the sum of the cross-entropy losses of the predicted probabilities of all possible classes of this sample x t under the current parameters of the model; K is the total number of sample classes in the dataset, is the probability value that the model predicts the sample x t belongs to the class k; L divRepresents the diversity maximization part of the IMLoss function, which is used to directly calculate each sample x t The KL divergence between the predicted probability distribution and the uniform distribution to promote the diversity of predictions, avoid preference for certain classes, and sum these values; where is the average probability that each sample is predicted as class k, is the expected probability of the uniform distribution. In this paper, it is assumed that the probability of selecting each class is equal.

[0216] After the above steps, while ensuring the diversity and average of classification labels, the confidence of samples that are not easily classified in the target domain is enhanced, the confidence of the model for micro-expression classification is enhanced, and the balanced recognition of various micro-expressions is ensured.

[0217] In traditional cross-entropy loss, all samples are treated equally, which easily leads to the model being dominated by a large number of easily classified samples, thus ignoring those few but difficult-to-classify samples. To solve this problem, an adjustable factor is introduced in the Focal Loss function, which can dynamically adjust the weight of each sample.

[0218] The Focal Loss function is defined as:

[0219]

[0220] where, represents the probability value that the model predicts x t as class k. The Focal Loss calculates the loss for each sample in the dataset independently; where k ∈ [k1, k2, k3], k1 is the positive class, k2 is the negative class, and k3 is the surprised class; α k is the preset weight parameter used to balance the weights of different classes, and this weight is set smaller for negative samples; γ is the adjustment parameter. By setting γ, the weight of difficult-to-classify samples can be increased, thus increasing the attention of the model to difficult-to-classify samples.

[0221] After the above steps, the problems caused by class imbalance can be alleviated, especially for those few but difficult-to-classify samples, enabling the model to pay more attention to those difficult-to-classify samples and making the model have better classification effects when facing various classes of samples.

[0222] To improve the model's recognition ability for difficult-to-classify samples and ensure the diversity and certainty of the output, a strategy of sample weighting based on the classification difficulty of sample classes is adopted. Specifically, the Focal Loss pays more attention to difficult-to-classify samples by dynamically adjusting the weights to address the class imbalance problem, while the IMLoss function increases the diversity of the model output while improving the model's certainty.

[0223] The Total Loss function is defined as:

[0224] L Total (x t ) = αL focal (p k (x t ) + βL LM (p k (x t )) (18)

[0225] where α and β are weight coefficients used to adjust the proportion of the two losses, represents the probability value that the model predicts x t as the class k, and L focal (p k (x t )) represents the Focal Loss function, and L IM (p k (x t )) represents the IM Loss function.

[0226] Through the above steps, the model can not only pay sufficient attention to hard samples at the overall level of the dataset, but also distinguish difficult-to-discriminate samples within a specific class, thereby improving the classification accuracy and robustness of the model.

[0227] Through steps S1 - S6, the micro-expression recognition model based on the three-branch network is trained and tested. Each time, all the micro-expressions of one subject are selected for testing, and the micro-expressions of other subjects are used as training data. After obtaining the trained model, the micro-expressions of the selected subject are then input into the model for testing. To ensure that the information of the tester is excluded during the training phase, iterative cross-validation is performed y times to end the test, where y is set according to the number of subjects.

[0228] In this embodiment, the evaluation criteria of unweighted average recall (UAR) and unweighted F1 score are used for evaluation. UF1 is also known as the macro-average F1 score and is determined by averaging the F1 scores of each class. Among them, the F1 score can be interpreted as the weighted average of precision and recall, reaching the best value when the F1 score is equal to 1 and the worst value when it is equal to 0. The relative contributions of precision (P) and recall (R) to the F1 score are the same. Therefore, UF1 is a good choice in multi-classification problems as it can equally emphasize rare classes. To calculate UF1, first calculate the F1 value of each class k, and then calculate the average of the F1 values of each class, which is UF1. The formula is:

[0229]

[0230] Among them, Q is the number of tags of micro-expressions, and F1 k is the F1 value of the classification result of class k, TP k , FP k and FN k respectively represent the number of true positives, false positives, and false negatives in the classification result of class k.

[0231] The unweighted average recall (UAR), also known as the "balanced accuracy", is a more reasonable evaluation criterion instead of the weighted average recall. UAR is obtained by adding the accuracies of each class and then taking the average over the number of classes. The formula is:

[0232]

[0233] Among them, TP c represents the number of true positive samples in the classification result of class c, P c is the predicted probability of class c, and M c represents the total number of samples in class c, that is, the number of samples actually belonging to class c in the dataset. It is used for weighted calculation of UAR to balance the problem of uneven class sample numbers.

[0234] The advantage of this evaluation method is that compared with the weighted average recall, its prediction result will not be biased towards the larger class in the data. In the MEGC2019 composite dataset, the number of samples in different classes is unbalanced. The ratio of the number of samples of the three types of expressions, surprise, positive, and negative, is approximately 1:1.3:3, with a serious problem of data class imbalance. Therefore, in this paper, the unweighted average recall (UAR) and unweighted F1-score (UF1) metrics are used to evaluate the performance of the method for three-class classification. These two metrics can well measure the performance of the model in terms of the accuracy and precision of positive example recognition. By treating each class equally, the ability of the model to correctly identify the less frequent classes is also taken into account, thus evaluating its performance more fairly.

[0235] In this embodiment, experiments are conducted on the datasets CASMEII, SMIC, SAMM, and the MEGC2019 composite dataset. The method of this embodiment is respectively compared with the method based on manually extracted features and the method based on deep learning for recognition. The comparison results are shown in Table 2.

[0236] Table 2 Comparison of micro-expression recognition performance of different methods on the three-class classification dataset

[0237]

[0238] According to Table 2, it can be seen that in all training and test datasets, compared with other methods, both UF1 and UAR of the recognition method in this embodiment have been significantly improved. The recognition results of the method in this embodiment on three categories are better than those of other methods, and it is more effective for three-class micro-expression recognition.

[0239] In addition, the method in this embodiment is also applied to multi-class micro-expression recognition. Using the CASME II and SAMM datasets, the classification recognition of five types of micro-expressions, namely "disgust", "happiness", "other", "repression", and "surprise", is carried out, and the recognition comparison is made with other methods. The recognition comparison results are shown in Table 3.

[0240] Table 3 Comparison of recognition rates of different methods for multi-class micro-expression recognition

[0241]

[0242] For multi-class micro-expression recognition, the UF1 and UAR results of the recognition method in this embodiment are also significantly higher than those of other methods. The recognition results of the method in this embodiment on multiple categories are better than those of other methods, and it also has a significant effect on multi-class micro-expression recognition.

[0243] This embodiment discloses a micro-expression recognition method based on a three-branch network. One deep convolutional neural network and two shallow 3D convolutional neural networks in parallel in the model are respectively used for the extraction of face features, optical flow features, and optical strain features. The full fusion of these different types of features can effectively reduce the loss of subtle face features and improve the accuracy of micro-expression recognition. A facial feature enhancement attention module is proposed in the shallow 3D CNN, and a depth-enhanced channel attention module is introduced in the deep convolutional neural network. This multi-attention mechanism enables the model to more specifically capture key face features and reduce the interference of redundant information. The constructed new loss function combining the IM Loss constraint term and the Focal Loss constraint term significantly improves the model's recognition ability and recognition accuracy for difficult samples.

[0244] In this embodiment, a subject is selected from the CASME II dataset for ablation experiments.

[0245] (1) Network architecture ablation

[0246] The first network branch in the model for extracting face features is denoted as N A , the second network branch for extracting optical flow features is denoted as N B , the third network branch for extracting optical strain features is denoted as N C , and the complete network model is denoted as N D, denote the depth-enhanced channel attention module ECANet in the first network branch as M E , denote the facial feature-enhanced attention module FEAM in the second and third network branches as M F .

[0247] Network architecture ablation experiment 1 is to remove the first face feature extraction branch from the complete model. Architecture experiment 2 is to remove the second optical flow feature extraction branch from the complete model. Architecture experiment 3 is to remove the third optical strain feature extraction branch from the complete model. Architecture experiment 4 is to remove the ECANet module from the complete model. Architecture experiment 5 is to remove the FEAM module from the second branch of the complete model. Architecture experiment 6 is to remove the FEAM module from the third branch of the complete model. Architecture experiment 7 is the complete model.

[0248] After the network architecture ablation experiment, the comparison results are shown in Table 4

[0249] Table 4 Results of network architecture ablation experiment

[0250]

[0251] It can be seen that removing any one of the feature extraction branches or attention modules will lead to a decline in the algorithm performance, indicating that the multi-feature fusion mechanism and multi-attention mechanism in this embodiment effectively improve the micro-expression recognition accuracy of the model.

[0252] (2) Ablation of loss function

[0253] Five ablation experiments were designed. Since in the Focal Loss function in the method of this embodiment, when its parameter γ is set to 0, it degenerates into the cross-entropy loss function commonly used in the existing micro-expression recognition methods. Therefore, ablation experiment 1 uses the most basic cross-entropy loss function, denoted as L CE , and use this as the baseline. Ablation experiment 2 uses the cross-entropy loss function L CE and the IMLoss function L IM , ablation experiment 3 only uses the Focal Loss function L focal , ablation experiment 4 only uses the IMLoss function L IM , ablation experiment 5 uses the Focal Loss function L focal and the IMLoss function L IM , that is, the complete loss function L of this article Total .

[0254] After the loss function ablation experiment, the comparison results are shown in Table 5

[0255] Table 5 Results of loss function ablation experiment

[0256]

[0257] It can be seen that in this embodiment, the total loss function composed of the Focal Loss function and the IM Loss function has the best effect.

[0258] Embodiment 2

[0259] This embodiment discloses a micro-expression recognition system based on a three-branch network, including:

[0260] A preprocessing module for obtaining a video sequence, preprocessing the video sequence to obtain an optical flow map, and further processing the optical flow map to obtain an optical strain map, where the video sequence includes video frames;

[0261] A feature extraction and fusion module for inputting video frames into a first branch network, where the first branch network extracts face features based on deep learning and depth-enhanced channel attention; inputting the optical flow map into a second branch network, where the second branch network extracts features of the optical flow map using a second shallow convolutional neural network and second facial feature enhancement attention; inputting the optical strain map into a third branch network, where the third branch network extracts features of the optical strain map using a third shallow convolutional neural network and third facial feature enhancement attention; and fusing the face features, features of the optical flow map, and features of the optical strain map;

[0262] An identification and classification module for classifying micro-expressions of the fused features to obtain a classification result of the micro-expressions.

[0263] Based on providing a micro-expression recognition system based on a three-branch network, the method steps in Embodiment 1 are implemented.

[0264] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. The present invention is not limited to any specific combination of hardware and software.

[0265] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that on the basis of the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A micro-expression recognition method based on a three-branch network, characterized in that, Including: Obtain a video sequence, preprocess the video sequence to obtain an optical flow map, and further process the optical flow map to obtain an optical strain map. The video sequence includes video frames; Input the video frames into a first branch network, and the first branch network extracts face features based on deep learning and depth-enhanced channel attention; Input the optical flow map into a second branch network, and the second branch network extracts the features of the optical flow map by using a second shallow convolutional neural network and second facial feature enhancement attention; Input the optical strain map into a third branch network, and the third branch network extracts the features of the optical strain map by using a third shallow convolutional neural network and third facial feature enhancement attention; Fuse the face features, the features of the optical flow map, and the features of the optical strain map; Perform micro-expression classification on the fused features to obtain the classification result of the micro-expression; The second branch network and the third branch network have the same structure, both including: three parallel streams, a channel layer, a facial feature enhancement attention module, a second fusion module, and a fully connected layer; In the second branch network, each parallel stream uses convolutional kernels with different parameters to extract features of different sizes. After the parallel streams, the features in the three parallel streams are concatenated together at the channel level to obtain a first fused feature; The first facial feature enhancement attention module combines and performs weighted averaging on the first fused feature; Further fuse the features obtained after the first facial feature enhancement attention module with the first fused feature to obtain a second fused feature; Process the second fused feature through the fully connected layer and output the features of the optical flow map.

2. The micro-expression recognition method based on a three-branch network according to claim 1, characterized in that, The video frame is a vertex frame, and the vertex frame refers to the frame in which the intensity or manifestation of the micro-expression is most obvious, or the intermediate frame sequence composed of the vertex frame and its front and back frames.

3. The micro-expression recognition method based on a three-branch network according to claim 1, characterized in that, Preprocess the video sequence, including locating the vertex frame to determine the position of the vertex frame in the micro-expression sequence for extracting the time information before and after the vertex frame; And process the vertex frame and the starting frame to obtain an optical flow map representing the dynamic change of the micro-expression, and process the optical flow map to obtain an optical strain map.

4. The micro-expression recognition method based on a three-branch network according to claim 1, characterized in that, The facial feature enhancement attention module includes a convolutional attention module, a double pooling layer, and a fully connected layer. The convolutional attention module includes a channel attention module and a spatial channel attention module; Perform max pooling on the first fused feature to obtain the output feature of the max pooling layer; Perform average pooling on the first fused feature to obtain the output feature of the average pooling layer; Merge the output feature of the max pooling layer and the output feature of the average pooling layer and input them into the fully connected layer to obtain the output feature of the fully connected layer; Input the output feature of the fully connected layer into the channel attention module to obtain the channel attention feature; Perform channel attention weighting on the channel attention feature and the output feature of the fully connected layer to obtain the channel attention weighted feature; Input the channel attention weighted feature into the spatial attention module to obtain the spatial attention feature; Perform spatial attention weighting on the spatial attention feature and the channel attention weighted feature to obtain the spatial attention weighted feature; Fuse the spatial attention feature and the first fused feature to obtain the second fused feature.

5. The micro-expression recognition method based on a three-branch network according to claim 1, characterized in that Process the second fused feature through the fully connected layer. The fully connected layer includes three layers, and the processing includes: The first two fully-connected layers respectively map the fused second fused features to a higher-level abstract representation and gradually reduce the feature dimension; The third fully-connected layer then converts the features output by the first two fully-connected layers into a dimension suitable for the output of the second branch network.

6. The micro-expression recognition method based on a three-branch network according to claim 1, characterized in that Before the optical flow map is input into the second branch network and the optical strain map is input into the third branch network, resampling is performed respectively.

7. The micro-expression recognition method based on a three-branch network according to claim 1, characterized in that, The first branch network includes a deep convolutional neural network and a channel attention module; the channel attention module includes a global average pooling module and a max pooling module; After the video frame is input into the deep convolutional neural network, the original face features are extracted; The original face features are input into adaptive average pooling, average pooling is performed on each channel, and the average pooling channel attention weights are obtained through convolution and function activation processing; The average pooling channel attention weights are multiplied by the original face features to obtain the global average pooling output features; The original face features are input into adaptive max pooling for max pooling, and the max pooling channel attention weights are obtained through convolution and function activation processing; The max pooling channel attention weights are multiplied by the original face features to obtain the max pooling output features; The global average pooling output features and the max pooling output features are fused to obtain the face features.

8. The micro-expression recognition method based on a three-branch network according to claim 1, characterized in that, After obtaining the micro-expression classification result, the total loss function is constructed through two loss functions, IM Loss and Focal Loss, to optimize the model.

9. A micro-expression recognition system based on a three-branch network, characterized in that, Implement a micro-expression recognition method based on a three-branch network according to any one of claims 1-8, including: A preprocessing module, configured to obtain a video sequence, preprocess the video sequence to obtain an optical flow map, and further process the optical flow map to obtain an optical strain map, where the video sequence includes video frames; A feature extraction and fusion module, configured to input the video frames into the first branch network, and the first branch network extracts face features based on deep learning and depth-enhanced channel attention; input the optical flow map into the second branch network, and the second branch network extracts the features of the optical flow map by using a second shallow convolutional neural network and a second facial feature enhancement attention; input the optical strain map into the third branch network, and the third branch network extracts the features of the optical strain map by using a third shallow convolutional neural network and a third facial feature enhancement attention; fuse the face features, the features of the optical flow map, and the features of the optical strain map; An identification and classification module, configured to perform micro-expression classification on the fused features to obtain the classification result of the micro-expression.

Citation Information

Patent Citations

  • Micro-expression recognition algorithm

    CN110287801A

  • Cross-domain facial expression recognition method irrelevant to source domain data

    CN114973350A

  • Micro-expression recognition method based on face key point and optical flow feature fusion

    CN118781636A