Micro-expression recognition method based on multiple modes

By combining the multimodal recognition method with optical flow maps and depth difference features, a dual-stream network of SwimTransformer is constructed, which solves the problem of low accuracy in micro-expression recognition and achieves higher recognition accuracy and robustness.

CN120599679APending Publication Date: 2025-09-05HEBEI UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510719717.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The accuracy of micro-expression recognition in existing technologies is low, deep learning methods have problems of unstable performance and overfitting when processing micro-expressions, and single-modal recognition methods fail to fully utilize multi-scale information.

Method used

A multimodal recognition method is adopted, combining facial images and depth images, and data preprocessing is performed through optical flow maps and depth difference features. The training data is expanded using a data enhancement algorithm, and a two-stream backbone network based on SwimTransformer is constructed for feature extraction and fusion to output micro-expression classification results.

Benefits of technology

It improves the accuracy and robustness of micro-expression recognition, enhances the ability to perceive dynamic changes in facial micro-expressions, overcomes the limitations of single-modality methods, and improves the accuracy and stability of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599679A_ABST
    Figure CN120599679A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-mode-based micro-expression recognition method, which comprises the steps of acquiring data to obtain a face image and a depth image, preprocessing the data to remove redundant information and obtain an optical flow graph and depth difference information, training a model, deploying the model and the like. The data acquisition module is also used for acquiring data of a to-be-recognized object and inputting the data into the trained micro-expression recognition model; and finally, the identification model outputs a prediction result of the to-be-identified data. According to the multi-modal-based micro-expression recognition system method, analysis can be carried out according to image data of the test object, the micro-expression state of the test object can be automatically classified and recognized, and compared with a single-modal recognition system and method, the performance is better.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The present invention belongs to the technical field of data processing and image recognition, and particularly relates to a micro-expression recognition method based on multimodality.

[0002] Facial microexpressions are transient representations of human emotional changes, typically appearing briefly (0.04 to 0.2 seconds) with a small range of motion. Unlike regular facial expressions, microexpressions are extremely short-lived and have limited range of motion, making them valuable for applications in psychological diagnosis, security monitoring, criminal investigation, business negotiations, and human-computer interaction. However, due to their subtle features and rapid changes, automatic microexpression recognition technology faces a series of challenges, including the subtle nature of the features, the scarcity of samples, and individual differences.

[0003] In recent years, with the development of deep learning technology, an increasing number of micro-expression recognition studies have begun to adopt deep learning methods. These methods automatically learn and extract features through multi-layer neural networks, effectively improving the performance of micro-expression recognition. Compared with traditional feature engineering methods, deep learning methods have higher recognition speed and accuracy, significantly improving the efficiency of micro-expression recognition. However, deep learning methods also have limitations when dealing with micro-expression recognition. As the network depth increases, subtle motion features may be weakened or even completely lost by the complex structure of the network, resulting in unstable performance of the model in recognizing micro-expressions. In addition, most deep learning networks use techniques such as residual connections and same-layer feature fusion to improve recognition performance. However, these methods fail to fully utilize the advantages of multi-scale information and cannot effectively capture the changes in facial micro-expressions at different scales.

[0004] Furthermore, micro-expression databases generally have smaller sample sizes. Compared to macro-expression datasets, their scarcity limits the effectiveness of model training. Insufficient data samples can lead to overfitting in deep learning models, where the model performs well on the training data but generalizes poorly to unseen data, further impacting the model's practical application. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a micro-expression recognition system based on multimodality to solve the problem of low accuracy of micro-expression recognition using only a single modality in the prior art.

[0006] The technical solution adopted in the present invention is:

[0007] A multimodal micro-expression recognition method comprises the following steps:

[0008] S01: Multimodal training data acquisition, using a camera to collect face images and depth images;

[0009] S02: Preprocess multimodal training data to remove redundant information and obtain discriminative input features, including optical flow maps and depth differences, and use data augmentation algorithms to expand training data;

[0010] S03: Build and train a multimodal micro-expression recognition model, train the pre-processed image data and its attached optical flow and depth information, and output the prediction results;

[0011] S04: Deploy the trained model to the computer;

[0012] S05: Acquire the micro-expression to be recognized using the steps of S01 and perform pre-processing through S02;

[0013] S06: The pre-processed micro-expression data to be recognized is sent to the model in S03 for recognition, and the corresponding category of the micro-expression is output.

[0014] Furthermore, the specific steps of S02 include:

[0015] S201: normalizing and cropping the face to obtain basic image data with background removed and face aligned, meeting the network size input requirements;

[0016] S202: Optical flow is a discriminative feature that contains two-dimensional pixel motion information. The TV-L1 algorithm is used to calculate optical flow using the start and peak frames of the normalized and cropped RGB image. Horizontal and vertical optical flows are obtained, and optical strain is calculated. The channels are then stitched together to obtain an optical flow map.

[0017] S203: The depth feature is a supplement to the optical flow feature, which contains the motion information of the facial area in depth. The depth difference is calculated using the depth images of the normalized cropped starting frame and the peak frame.

[0018] S204: Using data enhancement methods, the optical flow map and the depth difference are processed in pairs, and a random combination of multiple image transformations is adopted to generate more sample data and enrich the training set.

[0019] Furthermore, in S202, the specific method for obtaining the optical flow map includes:

[0020] The starting frame and peak frame of the original image are cropped after face normalization;

[0021] The TV-L1 algorithm is used to calculate the optical flow to obtain the horizontal and vertical optical flows. The horizontal and vertical optical flows are used to calculate the optical strain according to the Cauchy tensor theorem.

[0022] The horizontal optical flow, vertical optical flow and optical strain constitute a three-channel RGB image, namely the optical flow map.

[0023] Furthermore, normalized cropping is performed using the following method:

[0024] The original image data uses MediaPipe to obtain facial landmarks, where the left and right eyes are used for face alignment, and the closed curve formed by the points outside the landmarks is used to obtain the facial contour;

[0025] Calculate the rotation angle based on the left and right eye landsmarks, rotate the image data, and get the aligned face;

[0026] Since the closed curve formed by the points outside the landmarks is the facial contour, cv2.convexHull is used to calculate the convex hull corresponding to the landsmarks point set to obtain the facial area, and then the facial mask can be obtained;

[0027] The face mask is ANDed with the original image to obtain the face image with the background removed. Then cv2.boundingRect is used to obtain the minimum rectangle containing the face mask. The image is cropped within this rectangular range to obtain the cropped image.

[0028] The cropped image size does not meet the network input requirements and must be resized to 224×224 pixels using cv2.resize. This completes the face normalization and cropping steps.

[0029] Furthermore, in S03, a dual-stream backbone network based on SwimTransformer is used to extract multi-scale multimodal features from the optical flow input and the depth difference input respectively, and a fusion branch is used to realize the interactive fusion of multimodal features, and output the multimodal fusion features.

[0030] Furthermore, the dual-stream backbone network includes:

[0031] An input module includes an optical flow map input module located in an optical flow branch and a depth difference input module located in a depth branch. The optical flow map input module receives optical flow information representing dynamic changes in the face and captures temporal changes in facial expressions. The depth difference input module receives a depth image representing changes in the spatial structure of the face and provides spatial information of the facial expression.

[0032] The SwimTransformer backbone network is set up in the optical flow branch and the depth branch respectively. The one in the optical flow branch is used to process the optical flow map and extract the dynamic change characteristics of facial expressions through the self-attention mechanism. The one in the depth branch is used to process the depth difference and extract features related to the spatial structure of the face.

[0033] The FusionBranch fusion branch is used to interact and fuse the features extracted by the two SwimTransformer branches to generate a multimodal feature representation; and

[0034] The Head classification head is used to receive the two features and fusion features from the backbone network, process them, and finally output the classification results of micro-expressions.

[0035] Furthermore, the optical flow branch is composed of a 4-stage backbone network consisting of SwimTransformer, which extracts multi-scale features from the optical flow map. and The depth branch also consists of a 4-stage backbone network composed of SwimTransformer, which extracts multi-scale features from the depth difference. and FusionBranch is composed of convolutional neural networks. and Fusion to obtain F M , and then and F M Fusion to obtain F S , the classification head is used for classification, and F S Flatten and concatenate, and finally use the fully connected layer to achieve classification mapping.

[0036] Furthermore, the fusion branch is implemented using a fusion module and The feature fusion is achieved by using average pooling to reduce the dimensionality of the fusion features and obtain F M ; Implemented using fusion modules and F M The ConvBNAct module is used to achieve dimensionality reduction of fusion features and obtain F S .

[0037] Furthermore, the fusion module includes a ConvBNAct module, an IncepResBlock module, and a Concat operation, wherein Conv is a convolution operation for extracting local features; BN is a batch normalization operation for accelerating training and improving model performance; and Act is an activation function for introducing nonlinearity into the network.

[0038] The input of the fusion module is multiple features of the same size, which are spliced ​​in the channel dimension using Concat. The features are split according to the channel, half of which are shallowly processed using ConvBNAct, and the other half are deeply processed using ConvBNAct and multiple IncepResBlocks.

[0039] The shallow processing features and deep processing features are concatenated in the channel dimension through Concat;

[0040] The ConvBNAct module is used to achieve feature dimensionality reduction and obtain fusion features.

[0041] Furthermore, the IncepResBlock module introduces the InceptionV2 structure to implement multi-branch and multi-size convolution kernels for more robust feature extraction, and introduces residual connections to retain the details and semantic information in the original input data.

[0042] The positive effects of the present invention are:

[0043] This invention incorporates depth information, expanding the micro-expression recognition model into a multimodal recognition system. By incorporating depth information, it not only provides a richer representation of facial features but also enhances the ability to capture details in subtle facial motion features, effectively improving recognition accuracy. The combination of multimodal data, particularly the fusion of optical flow and depth information, enhances the perception of dynamic changes in facial micro-expressions, thereby overcoming the limitations of traditional single-modality approaches and improving the robustness and accuracy of micro-expression recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a flow chart of the present invention;

[0045] Figure 2 It is a flow chart of the data preprocessing algorithm of the present invention;

[0046] Figure 3 This is a flow chart of the normalized face cropping algorithm of the present invention;

[0047] Figure 4 This is a flowchart of obtaining an optical flow map of the present invention;

[0048] Figure 5 This is a simplified diagram of the multimodal dual-stream network architecture of the present invention;

[0049] Figure 6 It is a detailed diagram of the multimodal dual-stream network architecture of the present invention;

[0050] Figure 7 This is a diagram of the fusion branch architecture of the present invention;

[0051] Figure 8 This is a diagram of the fusion module architecture of the present invention;

[0052] Figure 9 This is the IncepResBlock module architecture diagram of the present invention;

[0053] Figure 10 It is a schematic diagram of the system of the present invention; DETAILED DESCRIPTION

[0054] See also Figure 1 , which shows a flowchart of a multimodal micro-expression recognition method according to an embodiment of the present invention, which is described in detail as follows, including:

[0055] S01: Collection of multimodal training data, using a camera to collect face images and depth images.

[0056] S02: Multimodal training data preprocessing, using Figure 2 The preprocessing algorithm flow shown implements preprocessing to remove irrelevant redundant information such as background, and obtains input features with strong identification ability - optical flow map and depth difference, and uses data enhancement algorithm to expand training data.

[0057] S03: Build and train a multimodal micro-expression recognition model. Use a dual-stream backbone network based on SwimTransformer to extract multi-scale multimodal features from optical flow input and depth difference input respectively. Use fusion branches to achieve interactive fusion of multimodal features and output multimodal fusion features. The model includes the backbone network (BackBone) of the open source SwimTransformer model, the self-built fusion branch (FusionBranch) and the classification head (Head). For detailed architecture, see Figure 3 .

[0058] S04: Deploy the trained model to your computer.

[0059] S05: Acquiring the micro-expression to be detected. Similarly, a camera is used to collect a face image and a depth image. The raw data of the micro-expression to be recognized needs to be pre-processed in the same way as S102.

[0060] S06: The pre-processed micro-expression data to be recognized is sent to the model in S03 for recognition, and the model can output the corresponding category of the micro-expression.

[0061] See also Figure 2 , which shows a data preprocessing algorithm flow chart of a multimodal micro-expression recognition method provided by an embodiment of the present invention, which is detailed as follows, including:

[0062] The original image data contains a lot of redundant information, and directly feeding it into the network is not conducive to feature extraction, so it needs to be preprocessed.

[0063] S201: Normalized face cropping can obtain basic image data with background removed, face aligned, and meeting the network size input requirements. Figure 3 .

[0064] S202: Optical flow feature is a discriminative feature containing two-dimensional pixel motion information. The TV-L1 algorithm is used to calculate the optical flow using the start frame and peak frame of the normalized cropped RGB image to obtain horizontal and vertical optical flows, and calculate the optical strain. The optical flow map can be obtained by channel splicing. The specific algorithm flow is shown in Figure 4 .

[0065] S203: The depth feature is a supplement to the optical flow feature, which contains the motion information of the facial area in depth. The depth difference is calculated using the depth images of the normalized cropped starting frame and the peak frame.

[0066] S204: Use data augmentation methods to process the optical flow map and depth difference in pairs. This includes, but is not limited to, random combinations of image transformations such as image flipping, blurring, cropping, rotation, translation, adding noise, and brightness adjustment to generate more sample data, thereby enriching the training set, increasing the model's generalization ability, and reducing the risk of overfitting.

[0067] See also Figure 3 , which shows a flow chart of a normalized face cropping algorithm provided by an embodiment of the present invention, and is described in detail as follows, including:

[0068] The original image data uses MediaPipe to obtain facial landmarks, where the left and right eyes can be used for face alignment, and the closed curve formed by the points outside the landmarks can be used to obtain the facial contour.

[0069] The rotation angle is calculated based on the left and right eye landsmarks, and the image data is rotated to obtain aligned faces.

[0070] Since the closed curve formed by the points outside the landmarks is the facial contour, cv2.convexHull is used to calculate the convex hull corresponding to the landsmarks point set to obtain the facial area, and then the facial mask can be obtained.

[0071] The face mask is ANDed with the original image to obtain the face image with the background removed. Then cv2.boundingRect is used to obtain the minimum rectangle containing the face mask. The image is cropped within this rectangular range to obtain the cropped image.

[0072] The cropped image size does not meet the network input requirements and must be resized to 224×224 pixels using cv2.resize. This completes the face normalization and cropping steps.

[0073] Use the facial key point algorithm to mark the points, then obtain the facial convex hull, and only retain the facial convex hull area. Therefore, the obtained optical flow and depth difference are free of background interference and are more effective as network input.

[0074] See also Figure 4 , which shows a flowchart of obtaining an optical flow map for removing background provided by an embodiment of the present invention, and is described in detail as follows, including:

[0075] The starting frame and peak frame of the original image are cropped after face normalization. The specific process is shown in Figure 3 .

[0076] The TV-L1 algorithm is used to calculate the optical flow to obtain the horizontal and vertical optical flows. The horizontal and vertical optical flows are used to calculate the optical strain according to the Cauchy tensor theorem.

[0077] Horizontal optical flow, vertical optical flow, and optical strain constitute a three-channel RGB image - the optical flow map.

[0078] See also Figure 5 , which shows a multimodal dual-stream backbone network architecture diagram provided by an embodiment of the present invention, as detailed below, including:

[0079] Network Architecture Overview

[0080] This multi-branch network architecture is designed to process and analyze multimodal data for micro-expression recognition. The network structure consists of two main SwimTransformer branches, each processing different input data (optical flow map and depth difference). The output features of the two branches are fused through the FusionBranch module and finally passed through the Head module for classification output.

[0081] Network module detailed description

[0082] Input module. The network inputs two sources: optical flow maps and depth differences. The optical flow input module receives optical flow information representing facial dynamics, helping to capture temporal variations in facial expressions. The depth difference input receives depth maps representing variations in facial spatial structure, providing spatial information about facial expressions.

[0083] The SwimTransformer backbone network consists of two main networks, each using a SwimTransformer to process different data sources. The optical flow branch processes the optical flow graph and extracts dynamic features of facial expressions through a self-attention mechanism. The depth branch processes depth differences and extracts features related to the spatial structure of the face.

[0084] FusionBranch. A key branch, FusionBranch is responsible for fusing the features extracted by the two SwimTransformer branches. Specifically, FusionBranch interacts and fuses features from optical flow and depth to generate a multimodal feature representation. This fusion approach helps enhance the model's ability to recognize micro-expressions, leveraging the complementarity of optical flow and depth information to improve model accuracy and robustness.

[0085] The Head receives two features from the backbone and the fusion feature, a total of three features for processing, and finally outputs the classification result of the micro-expression.

[0086] Input and network architecture advantages

[0087] By incorporating optical flow and depth information as multimodal inputs, this network can simultaneously process both the temporal and spatial structural characteristics of facial expressions. This multimodal fusion has the potential to improve the accuracy of micro-expression recognition. The use of the SwimTransformer enables the network to effectively capture multi-scale information, while the FusionBranch enhances the model's robustness by deeply fusing information from different modalities. Ultimately, through the efficient classification of the Head module, the system is able to accurately recognize micro-expressions.

[0088] See also Figure 6 , which shows a detailed diagram of the multimodal dual-stream network architecture provided by an embodiment of the present invention, as detailed below, including:

[0089] Optical flow branch BackBone optical A 4-stage backbone network composed of SwimTransformer extracts multi-scale features from optical flow maps and Among them, B stands for BackBone, the subscript o stands for optical flow, and the superscripts {L, M, S} represent three different sizes: large, middle, and small, respectively.

[0090] Deep Branch BackBone depth The 4-stage backbone network, also composed of SwimTransformer, extracts multi-scale features from depth differences. and Among them, B stands for BackBone, the subscript d stands for depth, and the superscripts {L, M, S} represent three different sizes: large, middle, and small, respectively.

[0091] The fusion branch is composed of a convolutional neural network and is used for interactive fusion of multi-modal and multi-scale features. and Fusion to obtain F M , and then and F M Fusion to obtain F S , where F stands for fusion feature, the superscript M stands for middle size, and the superscript S stands for small size.

[0092] The classification head is used for classification. and F S Flatten and concatenate, and finally use the fully connected layer to achieve classification mapping.

[0093] If only the final features extracted by the backbone network are used for classification, the correlation of multimodal features will be ignored. Adding a fusion branch can achieve staged fusion of multimodal features, extracting complex related information, which may be beneficial to the classification and recognition of the model.

[0094] See also Figure 7 , which shows a diagram of the fusion branch architecture provided by an embodiment of the present invention, and is described in detail as follows, including:

[0095] In response to the needs of multi-modal and multi-scale feature fusion, a fusion module is constructed and pooling and convolution are used to achieve dimensionality reduction, thus realizing the architecture of the fusion branch.

[0096] Implemented using FusionBlock and feature fusion.

[0097] Use average pooling to achieve dimensionality reduction of fusion features and get F M .

[0098] Implemented using FusionBlock and F M feature fusion.

[0099] Use the ConvBNAct module to achieve dimensionality reduction of fusion features and obtain F S .

[0100] See also Figure 8 , which shows the fusion module architecture provided by an embodiment of the present invention, which is described in detail as follows, including:

[0101] ConvBNAct module, IncepResBlock module, and Concat operation. Conv stands for convolution, which is used to extract local features; BN stands for batch normalization, a technique used to accelerate training and improve model performance by normalizing the output of each layer to reduce internal covariate shift; Act is an activation function used to introduce nonlinearity to the network, enabling it to learn complex patterns.

[0102] The input of the fusion module is multiple features of the same size, which are concatenated in the channel dimension using Concat.

[0103] The features are segmented by channel, with half fed into the left branch (shallow path) and the other half fed into the right branch (deep path). The left branch uses ConvBNAct for shallow processing, while the right branch uses ConvBNAct and multiple IncepResBlocks for deep processing.

[0104] The features of the shallow path and the deep path are concatenated in the channel dimension through Concat.

[0105] The ConvBNAct module is used to achieve feature dimensionality reduction and obtain fusion features.

[0106] See also Figure 9 , which shows the IncepResBlock module architecture provided by an embodiment of the present invention, as detailed below, including:

[0107] This module introduces the classic InceptionV2 structure to implement multi-branch and multi-size convolution kernels for more robust feature extraction.

[0108] This module introduces residual connections to preserve the details and semantic information in the original input data.

[0109] See also Figure 10 , which shows a system schematic diagram of an embodiment of the present invention, and is described in detail as follows:

[0110] System Overview

[0111] The present invention also discloses a multimodal micro-expression recognition system. This system captures micro-expression data using a data acquisition device, processes and analyzes the data using a computer program, and ultimately displays the recognition results on a display. The system utilizes a modular design to optimize data processing efficiency and improve recognition accuracy. It also offers flexible scalability and is applicable to a variety of fields, including security, healthcare, criminal investigation, commerce, and human-computer interaction.

[0112] System composition and functions

[0113] according to Figure 10, this system mainly includes the following modules: The data acquisition device is used to collect the user's micro-expression multimodal data. Exemplary implementation methods include but are not limited to: RGB camera, depth camera, RGB-D camera, etc. The computer runs the micro-expression recognition algorithm and is responsible for computing tasks such as data preprocessing, model training, classification and recognition. It can be implemented using CPU, GPU, FPGA or dedicated AI acceleration chip (such as NPU). The memory stores the collected raw data, intermediate processing results and final recognition results, and may include different storage media such as RAM, ROM, SSD, HDD or cloud storage. The processor performs computing tasks, including data cleaning, feature extraction, pattern matching, deep learning reasoning, etc., and may use multi-core CPU, GPU cluster, edge computing device or distributed computing architecture. The display is used to display the micro-expression recognition results, which can be an LCD, OLED screen, AR / VR device or remote terminal. The computer program includes the aforementioned multimodal micro-expression recognition method, data preprocessing algorithms, data augmentation algorithms, optical flow algorithms (such as TV-L1 and Lucas-Kanade), and deep learning models (such as CNN, Transformer, SwimTransformer, and InceptionV2). It can be implemented in various programming languages ​​and frameworks, including Python, C++, TensorFlow, and PyTorch. Data includes raw image frames, preprocessed data, and classification labels.

[0114] System workflow

[0115] Data Collection: Collects user micro-expression data using multimodal sensors. Data Preprocessing: Performs background removal, alignment, and resizing. Model Training: Uses preprocessed data to train a deep learning model. Classification and Recognition: The data to be recognized is fed into the trained model for classification, outputting the micro-expression category (e.g., anger, surprise, sadness, etc.). Result Display: Displays the recognition results on a monitor in real time or offline.

[0116] The above-mentioned multimodal micro-expression recognition method is described in detail below in conjunction with specific embodiments.

[0117] The operating system used is Windows 11, the CPU is Intel(R) Core(TM) i5-12400F, the graphics card is NVIDIA GeForce RTX 4060Ti with 8GB of video memory, the Python version is 3.12.3, the PyTorch version is 2.2.2+cu121, the learning rate is fixed at 0.0005, and the batch size is set to 64.

[0118] Use the cross entropy loss function (CrossEntropyLoss) for training, and its formula is:

[0119]

[0120] Among them, N is the number of training samples, C is the number of categories, and y i,c and They represent the true label and predicted probability of the i-th sample respectively.

[0121] The evaluation indicators are unweighted average recall (UAR), unweighted F1-score (UF1) and accuracy (ACC).

[0122] This paper uses the Leave-One-Subject-Out (LOSO) cross-validation method for evaluation. Specifically, samples from one subject are selected at a time as the test set, and the data from all other subjects are used as the training set. Assuming there are N subjects in the dataset, N experiments are required. This method can effectively mitigate the effects of the small sample size of micro-expression datasets and the large individual differences among subjects.

[0123] The embodiment of the present invention uses CASME 3 The PartC sub-dataset of the CASME dataset is used for comparative experiments. 3 Part C of the dataset is a multimodal dataset that includes RGB images, depth images, and ECG images. Part C contains 165 multimodal micro-expression samples from 31 subjects, with four labels: Negative, Positive, Surprise, and Other.

[0124] The multimodal micro-expression recognition method provided by the embodiment of the present invention is compared with the baseline method (baseline) using AlexNet as input of the original RGB image (C) and the depth image (D) in CASME 3 The comparison is conducted on the PartC sub-dataset of the dataset. For detailed information, see Table 1.

[0125] Table 1 Comparison of methods

[0126]

[0127] “-” means that the data is not provided

[0128] The results in Table 1 show that the multimodal two-stream network approach achieves 22.09% and 22.79% improvements in UAR and UF1, respectively, compared to the baseline, and reaches an ACC of 0.7091. This demonstrates that by combining preprocessing steps such as background removal, optical flow calculation, data augmentation, and depth difference calculation with a multimodal two-stream network, the method of this embodiment achieves significant improvements over the baseline.

[0129] In summary, this paper proposes a multimodal micro-expression recognition method and system that significantly improves the accuracy and robustness of micro-expression recognition. This achievement demonstrates that through effective data preprocessing and multimodal feature fusion, the method not only improves the accuracy of micro-expression recognition but also enhances the model's robustness and ability to recognize complex emotional features. Compared with traditional methods, the system of this invention is more adaptable to different application scenarios and has greater practical application value, especially in fields such as security, healthcare, and human-computer interaction.

[0130] Those skilled in the art should understand that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0131] Those skilled in the art should also understand that the above modules and implementation methods are only illustrative displays, and different technical solutions may be adopted in actual implementation, including but not limited to: Hardware integration: multiple modules (such as processor + memory) can be integrated into a single chip (such as SoC) to improve efficiency. Software optimization: the algorithm can adopt different models (such as ResNet, ViT) or hybrid architecture (traditional CV + deep learning). Data source expansion: in addition to the combination of visual and depth data, it can also be combined with multimodal inputs such as voice and physiological signals (heart rate, skin electricity). Deployment method: The system can be deployed on local devices, edge servers or the cloud to adapt to different application scenarios. Therefore: all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A multimodal micro-expression recognition method, characterized in that It includes the following steps: S01: Multimodal training data acquisition, using a camera to collect face images and depth images; S02: Preprocess multimodal training data to remove redundant information and obtain discriminative input features, including optical flow maps and depth differences, and use data augmentation algorithms to expand training data; S03: Build and train a multimodal micro-expression recognition model, train the pre-processed image data and its attached optical flow and depth information, and output the prediction results; S04: Deploy the trained model to the computer; S05: Acquire the micro-expression to be recognized using the steps of S01 and perform pre-processing through S02; S06: The pre-processed micro-expression data to be recognized is sent to the model in S03 for recognition, and the corresponding category of the micro-expression is output.

2. A multimodal micro-expression recognition method according to claim 1, characterized in that The specific steps of S02 include: S201: normalizing and cropping the face to obtain basic image data with background removed and face aligned, meeting the network size input requirements; S202: Optical flow is a discriminative feature that contains two-dimensional pixel motion information. The TV-L1 algorithm is used to calculate optical flow using the start and peak frames of the normalized and cropped RGB image. Horizontal and vertical optical flows are obtained, and optical strain is calculated. The channels are then stitched together to obtain an optical flow map. S203: The depth feature is a supplement to the optical flow feature, which contains the motion information of the facial area in depth. The depth difference is calculated using the depth images of the normalized cropped starting frame and the peak frame. S204: Using data enhancement methods, the optical flow map and the depth difference are processed in pairs, and a random combination of multiple image transformations is adopted to generate more sample data and enrich the training set.

3. A multimodal micro-expression recognition method according to claim 2, characterized in that In S202, the specific method for obtaining the optical flow map includes: The starting frame and peak frame of the original image are cropped after face normalization; The TV-L1 algorithm is used to calculate the optical flow to obtain the horizontal and vertical optical flows. The horizontal and vertical optical flows are used to calculate the optical strain according to the Cauchy tensor theorem. The horizontal optical flow, vertical optical flow and optical strain constitute a three-channel RGB image, namely the optical flow map.

4. A multimodal micro-expression recognition method according to claim 2 or 3, characterized in that Normalized cropping uses the following method: The original image data uses MediaPipe to obtain facial landmarks, where the left and right eyes are used for face alignment, and the closed curve formed by the points outside the landmarks is used to obtain the facial contour; Calculate the rotation angle based on the left and right eye landsmarks, rotate the image data, and get the aligned face; Since the closed curve formed by the points outside the landmarks is the facial contour, cv2.convexHull is used to calculate the convex hull corresponding to the landsmarks point set to obtain the facial area, and then the facial mask can be obtained; The face mask is ANDed with the original image to obtain the face image with the background removed. Then cv2.boundingRect is used to obtain the minimum rectangle containing the face mask. The image is cropped within this rectangular range to obtain the cropped image. The cropped image size does not meet the network input requirements and must be adjusted to 224×224 pixels using cv2.resize. This completes the face normalization and cropping steps.

5. A multimodal micro-expression recognition method according to claim 1, characterized in that In S03, a dual-stream backbone network based on SwimTransformer is used to extract multi-scale multimodal features from the optical flow input and depth difference input respectively, and a fusion branch is used to realize the interactive fusion of multimodal features and output the multimodal fusion features.

6. A multimodal micro-expression recognition method according to claim 5, characterized in that The dual-stream backbone network includes: An input module includes an optical flow map input module located in an optical flow branch and a depth difference input module located in a depth branch. The optical flow map input module receives optical flow information representing dynamic changes in the face and captures temporal changes in facial expressions. The depth difference input module receives a depth image representing changes in the spatial structure of the face and provides spatial information of the facial expression. The SwimTransformer backbone network is set up in the optical flow branch and the depth branch respectively. The one in the optical flow branch is used to process the optical flow map and extract the dynamic change characteristics of facial expressions through the self-attention mechanism. The one in the depth branch is used to process the depth difference and extract features related to the spatial structure of the face. The FusionBranch fusion branch is used to interact and fuse the multi-scale features extracted by the two SwimTransformer branches to generate a multimodal feature representation; and The Head classification head is used to receive the two features and fusion features from the backbone network, process them, and finally output the classification results of micro-expressions.

7. A multimodal micro-expression recognition method according to claim 6, characterized in that The optical flow branch includes a 4-stage backbone network composed of SwimTransformer, which extracts multi-scale features from the optical flow map. and B stands for BackBone, the subscript o stands for optical flow, and the superscripts {L, M, S} represent the three different sizes of large, middle, and small respectively. The depth branch includes a 4-stage backbone network composed of SwimTransformer, which extracts multi-scale features from the depth difference. and The subscript d stands for depth; FusionBranch is composed of convolutional neural networks. and Fusion to obtain F M , and then and F M Fusion to obtain F S , where F stands for fusion feature, i.e. fusion feature; The superscript M represents the middle size, and the superscript S represents the small size; The classification head is used for classification. and F S Flatten and concatenate, and finally use the fully connected layer to achieve classification mapping.

8. A multimodal micro-expression recognition method according to claim 7, characterized in that The fusion branch is implemented using a fusion module and The feature fusion of ; use average pooling to achieve dimensionality reduction of fusion features, and get F M ; Implemented using fusion modules and F M The ConvBNAct module is used to achieve dimensionality reduction of fusion features and obtain F S .

9. A multimodal micro-expression recognition method according to claim 8, characterized in that The fusion module includes a ConvBNAct module, an IncepResBlock module, and a Concat operation, wherein Conv is a convolution for extracting local features; BN is batch normalization, a technology used to accelerate training and improve model performance; Act is an activation function used to introduce nonlinearity into the network; The input of the fusion module is multiple features of the same size, and Concat is used to perform channel-dimensional splicing; Feature segmentation is performed based on the channels, with half processed shallowly using ConvBNAct and the other half processed deeply using ConvBNAct and multiple IncepResBlocks. The shallow processing features and deep processing features are concatenated in the channel dimension through Concat; The ConvBNAct module is used to achieve feature dimensionality reduction and obtain fusion features.

10. A multimodal micro-expression recognition method according to claim 9, characterized in that The IncepResBlock module introduces the InceptionV2 structure to implement multi-branch and multi-size convolution kernels for more robust feature extraction, and introduces residual connections to retain the details and semantic information in the original input data.

Citation Information

Cited By

  • High-precision micro-expression recognition method and device, electronic equipment and storage medium

    CN122266038A