Micro-expression recognition method, system and equipment based on multi-scale mixed channel and medium
Through the combination of the multi-scale hybrid channel module and the feature purification module, the problem of insufficient fusion of multi-scale feature in micro-expression recognition is solved, high-precision micro-expression recognition is achieved, and the recognition ability of the model is improved.
Patent Information
- Application Number
- CN202510566460.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The existing micro-expression recognition methods have shortcomings in dealing with multi-scale feature fusion, resulting in limited recognition accuracy, and when using video sequences as input, they are easily affected by changes in sequence length and frame rate, and ignore facial identity information.
The multi-scale hybrid channel module is used to extract micro-expression features from global to local, and the feature purification module is integrated and purified in the channel dimension, and classified in the full connection layer to build a micro-expression recognition model.
The accuracy of micro-expression recognition is significantly improved, especially the UF1 and UAR values on the public data set reach 0.8832 and 0.8968, which is better than the existing methods and proves the effectiveness and generalization of the model.
Smart Images

Figure CN120388409A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of facial expression recognition, and particularly relates to a micro-expression recognition method, system, device and medium based on multi-scale hybrid channels. Background Art
[0002] In the field of micro-expression recognition, existing methods can be mainly classified into methods that use manually labeled key frames as input and methods that use video sequences as input according to the input data. Although the method using video sequences as input can capture spatio-temporal features, it has a weak modeling ability for the dependence relationship of long sequences, is easily affected by the changes of sequence length and frame rate, and is prone to the problem of excessive redundant input information. The method using key frames as input usually relies on calculating the motion features of the starting frame and the peak frame for recognition. This type of method retains spatial features and temporal features while ignoring the face identity information.
[0003] In addition, the features of micro-expressions are usually very subtle and are easily mixed with noise at small scales, which further increases the difficulty of recognition. Existing methods have problems in dealing with multi-scale feature fusion, resulting in limited recognition accuracy. Therefore, researching a micro-expression recognition method that can effectively extract multi-scale features and fuse overall information has become a core problem to be solved urgently. Summary of the Invention
[0004] The purpose of the present invention is to provide a micro-expression recognition method, system, device and medium based on multi-scale hybrid channels to solve the problems existing in the above-mentioned prior art.
[0005] To achieve the above purpose, the present invention provides a micro-expression recognition method based on multi-scale hybrid channels, including:
[0006] Obtain an image sequence corresponding to the micro-expression video;
[0007] Extract an effective face region from the image sequence to obtain preprocessed micro-expression image data;
[0008] Input the preprocessed micro-expression image data into a micro-expression recognition model for classification prediction to obtain a recognition result; wherein, the micro-expression recognition model includes a multi-scale hybrid channel module, a feature purification module and a fully connected layer connected in sequence.
[0009] Optionally, the extracting an effective face region from the image sequence specifically includes:
[0010] Perform face detection on the image sequence based on the MTCNN network, extract the effective face region, and use the first frame image in the image sequence as the starting frame, and calculate the optical flow change between the starting frame and all subsequent frames according to the TVL1 method to extract motion features.
[0011] Optionally, the training process of the micro-expression recognition model specifically includes:
[0012] Obtain training data, where the training data includes micro-expression image training data and corresponding expression recognition labels;
[0013] Input the training data into the micro-expression recognition model for classification prediction, and train according to the target loss function to obtain a trained micro-expression recognition model.
[0014] Optionally, the processing process of the micro-expression recognition model specifically includes:
[0015] Input the preprocessed micro-expression image sequence data into the multi-scale hybrid channel module in sequence to extract micro-expression motion features from global to local;
[0016] Integrate the micro-expression motion features into the channel dimension through the feature purification module, and perform feature extraction and feature purification processing to obtain purified micro-expression motion features;
[0017] Input the purified micro-expression motion features into the fully connected layer for classification prediction, and output the recognition result.
[0018] Optionally, the processing process of the multi-scale hybrid channel module specifically includes:
[0019] Gradually divide the preprocessed micro-expression image data into blocks, and extract micro-expression features along the channel dimension after each block. At the same time, gradually reduce the size of the input tensor and increase the channel dimension to obtain block features;
[0020] Process the block features based on the rearrangement operation and the position-aware circular convolution module, and retain the relatively global facial motion information through residual connection to complete the processing of the multi-scale hybrid channel module and output the micro-expression motion features.
[0021] A micro-expression recognition system based on a multi-scale hybrid channel includes:
[0022] A data acquisition module, configured to obtain an image sequence corresponding to a micro-expression video; extract an effective face region from the image sequence to obtain preprocessed micro-expression image data;
[0023] A micro-expression recognition module, configured to input the preprocessed micro-expression image data into a micro-expression recognition model for classification prediction to obtain a recognition result; wherein, the micro-expression recognition model includes a multi-scale hybrid channel module, a feature purification module, and a fully connected layer connected in sequence.
[0024] An electronic device includes a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the described micro-expression recognition method based on multi-scale hybrid channels.
[0025] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the described micro-expression recognition method based on multi-scale hybrid channels.
[0026] The technical effects of the present invention are as follows:
[0027] Through the multi-scale hybrid channel module (MC), the present invention gradually extracts micro-expression features of different scales from the global to the local, effectively capturing the subtle changes of micro-expressions; through the feature purification module (FPM), it further focuses on micro-expression features in the channel dimension, reducing the model's attention to noise and enhancing the model's recognition ability; through the MLP layer for classifying features, the overall accuracy of the model is significantly improved; on the mixed dataset composed of public datasets (CASMEII, SMIC, and SAMM), the UAR and UF1 reach 0.8832 for UAR and 0.8968 for UF1 respectively. On the CASME II dataset, the present method achieves a UF1 value of 96.89% and a UAR value of 96.88%, significantly superior to existing methods. Description of the Drawings
[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0029] The drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0030] Figure 1 It is the implementation flowchart in the embodiments of the present invention.
[0031] Figure 2 It is the structural schematic diagram of the multi-scale hybrid channel module (MC) and the overall network structure diagram in the embodiments of the present invention;
[0032] Figure 3 It is the structural schematic diagram of the feature purification module (FPM) in the embodiments of the present invention. Detailed Embodiments
[0033] The various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as a limitation of the present invention, but rather as a more detailed description of certain aspects, features, and implementation manners of the present invention.
[0034] It should be understood that the terms described in the present invention are only for describing specific embodiments and are not used to limit the present invention. Additionally, for the numerical ranges in the present invention, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Each intermediate value within any stated value or stated range, as well as each smaller range between any other stated value or intermediate value within the stated range, is also included in the present invention. The upper and lower limits of these smaller ranges may be independently included or excluded from the range.
[0035] Without departing from the scope or spirit of the present invention, various improvements and changes can be made to the specific embodiments of the description of the present invention, which are obvious to those skilled in the art. Other embodiments obtained from the description of the present invention are obvious to those skilled in the art. The description and examples of this application are merely exemplary.
[0036] Regarding the use of "comprising", "including", "having", "containing", etc. in this article, they are all open-ended terms, meaning including but not limited to.
[0037] It should be noted that, without conflict, the embodiments and features in the embodiments of this application can be combined with each other. The following will refer to the drawings and combine with the embodiments to detail this application.
[0038] As Figures 1-3 shown, in this embodiment, a micro-expression recognition method based on a multi-scale hybrid channel is provided, including: obtaining an image sequence corresponding to a micro-expression video; extracting an effective face region from the image sequence to obtain preprocessed micro-expression image data; inputting the preprocessed micro-expression image data into a micro-expression recognition model for classification prediction to obtain a recognition result; wherein, the micro-expression recognition model includes a multi-scale hybrid channel module, a feature purification module, and a fully connected layer connected in sequence.
[0039] This embodiment aims to provide a micro-expression recognition method based on local-global temporal relationships, including the following steps: This embodiment aims to provide a micro-expression recognition method based on multi-scale hybrid channels, including the following steps: Obtain a micro-expression video and decompose it into an image sequence; Preprocess the micro-expression image sequence, perform face detection using MTCNN, and implement affine transformation based on facial feature points to achieve face alignment, and extract the effective face region; Construct a multi-scale hybrid channel module (MC): Extract micro-expression motion features from a global to local perspective, gradually divide the input image into blocks, and extract micro-expression features along the channel dimension after each block, while ensuring that the model continuously focuses on information from large-scale features through residual connections; Construct a feature purification module (FPM): Integrate the remaining spatial features into the channel dimension through a combination of convolutional and pooling layers, and use a 11-core convolutional operation and a transformer encoder to further extract and purify micro-expression features; Feed the outputs of MC and FPM into an MLP layer for classification, and the MLP layer includes a LayerNorm (LN) layer, a Reduce layer, and a Linear layer; Train the multi-scale hybrid channel network and save the weight model with the best performance; Preprocess the micro-expression video to be recognized, input it into the trained multi-scale hybrid channel network, and identify the micro-expression category.
[0040] The specific implementation process of this embodiment includes:
[0041] Step 1: Obtain a batch of micro-expression videos and decompose them into image sequences;
[0042] Step 2: Preprocess the micro-expression image sequence, perform face detection using MTCNN, and implement affine transformation based on facial feature points to achieve face alignment, and extract the effective face region; and use the first frame image as the starting frame and all subsequent frames as peak frames, and calculate the optical flow change using the TVL1 method.
[0043] Step 3: Construct a multi-scale hybrid channel module (MC): Extract micro-expression motion features from a global to local perspective, gradually divide the input image into blocks, and extract micro-expression features along the channel dimension after each block, while ensuring that the model continuously focuses on information from large-scale features through residual connections, and the feature processing process can be expressed as:
[0044]
[0045] 1 ≤ x1 ≤ 2, 1 ≤ y1 ≤ 2, k ∈ [1, 4]
[0046] Where In the formula, \(x1\in[1,2]\), \(y1\in[1,2]\), \(k1\in[1,4]\), which represents the order of the feature blocks of the I1 feature from the upper left corner to the lower right corner. Here, \(k1\in[1,4]\) is the rearranged digital label, \(P(.)\) represents the image processed by the ParC module, and \(R0\) represents the previous rearrangement layer.
[0047] Step 4: Construct the Feature Purification Module (FPM): Integrate the remaining spatial features into the channel dimension through a combination of convolutional and pooling layers, and use a 11-core convolutional operation and a transformer encoder to further extract and purify the micro-expression features;
[0048] Step 5: Feed the outputs of MC and FPM into the MLP layer for classification. The MLP layer includes a LayerNorm (LN) layer, a Reduce layer, and a Linear layer;
[0049] Step 6: Train the multi-scale hybrid channel network and save the weight model with the best performance;
[0050] Step 7: Preprocess the micro-expression video to be recognized, input it into the trained multi-scale hybrid channel network, and identify the micro-expression category.
[0051] Implementable, the preprocessing described in step 2 includes:
[0052] Step 2.1: Perform face detection and facial feature point detection on each frame of the image;
[0053] Step 2.2: Align the detected facial features using affine transformation;
[0054] Step 2.3: Crop out a valid face area with a consistent size from the inner corner of the left eye of each frame of the face.
[0055] Step 2.4: Use the TVL1 method to calculate the change in motion features between the first image of the sequence as the starting frame and all subsequent images respectively.
[0056] Implementable, step 3 specifically includes:
[0057] Step 3.1: Gradually divide the input image into blocks, extract micro-expression features along the channel dimension after each block division, and gradually reduce the size of the input tensor while increasing the channel dimension;
[0058] Step 3.2: Process the block-divided features through a rearrangement operation and a Position-Aware Ring Convolution (ParC) module to ensure that the attention to the feature scale does not drop too quickly;
[0059] Step 3.3: Retain relatively global facial motion information through residual connections while learning deeper features in the channel dimension.
[0060] Implementable, step 4 specifically includes:
[0061] Step 4.1: Integrate the remaining spatial features into the channel dimension through a combination of convolutional and pooling layers, and select the useful parts in the feature map;
[0062] Step 4.2: Adopt a convolutional operation with a 1×1 kernel and a transformer encoder to further extract and purify micro-expression features and enhance the global feature extraction ability;
[0063] Step 4.3: Convert the output of the FPM into a text state, make the height (H) and width (W) equal to 1, and further focus on the micro-expression feature information through the transformer encoder.
[0064] Implementable, step 5 specifically includes:
[0065] Step 5.1: Input the output features of the MC and FPM into the MLP layer, and the MLP layer includes a LayerNorm (LN) layer, a Reduce layer, and a Linear layer;
[0066] Step 5.2: Normalize the features through the LayerNorm layer, reduce the feature dimension through the Reduce layer, and complete the final classification task through the Linear layer;
[0067] Step 5.3: Use the Softmax function to complete the final classification of micro-expression categories.
[0068] A micro-expression recognition method based on multi-scale hybrid channels provided by this embodiment has the following advantages in view of the deficiencies of existing methods in multi-scale feature extraction and global-local feature fusion: through the multi-scale hybrid channel module (MC), micro-expression features are gradually extracted from the global to the local, effectively capturing the subtle changes of micro-expressions; through the feature purification module (FPM), features are further optimized and fused to enhance the recognition ability of the model; the features are classified through the MLP layer, significantly improving the recognition accuracy; on the mixed dataset composed of public datasets (CASMEII, SMIC, and SAMM), the UAR and UF1 respectively reach the UAR value of 0.8832 and the UF1 value of 0.8968. On the CASME II dataset, this embodiment achieves a UF1 value of 96.89% and a UAR value of 96.88%, significantly superior to existing methods. On the five-class mixed dataset composed of SAMM and CASME II, the overall UAR and UF1 of MMCN reach 0.8242 and 0.8476 respectively. Among them, on CASMEII, the effects of 0.9496 (UAR) and 0.9346 (UF1) are achieved. On the three-class classification task of the CASME 3 dataset, MMCN achieves excellent performance with a UAR value of 0.5679 and a UF1 value of 0.6125, exceeding other models in the field, proving the generalization and effectiveness of the model.
[0069] Implementable, this embodiment also provides a micro-expression recognition system based on multi-scale hybrid channels, including:
[0070] A data acquisition module, configured to obtain an image sequence corresponding to a micro-expression video; extract an effective face region from the image sequence to obtain preprocessed micro-expression image data;
[0071] A micro-expression recognition module, configured to input the preprocessed micro-expression image data into a micro-expression recognition model for classification prediction to obtain a recognition result; wherein, the micro-expression recognition model includes a multi-scale hybrid channel module, a feature purification module, and a fully connected layer connected in sequence.
[0072] Implementable, this embodiment also provides an electronic device, including a memory and a processor, where the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the described micro-expression recognition method based on multi-scale hybrid channels.
[0073] Implementable, this embodiment also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the described micro-expression recognition method based on multi-scale hybrid channels.
[0074] As described above, it is only the preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A micro-expression recognition method based on multi-scale hybrid channels, characterized in that, including: Obtain an image sequence corresponding to the micro-expression video; Extract an effective face region from the image sequence to obtain preprocessed micro-expression image data; Input the preprocessed micro-expression image data into a micro-expression recognition model for classification prediction to obtain a recognition result; wherein, the micro-expression recognition model includes a multi-scale hybrid channel module, a feature purification module, and a fully connected layer connected in sequence.
2. The micro-expression recognition method based on multi-scale hybrid channels according to claim 1, wherein The extracting an effective face region from the image sequence specifically includes: Perform face detection on the image sequence based on the MTCNN network, extract an effective face region, and use the first frame image in the image sequence as the starting frame, calculate the optical flow change between the starting frame and all subsequent frames according to the TVL1 method, and extract motion features.
3. A micro-expression recognition method based on multi-scale hybrid channels according to claim 1, characterized in that, The training process of the micro-expression recognition model specifically includes: Obtain training data, where the training data includes micro-expression image training data and corresponding expression recognition labels; Input the training data into the micro-expression recognition model for classification prediction, and train according to the target loss function to obtain a trained micro-expression recognition model.
4. A micro-expression recognition method based on multi-scale hybrid channels according to claim 1, characterized in that The processing process of the micro-expression recognition model specifically includes: Input the preprocessed micro-expression image sequence data into the multi-scale hybrid channel module in sequence to extract micro-expression motion features from global to local; Integrate the micro-expression motion features into the channel dimension through the feature purification module, and perform feature extraction and feature purification processing to obtain purified micro-expression motion features; Input the purified micro-expression motion features into the fully connected layer for classification prediction and output the recognition result.
5. The micro-expression recognition method based on multi-scale hybrid channels according to claim 4, characterized in that The processing process of the multi-scale hybrid channel module specifically includes: Gradually divide the preprocessed micro-expression image data into blocks, extract micro-expression features along the channel dimension after each block, and gradually reduce the size of the input tensor and increase the channel dimension to obtain block features; Process the block features based on rearrangement operations and a position-aware circular convolution module, and retain relatively global facial motion information through residual connections to complete the processing of the multi-scale hybrid channel module and output micro-expression motion features.
6. A micro-expression recognition system based on multi-scale hybrid channels, characterized in that, including: A data acquisition module for obtaining an image sequence corresponding to the micro-expression video; Extract an effective face region from the image sequence to obtain preprocessed micro-expression image data; A micro-expression recognition module for inputting the preprocessed micro-expression image data into a micro-expression recognition model for classification prediction to obtain a recognition result; wherein, the micro-expression recognition model includes a multi-scale hybrid channel module, a feature purification module, and a fully connected layer connected in sequence.
7. An electronic device, characterized in that, including a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute a micro-expression recognition method based on a multi-scale hybrid channel according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed by the processor, it implements a micro-expression recognition method based on a multi-scale hybrid channel according to any one of claims 1-5.
Citation Information
Patent Citations
Micro-expression recognition method and system based on channel attention mechanism
CN112001241A
Model training method, micro-expression recognition method and model training device
CN117423145A
Small target detection system and method based on improved YOLOv5
CN117523267A
Micro-expression recognition method and system based on three-branch network
CN119625805A
Fine-grained image classification method, equipment and medium
CN119785111A