Micro-expression recognition method based on multi-motion feature fusion
By adopting a multi-motor feature fusion method in micro-expression recognition, combined with LSTM and attention mechanism, the high computing resource consumption caused by insufficient capture of subtle muscle motion features and model complexity in the prior art is solved, and efficient and accurate micro-expression recognition effect is achieved.
Patent Information
- Application Number
- CN202510016099.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-06
AI Technical Summary
The prior art has shortcomings in processing subtle muscle motion characteristics on two-dimensional images. Three-dimensional reconstruction cannot fully capture the instantaneous changes in micro-expressions. The improved Inception network module is not sensitive enough to process short-term and local changes when processing optical flow characteristics. The complex model structure leads to high demand for computing resources, especially in real-time applications.
The micro-expression recognition method based on multi-motion feature fusion is adopted to detect the micro-expression time period through convolutional neural network, and multi-motion features including optical flow field, LBP-TOP, facial key point trajectory and inter-frame motion vector are extracted. The feature fusion model is used based on LSTM, combining attention mechanism and PCA dimensionality reduction technology to simplify the model structure and improve the recognition efficiency.
Through the multi-motor feature fusion method, the instantaneous changes of micro-expressions and the ability to capture subtle muscle movements are improved, the recognition accuracy and robustness are improved, the model structure is simplified, the calculation complexity is reduced, and the feasibility of real-time applications is enhanced.
Smart Images

Figure CN119942613A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a micro-expression recognition method based on multi-motion feature fusion. Background Art
[0002] With the development of computer vision and machine learning technologies, micro-expression recognition has gradually become an important research field, which aims to capture and analyze these subtle facial changes through automated methods to accurately identify people's true emotional states.
[0003] After searching, the invention patent with Chinese patent number CN114333002A discloses a micro-expression recognition method based on graph deep learning and three-dimensional reconstruction of human face, including the following steps: constructing a graph feature learning module, performing graph feature analysis to obtain a one-dimensional feature vector; constructing an optical flow feature learning module, obtaining a one-dimensional feature vector through optical flow feature extraction; constructing a three-dimensional detail reconstruction module to obtain a one-dimensional feature vector; constructing a multi-stream OGC-FL network model structure, and obtaining micro-expression recognition classification results through multi-stream fusion. Compared with a single strategy, the multi-strategy generation of optical flow features in the present invention can screen out the most favorable generation strategy for micro-expression recognition tasks; the multi-stream OGC-FL network model structure of the invention patent with Chinese patent number CN114333002A finds the consistency of facial key point information and dense image information in recognizing micro-expressions. The sparse spatial information of key points can judge the general state of micro-expressions through GFL, while the dense image information highlights the subtle muscle movements of the face, extracting more detailed information for MER.
[0004] However, in actual use, the above invention uses three-dimensional reconstruction and graph feature extraction to capture the sparse spatial information of facial key points, but has shortcomings in processing subtle muscle movement features on two-dimensional images, and three-dimensional reconstruction cannot fully capture the instantaneous changes of micro-expressions; although the improved Inception network module can capture motion information by processing optical flow features, it is not sensitive enough when processing short time series and local changes, especially its robustness under different lighting conditions is poor; in addition, the model structure is relatively complex, resulting in high computing resource requirements, especially in real-time applications, which may face the problem of increased latency.
[0005] Therefore, a micro-expression recognition method based on multi-motion feature fusion is proposed. Summary of the invention
[0006] The purpose of the present invention is to solve the shortcomings of the prior art in processing subtle muscle movement features on two-dimensional images, and the three-dimensional reconstruction cannot fully capture the instantaneous changes of micro-expressions; although the improved Inception network module can capture motion information when processing optical flow features, it is not sensitive enough when processing short time series and local changes, especially the robustness under different lighting conditions is poor; in addition, the model structure is relatively complex, resulting in high computing resource requirements, especially in real-time applications may face the disadvantage of increased delay, and a micro-expression recognition method based on multi-motion feature fusion is proposed.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] The micro-expression recognition method based on multi-motion feature fusion includes:
[0009] S1: Get the video sequence to be analyzed;
[0010] S2: inputting the video sequence into a pre-built micro-expression detection model, and detecting the time period in the video sequence containing micro-expressions through the micro-expression detection model, wherein the micro-expression detection model is trained using a convolutional neural network (CNN), and the training data set includes video clips with the time when micro-expressions occur marked;
[0011] S3: extracting multiple motion features from the time period, the multiple motion features including an optical flow field, a local binary pattern histogram LBP-TOP, a facial key point trajectory, and an inter-frame motion vector, the optical flow field being calculated by an Endo optical flow algorithm;
[0012] S4: inputting the multi-motion features into a pre-trained feature fusion model, and obtaining a fusion feature vector through the feature fusion model, wherein the feature fusion model is a deep learning model based on a long short-term memory network (LSTM), and an attention mechanism is introduced during the training process to enhance the recognition of key frames;
[0013] S5: identifying the micro-expression category within the time period according to the fused feature vector and using a pre-established micro-expression category database, wherein the micro-expression category database includes a plurality of micro-expression types and their corresponding typical feature vectors;
[0014] S6: Post-process the recognition results.
[0015] The above technical solution further includes:
[0016] Preferably, the micro-expression detection model in step S2 is obtained by training through the following steps:
[0017] Collect a standard micro-expression video dataset, which contains micro-expression video clips in various emotional states, and each video clip is annotated by experts to indicate whether there is a micro-expression and the time period in which it occurs;
[0018] The micro-expression detection model is trained using the annotated data set, and the model parameters are optimized using a cross entropy loss function during the training process until the detection accuracy reaches more than 95%;
[0019] In the process of training the micro-expression detection model, data enhancement techniques are used, including random cropping, flipping, and adding Gaussian noise;
[0020] After training, the cross-validation method is used to evaluate the generalization ability of the model and adjust the hyperparameters to optimize the model performance.
[0021] Preferably, the extraction of multiple motion features in step S3 includes:
[0022] Endo optical flow algorithm is used to calculate the optical flow field between frames, and multi-scale analysis is used to obtain the optical flow dynamics at different scales;
[0023] The local binary pattern histogram LBP-TOP algorithm is used to extract facial texture features;
[0024] Use facial key point detection algorithm to extract facial key points and track their trajectory changes;
[0025] The facial contour and texture features are extracted using HOG feature descriptor.
[0026] Preferably, the feature fusion model in step S4 adopts a deep learning architecture including:
[0027] In the feature fusion stage, a multi-layer LSTM structure is used to capture long-term dependencies;
[0028] Introducing the attention mechanism to dynamically assign importance weights to different features and improve recognition accuracy;
[0029] During the training process, enhancement techniques such as random masking and data perturbation are used to enhance the robustness of the model;
[0030] Residual connections are added to the feature fusion model to alleviate the gradient vanishing problem.
[0031] Preferably, the micro-expression category database in step S5 is constructed by the following steps:
[0032] Collect micro-expression samples from different individuals, and each sample is annotated by experts with the specific micro-expression type and its feature vector;
[0033] Establish the mapping relationship between micro-expression categories and fused feature vectors, and use principal component analysis (PCA) to reduce the dimension of feature vectors;
[0034] After the database is built, the newly extracted feature vectors are classified using the K nearest neighbor algorithm or the support vector machine SVM classifier;
[0035] By continuously updating the database, new micro-expression samples are included;
[0036] A clustering algorithm is used to cluster the feature vectors in the database to discover potential micro-expression types.
[0037] Preferably, the step S1 includes:
[0038] Preprocess the video sequence;
[0039] Use Gaussian filter to denoise the video sequence;
[0040] Use histogram equalization technology to normalize the brightness of video frames;
[0041] Scale the video frames to unify the frame size;
[0042] Perform color space conversion.
[0043] Preferably, the micro-expression detection model in step S2 includes:
[0044] During the training process, we use transfer learning to first pre-train on a large-scale facial expression dataset, and then fine-tune on a specific micro-expression dataset;
[0045] During the training process, a class balancing strategy is used to ensure sufficient representation of various types of micro-expression samples;
[0046] Use a multi-task learning framework to simultaneously optimize micro-expression detection and classification tasks.
[0047] Preferably, the feature fusion model in step S4 includes:
[0048] By monitoring the performance indicators on the validation set, training is stopped when the performance stops improving to prevent overfitting.
[0049] During the training process, the best performing model version is saved regularly for subsequent use;
[0050] Adopt a multi-task learning framework to optimize micro-expression detection and classification tasks simultaneously;
[0051] During the training process, the dropout technique is used to reduce the risk of overfitting of the model.
[0052] Preferably, the post-processing in step S6 includes:
[0053] Kalman filtering technology is used to smooth the recognition results;
[0054] Set a threshold and filter the recognition results below the threshold;
[0055] Use morphological operations to further optimize the recognition results and remove isolated points and small areas;
[0056] The recognition results of consecutive frames are checked for consistency using the time window smoothing technique.
[0057] Preferably, including:
[0058] In the feature fusion model, an adaptive learning rate adjustment strategy is introduced to accelerate model convergence;
[0059] In the post-processing step, a threshold filtering mechanism is used to filter the recognition results to improve the accuracy of the recognition results.
[0060] In the feature extraction step, an environmental factor correction algorithm is used, including corrections for changes in lighting conditions and non-uniform backgrounds;
[0061] In the feature extraction step, a head posture correction algorithm is used to handle problems such as head tilt and occlusion;
[0062] In the feature extraction step, a lightweight processing algorithm is used to deal with the problem of limited low pixels of video acquisition equipment and poor computer processing capabilities.
[0063] The present invention has the following beneficial effects:
[0064] 1. In the present invention, by combining multiple feature extraction methods such as optical flow field, LBP-TOP, facial key point trajectory and HOG features, it is ensured that the instantaneous changes of micro-expressions and subtle muscle movements are captured, and the comprehensiveness and accuracy of feature extraction are improved. Secondly, a variety of data enhancement techniques are used to improve the robustness of the model under different conditions, so that it can still maintain a high recognition accuracy under different lighting conditions, facial expression changes and other factors. In addition, by combining LSTM and attention mechanism with PCA dimensionality reduction technology, the model structure is simplified, the computational complexity is reduced, the feasibility of real-time application is improved, and the training time and deployment cost are also reduced.
[0065] 2. In the present invention, by ensuring that the data set contains a sufficient number of micro-expression samples and covers a variety of emotional states, a class balancing strategy is used to ensure that various types of micro-expression samples are fully represented, thereby improving the generalization ability of the model. In addition, the attention mechanism, early stopping strategy and dropout technology are introduced to prevent overfitting, improve the performance of the model on new data, and enhance the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 A flowchart of a micro-expression recognition method based on multi-motion feature fusion proposed by the present invention;
[0067] Figure 2 This is a multi-feature fusion framework diagram in the present invention;
[0068] Figure 3 This is a framework diagram for processing interference factors in the present invention. DETAILED DESCRIPTION
[0069] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0070] like Figure 1 As shown, the micro-expression recognition method based on multi-motion feature fusion proposed by the present invention includes:
[0071] S1: Get the video sequence to be analyzed;
[0072] S2: inputting the video sequence into a pre-built micro-expression detection model, and detecting the time period in the video sequence containing micro-expressions through the micro-expression detection model, wherein the micro-expression detection model is trained using a convolutional neural network (CNN), and the training data set includes video clips with the time when micro-expressions occur marked;
[0073] S3: extracting multiple motion features from the time period, the multiple motion features including an optical flow field, a local binary pattern histogram LBP-TOP, a facial key point trajectory, and an inter-frame motion vector, the optical flow field being calculated by an Endo optical flow algorithm;
[0074] S4: inputting the multi-motion features into a pre-trained feature fusion model, and obtaining a fusion feature vector through the feature fusion model, wherein the feature fusion model is a deep learning model based on a long short-term memory network (LSTM), and an attention mechanism is introduced during the training process to enhance the recognition of key frames;
[0075] S5: identifying the micro-expression category within the time period according to the fused feature vector and using a pre-established micro-expression category database, wherein the micro-expression category database includes a plurality of micro-expression types and their corresponding typical feature vectors;
[0076] S6: Post-process the recognition results.
[0077] In one embodiment, the video sequence to be analyzed is obtained.
[0078] First, you need to obtain the video sequences to be analyzed. These video sequences can come from camera recordings, storage files, or other video sources. After obtaining the video sequences, perform the following preprocessing steps:
[0079] Denoising: Use a Gaussian filter to denoise the video frame to remove high-frequency noise. The specific steps are as follows:
[0080] Use the `cv2.GaussianBlur()` function from the OpenCV library with a filter size of 3×3 and a standard deviation of 1.5.
[0081] Apply a Gaussian filter to each frame of video until all video frames have been processed.
[0082] Brightness standardization: Use histogram equalization technology to standardize the brightness of video frames to ensure the consistency of feature extraction. The specific steps are as follows:
[0083] Use the `cv2.equalizeHist()` function in the OpenCV library to perform histogram equalization on each frame of video.
[0084] Ensure that the histogram distribution of each frame image after processing is uniform.
[0085] Color space conversion: Convert the color space of the video frame from RGB to HSV to enhance the expression of color features. The specific steps are as follows:
[0086] Convert the RGB image to HSV format using `cv2.cvtColor()` function from OpenCV library.
[0087] Extract the V channel (brightness) as part of subsequent feature extraction.
[0088] Scaling: Scaling the video frames to unify the frame size to meet the needs of subsequent feature extraction. The specific steps are as follows:
[0089] Use the `cv2.resize()` function in the OpenCV library to uniformly resize all video frames to a fixed size, such as 640×480 pixels.
[0090] In one embodiment, micro-expression detection
[0091] Next, the preprocessed video sequence is input into a pre-built micro-expression detection model, and the model is used to detect the time periods in the video sequence that contain micro-expressions.
[0092] Data collection: A standard micro-expression video dataset was collected, and each video clip was annotated by experts to indicate whether there was a micro-expression and the time period in which it occurred. The specific steps are as follows:
[0093] Select appropriate data from public datasets, such as SAMM or CASME II, and ensure that the dataset contains at least 1,000 micro-expression samples covering different emotional states.
[0094] Data annotation: Use specialized annotation tools (such as Labelbox or VGG Image Annotator) to annotate the video dataset. The specific steps are as follows:
[0095] Go through the video frame by frame and mark the time periods that contain micro-expressions.
[0096] Label the specific type of each micro-expression (such as surprise, disgust, happiness, etc.).
[0097] Data enhancement: Use data enhancement techniques, including random cropping, flipping, and adding Gaussian noise, to improve the robustness of the model. The specific steps are as follows:
[0098] Using random cropping technology, different areas of the video frame are cropped each time, and the cropping ratio is 80%-120% of the original image.
[0099] Use horizontal flipping technology to increase the diversity of the data set.
[0100] Add Gaussian noise with a standard deviation of 0.1.
[0101] Model training: Use convolutional neural network (CNN) as the basic model architecture. The specific steps are as follows:
[0102] Use TensorFlow or PyTorch framework to build a CNN model, which includes 3 convolutional layers, 2 maximum pooling layers, and 2 fully connected layers.
[0103] The convolution kernel sizes of the convolutional layers are 3×3, 3×3, and 3×3, respectively, with a stride of 1 and the same padding.
[0104] The pooling kernel sizes of the maximum pooling layer are 2×2 and 2×2, respectively, with a stride of 2.
[0105] The number of nodes in the fully connected layers is 128 and 64 respectively, and the ReLU activation function is used.
[0106] The output layer uses Softmax as the activation function.
[0107] The micro-expression detection model is trained using the labeled dataset, with the batch size set to 32, the learning rate set to 0.001, and the number of iterations set to 100.
[0108] The model parameters were optimized using the cross entropy loss function until the detection accuracy reached above 95%.
[0109] Transfer learning: During the training process of the micro-expression detection model, transfer learning technology is used to first pre-train on a large-scale facial expression dataset, and then fine-tune on a specific micro-expression dataset to improve the generalization ability of the model.
[0110] Model evaluation: After training is completed, use the cross-validation method to evaluate the generalization ability of the model and adjust the hyperparameters to optimize the model performance. The specific steps are as follows:
[0111] The dataset is divided into training set (70%), validation set (15%) and test set (15%).
[0112] The model performance was evaluated using 5-fold cross validation.
[0113] Tune hyperparameters based on performance on the validation set.
[0114] In one embodiment, multiple motion feature extraction
[0115] Multiple motion features are extracted from the time period containing micro-expressions detected, mainly including optical flow field, local binary pattern histogram (LBP-TOP), facial key point trajectory and inter-frame motion vector.
[0116] Optical flow calculation: Use the Farneback optical flow algorithm to calculate the inter-frame optical flow field, and obtain the optical flow dynamics at different scales through multi-scale analysis. The specific steps are as follows:
[0117] Use the `cv2.calcOpticalFlowFarneback()` function in the OpenCV library to calculate the optical flow field.
[0118] Perform multi-scale analysis on the optical flow field to extract features at different scales.
[0119] The formula represents the calculation of optical flow field:
[0120] F=calcOpticalFlowFarneback(I t ,I t+1 ,pyr s cale=0.5,levels=3,winsize
[0121] =15,iterations=3,poly n =5,poly s igma=1.2,flags=0)
[0122] Among them, I t and I t+1 Represent two consecutive frames of images respectively.
[0123] Local Binary Pattern Histogram (LBP-TOP): Apply LBP-TOP algorithm to extract facial texture features. The specific steps are as follows:
[0124] Use the `localBinaryPatternsHistograms` function in the OpenCV library to extract LBP features.
[0125] Construct LBP-TOP histogram to capture facial texture information.
[0126] The formula represents LBP feature extraction:
[0127]
[0128] Where p is the number of neighbor points, R is the circle radius, and s is the sign function.
[0129] Facial key point trajectory: Use the facial key point detection algorithm to extract the trajectory changes of facial key points (such as eyebrows, eyes, nose, mouth, cheeks, etc.) and track their trajectory changes. The specific steps are as follows:
[0130] Use the `face_alignment` tool from the OpenCV library to detect facial landmarks.
[0131] Track the trajectory changes of key points in video sequences and extract motion features.
[0132] Use the Dense Optical Flow method to track key point trajectories.
[0133] HOG feature: Use HOG feature descriptor to extract facial contour and texture features. The specific steps are as follows:
[0134] Use the `HOGDescriptor` class in the OpenCV library to extract HOG features.
[0135] Set appropriate cell size and block size to extract HOG features of the facial area.
[0136] The cell size is 8 × 8 pixels and the block size is 2 × 2 cells.
[0137] Environmental factors correction:
[0138] Lighting correction: Use histogram equalization technology to normalize the brightness of video frames to reduce the impact of lighting changes on feature extraction.
[0139] Background correction: Use background subtraction technology or image segmentation technology to remove background interference and improve the accuracy of feature extraction.
[0140] Head posture correction:
[0141] Head tilt correction: Use geometric transformation techniques affine transformation or perspective transformation to correct head tilt.
[0142] Occlusion handling: Use the deep learning model Mask R-CNN to detect and repair occluded areas to ensure the integrity of feature extraction.
[0143] Lightweight handling:
[0144] Low-resolution processing: Use super-resolution technology ESRGAN to improve the clarity of low-resolution videos.
[0145] Lightweight model: Use lightweight neural network architectures such as MobileNet or ShuffleNet to reduce the computational complexity of the model and improve processing speed.
[0146] In one embodiment, feature fusion
[0147] The above multi-motion features are input into the pre-trained feature fusion model, and the fused feature vector is obtained through the feature fusion model.
[0148] Feature fusion model: A deep learning architecture based on the long short-term memory network (LSTM) is used, and an attention mechanism is introduced during the training process to enhance the recognition of key frames. The specific steps are as follows:
[0149] Use TensorFlow or PyTorch framework to build an LSTM model, which includes two LSTM layers, and the number of hidden units in each LSTM layer is 128.
[0150] The attention mechanism is introduced into the LSTM model, and the Softmax function is used to calculate the attention weight.
[0151] Set the appropriate hidden layer size and number of layers to optimize the model structure.
[0152] The formula represents the calculation of the LSTM layer:
[0153] h t =LSTM(h t-1 ,x t )
[0154] Among them, ht is the hidden state at the current moment, h t-1 is the hidden state at the previous moment, x t is the input feature at the current moment.
[0155] Data enhancement: During the training process, enhancement techniques such as random masking and data perturbation are used to enhance the robustness of the model. The specific steps are as follows:
[0156] Randomly mask a portion of the data in the input features, with a ratio of 20%.
[0157] Use data perturbation techniques such as adding random noise with a standard deviation of 0.05.
[0158] Residual connection: Add residual connection to the feature fusion model to alleviate the gradient vanishing problem. The specific steps are as follows:
[0159] Add residual blocks to the LSTM model and use skip connections.
[0160] Set the appropriate number of residual blocks to optimize model performance.
[0161] Adaptive learning rate adjustment strategy: In the feature fusion model, the adaptive learning rate adjustment strategy Adam optimizer is introduced to accelerate model convergence
[0162] In one embodiment, micro-expression category recognition
[0163] According to the fused feature vector, the micro-expression category within the time period is identified using a pre-established micro-expression category database.
[0164] Database construction: Collect micro-expression samples of different individuals, and each sample is annotated by experts with the specific micro-expression type and its feature vector. The specific steps are as follows:
[0165] Use the labeled dataset to build an initial database containing at least 1,000 micro-expression samples.
[0166] Make sure your database contains a variety of micro-expression samples.
[0167] Feature mapping: Establish the mapping relationship between micro-expression categories and fused feature vectors, and use principal component analysis (PCA) to reduce the dimension of feature vectors to reduce computational complexity. The specific steps are as follows:
[0168] The PCA algorithm is used to reduce the dimension of the fused feature vector and retain the first 50 principal components.
[0169] Construct a mapping relationship between feature vectors and micro-expression categories.
[0170] The formula represents PCA dimensionality reduction:
[0171] Y=XW
[0172] Among them, X is the original feature matrix, W is the weight matrix after dimensionality reduction, and Y is the feature matrix after dimensionality reduction.
[0173] Classification verification: After the database is built, use a classifier such as the K nearest neighbor algorithm or support vector machine (SVM) to classify the newly extracted feature vectors to verify the validity of the database. The specific steps are as follows:
[0174] Train a classifier using the `KNeighborsClassifier` or `SVC` class from the Scikit-learn library.
[0175] Set appropriate hyperparameters to optimize classifier performance.
[0176] The classifier performance was evaluated using the cross-validation method.
[0177] Continuous updating: By continuously updating the database and incorporating new micro-expression samples, the database can be kept up to date and accurate. The specific steps are as follows:
[0178] Regularly collect new micro-expression samples and update the database.
[0179] Retrain the classifier using new data, maintaining classification performance.
[0180] In one embodiment, post-processing
[0181] Post-process the recognition results, including but not limited to removing false positives and smoothing the recognition results, to improve the reliability of the final recognition results.
[0182] Smoothing: Kalman filtering technology is used to smooth the recognition results to reduce the impact of noise. The specific steps are as follows:
[0183] The Kalman filter algorithm is used to smooth the recognition results.
[0184] Set appropriate filtering parameters to optimize the smoothing effect.
[0185] The formula represents Kalman filtering:
[0186]
[0187] in, is the state estimate at the current moment, A is the state transfer matrix, K t is the Kalman gain, z t is the observation value, and H is the observation matrix.
[0188] Threshold filtering: Set a threshold and filter the recognition results below the threshold to remove false positives. The specific steps are as follows:
[0189] Set an appropriate recognition threshold, usually 0.5 or higher.
[0190] Results below the threshold are filtered out, and only recognition results with high confidence are retained.
[0191] Morphological operation: Use morphological operation to further optimize the recognition results and remove isolated points and small areas. The specific steps are as follows:
[0192] Morphological operations are performed using the `morphologyEx` function from the OpenCV library.
[0193] Set the appropriate kernel size to optimize the results.
[0194] Use erosion followed by dilation to remove isolated points.
[0195] Use threshold filtering mechanism to filter the recognition results to further improve the accuracy of the recognition results
[0196] Time window smoothing: Use time window smoothing technology to check the consistency of recognition results of consecutive frames. The specific steps are as follows:
[0197] The consistency of the recognition results is checked in consecutive time windows.
[0198] Remove inconsistent results to maintain the consistency of recognition results.
[0199] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A micro-expression recognition method based on multi-motion feature fusion, characterized in that: include: S1: Get the video sequence to be analyzed; S2: inputting the video sequence into a pre-built micro-expression detection model, and detecting the time period in the video sequence containing micro-expressions through the micro-expression detection model, wherein the micro-expression detection model is trained using a convolutional neural network (CNN), and the training data set includes video clips with the time when micro-expressions occur marked; S3: extracting multiple motion features from the time period, the multiple motion features including an optical flow field, a local binary pattern histogram LBP-TOP, a facial key point trajectory, and an inter-frame motion vector, the optical flow field being calculated by an Endo optical flow algorithm; S4: inputting the multi-motion features into a pre-trained feature fusion model, and obtaining a fusion feature vector through the feature fusion model, wherein the feature fusion model is a deep learning model based on a long short-term memory network (LSTM), and an attention mechanism is introduced during the training process to enhance the recognition of key frames; S5: identifying the micro-expression category within the time period according to the fused feature vector and using a pre-established micro-expression category database, wherein the micro-expression category database includes a plurality of micro-expression types and their corresponding typical feature vectors; S6: Post-process the recognition results.
2. The micro-expression recognition method based on multi-motion feature fusion according to claim 1, characterized in that: The micro-expression detection model in step S2 is obtained by training through the following steps: Collect a standard micro-expression video dataset, which contains micro-expression video clips in various emotional states, and each video clip is annotated by experts to indicate whether there is a micro-expression and the time period in which it occurs; The micro-expression detection model is trained using the annotated data set, and the model parameters are optimized using a cross entropy loss function during the training process until the detection accuracy reaches more than 95%; In the process of training the micro-expression detection model, data enhancement techniques are used, including random cropping, flipping, and adding Gaussian noise; After training, the cross-validation method is used to evaluate the generalization ability of the model and adjust the hyperparameters to optimize the model performance.
3. The micro-expression recognition method based on multi-motion feature fusion according to claim 1, characterized in that: The extraction of multiple motion features in step S3 includes: Endo optical flow algorithm is used to calculate the optical flow field between frames, and multi-scale analysis is used to obtain the optical flow dynamics at different scales; The local binary pattern histogram LBP-TOP algorithm is used to extract facial texture features; Use facial key point detection algorithm to extract facial key points and track their trajectory changes; The facial contour and texture features are extracted using HOG feature descriptor.
4. The micro-expression recognition method based on multi-motion feature fusion according to claim 1, characterized in that: The feature fusion model in step S4 adopts a deep learning architecture including: In the feature fusion stage, a multi-layer LSTM structure is used to capture long-term dependencies; Introducing the attention mechanism to dynamically assign importance weights to different features and improve recognition accuracy; During the training process, enhancement techniques such as random masking and data perturbation are used to enhance the robustness of the model; Residual connections are added to the feature fusion model to alleviate the gradient vanishing problem.
5. The micro-expression recognition method based on multi-motion feature fusion according to claim 1, characterized in that: The micro-expression category database in step S5 is constructed by the following steps: Collect micro-expression samples from different individuals, and each sample is annotated by experts with the specific micro-expression type and its feature vector; Establish the mapping relationship between micro-expression categories and fused feature vectors, and use principal component analysis (PCA) to reduce the dimension of feature vectors; After the database is built, the newly extracted feature vectors are classified using the K nearest neighbor algorithm or the support vector machine SVM classifier; By continuously updating the database, new micro-expression samples are included; A clustering algorithm is used to cluster the feature vectors in the database to discover potential micro-expression types.
6. The micro-expression recognition method based on multi-motion feature fusion according to claim 1, characterized in that: The step S1 includes: Preprocess the video sequence; Use Gaussian filter to denoise the video sequence; Use histogram equalization technology to normalize the brightness of video frames; Scale the video frames to unify the frame size; Perform color space conversion.
7. The micro-expression recognition method based on multi-motion feature fusion according to claim 1, characterized in that: The micro-expression detection model in step S2 includes: During the training process, we use transfer learning to first pre-train on a large-scale facial expression dataset, and then fine-tune on a specific micro-expression dataset; During the training process, a class balancing strategy is used to ensure sufficient representation of various types of micro-expression samples; Use a multi-task learning framework to simultaneously optimize micro-expression detection and classification tasks.
8. The micro-expression recognition method based on multi-motion feature fusion according to claim 1, characterized in that: The feature fusion model in step S4 includes: By monitoring the performance indicators on the validation set, training is stopped when the performance stops improving to prevent overfitting. During the training process, the best performing model version is saved regularly for subsequent use; Adopt a multi-task learning framework to optimize micro-expression detection and classification tasks simultaneously; During the training process, the dropout technique is used to reduce the risk of overfitting of the model.
9. The micro-expression recognition method based on multi-motion feature fusion according to claim 1, characterized in that: The post-processing in step S6 includes: Kalman filtering technology is used to smooth the recognition results; Set a threshold and filter the recognition results below the threshold; Use morphological operations to further optimize the recognition results and remove isolated points and small areas; The recognition results of consecutive frames are checked for consistency using the time window smoothing technique.
10. The micro-expression recognition method based on multi-motion feature fusion according to claim 1, characterized in that: include: In the feature fusion model, an adaptive learning rate adjustment strategy is introduced to accelerate model convergence; In the post-processing step, a threshold filtering mechanism is used to filter the recognition results to improve the accuracy of the recognition results. In the feature extraction step, an environmental factor correction algorithm is used, including corrections for changes in lighting conditions and non-uniform backgrounds; In the feature extraction step, a head posture correction algorithm is used to handle problems such as head tilt and occlusion; In the feature extraction step, a lightweight processing algorithm is used to deal with the problem of limited low pixels of video acquisition equipment and poor computer processing capabilities.
Citation Information
Patent Citations
Micro-expression recognition method based on graph deep learning and face three-dimensional reconstruction
CN114333002A
Multi-modal face emotion recognition method and device
CN114399818A
Micro-expression recognition method and device
CN116543440A
Micro-expression recognition method based on face key point and optical flow feature fusion
CN118781636A
Micro-expression recognition method based on multi-modal fusion
CN119007270A
Cited By
Video packet loss synchronous compensation method and system based on deep learning
CN120640087A