Micro-expression recognition method based on active three-dimensional imaging
By collecting multimodal data through an active three-dimensional camera and combining spatiotemporal self-attention and cross-modal cross-attention mechanisms, the problems of insufficient generalization ability and video trimming of traditional micro-expression recognition under complex lighting conditions are solved, and efficient micro-expression detection and recognition are achieved.
Patent Information
- Application Number
- CN202510800373.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-16
AI Technical Summary
Traditional micro-expression recognition methods lack generalization capabilities under complex lighting conditions, rely on manual video trimming, resulting in information loss, have difficulty processing multi-emotional clips and nesting in long videos, and ignore the temporal information of dynamic changes in micro-expressions.
An active three-dimensional camera is used to collect multimodal data, and end-to-end detection and classification are performed through spatiotemporal self-attention and cross-modal cross-attention mechanisms to construct a cross-modal micro-expression dataset. The infrared modality is used to eliminate ambient light interference, the depth modality is used to record 3D structure dynamics, and the RGB modality is used to retain apparent texture, achieving lighting robustness and eliminating the need for video trimming.
It improves the feature robustness and adaptability of the micro-expression detection model under complex lighting conditions, effectively captures the instantaneous dynamic features and cross-modal complementary information of micro-expressions, breaks through the traditional method's reliance on manually trimmed videos and a single modality, and improves detection accuracy and robustness.
Smart Images

Figure CN120635966A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of micro-expression recognition, and in particular to a micro-expression recognition method based on active three-dimensional imaging. Background Art
[0002] The short duration and slight facial deformation of micro-expressions make it difficult for traditional visual perception methods to effectively capture them, and place extremely high demands on the spatiotemporal resolution and feature sensitivity of the detection algorithm.
[0003] Current mainstream micro-expression datasets (such as SAMM and CASME) are primarily collected in controlled laboratory lighting environments, relying on uniform white light illumination and fixed camera angles. While these datasets provide standardized benchmarks for early algorithm validation, they have significant limitations for real-world applications. On the one hand, lighting variations in real scenes (such as low illumination, strong reflections, and color temperature differences) can severely disrupt the color consistency and texture clarity of RGB images, resulting in insufficient generalization capabilities for detection models based on surface features. On the other hand, the natural occurrence of micro-expressions is sporadic and hidden, and the expressiveness of expressions during artificially induced collection deviates from real-world scenarios. Furthermore, labeling the start and end frames and categories of expressions frame by frame requires a significant amount of specialized manpower, resulting in the generally small size of public datasets, making it difficult to support the training needs of complex deep learning models.
[0004] The micro-expression analysis process is typically divided into two core stages: micro-expression localization and micro-expression recognition. Early micro-expression localization methods relied on manually designed spatiotemporal features, using a sliding window to detect segments with significant differences between frames.
[0005] With the development of deep learning, temporal modeling methods based on three-dimensional convolutional neural networks (3DCNNs) and long short-term memory (LSTM) recurrent neural networks have been proposed. However, existing methods suffer from two major drawbacks: most models assume that micro-expressions exist in a single segment, making it difficult to handle the nesting and overlap of multiple emotional segments in long videos; and they rely on manually trimmed short clips as input, making them unable to directly process original long video sequences. This necessitates pre-segmentation in practical applications, increasing preprocessing costs and the risk of information loss. In recognition tasks, traditional methods extract static surface features from peak frames for classification, ignoring the temporal information involved in the dynamic evolution of micro-expressions. In recent years, end-to-end models based on spatiotemporal networks have attempted to exploit inter-frame motion information in video sequences. However, they still face numerous challenges in real-world scenarios, such as a strong dependence on laboratory lighting conditions and difficulty handling invalid background frames in untrimmed videos. Summary of the Invention
[0006] In order to overcome the defects of the above-mentioned existing technologies, the present invention provides a micro-expression recognition method based on active three-dimensional imaging, which adopts an active three-dimensional camera to collect multimodal data, and realizes end-to-end detection and classification through spatiotemporal self-attention and cross-modal cross-attention mechanisms. It has the characteristics of good illumination robustness and no need for video trimming.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is:
[0008] A micro-expression recognition method based on active three-dimensional imaging comprises the following steps:
[0009] Step 1: Use active 3D cameras to capture images and construct a micro-expression dataset with illumination robustness and multimodal features; active 3D cameras include time-of-flight cameras and structured light cameras;
[0010] Step 2: Input the micro-expression dataset into the 3D convolutional network to obtain the initial fusion features;
[0011] Step 3: Input the initial fusion features into the spatiotemporal self-attention mechanism module, use the multi-head attention mechanism to perform spatiotemporal joint modeling on the feature sequence, capture the long-range temporal dependency and dynamic change areas of micro-expressions; introduce the cross-attention mechanism to obtain the fusion features, and compress the fusion features into compressed features F through global average pooling. global , retaining key spatiotemporal information;
[0012] Step 4: Input the compressed features into a two-branch network to simultaneously achieve temporal localization and emotion classification of micro-expression segments.
[0013] In step 1, during the dataset construction phase, an active 3D camera Intel Real Sense LiDARL515 is used to collect active infrared light intensity maps, depth maps, and RGB images.
[0014] By setting up multiple scenarios – low illumination (<10 lux in a dark room) and normal indoor illumination (300-500 lux), we invited subjects of different ages, genders, and skin colors to watch emotionally impactful videos and naturally induce micro-expressions to ensure data diversity.
[0015] Analyze the video frame by frame, annotate seven micro-expression categories (angry, disgusted, fearful, happy, surprised, sad, other) and time boundaries (start frame, peak frame, end frame);
[0016] During preprocessing, the active infrared light intensity map is enhanced by histogram equalization, the depth map is denoised by median filtering, the RGB image is white-balanced, and multimodal data is aligned in time by analyzing timestamps. The video sequence is stabilized based on the optical flow method to reduce position deviations caused by shaking or jitter.
[0017] In the step 2:
[0018] Based on the duration characteristics of micro-expressions, the window size is set to s frames and slid with a fixed step size θ (taking into account both real-time detection and accuracy, ensuring that adjacent windows overlap to avoid missed detections). Continuous frame subsequences of active infrared light intensity maps, depth maps, and RGB images are extracted and used as input to the 3D convolutional network branch to extract the spatiotemporal features of each modality.
[0019] The RGB modality branch of the RGB image is input with 3 channels and the spatiotemporal features of the apparent texture are extracted through a 3D convolutional network;
[0020] The infrared modality branch of the active infrared light intensity map is input as a single channel; the spatiotemporal features of the intensity changes are extracted through a 3D convolutional network;
[0021] The depth modality branch of the depth map takes 1 channel as input and extracts the spatiotemporal features of the facial 3D structure displacement;
[0022] The output features of each branch are spliced through the channel to form the initial fusion feature F init .
[0023] The step 2 is specifically as follows:
[0024] The pre-processed multimodal sequence data (RGB, active infrared, depth map) is normalized and input into the modality-specific feature extraction module;
[0025] Design independent 3D convolutional networks for different modal characteristics:
[0026] The first layer of the RGB modality branch is a 3D convolutional layer (kernel size 3×3×3, stride 1×1×1, padding 1×1×1) with 64 convolution kernels, followed by a batch normalization (BN) layer and a rectified linear unit (ReLU) activation function to extract spatiotemporal features of the apparent texture (such as the spatiotemporal gradient of facial muscle texture changes).
[0027] The first 3D convolutional layer of the infrared / depth modality branch (kernel size 3×3×3, stride 1×1×1, padding 1×1×1) is configured with 32 convolution kernels. Subsequent processing is the same as the RGB branch, extracting the spatiotemporal features of intensity changes (infrared) and facial 3D structure displacement (depth), respectively.
[0028] After two layers of 3D convolution, the output feature tensor size of each branch is unified to T×64×64×C (T is the sequence length, C is the number of channels, C=64 for RGB branch, C=32 for infrared / depth branch), and the features of three different forms are spliced through channels to form the initial fusion feature.
[0029] The step three is specifically as follows:
[0030] The spatiotemporal self-attention mechanism module is used to capture the long-range temporal dependencies and dynamic change areas of the initial fusion features:
[0031] First, the spatiotemporal position coding is added to the frames in each segment of the initial fusion feature, and the time coding adopts the sine function Where k is the dimension index, D is the feature dimension; t is the length of the time segment, x and y are the width and height of the video respectively; by the ratio of the dimension index k to the total dimension D As an index, it makes the position encoding of different dimensions have different frequencies, thereby distinguishing the relative position information of time and space in the spatiotemporal self-attention mechanism and capturing the long-range temporal dependency and spatial structure characteristics of micro-expressions;
[0032] Then, a multi-head attention mechanism is used to perform spatiotemporal joint modeling on the feature sequence after time encoding, and the output is a feature F containing the temporal dynamic priority. temp ; Each encoder layer contains multi-head self-attention and feedforward network (FFN);
[0033] Multi-head self-attention projects the feature sequence into query Q, key K, value V, and calculates the attention weight by scaling the dot product Among them, Q, K, and V are generated by linear mapping of input features, d k is the dimension scaling factor; the activation function is GELU, followed by residual connection and layer normalization (LN) to prevent gradient disappearance; after being processed by the encoder, the output contains the feature F of temporal dynamic priority temp , where the areas with high attention weights correspond to micro-expression key frames (such as starting frame and peak frame);
[0034] Design a cross-modal attention mechanism to achieve inter-modal information complementarity, using RGB modality feature F r For the benchmark query Q r , infrared modal characteristics F i With deep modal features F d As key-value pairs (K i ,V i )、(K d ,V d ), unify the dimension to d through linear transformation k =64;
[0035] The fusion between modalities is achieved by calculating the correlation weights between modalities, and the correlation weights between modalities are calculated through a learnable parameter matrix. where β r,iIndicates the degree of dependence of the RGB modality on the infrared modality. In low-light illumination scenes, the weight automatically increases to enhance the compensation of infrared features for apparent noise. Finally, the gating parameter is introduced to dynamically balance the cross-modal fusion and intrinsic modal features to avoid information overload. The fused feature is F fusion =γ(β r,i V i +β r,d V d )+(1-γ)F r , and then the fused features are compressed into Preserve key spatiotemporal information.
[0036] In step 4, one branch passes through two fully connected layers followed by two regression layers to output standardized micro-expression start and end frames, and uses a smooth L1 loss function to calculate the positioning loss; the other branch passes through three fully connected layers, and the last layer uses a SoftMax function to output the probability distribution of seven categories of emotions. To address the problem of class imbalance, a weighted cross-entropy loss function is used to calculate the classification loss; the positioning loss and the classification loss are jointly optimized, and after the loss is obtained through training, the loss is continuously reduced through a backpropagation algorithm, thereby improving the positioning accuracy of micro-expression fragments and the recognition accuracy of emotional categories.
[0037] The step 4 is specifically as follows:
[0038] Input the fused features into the dual-branch network to achieve joint detection and classification:
[0039] The dual-branch network uses two fully connected layers (128→256→128), followed by two regression layers (128→64→2), and outputs a standardized starting frame. and end frame (range [0,1], through Mapped to actual frame number);
[0040] In order to deal with the ambiguity of micro-expression boundaries, smooth L1 loss is used as the positioning loss: in is the frame difference, is the indicator function;
[0041] The other path of the dual-branch network passes through three fully connected layers (128→512→256→8), and the last layer SoftMax outputs the probability distribution of 7 categories of emotions;
[0042] To address the class imbalance problem in the micro-expression dataset where the number of samples of each emotion category varies significantly, a weighted cross entropy loss is used:
[0043] Weight ω c The weight is calculated inversely according to the number of category samples (for example, if the proportion of neutral samples is high, the weight is reduced) to balance the gradient contribution of each emotion category during training.
[0044] The overall network model adopts end-to-end training, and the joint loss function is: The optimizer selected is AdamW (weight decay 0.01), and the initial learning rate is 1e -4 , using cosine annealing learning rate scheduling (200 rounds, minimum learning rate 1e -6 .
[0045] Beneficial effects of the present invention:
[0046] By simultaneously collecting active infrared light intensity maps, depth maps, and RGB images, a cross-modal micro-expression dataset was constructed. Leveraging the complementary properties of infrared modality to eliminate ambient light interference, depth modality to record 3D structural dynamics, and RGB modality to preserve apparent texture, this method effectively addresses the feature degradation problem of traditional single-modal RGB data under complex lighting conditions. In low-light scenarios, feature robustness is significantly enhanced, significantly improving the detection model's adaptability to real-world environments.
[0047] A cross-modal micro-expression dataset is constructed by simultaneously collecting active infrared light intensity maps, depth maps, and RGB images. The principle behind this approach is to leverage the physical properties of different modalities to achieve environmental adaptability and feature complementarity. The active infrared modality utilizes the thermal radiation principle of 860nm near-infrared light to avoid interference from ambient visible light and stably record the dynamic facial temperature field, achieving illumination invariance. The depth modality, based on the time-of-flight (ToF) method, calculates the 3D facial coordinates by measuring the round-trip time of laser pulses, recording the spatial deformation trajectory of micro-expressions. The RGB modality, based on the reflectivity of visible light, preserves facial surface texture and color information. The cross-modal complementarity mechanism dynamically assigns weights and strengthens semantic associations through spatiotemporal self-attention and cross-modal cross-attention, forming a composite feature that combines "illumination robustness, spatial structure, and surface details." This addresses the feature degradation problem of traditional single-modal RGB under complex illumination at the physical signal level. The proposed spatiotemporal joint detection and classification model breaks away from the traditional "localization-then-recognition" staged framework. It automatically focuses on micro-expression keyframes through a spatiotemporal self-attention mechanism and combines multimodal cross-attention for feature fusion, enabling direct processing of raw long video sequences without manual video trimming. This method effectively reduces positioning errors and improves the model's detection efficiency for micro-expressions that appear randomly and have non-fixed durations in natural scenes.
[0048] By synergistically optimizing spatiotemporal self-attention and multimodal cross-attention, the model effectively captures the instantaneous dynamic characteristics of micro-expressions and their cross-modal complementary information, breaking through the existing methods' reliance on manually trimmed videos and a single modality. This technical solution not only improves detection accuracy but also significantly enhances the model's robustness and generalization capabilities, providing an efficient solution for micro-expression detection in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 Schematic diagram comparing the present invention with other micro-expression recognition methods.
[0050] Figure 2 The seven emotional distributions of micro-expressions and macro-expressions in the micro-expression dataset constructed by the present invention are shown.
[0051] Figure 3 This is a distribution information diagram of the duration of expression segments in the micro-expression dataset videos collected by the present invention.
[0052] Figure 4 is a schematic diagram of the sliding window mechanism used in the present invention.
[0053] Figure 5 This is the entire process of frame processing of expression segments in micro-expression detection of the present invention.
[0054] Figure 6 A comparison chart of the confusion matrices of the micro-expression detection models under different attention mechanism configurations in the present invention. DETAILED DESCRIPTION
[0055] The present invention will be described in further detail below with reference to the accompanying drawings.
[0056] A micro-expression recognition method based on active three-dimensional imaging comprises the following steps:
[0057] 1. Dataset construction:
[0058] In order to construct a micro-expression dataset with illumination robustness and multimodal features, the present invention uses the Intel Real Sense LiDAR L515 ToF laser radar depth camera device for data acquisition. The device can synchronously collect active infrared light intensity maps, depth maps and RGB images, providing rich modal information for the dataset. The active infrared wavelength of the camera's laser radar module is 860nm, and the depth acquisition principle is the time-of-flight method (ToF). The target distance is calculated by measuring the round-trip time of the laser pulse. The depth resolution is 1024×768 pixels, the accuracy is ±0.5mm (within the range of 0.5-5m), and the frame rate is 30fps (the depth map is synchronized with the infrared map). The image resolution of the RGB camera captured by the color CMOS sensor is 1920×1080 pixels, the spectral response range is 400-700nm (visible light band), the frame rate is 30fps, and it is synchronized with the laser radar module through hardware triggering.
[0059] During the data collection process, multiple lighting scenarios were created to simulate various real-world lighting conditions. These included low-light scenarios (such as a darkroom with less than 10 lux) and normal indoor lighting scenarios (approximately 300-500 lux). Multiple subjects were invited to participate in the micro-expression data collection under each lighting scenario. The subjects encompassed a diverse age, gender, and skin color demographic to ensure a diverse and representative dataset. Prior to data collection, the purpose and process of micro-expression data collection were explained to the subjects to ensure their understanding and consent to participate. Subjects were instructed to sit in a fixed position, with their faces roughly centered within the camera's field of view and free of obstructions (such as hair or hats) to ensure image quality. A naturalistic approach was used to elicit micro-expressions from the subjects. Subjects were instructed to watch emotionally charged video clips, encouraging them to produce micro-expressions in a natural manner. As the subjects produced micro-expressions, the device simultaneously captured active infrared light intensity maps, depth maps, and RGB images, recording the complete dynamic process of the micro-expressions.
[0060] During the data standardization process, the annotation team analyzes the collected video sequences frame by frame to determine the categories of micro-expressions (including anger, disgust, fear, happiness, surprise, sadness, and others). Based on the facial movement characteristics and emotional expression of micro-expressions, combined with relevant knowledge and experience in micro-expression recognition, the annotation team accurately determines the category of each micro-expression segment. They also annotate temporal boundaries: marking the start frame (Onset), peak frame (Apex), and end frame (Offset) of the micro-expression to determine the temporal boundaries of the micro-expression within the video sequence.
[0061] Because active infrared light intensity images may be affected by environmental factors, they can lead to low image contrast and unclear details. Histogram equalization and contrast-limited adaptive histogram equalization are used as image enhancement algorithms to improve the contrast and clarity of active infrared light intensity images, enhancing the discernibility of facial thermal radiation features. Depth images may contain a certain degree of noise and rough areas, which may affect the subsequent analysis of muscle deformation and depth changes. A median filter is used to smooth the depth image, removing noise points and making it smoother and more natural, while preserving key details of facial muscle deformation. Color correction is performed to address color deviations that may occur in RGB images under different lighting conditions. A white balance correction algorithm is used to adjust the color balance of the RGB image, ensuring that it presents more accurate color information in different lighting scenarios, improving the visual quality and feature consistency of the RGB image.
[0062] Since data from different modalities (active infrared light intensity map, depth map and RGB image) may have slight timing differences during the acquisition process, they need to be time-aligned. By analyzing the timestamp information of data from different modalities, the data from different modalities are time-calibrated to ensure their consistency in the time dimension, so that the relevant information in the multimodal data can be accurately fused and analyzed later. At the same time, the time-aligned video sequence is stabilized to reduce image position deviations caused by factors such as slight shaking of the subject's head and camera shake. Using an optical flow-based approach, the optical flow of the nose tip of the face is subtracted from the global optical flow, and each frame in the video sequence is aligned with the reference frame (such as the first frame). The displacement and rotation parameters between the images are calculated, and the subsequent frames are subjected to corresponding geometric transformations to keep the entire video sequence stable in spatial position, thereby improving the quality and availability of the data.
[0063] 2. Spatiotemporal Detection Model Training
[0064] The preprocessed multimodal sequence data (RGB, active infrared, and depth images) is normalized and then fed into the modality-specific feature extraction module. Independent 3D convolutional networks are designed for different modal characteristics: the RGB modality branch uses a three-channel input. The first layer is a 3D convolutional layer (kernel size 3×3×3, stride size 1×1×1, padding 1×1×1) with 64 convolution kernels, followed by a batch normalization (BN) layer and a rectified linear unit (ReLU) activation function. This extracts spatiotemporal features of apparent texture (such as the spatiotemporal gradient of facial muscle texture changes). The infrared / depth modality branch uses a single-channel input. The first layer is a 3D convolutional layer (kernel size 3×3×3, stride size 1×1×1, padding 1×1×1) with 32 convolution kernels. Subsequent processing is similar to the RGB branch, extracting spatiotemporal features of intensity change (infrared) and 3D facial structure displacement (depth), respectively. After two layers of 3D convolution, the output feature tensor size of each branch is unified to T×64×64×C (T is the sequence length, C is the number of channels, RGB branch C=64, infrared / depth branch C=32), and the initial fusion feature is formed by channel splicing.
[0065] To capture the long-range temporal dependencies and dynamic change regions of micro-expressions, the model uses a spatiotemporal self-attention mechanism: the time dimension T is divided into non-overlapping segments of length s = 16 (based on the average duration of micro-expressions of 30 frames, allowing for coverage of the entire expression cycle) to avoid loss of boundary information. Spatiotemporal position encoding is added to the frames within each segment, and the time encoding uses a sine function. Where k is the dimension index and D is the feature dimension. A multi-head attention mechanism is used to perform spatiotemporal joint modeling of feature sequences. Each encoder layer contains multi-head self-attention and a feed-forward network (FFN). Multi-head self-attention projects features into query Q, key K, and value V, and calculates attention weights by scaling dot products. Among them, Q, K, and V are generated by linear mapping of input features, d k is the dimension scaling factor. The activation function is GELU, followed by residual connection and layer normalization (LayerNorm, LN) to prevent gradient disappearance. After being processed by the encoder, the output contains the feature F of temporal dynamic priority temp , where the areas with high attention weights correspond to micro-expression key frames (such as starting frame and peak frame).
[0066] Design a cross-modal attention mechanism to achieve inter-modal information complementarity, using RGB modality feature F r For the benchmark query Q r , infrared modal characteristics F i With deep modal features F dAs key-value pairs (K i ,V i )、(K d ,V d ), unify the dimension to d through linear transformation k = 64. Calculate the inter-modal correlation weights through the learnable parameter matrix where β r,i Indicates the degree of dependence of the RGB modality on the infrared modality. In low-light scenes, the weight automatically increases to enhance the compensation of infrared features for apparent noise. Finally, a gating parameter is introduced to dynamically balance cross-modal fusion and intrinsic modal features to avoid information overload. The fused feature is F fusion =γ(β r,i V i +β r,d V d )+(1-γ)F r , and then the fused features are compressed into Preserve key spatiotemporal information.
[0067] The fusion features are input into the dual-branch network to achieve joint detection and classification: two layers of fully connected layers (128→256→128) are used, followed by two layers of regression layers (128→64→2), and the standardized starting frame is output. and end frame (range [0,1], through To handle the ambiguity of micro-expression boundaries, smooth L1 loss is used as the positioning loss: in is the frame difference, is the indicator function. Through three layers of fully connected layers (128→512→256→8), the final SoftMax layer outputs 7 categories of emotional probability distribution. To address the problem of class imbalance, weighted cross entropy loss is used: Weight ω c The weight is calculated inversely according to the number of category samples (for example, if the proportion of neutral samples is high, the weight is reduced) to balance the gradient contribution of each emotion category during training.
[0068] The model adopts end-to-end training, and the joint loss function is: The optimizer selected is AdamW (weight decay 0.01), and the initial learning rate is 1e -4 , using cosine annealing learning rate scheduling (200 rounds, minimum learning rate 1e -6 .
[0069] Figure 1The figure shows a comparison between the present invention and other micro-expression recognition methods. Conventional methods for emotion classification of micro-expression videos involve manually segmenting the original video into multiple video clips containing micro-expression instances, and then using a video classification model to identify the emotions in each clip. The present invention uses a temporal emotion detection method, directly employing a temporal emotion detection model to perform emotion detection on the entire, un-cropped video. This method not only allows for temporal localization of multiple emotions in the video, but also allows for simultaneous identification of the category of each emotion.
[0070] Figure 2 The seven emotional distributions of micro-expressions and macro-expressions in the micro-expression dataset constructed by the present invention are shown, including two radar charts, which are compared from the dimensions of quantity and proportion respectively. The radar chart on the left is displayed in terms of quantity: the blue part in the figure represents macro-expressions, the red part represents micro-expressions, and the seven emotion categories (fear, disgust, anger, other, surprise, sadness, and happiness) are evenly distributed on each axis of the radar chart. In comparison, the number of micro-expressions is relatively small overall. This figure can be used to intuitively compare the absolute number differences between macro-expressions and micro-expressions under each emotion. The radar chart on the right is displayed in terms of proportion: the blue part also represents macro-expressions, and the red part represents micro-expressions. This figure clearly shows the relative proportional relationship between macro-expressions and micro-expressions in the overall data for each emotion, which is convenient for analyzing the distribution characteristics of different emotions in the data set.
[0071] Figure 3 This is a distribution information diagram of the duration of expression segments in the micro-expression dataset video collected by the present invention. The horizontal axis in the figure represents the segment length (unit: frame), ranging from 0 to 120 frames, corresponding to an actual duration of 0 to 4 seconds (the video frame rate is 30 frames / second); the vertical axis represents the distribution, reflecting the frequency of occurrence of expression segments of different lengths in the dataset. According to the time definition of micro-expressions and macro-expressions, the duration of micro-expressions does not exceed 0.5 seconds, and the corresponding segment length does not exceed 15 frames. The duration of macro-expressions is 0.5 seconds to 4 seconds, and the corresponding segment length is 15 frames to 120 frames. The frame length intervals of the two types of expressions are clearly reflected in the figure. From the distribution characteristics, there are differences in the frequency of segments in each length interval. The figure intuitively presents the distribution law of the length of expression segments in a visual way. It not only provides a key basis for distinguishing micro-expressions from macro-expressions and analyzing the composition of the dataset, but also lays a data foundation for the subsequent targeted extraction of typical expression features and optimization of micro-expression detection algorithms through the identification of high-frequency intervals.
[0072] Figure 4This is a schematic diagram of the sliding window mechanism used in the present invention. A sliding window is used to detect emotion instances within a video. The sliding window size is set to s frames, which is determined based on the duration of micro-expressions to ensure that the entire micro-expression is captured. The sliding window slides from left to right across the image frame sequence with a fixed step size θ. The step size is set to take into account the real-time and accuracy requirements of detection, ensuring a certain overlap between adjacent windows to avoid missing emotion instances while improving processing efficiency.
[0073] Figure 5 This is the entire frame processing flow for expression segments in the micro-expression detection of the present invention. A series of consecutively arranged image frames are presented at the top of the figure, where the green-framed 3D area represents a specific example of an expression segment. Three key frames are clearly marked within the segment: #0138 is the start frame, marking the beginning of the micro-expression; #0143 is the peak frame, representing the moment when the micro-expression intensity reaches its peak; and #0148 is the end frame, meaning the end of the micro-expression. These three key frames clearly define the time range of the expression segment. Below the key frame annotations, the true value section defines the standard positions of the start, peak, and end frames of the expression segment, providing a reference benchmark for subsequent processing and ensuring a measurable basis for the accuracy of the detection results. The label section is intuitively displayed through frame-level labels. Rectangles filled with dark green represent frames belonging to expression segments, while rectangles filled with light green represent frames that are not expression segments. This visualization method clearly presents the label assignment for each frame, making it easier to understand the model's preliminary classification of frames. The probability component is presented as a segment-level probability bar chart. Each bar corresponds to a frame, and the bar height reflects the probability that the frame belongs to an expression segment. This design intuitively presents the model's prediction confidence for each frame, with higher values indicating a greater likelihood of the frame belonging to an expression segment. Post-processing optimizes the probability results to correct for potential errors or unreasonable intervals, further improving detection accuracy. The final prediction results, represented by black lines, represent the expression segment detection intervals determined after a series of processing steps. This diagram integrates multiple steps, including ground truth reference, frame-level labels, segment-level probabilities, and post-processing optimization, to produce an accurate and reliable detection range. This diagram comprehensively and meticulously illustrates the key micro-expression detection process, from raw frame annotation, label definition, probability calculation, post-processing, to the final prediction results. It clearly demonstrates how these steps work together to determine the temporal location of expression segments, providing an intuitive and comprehensive visual aid for understanding the mechanisms of micro-expression detection.
[0074] Figure 6This is a comparison diagram of the confusion matrices of the micro-expression detection model under different attention mechanism configurations in the present invention. The main framework of the micro-expression detection model of the present invention is: first, the self-attention mechanism is used to extract spatiotemporal features for videos of different modalities, and then the cross-attention mechanism is used to extract the difference information between modalities. The figure shows the emotion prediction confusion matrices of the complete model (including self-attention and cross-attention mechanisms), the model without self-attention mechanism, and the model without cross-attention mechanism from left to right. The horizontal axis is the predicted emotion, and the vertical axis is the real emotion. The emotion categories include anger, disgust, fear, happiness, sadness, surprise and others. Through comparison, it can be seen that the complete model has more advantages in the prediction performance of various emotions by virtue of the effective extraction of spatiotemporal features by the self-attention mechanism and the deep mining of the difference information between modalities by the cross-attention mechanism. It intuitively reflects the key role of the two attention mechanisms in improving the accuracy of model emotion recognition, and provides a visual basis for evaluating the effectiveness of the attention mechanism in the model framework of the present invention. It is an important graphical content for analyzing model performance in the technical solution of this patent.
[0075] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings:
[0076] To further illustrate the objectives, technical solutions, and core advantages of this solution, the present invention is further described in detail below with reference to the accompanying drawings and examples. Please note that the following specific examples are for illustrative purposes only and are not intended to limit the present invention. Furthermore, the technical features involved in different implementations of the embodiments may be combined with each other as long as they do not conflict with each other.
[0077] Step 1: Reference Figure 2 , Figure 3 As shown, in the data set construction link, the present invention uses an active three-dimensional camera IntelRealSense LiDAR L515 ToF to collect active infrared light intensity maps, depth maps and RGB images. By building multiple scenes such as low illumination (such as a dark room with less than 10 lux) and normal indoor lighting (about 300-500 lux), subjects of different ages, genders, and skin colors are invited to watch videos with emotional impact to naturally induce micro-expressions and ensure data diversity. The annotation team analyzes the video frame by frame, annotates the micro-expression categories (angry, disgusted, fearful, happy, surprised, sad, and other) and time boundaries (starting frame, peak frame, and ending frame). During preprocessing, the active infrared light intensity map is enhanced by histogram equalization, the depth map is denoised by median filtering, the RGB image is white-balanced, and the multimodal data is aligned in time by analyzing the timestamps. The video sequence is stabilized based on the optical flow method to reduce position deviations caused by shaking or jitter.
[0078] Step 2: Reference Figure 1 , Figure 4As shown, the spatiotemporal detection model training is performed. Figure 4 The sliding window mechanism sets the window size to s frames according to the duration characteristics of micro-expressions, and slides with a fixed step size θ (taking into account both real-time detection and accuracy, ensuring that adjacent windows overlap to avoid missed detections), and extracts continuous frame subsequences. The RGB modality branch takes 3 channels as input and extracts the spatiotemporal features of the apparent texture through 3D convolution; the infrared / depth modality branch takes 1 channel as input and extracts the spatiotemporal features of the intensity change and the facial 3D structure displacement respectively after similar processing. The output features of each branch are spliced through the channels to form the initial fusion feature F. init .
[0079] Step 3: Reference Figure 5 The model adopts a spatiotemporal self-attention mechanism to divide the time dimension T into non-overlapping segments of length s=16, and adds spatiotemporal position encoding to the frames within the segments. Where k is the dimension index, D is the feature dimension, t is the time segment length, x and y are the width and height of the video respectively. The feature sequence is modeled in time and space through a multi-head attention mechanism, and the encoder processes the feature F with temporal dynamic priority. temp , highlighting the key frames of micro-expressions (such as the starting frame and the peak frame). Design a cross-modal attention mechanism to use the RGB modal feature F r For the benchmark query Q r , infrared modal characteristics F i With deep modal features F d As key-value pairs (K i ,V i )、(K d ,V d ), calculate the inter-modal correlation weight after linear transformation and unified dimension where d k is the scaling factor to avoid gradient disappearance. The gating parameter is introduced to dynamically balance the cross-modal fusion and intrinsic modal features. The fusion features are compressed to F by global average pooling. global , retaining key spatiotemporal information.
[0080] Step 4: Input the fused features into a two-branch network, one of which passes through two fully connected layers followed by two regression layers, and outputs a standardized micro-expression start frame and end frame Smooth L1 loss is used as the positioning loss; secondly, through three layers of fully connected layers, the last layer SoftMax outputs 7 categories of emotion probability distribution, and weighted cross entropy loss is used to address category imbalance. The model is trained end-to-end, and the joint loss function and They are classification loss and positioning loss respectively. The optimizer is AdamW (weight decay is 0.01) and the initial learning rate is 1e-4 , using cosine annealing learning rate scheduling (200 rounds, minimum learning rate 1e -6 After the loss is obtained through training, the back propagation algorithm is used to continuously reduce the loss and improve the accuracy of positioning and recognition.
[0081] The multimodal dataset (including active infrared, depth, and RGB modalities) constructed in Step 1 serves as the foundation for training the spatiotemporal detection model in Step 2. The "training the spatiotemporal detection model on the micro-expression dataset" in Step 2 directly relies on the multimodal data collected and preprocessed in Step 1. The dataset's illumination robustness and diversity determine the effectiveness of model training. The initial fused features output in Step 2 serve as the input to the spatiotemporal self-attention mechanism in Step 3. The "inputting the initial fused features into the spatiotemporal self-attention module" in Step 3 is based on the spatiotemporal features extracted from each modality through 3D convolution in Step 2. These initial fused features, formed after channel concatenation, provide the multimodal information foundation for subsequent spatiotemporal modeling. The compressed features generated in Step 3 serve as the input to the dual-branch network in Step 4. The "fused features" in Step 4 are the features processed through spatiotemporal self-attention and cross-modal cross-attention in Step 3. The key spatiotemporal information retained by the dual-branch network directly impacts the accuracy of micro-expression localization and classification.
[0082] In step 1 (dataset construction), the Intel RealSense LiDAR L515 ToF camera is used to synchronously collect active infrared light intensity maps, depth maps, and RGB images. Its 860nm near-infrared light and ToF technology respectively achieve immunity to ambient light interference and facial 3D structure measurement, thereby constructing a multimodal dataset. Depth map denoising and temporal alignment in preprocessing rely on camera characteristics.
[0083] Step 2 (model training): The infrared / depth modality branch directly adapts the ToF output data, extracts the intensity change and 3D structure displacement features, and splices them with the RGB features to form the initial fusion features.
[0084] Steps 3 and 4 (modeling and detection): The cross-modal attention mechanism enhances the interaction between the ToF modality and RGB, increases the infrared weight in low light conditions, and uses depth to locate spatial dynamics. The dual-branch network uses ToF data to improve detection robustness in complex scenarios.
[0085] This paper uses the Intel RealSense LiDAR L515 active three-dimensional imaging device to simultaneously collect active infrared light intensity maps, depth maps, and RGB images to construct a cross-modal micro-expression dataset. The active infrared modality utilizes an 860nm near-infrared light source to eliminate the effects of ambient light and preserve stable facial thermal radiation characteristics. The depth modality uses a structured light camera to acquire three-dimensional facial coordinate information and record the depth changes in muscle deformation during micro-expressions. The RGB modality provides traditional surface texture features. The complementarity of the three modalities effectively improves the data's adaptability to complex lighting conditions. In scenarios with a variety of lighting intensities, the multimodal fusion features can still fully preserve the dynamic details of facial expressions, providing key data support for the development of lighting robustness algorithms.
[0086] At the same time, the present invention proposes a spatiotemporal joint detection and classification model that does not require video trimming, achieving unified modeling from raw video input to emotion category output. By designing a temporal attention mechanism, the model automatically focuses on time intervals with significant expression changes in the video sequence. At the same time, combined with a multi-stage positioning regression network, it accurately predicts the start and end frame positions of micro-expressions and jointly optimizes the segment positioning loss and classification loss. Compared with the traditional two-stage framework of "first positioning and then recognition", this method achieves the joint optimization of positioning and recognition, avoiding the accumulation of errors in the intermediate links, and is particularly suitable for processing raw long videos containing complex emotional dynamics.
Claims
1. A micro-expression recognition method based on active three-dimensional imaging, characterized in that: The following steps are included: Step 1: Use active 3D cameras to capture images and construct a micro-expression dataset with illumination robustness and multimodal features. Active 3D cameras include time-of-flight cameras and structured light cameras. Step 2: Input the micro-expression dataset into the 3D convolutional network to obtain the initial fusion features; Step 3: Input the initial fusion features into the spatiotemporal self-attention mechanism module, use the multi-head attention mechanism to perform spatiotemporal joint modeling on the feature sequence, capture the long-range temporal dependency and dynamic change areas of micro-expressions; introduce the cross-attention mechanism to obtain the fusion features, and compress the fusion features into compressed features F through global average pooling. global , retaining key spatiotemporal information; Step 4: Input the compressed features into a two-branch network to simultaneously achieve temporal localization and emotion classification of micro-expression segments.
2. The micro-expression recognition method based on active three-dimensional imaging according to claim 1, characterized in that: In step 1, during the dataset construction phase, an active 3D camera Intel Real Sense LiDAR L515 ToF is used to collect active infrared light intensity maps, depth maps, and RGB images. By setting up multiple scenarios – low illumination (<10 lux in a dark room) and normal indoor illumination (300-500 lux), we invited subjects of different ages, genders, and skin colors to watch emotionally impactful videos and naturally induce micro-expressions to ensure data diversity. Analyze the video frame by frame and annotate seven micro-expression categories and their temporal boundaries; During preprocessing, the active infrared light intensity map is enhanced by histogram equalization, the depth map is denoised by median filtering, the RGB image is white-balanced, and multimodal data is aligned in time by analyzing timestamps. The video sequence is stabilized based on the optical flow method to reduce position deviations caused by shaking or jitter.
3. The micro-expression recognition method based on active three-dimensional imaging according to claim 2, characterized in that: In the step 2: According to the duration characteristics of micro-expressions, the window size is set to s frames, and the sliding step is fixed at θ to extract continuous frame subsequences of active infrared light intensity map, depth map and RGB image. These subsequences are used to input into the 3D convolutional network branch to extract the spatiotemporal features of each modality. The RGB modality branch of the RGB image is input with 3 channels and the spatiotemporal features of the apparent texture are extracted through a 3D convolutional network; The infrared modality branch of the active infrared light intensity map is input as a single channel; the spatiotemporal features of the intensity changes are extracted through a 3D convolutional network; The depth modality branch of the depth map takes 1 channel as input; it extracts the spatiotemporal features of the facial 3D structure displacement; The output features of each branch are spliced through the channel to form the initial fusion feature F init .
4. The micro-expression recognition method based on active three-dimensional imaging according to claim 3, characterized in that: The step 2 is specifically as follows: Design independent 3D convolutional networks for different modal characteristics: The first layer of the RGB modality branch is a 3D convolution layer with 64 convolution kernels, followed by a batch normalization layer and a linear rectification activation function to extract the spatiotemporal features of the apparent texture. The first 3D convolutional layer of the infrared modality branch is configured with 32 convolution kernels, followed by a batch normalization layer and a linear rectification activation function to extract the spatiotemporal features of infrared changes and facial 3D depth displacement; The first 3D convolutional layer of the deep modality branch is configured with 32 convolution kernels, followed by a batch normalization layer and a linear rectification activation function to extract the spatiotemporal features of intensity changes and facial 3D structural displacements; After two layers of 3D convolution, the output feature tensor size of each branch is unified to T×64×64×C, where T is the sequence length and C is the number of channels. The RGB branch has C=64 and the infrared / depth branch has C=32. The features of the three different forms are spliced through the channels to form the initial fusion feature.
5. The micro-expression recognition method based on active three-dimensional imaging according to claim 4, characterized in that: The step three is specifically as follows: The spatiotemporal self-attention mechanism module is used to capture the long-range temporal dependencies and dynamic change areas of the initial fusion features: First, the spatiotemporal position coding is added to the frames in each segment of the initial fusion feature, and the time coding adopts the sine function Where k is the dimension index, D is the feature dimension; t is the length of the time segment, x and y are the width and height of the video respectively; by the ratio of the dimension index k to the total dimension D As an index, it makes the position encoding of different dimensions have different frequencies, thereby distinguishing the relative position information of time and space in the spatiotemporal self-attention mechanism and capturing the long-range temporal dependency and spatial structure characteristics of micro-expressions; Then, a multi-head attention mechanism is used to perform spatiotemporal joint modeling on the feature sequence after time encoding to output the feature F containing the temporal dynamic priority. temp ; Design a cross-modal attention mechanism to achieve inter-modal information complementarity, using RGB modality feature F r For the benchmark query Q r , infrared modal characteristics F i With deep modal features F d As key-value pairs (K i ,V i )、(K d ,V d ), unify the dimensions through linear transformation; The fusion between modalities is achieved by calculating the correlation weights between modalities, and the correlation weights between modalities are calculated through a learnable parameter matrix. where β r,i Indicates the degree of dependence of the RGB modality on the infrared modality. In low-light scenes, this weight automatically increases to enhance the compensation of infrared features for apparent noise. Finally, the gating parameter is introduced to dynamically balance the cross-modal fusion and intrinsic modal features. The fused feature is F fusion =γ(β r,i V i +β r,d V d )+(1-γ)F r , and then the fused features are compressed into Preserve key spatiotemporal information.
6. The micro-expression recognition method based on active three-dimensional imaging according to claim 5, characterized in that: Multi-head self-attention projects the feature sequence into query Q, key K, value V, and calculates the attention weight by scaling the dot product Among them, Q, K, and V are generated by linear mapping of input features, d k is the dimension scaling factor; the activation function is GELU, followed by residual connection and layer normalization to prevent gradient disappearance; after being processed by the encoder, the output contains the feature F of temporal dynamic priority temp , where the areas with high attention weights correspond to micro-expression keyframes.
7. The micro-expression recognition method based on active three-dimensional imaging according to claim 6, characterized in that: In step 4, one branch passes through two fully connected layers followed by two regression layers to output standardized micro-expression start and end frames, and uses a smooth L1 loss function to calculate the positioning loss; the other branch passes through three fully connected layers, and the last layer uses a SoftMax function to output the probability distribution of seven categories of emotions. To address the problem of class imbalance, a weighted cross-entropy loss function is used to calculate the classification loss; the positioning loss and the classification loss are jointly optimized, and after the loss is obtained through training, the loss is continuously reduced through a backpropagation algorithm to improve the positioning accuracy of micro-expression fragments and the recognition accuracy of emotional categories.
8. The micro-expression recognition method based on active three-dimensional imaging according to claim 7, characterized in that: The step 4 is specifically as follows: Input the fused features into the dual-branch network to achieve joint detection and classification: The dual-branch network uses two fully connected layers (128→256→128), followed by two regression layers (128→64→2), and outputs a standardized starting frame. and end frame Range [0,1], by Mapped to actual frame number; Smooth L1 loss is used as the positioning loss to deal with the ambiguity of micro-expression boundaries: in is the frame difference, is the indicator function; The other path of the dual-branch network passes through three fully connected layers (128→512→256→8), and the last layer SoftMax outputs the probability distribution of 7 categories of emotions; To address the class imbalance problem in the micro-expression dataset where the number of samples of each emotion category varies significantly, a weighted cross entropy loss is used: Weight ω c Calculate the gradient contribution of each emotion category during training by the inverse of the number of category samples.
Citation Information
Patent Citations
Micro-expression recognition method based on multi-mode double-branch space-time motion feature fusion
CN118262397A
Micro-expression classification method based on multi-level double-branch space-time attention architecture
CN118366195A
English classroom real-time micro-expression recognition method based on double-branch continuous attention
CN118887717A
Micro-expression recognition method based on fast and slow motion double-branch network
CN119399819A
Facial expression recognition method and system combined with attention mechanism
US20230298382A1
Cited By
Physiological and emotional state recognition method for diversified people
CN122320548B