Micro-expression analysis method and system based on adaptive pseudo-label and attention mechanism
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANCHANG UNIV
- Filing Date
- 2025-04-23
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]现有技术中,微表情分析任务的核心在于精确捕捉并解读个体的真实情感,而伪标记的准确性和有效性是此过程中至关重要的一环,但逐帧进行微表情标记不仅是一项极为耗时且劳动强度大的工作,而且在实际操作中效率低下,难以满足大规模数据分析的需求,由于现有的微表情分析方法通常只聚焦于单个微表情帧的识别,而忽视了微表情在整个时间序列中的分布比例和连续性特征,导致伪标记结果往往无法全面、准确地反映微表情的真实情况,同时,微表情因为出现时间短暂且变化微小,因此难以准确定位和识别,进一步的影响了分析的准确性
附图说明
Smart Images

Figure CN120544245B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing, and in particular to a micro-expression analysis method and system based on adaptive pseudo-labeling and attention mechanisms. Background Technology
[0002] Microexpressions refer to unconscious facial expressions that last between 1 / 25 and 1 / 3 of a second in a specific context. They often contain genuine emotions that people try to conceal and are not easily faked. They are a form of non-verbal communication. However, even professionals trained in microexpression recognition can only effectively identify microexpressions with an accuracy rate of less than 47%. Therefore, highly accurate microexpression analysis methods are crucial.
[0003] In existing technologies, the core of micro-expression analysis lies in accurately capturing and interpreting an individual's true emotions. The accuracy and effectiveness of pseudo-labels are crucial in this process. However, frame-by-frame micro-expression labeling is not only extremely time-consuming and labor-intensive, but also inefficient in practice, making it difficult to meet the needs of large-scale data analysis. Because existing micro-expression analysis methods usually focus only on the recognition of a single micro-expression frame, ignoring the distribution ratio and continuity characteristics of micro-expressions throughout the entire time series, the pseudo-labeling results often fail to fully and accurately reflect the true situation of micro-expressions. At the same time, because micro-expressions appear briefly and change very little, they are difficult to locate and identify accurately, further affecting the accuracy of the analysis.
[0004] Therefore, how to design a micro-expression analysis method to improve the accuracy of micro-expression localization and recognition has become an urgent problem to be solved. Summary of the Invention
[0005] Based on this, this invention proposes a micro-expression analysis method and system based on adaptive pseudo-labeling and attention mechanisms. To more accurately and efficiently locate micro-expression intervals and identify emotion categories in videos, an adaptive pseudo-labeling algorithm and a micro-expression analysis network based on a fusion attention mechanism are designed. Addressing the problem of large pseudo-labeling errors in existing methods, the adaptive pseudo-labeling algorithm uses the proportion of micro-expressions throughout the entire video interval and combines a sliding window for adaptive labeling to reduce labeling errors and improve the accuracy of micro-expression analysis. For the problem of existing methods struggling to accurately capture subtle changes in micro-expressions, this invention designs a micro-expression analysis network based on a fusion attention mechanism to seamlessly complete micro-expression localization and recognition tasks. This includes a shared sub-network, a localization sub-network, and a recognition sub-network. The micro-expression analysis network uses a two-layer convolutional neural network to extract features and introduces a dual attention mechanism to capture channel and spatial information respectively, focusing on salient features and improving the accuracy of micro-expression localization and recognition. This invention improves the accuracy of micro-expression analysis.
[0006] This invention proposes a micro-expression analysis method based on adaptive pseudo-labeling and attention mechanisms, comprising: The target face video frame image is acquired and preprocessed. Optical flow features are extracted from the preprocessed target face video frame image, and optical flow triples are made based on the optical flow features. The optical flow feature extraction is based on the TV-L1 algorithm, and the optical flow triples include optical flow features and optical strain. The key regions in the target face video frame image are resampled, and micro-expression frames are labeled according to the adaptive pseudo-labeling algorithm to obtain pseudo-labels. The key regions are face regions containing micro-expression information. The adaptive pseudo-labeling algorithm dynamically adjusts the comparison parameters according to the length of the sliding window and the real interval. The optical flow triplet and the pseudo-label are respectively input into the micro-expression analysis network. The micro-expression analysis network includes a shared sub-network, a localization sub-network, and a recognition sub-network. The micro-expression analysis network is based on a fusion attention mechanism and has a dual three-flow neural network structure. The shared subnetwork adjusts the attention weights of the optical flow triples according to a dual attention mechanism to obtain attention-enhanced optical flow features. The dual attention mechanism is based on channel attention mechanism and spatial attention mechanism. The localization subnetwork performs feature processing on the pseudo-labels and the attention-enhanced optical flow features to obtain confidence scores. The feature processing is used to reduce the feature dimension of the attention-enhanced optical flow features and reduce the data size. The recognition subnetwork performs micro-expression emotion recognition based on the attention-enhanced optical flow features to obtain the predicted emotion probability.
[0007] In summary, based on the aforementioned micro-expression analysis method using adaptive pseudo-labeling and attention mechanisms, this invention aims to more accurately and efficiently locate micro-expression intervals and identify emotion categories in videos. By designing an adaptive pseudo-labeling algorithm and a micro-expression analysis network based on a fusion attention mechanism, the method addresses the issue of large pseudo-labeling errors in existing methods. The adaptive pseudo-labeling algorithm uses the proportion of micro-expressions across the entire video interval and combines a sliding window for adaptive labeling, thereby reducing labeling errors and improving the accuracy of micro-expression analysis. Furthermore, to address the difficulty of accurately capturing subtle changes in micro-expressions using existing methods, this invention designs a micro-expression analysis network based on a fusion attention mechanism to seamlessly complete micro-expression localization and recognition tasks. This network includes a shared sub-network, a localization sub-network, and a recognition sub-network. The micro-expression analysis network employs a two-layer convolutional neural network to extract features and introduces a dual attention mechanism to capture channel and spatial information respectively, focusing on salient features and improving the accuracy of micro-expression localization and recognition. Therefore, this invention enhances the accuracy of micro-expression analysis. Specifically, the process involves acquiring and preprocessing target face video frame images, extracting optical flow features from the preprocessed images, and creating optical flow triples based on these features. The optical flow feature extraction is based on the TV-L1 algorithm, and the optical flow triples include optical flow features and optical strain. Key regions in the target face video frame images are resampled, and micro-expression frames are labeled using an adaptive pseudo-labeling algorithm to obtain pseudo-labels. The key regions are facial regions containing micro-expression information. The adaptive pseudo-labeling algorithm dynamically adjusts comparison parameters based on the length of the sliding window and the true interval. Addressing the issue of large pseudo-labeling errors in existing methods, the adaptive pseudo-labeling algorithm uses the proportion of micro-expressions throughout the entire video interval, combined with a sliding window, to adaptively label, thereby reducing labeling errors and improving the accuracy of micro-expression analysis. The optical flow triples and pseudo-labels are then input into the micro-expression analysis network. The micro-expression analysis network includes a shared sub-network, a localization sub-network, and a recognition sub-network. Based on a fusion attention mechanism, the shared sub-network adjusts the attention weights of the optical flow triples according to a dual attention mechanism to obtain attention-enhanced optical flow features. The dual attention mechanism, based on channel attention and spatial attention, uses a two-layer convolutional neural network to extract features and introduces a dual attention mechanism to capture channel and spatial information respectively, focusing on salient features and improving the accuracy of micro-expression localization and recognition. The localization sub-network performs feature processing on the pseudo-labels and the attention-enhanced optical flow features to obtain confidence scores. This feature processing reduces the feature dimension of the attention-enhanced optical flow features and decreases the data size. The recognition sub-network performs micro-expression emotion recognition based on the attention-enhanced optical flow features to obtain predicted emotion probabilities. This invention improves the accuracy of micro-expression analysis.
[0008] Furthermore, the steps of acquiring and preprocessing the target face video frame image, extracting optical flow features from the preprocessed target face video frame image, and creating optical flow triples based on the optical flow features specifically include: The target face video frame image is acquired and the face is cropped. The face cropping uses the first frame image after face detection as the cropping standard and marks the key facial region points in the target face video frame image. Optical flow features are extracted using the V-L1 algorithm to obtain horizontal and vertical optical flow features, respectively. The specific formula for optical flow feature extraction is as follows: , , in, Indicates the characteristics of horizontal optical flow. Indicates the characteristics of vertical optical flow. and These represent the positive horizontal direction and the positive vertical direction, respectively. Indicates the number of time frames; The optical strain and optical strain amplitude values are calculated to obtain the facial deformation intensity. The specific formulas for calculating the optical strain and optical strain amplitude values are as follows: , , in, Indicates optical strain. , , , These represent optical strain in different directions. Indicates the optical strain amplitude value; Based on the horizontal optical flow characteristics, vertical optical flow characteristics, and optical strain, an optical flow triple (u, v, ϵ) is constructed and input into the micro-expression analysis network.
[0009] Furthermore, the step of resampling key regions in the target face video frame image and annotating micro-expression frames according to an adaptive pseudo-labeling algorithm to obtain pseudo-labels specifically includes: ROI selection and resampling are performed on key regions in the target face video frame image. The key regions include the left eye and left eyebrow region, the right eye and right eyebrow region, and the mouth region. Micro-expression frames are annotated based on key regions in the target face video frame image according to the adaptive pseudo-labeling algorithm to obtain pseudo-labels. The specific formula of the adaptive pseudo-labeling algorithm is as follows: , , , in, This represents half the average length of a micro-expression frame, where 𝑁 represents the average length of a micro-expression frame. This indicates the total number of micro-expression frames. Indicates the frame number of micro-expression. and These represent the end frame and the beginning frame of the micro-expression, respectively. IoU This represents the intersection-over-union ratio (IoU) between the sliding window and the actual interval, where the sliding window represents a range starting from any frame and having a length of [value missing]. The frame sequence, wherein the real interval is the frame sequence from the micro-expression start frame to the micro-expression end frame. Indicates a sliding window. Represents the true interval.
[0010] Furthermore, the shared sub-network adjusts the attention weights of the optical flow triples according to a dual-attention mechanism to obtain attention-enhanced optical flow features, specifically including: After the optical flow triples are divided, they are input into different parallel channels of the shared sub-network. Each parallel channel of the shared sub-network includes a convolutional layer, an attention enhancement layer, a max pooling layer, and a splicing layer. The shared subnetwork is connected to the positioning subnetwork and the identification subnetwork respectively, and the positioning subnetwork and the identification subnetwork are connected through a shared layer; After convolution processing of the optical flow triplet, channel attention enhancement and spatial attention enhancement are performed according to the dual attention mechanism to obtain attention enhancement features. The attention enhancement features are then subjected to convolution processing and max pooling processing again to enhance the attention enhancement features by saliency feature selection. The attention-enhanced features output from all parallel channels are concatenated to obtain attention-enhanced optical flow features.
[0011] Furthermore, the step of the localization subnetwork performing feature processing on the pseudo-labels and the attention-enhanced optical flow features to obtain a confidence score specifically includes: The localization subnetwork performs convolution and max pooling on the attention-enhanced optical flow features based on pseudo-labels to reduce feature dimensionality and data size. The pseudo-labels are used as weak supervision signals. In the convolution process, the attention-enhanced optical flow features are focused based on the micro-expression regions marked in the pseudo-labels. The target of the feature focusing is the small deformation region of facial muscles in the attention-enhanced optical flow features. Micro-expression dynamic trajectory enhancement is performed based on the feature focusing. The attention-enhanced optical flow features are then tiled to convert them into one-dimensional data. The location reliability score is predicted based on the pseudo-labels, and the specific formula for the location reliability score is as follows: , in, This indicates the location reliability score. This represents the activation function value. represents the activation function value weight, and b represents the bias; Then, the positioning loss is calculated based on the positioning confidence score. The specific formula for the positioning loss is as follows: , in, Indicates location loss. This indicates the total number of micro-expression frames. Indicates the frame number of micro-expression. This indicates location reliability.
[0012] Furthermore, the step of the recognition sub-network performing micro-expression emotion recognition based on the attention-enhanced optical flow features to obtain the predicted emotion probability specifically includes: According to the cross-task learning mechanism, the recognition subnetwork freezes the shared feature extraction layer between the localization subnetwork and the recognition subnetwork. The shared feature extraction layer is used to extract general features related to micro-expression localization in the localization subnetwork. Convolution and max pooling are performed on the attention-enhanced optical flow features to reduce feature dimension and data size. The attention-enhancing optical flow features are then tiled to calculate the predicted emotion probability. The specific formula for calculating the predicted emotion probability is as follows: , , Where p represents the probability of predicting sentiment. This represents the predicted score, where C represents the total number of categories and c represents the category ordinal number. j Represents a set of categories; The emotional category of micro-expressions is determined based on the predicted emotional probability; The recognition loss is calculated based on the predicted emotion probability, and the specific formula for calculating the recognition loss is as follows: , in, Indicates identification loss, This indicates the total number of micro-expression frames. Indicates the frame number of micro-expression.
[0013] Furthermore, the step of calculating the predicted emotion probability further includes: Post-processing is performed based on a peak detection mechanism to identify micro-expression regions by locating the maximum local peak and the segments surrounding the peak. The location confidence score of each individual video frame is smoothed, and a location threshold is calculated based on the average and maximum location confidence scores of each video frame. This location threshold is used to divide the score sequence into multiple micro-expression intervals. The specific formula for calculating the location threshold is as follows: , , in, This represents the smoothed location reliability score. This represents half the average length of micro-expression frames. Represents a video frame. j Represents a set of categories. Indicates the frame number of micro-expression. 𝑇 Indicates the location threshold. This represents the average location reliability score. This represents the maximum local reliability score. 𝑝 This indicates the preset positioning parameters.
[0014] This invention proposes a micro-expression analysis system based on adaptive pseudo-labeling and attention mechanisms, comprising: The preprocessing module is used to acquire and preprocess the target face video frame image, extract optical flow features from the preprocessed target face video frame image, and create optical flow triples based on the optical flow features. The optical flow feature extraction is based on the TV-L1 algorithm, and the optical flow triples include optical flow features and optical strain. An adaptive pseudo-labeling module is used to resample key regions in the target face video frame image and annotate micro-expression frames according to an adaptive pseudo-labeling algorithm to obtain pseudo-labels. The key regions are face regions containing micro-expression information. The adaptive pseudo-labeling algorithm dynamically adjusts the comparison parameters according to the length of the sliding window and the real interval. The micro-expression analysis network module is used to input the optical flow triplet and the pseudo-label into the micro-expression analysis network, which includes a shared sub-network, a localization sub-network and a recognition sub-network, and is based on a fusion attention mechanism. The shared subnetwork adjusts the attention weights of the optical flow triples according to a dual attention mechanism to obtain attention-enhanced optical flow features. The dual attention mechanism is based on channel attention mechanism and spatial attention mechanism. The localization subnetwork performs feature processing on the pseudo-labels and the attention-enhanced optical flow features to obtain confidence scores. The feature processing is used to reduce the feature dimension of the attention-enhanced optical flow features and reduce the data size. The recognition subnetwork performs micro-expression emotion recognition based on the attention-enhanced optical flow features to obtain the predicted emotion probability.
[0015] The present invention also provides a storage medium that stores one or more programs, which, when executed by a processor, implement the micro-expression analysis method based on adaptive pseudo-tags and attention mechanisms as described above.
[0016] The present invention also provides a computer device, the computer device including a memory and a processor, wherein: The memory is used to store computer programs; When the processor executes the computer program stored in the memory, it implements the micro-expression analysis method based on adaptive pseudo-labels and attention mechanisms as described above. Attached Figure Description
[0017] Figure 1 The flowchart shows the micro-expression analysis method based on adaptive pseudo-labeling and attention mechanism proposed in the first embodiment of the present invention. Figure 2 This is a schematic diagram of the micro-expression analysis system based on adaptive pseudo-labeling and attention mechanism proposed in the second embodiment of the present invention; Figure 3 This is a schematic diagram of the adaptive pseudo-labeling algorithm of the present invention; Figure 4 This is a schematic diagram of the micro-expression analysis network logic of the present invention; Figure 5 This is a schematic diagram of the network structure of the shared subnetwork in the micro-expression analysis network of the present invention; Figure 6 This is a diagram showing the network structure relationship between the localization subnetwork and the recognition subnetwork in the micro-expression analysis network of this invention; The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation
[0018] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.
[0019] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0021] Please see Figure 1 The diagram shows a flowchart of the micro-expression analysis method based on adaptive pseudo-labeling and attention mechanism proposed in the first embodiment of the present invention. This micro-expression analysis method based on adaptive pseudo-labeling and attention mechanism includes steps S01 to S06, wherein: Step S01: Acquire the target face video frame image and preprocess it. Extract optical flow features from the preprocessed target face video frame image and create an optical flow triplet based on the optical flow features. It should be noted that in this embodiment, the optical flow feature extraction is based on the TV-L1 algorithm, the optical flow triplet includes optical flow features and optical strain, the target face video frame image is acquired and the face is cropped, the face cropping uses the first frame image after face detection as the cropping standard, and the key facial region points in the target face video frame image are marked. Optical flow features are extracted using the V-L1 algorithm to obtain horizontal and vertical optical flow features, respectively. The specific formula for optical flow feature extraction is as follows: , , in, Indicates the characteristics of horizontal optical flow. Indicates the characteristics of vertical optical flow. and These represent the positive horizontal direction and the positive vertical direction, respectively. Indicates the number of time frames; The optical strain and optical strain amplitude values are calculated to obtain the facial deformation intensity. The specific formulas for calculating the optical strain and optical strain amplitude values are as follows: , , in, Indicates optical strain. , , , These represent optical strain in different directions. Indicates the optical strain amplitude value; Based on the horizontal optical flow characteristics, vertical optical flow characteristics, and optical strain, an optical flow triple (u, v, ϵ) is constructed and input into the micro-expression analysis network.
[0022] Step S02: Resample the key regions in the target face video frame image and annotate the micro-expression frames according to the adaptive pseudo-labeling algorithm to obtain pseudo-labels; It should be noted that in this embodiment, the key region is the facial region containing micro-expression information. The adaptive pseudo-labeling algorithm dynamically adjusts the comparison parameters according to the length of the sliding window and the real interval, and performs ROI selection and resampling processing on the key regions in the target facial video frame image. The key regions include the left eye and left eyebrow region, the right eye and right eyebrow region, and the mouth region. For detailed instructions on the adaptive pseudo-labeling algorithm, please refer to [link / reference]. Figure 3 The four-pointed star region represents the actual micro-expression frame, the five-pointed star represents the labeling result of the adaptive pseudo-labeling algorithm of this invention, and the six-pointed star represents the labeling result of the existing pseudo-labeling algorithm. It can be clearly seen from the figure that this is a comparison of the labeling effect of the adaptive pseudo-labeling method proposed in this invention and the existing automatic pseudo-labeling method. The existing automatic pseudo-labeling method exaggerates the proportion of the micro-expression region in the entire video region. In contrast, the new labeling method proposed in this paper shows a stronger ability to cope with this situation and can effectively reduce the micro-expression localization error caused by labeling deviation. Micro-expression frames are annotated based on key regions in the target face video frame image according to the adaptive pseudo-labeling algorithm to obtain pseudo-labels. The specific formula of the adaptive pseudo-labeling algorithm is as follows: , , , in, This represents half the average length of a micro-expression frame, where 𝑁 represents the average length of a micro-expression frame. This indicates the total number of micro-expression frames. Indicates the frame number of micro-expression. and These represent the end frame and the beginning frame of the micro-expression, respectively. IoU This represents the intersection-over-union ratio (IoU) between the sliding window and the actual interval, where the sliding window represents a range starting from any frame and having a length of [value missing]. The frame sequence, wherein the real interval is the frame sequence from the micro-expression start frame to the micro-expression end frame. Indicates a sliding window. Represents the true interval.
[0023] Step S03: Input the optical flow triplet and pseudo-labels into the micro-expression analysis network respectively; It should be noted that in this embodiment, the micro-expression analysis network includes a sharing sub-network, a localization sub-network, and a recognition sub-network. The micro-expression analysis network is based on a fusion attention mechanism. For the specific logical flow of the micro-expression analysis network, please refer to [link / reference]. Figure 4 .
[0024] Step S04: The shared subnetwork adjusts the attention weights of the optical flow triples according to the dual attention mechanism to obtain attention-enhanced optical flow features; It should be noted that in this embodiment, the dual attention mechanism is based on the channel attention mechanism and the spatial attention mechanism. After the optical flow triple is divided, it is input into different parallel channels of the shared sub-network. Each parallel channel of the shared sub-network includes a convolutional layer, an attention enhancement layer, a max pooling layer and a splicing layer. The shared subnetwork is connected to the positioning subnetwork and the identification subnetwork respectively, and the positioning subnetwork and the identification subnetwork are connected through a shared layer; After convolution processing of the optical flow triplet, channel attention enhancement and spatial attention enhancement are performed according to the dual attention mechanism to obtain attention enhancement features. The attention enhancement features are then subjected to convolution processing and max pooling processing again to enhance the attention enhancement features by saliency feature selection. The attention-enhanced features output from all parallel channels are concatenated to obtain attention-enhanced optical flow features; Please refer to the specific network structure of the shared subnetwork. Figure 5 The detailed structural parameters of the shared subnetwork are shown in Table 1 below: Table 1. Shared Subnetwork Structure Parameters Wherein, Conv represents a convolutional layer, each of which uses ReLU as the activation function and the weight regularization is set to 0.001; CBAM represents an attention enhancement layer, based on a dual attention mechanism; Bn represents a batch normalization layer, each pool is a max pooling layer; and Dropout represents a dropout layer, with the proportion of neurons randomly dropped in each Dropout layer set to 0.3.
[0025] Step S05: The localization subnetwork performs feature processing on the pseudo-labels and attention-enhanced optical flow features to obtain confidence scores; It should be noted that in this embodiment, the feature processing is used to reduce the feature dimension of the attention-enhanced optical flow feature and reduce the data size. The localization subnetwork performs convolution and max pooling processing on the attention-enhanced optical flow feature according to the pseudo-label to reduce the feature dimension and data size. The pseudo-label is used as a weak supervision signal. In the convolution processing, the attention-enhanced optical flow feature is focused on the micro-expression region marked in the pseudo-label. The target of the feature focusing is the micro-deformation region of facial muscles in the attention-enhanced optical flow feature. Micro-expression dynamic trajectory enhancement is performed according to the feature focusing. In this embodiment, the addition of pseudo-labels suppresses interference from static facial backgrounds or irrelevant regions. The introduction of pseudo-labels effectively alleviates the problem of scarce and costly real labels in micro-expression datasets. The weak supervision signals generated by pseudo-labels expand the diversity of training data, enabling convolutional layers to learn more generalizable spatiotemporal features during feature extraction. For example, pseudo-labels guide convolutional layers to learn the dynamic features of muscle contraction speed and amplitude during micro-expression, thereby improving the robustness of the localization network under low signal-to-noise ratio conditions. The feature map generated by the pseudo-label-guided feature extraction process has significantly improved in terms of spatiotemporal resolution and dynamic response sensitivity, providing more suitable feature inputs for subsequent modules such as keypoint trajectory prediction and motion unit activation detection in the localization network, effectively improving the accuracy and stability of micro-expression localization. The attention-enhanced optical flow features are then tiled to convert them into one-dimensional data. The location reliability score is predicted based on the pseudo-labels, and the specific formula for the location reliability score is as follows: , in, This indicates the location reliability score. This represents the activation function value. represents the activation function value weight, and b represents the bias; Then, the positioning loss is calculated based on the positioning confidence score. The specific formula for the positioning loss is as follows: , in, Indicates location loss. This indicates the total number of micro-expression frames. Indicates the frame number of micro-expression. Indicates location reliability; The detailed structural parameters of the localization subnetwork are shown in Table 2 below: Table 2. Location Subnetwork Structure Parameter Table In this context, Shared represents a shared feature extraction layer, with the proportion of randomly dropped neurons in each Dropout layer set to 0.2; Maxpooling represents a max pooling layer; Flatten represents a flattening layer; and Dense represents a fully connected layer.
[0026] Step S06: The recognition sub-network performs micro-expression emotion recognition based on attention-enhanced optical flow features to obtain the predicted emotion probability; It should be noted that, in this embodiment, the structural relationship between the positioning subnetwork and the identification subnetwork is explained in the following reference: Figure 6 According to the cross-task learning mechanism, the recognition subnetwork freezes the shared feature extraction layer between the localization subnetwork and the recognition subnetwork. The shared feature extraction layer is used to extract general features related to micro-expression localization in the localization subnetwork. Convolution and max pooling are performed on the attention-enhanced optical flow features to reduce feature dimension and data size. This embodiment employs a cross-task learning mechanism, namely the ITL strategy. By freezing the shared feature extraction layer between the localization subnetwork and the recognition subnetwork, the features learned by the localization subnetwork are directly transferred to the recognition subnetwork. Specifically, the localization subnetwork focuses on micro-expression interval prediction, and the shared feature extraction layer has extracted general features related to micro-expression localization. When training the recognition subnetwork, the parameters of these shared feature extraction layers are fixed to avoid destroying the localization features due to retraining. At the same time, only the classification head specific to the recognition subnetwork is updated, enabling it to efficiently learn target category information based on shared features. This not only achieves effective transfer of localization knowledge and recognition ability, but also improves training efficiency by reducing the amount of parameter updates and reduces the risk of overfitting. The attention-enhancing optical flow features are then tiled to calculate the predicted emotion probability. The specific formula for calculating the predicted emotion probability is as follows: , , Where p represents the probability of predicting sentiment. This represents the predicted score, where C represents the total number of categories and c represents the category ordinal number. j Represents a set of categories; The emotional category of micro-expressions is determined based on the predicted emotional probability; The recognition loss is calculated based on the predicted emotion probability, and the specific formula for calculating the recognition loss is as follows: , in, Indicates identification loss, This indicates the total number of micro-expression frames. Indicates the frame number of micro-expression; Post-processing is performed based on a peak detection mechanism to identify micro-expression regions by locating the maximum local peak and the segments surrounding the peak. The location confidence score of each individual video frame is smoothed, and a location threshold is calculated based on the average and maximum location confidence scores of each video frame. This location threshold is used to divide the score sequence into multiple micro-expression intervals. The specific formula for calculating the location threshold is as follows: , , in, This represents the smoothed location reliability score. This represents half the average length of micro-expression frames. Represents a video frame. j Represents a set of categories. Indicates the frame number of micro-expression. 𝑇 Indicates the location threshold. This represents the average location reliability score. This represents the maximum local reliability score. 𝑝 Indicates the preset positioning parameters; The detailed parameters for identifying the sub-network structure are shown in Table 3 below: Table 3. Subnetwork Structure Parameter Table In this case, the proportion of neurons randomly dropped in each Dropout is set to 0.2.
[0027] In summary, based on the aforementioned micro-expression analysis method using adaptive pseudo-labeling and attention mechanisms, this invention aims to more accurately and efficiently locate micro-expression intervals and identify emotion categories in videos. By designing an adaptive pseudo-labeling algorithm and a micro-expression analysis network based on a fusion attention mechanism, the method addresses the issue of large pseudo-labeling errors in existing methods. The adaptive pseudo-labeling algorithm uses the proportion of micro-expressions across the entire video interval and combines a sliding window for adaptive labeling, thereby reducing labeling errors and improving the accuracy of micro-expression analysis. Furthermore, to address the difficulty of accurately capturing subtle changes in micro-expressions using existing methods, this invention designs a micro-expression analysis network based on a fusion attention mechanism to seamlessly complete micro-expression localization and recognition tasks. This network includes a shared sub-network, a localization sub-network, and a recognition sub-network. The micro-expression analysis network employs a two-layer convolutional neural network to extract features and introduces a dual attention mechanism to capture channel and spatial information respectively, focusing on salient features and improving the accuracy of micro-expression localization and recognition. Therefore, this invention enhances the accuracy of micro-expression analysis. Specifically, the process involves acquiring and preprocessing target face video frame images, extracting optical flow features from the preprocessed images, and creating optical flow triples based on these features. The optical flow feature extraction is based on the TV-L1 algorithm, and the optical flow triples include optical flow features and optical strain. Key regions in the target face video frame images are resampled, and micro-expression frames are labeled using an adaptive pseudo-labeling algorithm to obtain pseudo-labels. The key regions are facial regions containing micro-expression information. The adaptive pseudo-labeling algorithm dynamically adjusts comparison parameters based on the length of the sliding window and the true interval. Addressing the issue of large pseudo-labeling errors in existing methods, the adaptive pseudo-labeling algorithm uses the proportion of micro-expressions throughout the entire video interval, combined with a sliding window, to adaptively label, thereby reducing labeling errors and improving the accuracy of micro-expression analysis. The optical flow triples and pseudo-labels are then input into the micro-expression analysis network. The micro-expression analysis network comprises a shared sub-network, a localization sub-network, and a recognition sub-network. Based on a fusion attention mechanism, the shared sub-network adjusts the attention weights of the optical flow triples according to a dual attention mechanism to obtain attention-enhanced optical flow features. The dual attention mechanism, based on channel attention and spatial attention, employs a two-layer convolutional neural network to extract features and introduces a dual attention mechanism to capture channel and spatial information respectively, focusing on salient features and improving the accuracy of micro-expression localization and recognition. The localization sub-network performs feature processing on the pseudo-labels and the attention-enhanced optical flow features to obtain confidence scores. This feature processing reduces the feature dimension of the attention-enhanced optical flow features and decreases the data size. The recognition sub-network performs micro-expression emotion recognition based on the attention-enhanced optical flow features to obtain predicted emotion probabilities. This invention improves the accuracy of micro-expression analysis.
[0028] Please see Figure 3The figure shows a schematic diagram of the micro-expression analysis system based on adaptive pseudo-labeling and attention mechanism proposed in the third embodiment of the present invention. The system includes: The preprocessing module 10 is used to acquire and preprocess the target face video frame image, extract optical flow features from the preprocessed target face video frame image, and create an optical flow triplet based on the optical flow features. The optical flow feature extraction is based on the TV-L1 algorithm, and the optical flow triplet includes optical flow features and optical strain. The adaptive pseudo-labeling module 20 is used to resample the key regions in the target face video frame image and annotate the micro-expression frames according to the adaptive pseudo-labeling algorithm to obtain pseudo-labels. The key regions are face regions containing micro-expression information. The adaptive pseudo-labeling algorithm dynamically adjusts the comparison parameters according to the length of the sliding window and the real interval. The micro-expression analysis network module 30 is used to input the optical flow triplet and the pseudo-label into the micro-expression analysis network, which includes a shared sub-network, a localization sub-network and a recognition sub-network, and is based on a fusion attention mechanism. The shared subnetwork adjusts the attention weights of the optical flow triples according to a dual attention mechanism to obtain attention-enhanced optical flow features. The dual attention mechanism is based on channel attention mechanism and spatial attention mechanism. The localization subnetwork performs feature processing on the pseudo-labels and the attention-enhanced optical flow features to obtain confidence scores. The feature processing is used to reduce the feature dimension of the attention-enhanced optical flow features and reduce the data size. The recognition subnetwork performs micro-expression emotion recognition based on the attention-enhanced optical flow features to obtain the predicted emotion probability.
[0029] The present invention also proposes a computer storage medium storing one or more programs that, when executed by a processor, implement the aforementioned micro-expression analysis method based on adaptive pseudo-labeling and attention mechanisms.
[0030] The present invention also proposes a computer device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the above-mentioned micro-expression analysis method based on adaptive pseudo-labeling and attention mechanism.
[0031] Those skilled in the art will understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain stored, communicated, propagated, or transmitted programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0032] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0033] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0034] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0035] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A micro-expression analysis method based on adaptive pseudo-labeling and attention mechanism, characterized in that, include: The target face video frame image is acquired and preprocessed. Optical flow features are extracted from the preprocessed target face video frame image, and optical flow triples are made based on the optical flow features. The optical flow feature extraction is based on the TV-L1 algorithm, and the optical flow triples include optical flow features and optical strain. The key regions in the target face video frame image are resampled, and micro-expression frames are labeled according to the adaptive pseudo-labeling algorithm to obtain pseudo-labels. The key regions are face regions containing micro-expression information. The adaptive pseudo-labeling algorithm dynamically adjusts the comparison parameters according to the length of the sliding window and the real interval. The steps of resampling key regions in the target face video frame image and annotating micro-expression frames according to an adaptive pseudo-labeling algorithm to obtain pseudo-labels specifically include: ROI selection and resampling are performed on key regions in the target face video frame image. The key regions include the left eye and left eyebrow region, the right eye and right eyebrow region, and the mouth region. Micro-expression frames are annotated based on key regions in the target face video frame image according to the adaptive pseudo-labeling algorithm to obtain pseudo-labels. The specific formula of the adaptive pseudo-labeling algorithm is as follows: , , , in, This represents half the average length of micro-expression frames. This represents the average length of micro-expression frames. This indicates the total number of micro-expression frames. Indicates the micro-expression frame sequence number. and These represent the end frame and the beginning frame of the micro-expression, respectively. IoU This represents the intersection-over-union ratio (IoU) between the sliding window and the actual interval, where the sliding window represents a region starting from any frame and having a length of [missing information]. The frame sequence, wherein the real interval is the frame sequence from the micro-expression start frame to the micro-expression end frame. Indicates a sliding window. Represents the true interval; The optical flow triplet and the pseudo-label are respectively input into the micro-expression analysis network. The micro-expression analysis network includes a shared sub-network, a localization sub-network, and a recognition sub-network. The micro-expression analysis network is based on a fusion attention mechanism and has a dual three-flow neural network structure. The shared subnetwork adjusts the attention weights of the optical flow triples according to a dual attention mechanism to obtain attention-enhanced optical flow features. The dual attention mechanism is based on channel attention mechanism and spatial attention mechanism. The localization subnetwork performs feature processing on the pseudo-labels and the attention-enhanced optical flow features to obtain confidence scores. The feature processing is used to reduce the feature dimension of the attention-enhanced optical flow features and reduce the data size. The recognition subnetwork performs micro-expression emotion recognition based on the attention-enhanced optical flow features to obtain the predicted emotion probability.
2. The micro-expression analysis method based on adaptive pseudo-labeling and attention mechanism according to claim 1, characterized in that, The steps of acquiring and preprocessing the target face video frame image, extracting optical flow features from the preprocessed target face video frame image, and creating optical flow triples based on the optical flow features specifically include: The target face video frame image is acquired and the face is cropped. The face cropping uses the first frame image after face detection as the cropping standard and marks the key facial region points in the target face video frame image. Optical flow features are extracted using the TV-L1 algorithm to obtain horizontal and vertical optical flow features, respectively. The specific formula for optical flow feature extraction is as follows: , , in, Indicates the characteristics of horizontal optical flow. Indicates the characteristics of vertical optical flow. and These represent the positive horizontal direction and the positive vertical direction, respectively. Indicates the number of time frames; The optical strain and optical strain amplitude values are calculated to obtain the facial deformation intensity. The specific formulas for calculating the optical strain and optical strain amplitude values are as follows: , , in, Indicates optical strain. , , , These represent optical strain in different directions. Indicates the optical strain amplitude value; Based on the horizontal optical flow characteristics, vertical optical flow characteristics, and optical strain, an optical flow triple (u, v, ...) is constructed. The optical flow triplet is then input into the micro-expression analysis network.
3. The micro-expression analysis method based on adaptive pseudo-labeling and attention mechanism according to claim 1, characterized in that, The shared subnetwork adjusts the attention weights of the optical flow triples according to a dual-attention mechanism to obtain attention-enhanced optical flow features, specifically including: After the optical flow triples are divided, they are input into different parallel channels of the shared sub-network. Each parallel channel of the shared sub-network includes a convolutional layer, an attention enhancement layer, a max pooling layer, and a splicing layer. The shared subnetwork is connected to the positioning subnetwork and the identification subnetwork respectively, and the positioning subnetwork and the identification subnetwork are connected through a shared layer; After convolution processing of the optical flow triplet, channel attention enhancement and spatial attention enhancement are performed according to the dual attention mechanism to obtain attention enhancement features. The attention enhancement features are then subjected to convolution processing and max pooling processing again to enhance the attention enhancement features by saliency feature selection. The attention-enhanced features output from all parallel channels are concatenated to obtain attention-enhanced optical flow features.
4. The micro-expression analysis method based on adaptive pseudo-labeling and attention mechanism according to claim 1, characterized in that, The step of the localization subnetwork performing feature processing on the pseudo-labels and the attention-enhanced optical flow features to obtain a confidence score specifically includes: The localization subnetwork performs convolution and max pooling on the attention-enhanced optical flow features based on pseudo-labels to reduce feature dimensionality and data size. The pseudo-labels are used as weak supervision signals. In the convolution process, the attention-enhanced optical flow features are focused based on the micro-expression regions marked in the pseudo-labels. The target of the feature focusing is the small deformation region of facial muscles in the attention-enhanced optical flow features. Micro-expression dynamic trajectory enhancement is performed based on the feature focusing. The attention-enhanced optical flow features are then tiled to convert them into one-dimensional data. The location reliability score is predicted based on the pseudo-labels, and the specific formula for the location reliability score is as follows: , in, This indicates the location reliability score. This represents the activation function value. represents the activation function value weight, and b represents the bias; Then, the positioning loss is calculated based on the positioning confidence score. The specific formula for the positioning loss is as follows: , in, Indicates location loss. This indicates the total number of micro-expression frames. Indicates the micro-expression frame sequence number. This indicates location reliability.
5. The micro-expression analysis method based on adaptive pseudo-labeling and attention mechanism according to claim 1, characterized in that, The step of the recognition subnetwork performing micro-expression emotion recognition based on the attention-enhanced optical flow features to obtain the predicted emotion probability specifically includes: According to the cross-task learning mechanism, the recognition subnetwork freezes the shared feature extraction layer between the localization subnetwork and the recognition subnetwork. The shared feature extraction layer is used to extract general features related to micro-expression localization in the localization subnetwork. Convolution and max pooling are performed on the attention-enhanced optical flow features to reduce feature dimension and data size. The attention-enhancing optical flow features are then tiled to calculate the predicted emotion probability. The specific formula for calculating the predicted emotion probability is as follows: , , Where p represents the probability of predicting sentiment. This represents the predicted score, where C represents the total number of categories and c represents the category ordinal number. j Represents a set of categories; The emotional category of micro-expressions is determined based on the predicted emotional probability; The recognition loss is calculated based on the predicted emotion probability, and the specific formula for calculating the recognition loss is as follows: , in, Indicates identification loss, This indicates the total number of micro-expression frames. Indicates the frame number of micro-expression.
6. The micro-expression analysis method based on adaptive pseudo-labeling and attention mechanism according to claim 5, characterized in that, The step of calculating the probability of predicted emotion further includes: Post-processing is performed based on a peak detection mechanism to identify micro-expression regions by locating the maximum local peak and the segments surrounding the peak. The location confidence score of each individual video frame is smoothed, and a location threshold is calculated based on the average and maximum location confidence scores of each video frame. This location threshold is used to divide the score sequence into multiple micro-expression intervals. The specific formula for calculating the location threshold is as follows: , , in, This represents the smoothed location reliability score. This represents half the average length of micro-expression frames. Represents a video frame. j Represents a set of categories. Indicates the micro-expression frame sequence number. Indicates the location threshold. This represents the average location reliability score. This represents the maximum local reliability score. This indicates the preset positioning parameters.
7. A micro-expression analysis system based on adaptive pseudo-labeling and attention mechanism, characterized in that, include: The preprocessing module is used to acquire and preprocess the target face video frame image, extract optical flow features from the preprocessed target face video frame image, and create optical flow triples based on the optical flow features. The optical flow feature extraction is based on the TV-L1 algorithm, and the optical flow triples include optical flow features and optical strain. An adaptive pseudo-labeling module is used to resample key regions in the target face video frame image and annotate micro-expression frames according to an adaptive pseudo-labeling algorithm to obtain pseudo-labels. The key regions are face regions containing micro-expression information. The adaptive pseudo-labeling algorithm dynamically adjusts the comparison parameters according to the length of the sliding window and the real interval. The steps of resampling key regions in the target face video frame image and annotating micro-expression frames according to an adaptive pseudo-labeling algorithm to obtain pseudo-labels specifically include: ROI selection and resampling are performed on key regions in the target face video frame image. The key regions include the left eye and left eyebrow region, the right eye and right eyebrow region, and the mouth region. Micro-expression frames are annotated based on key regions in the target face video frame image according to the adaptive pseudo-labeling algorithm to obtain pseudo-labels. The specific formula of the adaptive pseudo-labeling algorithm is as follows: , , , in, This represents half the average length of micro-expression frames. This represents the average length of micro-expression frames. This indicates the total number of micro-expression frames. Indicates the micro-expression frame sequence number. and These represent the end frame and the beginning frame of the micro-expression, respectively. IoU This represents the intersection-over-union ratio (IoU) between the sliding window and the actual interval, where the sliding window represents a region starting from any frame and having a length of [missing information]. The frame sequence, wherein the real interval is the frame sequence from the micro-expression start frame to the micro-expression end frame. Indicates a sliding window. Represents the true interval; The micro-expression analysis network module is used to input the optical flow triplet and the pseudo-label into the micro-expression analysis network respectively. The micro-expression analysis network includes a shared sub-network, a localization sub-network, and a recognition sub-network. The micro-expression analysis network is based on a fusion attention mechanism and has a dual three-flow neural network structure. The shared subnetwork adjusts the attention weights of the optical flow triples according to a dual attention mechanism to obtain attention-enhanced optical flow features. The dual attention mechanism is based on channel attention mechanism and spatial attention mechanism. The localization subnetwork performs feature processing on the pseudo-labels and the attention-enhanced optical flow features to obtain confidence scores. The feature processing is used to reduce the feature dimension of the attention-enhanced optical flow features and reduce the data size. The recognition subnetwork performs micro-expression emotion recognition based on the attention-enhanced optical flow features to obtain the predicted emotion probability.
8. A storage medium, characterized in that, The storage medium stores one or more programs that, when executed by a processor, implement the micro-expression analysis method based on adaptive pseudo-labeling and attention mechanism as described in any one of claims 1-6.
9. A computer device, characterized in that, The computer device includes a memory and a processor, wherein: The memory is used to store computer programs; When the processor executes the computer program stored in the memory, it implements the micro-expression analysis method based on adaptive pseudo-labeling and attention mechanism as described in any one of claims 1-6.