Joint expression positioning and recognition method and device based on aggregation features and double anchoring
By building an end-to-end deep learning framework, combining local-global feature aggregation and dual anchoring mechanisms, the limitations of expression recognition in the existing technology are solved, and high-precision positioning and recognition of macro/micro-expressions are realized to adapt to the complex situation of multi-expression mixed scenarios.
Patent Information
- Application Number
- CN202510352352.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The prior art has problems in the recognition of facial expressions, insufficient local and global features coordination, limitations in multi-scale timing modeling, and lack of joint optimization frameworks, resulting in insufficient accuracy and stability of expression positioning and recognition.
Using a joint expression positioning and recognition method based on aggregation features and dual anchoring, we realize unified positioning and recognition of macro/micro-expressions by building an end-to-end deep learning framework, combining local-global feature aggregation, dual-stream 3D convolutional network and feature pyramid network. The method includes local-global feature aggregation, multi-scale spatiotemporal feature fusion, dual anchoring mechanism, and joint optimization of expression positioning and recognition networks.
It improves the detection sensitivity and recognition accuracy of macro/micro-expression, adapts to complex scenes that appear in the mixture of multiple expressions, has high robustness and real-timeness, and can efficiently handle multi-expression overlap in real scenes.
Smart Images

Figure CN120299073A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a joint expression localization and recognition method and device based on aggregated features and dual anchoring, belonging to the field of computer vision technology. Background Art
[0002] In the fields of interpersonal interaction analysis, mental state assessment, and intelligent security, etc., the accurate localization and recognition of facial expressions are the core technologies for understanding human emotional intentions. With the popularization of video acquisition devices and the breakthrough of deep learning technology, facial expression analysis technology based on computer vision has gradually moved towards practical applications. Especially in scenarios such as social robots, mental health assessment, and remote conference emotion analysis, the system needs to have the ability to accurately capture and understand expressions in natural interaction scenarios. Traditional psychological research divides facial expressions into macro-expressions with a relatively long duration (0.5 - 4 seconds) and significant movement amplitude, and micro-expressions that are short (<0.5 seconds) and of low intensity due to emotion suppression or concealment. There are significant differences in the activation patterns of action units (AUs) and temporal dynamic characteristics between the two types of expressions, but both can reflect an individual's true emotional state, which makes the joint detection and recognition of macro / micro expressions a frontier topic in the field of affective computing.
[0003] Early expression recognition methods were mainly based on manually designed features and shallow classifiers (such as SVM), and classified expressions by extracting local texture or geometric features of static images. However, such methods are sensitive to illumination changes and head poses, and it is difficult to model the temporal dynamic characteristics of expressions. With the development of deep learning, methods based on 2D / 3D convolutional neural networks (CNNs) have achieved joint optimization of feature representation and classification through end-to-end learning. Among them, 3D CNNs directly model the short-term motion patterns of video sequences through spatio-temporal convolutional kernels, while architectures based on Transformers use self-attention mechanisms to capture long-range dependencies. To further improve the fine-grained expression analysis ability, researchers introduced a multi-task learning framework, combined AU (facial action unit) detection with expression classification, and used the local muscle movement information of AUs to assist in expression recognition.
[0004] In terms of expression localization technology, the mainstream methods can be divided into two categories: candidate interval generation based on sliding windows and temporal detection based on anchors. The former traverses the video sequence through multi-scale windows, but has a high computational complexity and is difficult to process expression segments with significantly different lengths; the latter draws on the anchor box mechanism in object detection, predefines anchor intervals with different durations for regression optimization, but has insufficient adaptability to extremely short or extremely long expressions. In addition, existing methods mostly process expression localization and recognition as independent modules in stages, resulting in problems of temporal information fragmentation and error accumulation.
[0005] Although significant progress has been made in existing methods, in practical applications, there are still problems such as insufficient coordination between local and global features, limitations in multi-scale temporal modeling, and the lack of a joint optimization framework. Feature aggregation technology and multi-task joint learning have become important breakthroughs to solve the above bottlenecks. The Feature Pyramid Network (FPN) enhances the model's detection ability for objects of different sizes through multi-scale feature fusion, while the combination of the dynamic anchor mechanism and the anchor-free method can improve the flexibility of temporal detection. At the same time, models based on the two-stream architecture strengthen the spatio-temporal modeling ability through heterogeneous feature complementarity. In the field of micro-expression analysis, the use of AU detection as an intermediate supervision signal has begun to be explored to enhance the interpretability of expression recognition through fine-grained muscle movement encoding. Summary of the Invention
[0006] Object of the Invention: In order to overcome the deficiencies in the prior art, the present invention provides a joint expression localization and recognition method and device based on aggregated features and dual anchoring, which can collaboratively process macro / micro-expression localization in multi-expression overlapping scenarios through a dual anchoring mechanism, combine AU encoding and multi-scale spatio-temporal aggregation to improve the detection sensitivity of expressions with different durations, and ensure high-precision recognition and localization stability while achieving end-to-end joint optimization.
[0007] Technical Solution: To achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] A joint expression localization and recognition method based on aggregated features and dual anchoring constructs an end-to-end deep learning framework for the fine-grained expression analysis of a single face in a first-person perspective video to achieve unified localization and recognition of macro-expressions and micro-expressions; the method first passes the obtained single-face video data through a frame-by-frame AU recognition network that aggregates local and global features, combines the advantages of convolutional neural networks and vision transformers, extracts multi-scale features and generates AU encoding; then uses the two-stream 3D convolutional network assisted by AU encoding to fuse spatio-temporal features, combines the multi-scale sliding window strategy and the feature pyramid network, and proposes a dual anchoring mechanism to jointly locate the expression interval; finally, intercepts the video segment based on the localization result and completes the classification through the expression localization and recognition network. The method includes the following steps:
[0009] S01: Perform face tracking on the input first-person perspective video, extract the trajectory of each face in the first-person perspective video, and form each single-face video;
[0010] S02: Perform local-global feature aggregation on each single-face image frame in the single-face video to generate AU encoding. The local-global feature aggregation is achieved through a cascaded local feature aggregator and global feature aggregator. Among them, the local feature aggregator uses a multi-head convolution module and a multi-layer perceptron module to extract local high-frequency action features, and the global feature aggregator uses a multi-head self-attention module, a multi-head convolution module and a multi-layer perceptron module to extract global low-frequency face features, and suppresses the gradient dispersion of the local feature aggregator and the global feature aggregator through residual connection.
[0011] S03: Reshape and concatenate the time series of AU encoding, convert it into a single-channel temporal feature matrix I, extract features from the single-face video to obtain a single-channel temporal feature matrix II, and input the single-channel temporal feature matrix I and the single-channel temporal feature matrix II into a two-stream 3D convolutional network for multi-scale spatio-temporal feature fusion to generate fused expression features F (e) ;
[0012] S04: Based on the sliding window strategy, perform multi-scale sampling on the fused expression features F (e) Perform expression localization through a feature pyramid network combined with a dual localization mechanism with and without anchors, and screen the final expression localization through non-maximum suppression.
[0013] S05: Intercept the expression feature segment according to the start frame and end frame of the final expression localization, and input the expression feature segment into the expression localization and recognition network to predict the macro-expression category and micro-expression category.
[0014] S06: Train the first-person perspective fine-grained expression localization and recognition unified model composed of steps S01 - S05 by jointly optimizing the AU recognition loss, the dual-anchor localization loss and the expression recognition loss.
[0015] S07: Input the first-person perspective video with any given number of frames into the trained first-person perspective fine-grained expression localization and recognition unified model to recognize the macro-expression category and micro-expression category.
[0016] Specifically, in step S02, first input the single-face video into a 2D convolutional layer for pre-encoding to obtain a pre-encoded feature map, and then pass through N (lg) cascaded local-global feature aggregation modules to simultaneously capture the local high-frequency action features and global low-frequency face features of the single-face image, and then obtain the AU encoding F t(u) through a batch normalization layer and a global average pooling layer, and finally obtain the AU recognition probability through a fully connected layer The pre-coded feature map serves as the input to the first local-global feature aggregation module. The output of the previous local-global feature aggregation module serves as the input to the subsequent local-global feature aggregation module, and the output of the last local-global feature aggregation module.
[0017] Each local-global feature aggregation module is implemented using a sub-block embedding layer, N (l) cascaded local feature aggregators, and N (g) global feature aggregators. The sub-block embedding layer is used to embed and perform preliminary feature transformation on the input of the local-global feature aggregation module, and map it to a feature space suitable for subsequent processing. The local feature aggregator is composed of a multi-head convolution module and a multi-layer perceptron module, and is used to extract local high-frequency action features (reflecting subtle changes in the face). The global feature aggregator is composed of a 2D pointwise convolution layer, a batch normalization layer, a multi-head self-attention module, a multi-head convolution module, and a multi-layer perceptron module, and is used to extract global low-frequency face features (capturing the overall structure and pose changes of the face). The outputs of the local feature aggregator and the global feature aggregator both use residual connections to suppress the vanishing gradient problem.
[0018] The sub-block embedding layer embeds and performs preliminary feature transformation on the input of the local-global feature aggregation module, maps it to a feature space suitable for subsequent processing, and uses the resulting feature map as the input to the first local feature aggregator.
[0019] The implementation process of the local feature aggregator is as follows: In the multi-head convolution module, first divide the input feature map along the channel dimension into multiple sub-feature maps, and perform 2D convolution operations on each sub-feature map respectively. After concatenating the convolution results of each sub-feature map along the channel dimension, perform batch normalization, non-linear activation, and 2D pointwise convolution in sequence to further enhance the feature representation ability, and obtain the output of the multi-head convolution module. In the multi-layer perceptron module, the output of the multi-head convolution module is connected to the input as the input of the multi-layer perceptron module, and 2D convolution, non-linear activation, and dropout are performed in sequence to obtain the output of the multi-layer perceptron module, and the expression ability of local features is enhanced through the multi-layer perceptron module. The feature map obtained by connecting the input and output of the multi-layer perceptron module with a residual connection is used as the output of the local feature aggregator, and the output of the last local feature aggregator is used as the input of the first global feature aggregator.
[0020] The implementation process of the global feature aggregator is as follows: First, the number of input channels is adjusted through a 2D pointwise convolutional layer, and then the batch normalization layer and the multi-head self-attention module are sequentially passed through to capture the global low-frequency face features. The output of the multi-head self-attention module is connected to the input of the batch normalization layer through a residual connection as the input of the multi-head convolutional module. The output of the multi-head convolutional module is connected to the input through a residual connection as the input of the multi-layer perceptron module. The feature map obtained by connecting the input and output of the multi-layer perceptron module through a residual connection is used as the output of the global feature aggregator to accurately represent the global shape and dynamic features of the face; the output of the last global feature aggregator passes through the batch normalization layer and the global average pooling layer to obtain the AU encoding F t(u) 。
[0021] Specifically, in step S02, in order to improve the accuracy of AU recognition, the AU recognition probability uses the weighted binary cross-entropy loss L (u) for supervision:
[0022]
[0023] where: represents the AU recognition probability, M (u) represents the total number of AUs, N represents the total number of sample images in the training set, represents the total number of sample images in the training set where the j-th AU appears, represents the appearance probability of the j-th AU in the training set, w j represents the weight for suppressing the imbalance in the appearance frequencies between the j-th AU and other AUs, v j represents the weight for suppressing the too-low appearance frequency of the j-th AU (used for the first term of the cross-entropy, i.e., the appearance term), and respectively represent the true appearance probability and the predicted appearance probability of the j-th AU in the t-th frame of the sample image.
[0024] The weight calculation is to address the problem of the imbalance in the appearance frequencies of AUs in AU recognition. The form of the weight is:
[0025]
[0026] where: represents the total number of sample images in the training set where the j-th AU appears; the weight w j is beneficial for suppressing the problem of the imbalance in the appearance frequencies between different AUs. Low-frequency AUs will get higher weights, making the model more sensitive to rare AUs and avoiding high-frequency AUs dominating the model training; the weight v j further adjusts the appearance frequency of each AU and is beneficial for suppressing the problem that the appearance frequency of each AU is often lower than the non-appearance frequency.
[0027] Specifically, in step S03, each AU code is initially in the form of a vector. To make it compatible with the temporal data, the AU code of each frame is reshaped from a vector into a single-channel matrix form, and then these matrices are concatenated to form a single-channel temporal feature matrix I. The single-channel temporal feature matrix I represents the change of AU features in each frame of the single-face video over time; feature extraction is performed on the single-face video to obtain a single-channel temporal feature matrix II, and the single-channel temporal feature matrix II reflects the dynamic changes of the single-face video; the two-stream 3D convolutional network includes two inputs. The first input is the single-channel temporal feature matrix I, and the second input is the single-channel temporal feature matrix II. The first input extracts spatio-temporal feature I through an independent spatio-temporal feature extraction module, and the second input extracts spatio-temporal feature II through an independent spatio-temporal feature extraction module. After adding spatio-temporal feature I and spatio-temporal feature II, the fused expression feature F is obtained. (e) 。
[0028] The spatio-temporal feature extraction module extracts spatio-temporal features with different receptive fields through multiple multi-scale 3D convolutional modules, aiming to capture spatial and temporal information at different scales in the video. The spatio-temporal feature extraction module consists of cascaded 3D convolutional layers, max pooling layers, a group of multi-scale 3D convolutional modules, and a global average pooling layer; the multi-scale 3D convolutional module contains a group of branches arranged in parallel, and each branch uses 3D convolutional kernels of different sizes (including 3×3×3 and 5×5×5 convolutional kernels) to capture information with different receptive fields respectively (3D convolutional kernels of different sizes can extract features from different receptive fields, thus effectively capturing spatio-temporal changes in local details and global structures). The outputs of each branch are concatenated at the channel level as the output of the multi-scale 3D convolutional module. The output of the last multi-scale 3D convolutional module is dimensionally reduced through the global average pooling layer to reduce redundant information, and at the same time, spatio-temporal features with different receptive fields are extracted. Through this multi-scale feature extraction method, it is possible to simultaneously pay attention to local action features and global structure features in the single-face video, thereby enhancing the sensitivity to facial expression changes.
[0029] Specifically, in step S04, first, multi-scale sampling is performed on the fused expression feature F (e) using a sliding window strategy. Sliding windows of different lengths are set to cover the typical durations of macro-expressions and micro-expressions. There is partial overlap between two adjacent sliding windows in the time dimension. T window frames are evenly sampled from the single-face video segments within different sliding windows to obtain different window frame sequences.
[0030] Along the fused expression feature F (e)After applying a sliding window to the time dimension, first use a 1D convolutional layer and a max pooling layer to encode the window frame sequence, and then input it into a feature pyramid network with L levels. In each level of the feature pyramid network, a dual localization mechanism is executed in parallel, and a 1D convolutional layer is used to generate both anchor candidate intervals and anchor-free candidate intervals simultaneously;
[0031] The expression localization of the anchor candidate intervals is achieved through intersection over union (IoU) threshold screening and dynamic regression parameter adjustment: For an anchor candidate interval to be predicted, set the predefined anchor box as (c d , w d ), the predicted classification score is The IoU between the anchor candidate interval to be predicted and the closest true expression interval is p o , denote the dynamic regression parameters of the anchor candidate interval to be predicted as {Δ c , Δ w}, then the anchor candidate interval to be predicted represented using dynamic regression parameters is c = c d + αΔ c w d , w = w d exp(βΔ ω ); During training, only the accurately predicted anchor candidate intervals are used to learn the IoU p o and the dynamic regression parameters {Δ c , Δ ω}; For the N p(ab) generated anchor candidate intervals, first calculate the IoU, then screen out M p(ab) positive candidate intervals (samples matching the true expression intervals) according to the IoU threshold, and finally calculate the IoU loss based on the positive candidate intervals At the same time, calculate the classification loss based on both positive and negative candidate intervals and the regression loss based only on positive candidate intervals at each level of the feature pyramid network Regression loss based only on positive candidate intervals
[0032] where: c d and w d represent the center point and width of the predefined anchor box respectively; c and w represent the predicted center point and predicted width of the anchor candidate interval respectively; Δ c represents the offset of c relative to c d , Δ w represents the offset of w relative to w d ; represents the score that the anchor candidate interval is located as expression n, M (e) is the total number of expression categories, and the predicted confidence is α and β are scale parameters;
[0033] The expression localization of the anchor-free candidate intervals is achieved through foreground classification and dynamic regression parameters: In the feature pyramid network, the input of the bottommost level is a sequence of window frames consisting of T window frames. The output of the next level is used as the input of the upper level. Each level downsamples the input of this level (in the feature pyramid network, each time moving up one level, the time dimension usually shrinks by half), and maps the window frame positions in the window frame sequence to different positions in the feature map: Let j ∈ {0, 1, …, T} represent the index of a certain window frame in the window frame sequence, and the corresponding mapping position of the index in the i-th level is denoted as j i , based on the mapping of the feature map, if j i The corresponding window frame j falls into a certain expression interval, then j i is called the foreground, otherwise j i is called the background (which means the window frame does not belong to any expression interval); For an anchor-free candidate interval to be predicted, calculate the predicted classification score as The dynamic regression parameters are {Δ b , Δ e}; During the training process, calculate the classification errors of all N p(af) anchored candidate intervals to obtain the classification loss Optimize the offset for the M p(af) anchor-free candidate intervals predicted as the foreground to obtain the regression loss
[0034] where: represents the score of the anchor-free candidate interval being located as expression m, and the predicted confidence is max(s (af) ); Δ b = j - b, which is the start frame offset, representing the predicted distance from window frame j to the start frame b of the expression; Δ e = e - j, which is the end frame offset, representing the predicted distance from window frame, to the end frame b of the expression;
[0035] In each level of the feature pyramid network, each window frame has the expression localization results with and without anchors. Aggregate all the expression localization results of each level of the feature pyramid network and perform non-maximum suppression operation to obtain the final expression localization.
[0036] Specifically, in step S05, according to the final expression localization, intercept the expression feature segment along the time dimension of the fused expression feature F (e) , and input the intercepted expression feature segment into an expression localization and recognition network composed of a 2D convolutional layer, a temporal pooling layer, and a fully connected classifier to predict the macro-expression category and micro-expression category, and use the multi-class cross-entropy loss L ce to calculate the expression recognition loss.
[0037] Specifically, in the step S06, in the loss function of the unified model for fine-grained expression localization and recognition from the first perspective, the AU recognition loss is the weighted binary cross-entropy loss L (u) , and the dual-anchor localization loss includes a classification loss intersection over union loss regression loss classification loss and distance regression loss The expression recognition loss is the multi-class cross-entropy loss L ce .
[0038] An apparatus for implementing any one of the above-mentioned joint expression localization and recognition methods based on aggregated features and dual anchors includes a single-face video extraction unit, a local-global feature aggregation unit, a spatio-temporal feature fusion unit, a dual-anchor localization unit, an expression recognition unit, and a joint training and optimization unit;
[0039] The single-face video extraction unit performs face tracking on the input first-person perspective video, extracts the trajectory of each face in the first-person perspective video, and forms single-face videos that are temporally coherent;
[0040] The local-global feature aggregation unit first inputs the single-face video into a 2D convolutional layer for pre-encoding to obtain a pre-encoded feature map, and then captures the local high-frequency action features and global low-frequency face features of the single-face image through N (lg) cascaded local-global feature aggregation modules. Then, through a batch normalization layer and a global average pooling layer, the AU encoding F t(u) is obtained. Finally, through a fully connected layer, the AU recognition probability is obtained
[0041] The spatio-temporal feature fusion unit first extracts features from the time series of the single-face video and the AU encoding F t(u) respectively to obtain two single-channel temporal feature matrices. Then, through a two-stream 3D convolutional network, multi-scale spatio-temporal feature fusion is performed on the two single-channel temporal feature matrices to generate the fused expression feature F (e) ;
[0042] The dual-anchor localization unit, based on the multi-layer temporal features extracted by the feature pyramid network containing L levels, combines the sliding window strategy to generate predefined anchor boxes and dynamic regression parameters, selects the expression localization interval corresponding to the maximum confidence value through non-maximum suppression, and simultaneously realizes the classification and regression with and without anchors;
[0043] The expression recognition unit, according to the final expression localization result of the dual-anchor localization unit, along the fused expression feature F (e)Intercept the expression feature segments in the time dimension, compress the feature dimensions through the temporal pooling layer, and use a fully connected classifier to distinguish between macro-expression categories and micro-expression categories to achieve unified expression recognition;
[0044] The joint training and optimization unit synchronously optimizes the expression localization accuracy and expression recognition performance through a multi-task loss function, and introduces weight parameters to balance the gradients of the two types of losses; specifically, calculates the parameters and loss values of the first-perspective fine-grained expression localization and recognition unified model composed of the local-global feature aggregation unit, spatio-temporal feature fusion unit, dual-anchor localization unit, and expression recognition unit, and updates the parameters based on the gradient optimization method.
[0045] The present invention trains the first-perspective fine-grained expression localization and recognition unified model through an end-to-end method, including a frame-by-frame AU recognition network based on the local-global feature aggregation unit, a spatio-temporal feature extraction network based on the spatio-temporal feature fusion unit, and an expression localization and recognition network based on the dual-anchor localization unit and the expression recognition unit; first, trains the frame-by-frame AU recognition network to extract the AU encoding F t(u) as the input of the expression localization and recognition network; then, the expression localization and recognition network further learns the specific patterns and spatio-temporal correlation features of the AUs, and uses the spatio-temporal correlation features between the AUs to promote the localization and recognition of facial action units.
[0046] Specifically, the local-global feature aggregation module is implemented using a sub-block embedding layer, N (l) cascaded local feature aggregators, and N (g) global feature aggregators. The sub-block embedding layer is used to embed and perform preliminary feature transformation on the input, and map it to a feature space suitable for subsequent processing. The local feature aggregator (extracting facial local textures based on multi-head convolution) uses a multi-head convolution module and a multi-layer perceptron module to extract local high-frequency action features. The global feature aggregator (modeling global muscle correlations based on multi-head self-attention) uses a multi-head self-attention module, a multi-head convolution module, and a multi-layer perceptron module to extract global low-frequency face features; the gradient dispersion of the local feature aggregator and the global feature aggregator is suppressed through residual connections.
[0047] Specifically, for the two-stream 3D convolutional network, independent spatio-temporal feature extraction modules are respectively used to extract spatio-temporal features from the two single-channel temporal feature matrices, and the two extracted spatio-temporal features are added together to obtain the fused expression feature F (e) ;
[0048] The spatio-temporal feature extraction module is composed of cascaded 3D convolutional layers, max pooling layers, a set of multi-scale 3D convolutional modules, and a global average pooling layer; the multi-scale 3D convolutional module contains a set of branches arranged in parallel, each branch uses 3D convolutional kernels of different sizes to capture information of different receptive fields respectively, and the outputs of each branch are concatenated at the channel level as the output of the multi-scale 3D convolutional module, and the output of the last multi-scale 3D convolutional module is dimension-reduced by the global average pooling layer to generate corresponding spatio-temporal features.
[0049] Advantages: The joint expression localization and recognition method and device based on aggregated features and dual anchoring provided by the present invention have the following advantages compared with the prior art: 1. Aggregate high-frequency action and low-frequency structure information through the local-global feature perception module, effectively capture the fine-grained features of AUs, and combine the weighted loss function to alleviate the data imbalance problem and improve the per-frame AU recognition accuracy; 2. Adopt two-stream 3D convolution to fuse the original video and AU encoded features, use the dual anchoring strategy to take into account the localization requirements of expressions with different durations, and achieve efficient screening of multi-scale candidate intervals through the feature pyramid network and non-maximum suppression; 3. Optimize the localization and recognition tasks simultaneously through a unified framework, support parallel processing of macro / micro expressions, adapt to the complex situation of mixed multi-expressions in real scenarios, and have both high robustness and real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a schematic diagram of the implementation process of the method of the present invention;
[0051] Figure 2 is a schematic diagram of the structure of the per-frame AU recognition network based on local-global feature aggregation;
[0052] Figure 3 is a schematic diagram of the structures of the local feature aggregator and the global feature aggregator;
[0053] Figure 4 is a schematic diagram of the structure of the multi-scale 3D convolutional module;
[0054] Figure 5 is a schematic diagram of the structure of the prediction head;
[0055] Figure 6 is a schematic diagram of the structure of the entire unified model for fine-grained expression localization and recognition from the first perspective. DETAILED DESCRIPTION OF THE INVENTION
[0056] The present invention will be specifically introduced below in conjunction with the accompanying drawings and specific embodiments.
[0057] The present invention provides a joint expression localization and recognition method and device based on aggregated features and dual anchoring. In a frame-by-frame AU recognition network based on local-global feature aggregation, a cascaded local-global feature aggregation module is used to simultaneously capture local high-frequency action features focused by multi-head convolution and global low-frequency face features modeled by multi-head self-attention. Due to the fusion of fine-grained local and global information of the face, AU encoding can accurately characterize the semantics and spatial relevance of each AU. In the dual-anchored expression localization and recognition network, a dual-stream 3D convolutional network is used to extract multi-scale spatiotemporal features of single face video and AU encoding, combined with the dual anchoring strategy of the feature pyramid network, medium-duration expressions are adapted in the time domain through an anchor-based positioning module, and extremely short or extremely long expression fragments are captured through an anchor-free positioning module, and the start and end frames and categories of the final expression positioning are output. Due to the use of local-global feature aggregation and dual anchoring collaborative optimization, this case can adapt to complex situations in uncontrolled scenes where macro expressions and micro expressions appear mixedly and the duration difference is large, and it is expected to achieve high-precision expression recognition while improving the robustness of multi-expression positioning.
[0058] like Figure 1 As shown, it is a flowchart of the joint expression localization and recognition method based on aggregated features and dual anchoring, and each specific step is explained below.
[0059] S01: Perform face tracking on the input first-person perspective video, extract the trajectory of each face required for training in the first-person perspective video, and form each single face video.
[0060] For single face videos, in order to avoid the situation where too few frames are collected, making it difficult to learn the correct spatiotemporal association of action units, or too many frames are collected, making the learning time too long, it is necessary to selectively extract the video frames.
[0061] S02: Perform local-global feature aggregation on each frame of a single face image in a single face video to generate an AU code.
[0062] like Figure 2 As shown in Figure 2, for a given single face video, it is first input into a 2D convolutional layer for pre-coding to obtain a pre-coded feature map, and then (lg) A cascade of local-global feature aggregation modules is used to capture the local high-frequency motion features and global low-frequency face features of a single face image, followed by a batch normalization layer and a global average pooling layer to obtain the AU code F. t(u) , and finally a fully connected layer is used to obtain the AU recognition probability The pre-coded feature map serves as the input to the first local-global feature aggregation module. The output of the previous local-global feature aggregation module serves as the input to the subsequent local-global feature aggregation module, and the output of the last local-global feature aggregation module.
[0063] As Figure 3 shown, each local-global feature aggregation module is implemented using a sub-block embedding layer, N (l) cascaded local feature aggregators, and N (g) global feature aggregators. The sub-block embedding layer is used to embed and perform preliminary feature transformation on the input of the local-global feature aggregation module and map it to a feature space suitable for subsequent processing. The local feature aggregator is composed of a multi-head convolutional module and a multi-layer perceptron module and is used to extract local high-frequency action features. The global feature aggregator is composed of a 2D pointwise convolutional layer, a batch normalization layer, a multi-head self-attention module, a multi-head convolutional module, and a multi-layer perceptron module and is used to extract global low-frequency face features. The outputs of both the local feature aggregator and the global feature aggregator use residual connections.
[0064] The sub-block embedding layer embeds and performs preliminary feature transformation on the input of the local-global feature aggregation module, maps it to a feature space suitable for subsequent processing, and uses the resulting feature map as the input to the first local feature aggregator.
[0065] The implementation process of the local feature aggregator is as follows: In the multi-head convolutional module, first, each input feature map is divided into multiple sub-feature maps along the channel dimension, and 2D convolution operations are performed on each sub-feature map separately. After concatenating the convolution results of each sub-feature map along the channel dimension, batch normalization, non-linear activation, and 2D pointwise convolution are performed in sequence to obtain the output of the multi-head convolutional module. In the multi-layer perceptron module, the output of the multi-head convolutional module is connected to the input with a residual connection and used as the input to the multi-layer perceptron module. 2D convolution, non-linear activation, and dropout are performed in sequence to obtain the output of the multi-layer perceptron module, enhancing the expression ability of local features through the multi-layer perceptron module. The feature map obtained by connecting the input and output of the multi-layer perceptron module with a residual connection is used as the output of the local feature aggregator, and the output of the last local feature aggregator is used as the input to the first global feature aggregator.
[0066] The implementation process of the global feature aggregator is as follows: First, the number of input channels is adjusted through a 2D pointwise convolutional layer, and then the batch normalization layer and the multi-head self-attention module are sequentially passed through to capture the global low-frequency face features. The output of the multi-head self-attention module is connected to the input of the batch normalization layer through a residual connection and used as the input of the multi-head convolutional module. The output of the multi-head convolutional module is connected to the input through a residual connection and used as the input of the multi-layer perceptron module. The feature map obtained by connecting the input and output of the multi-layer perceptron module through a residual connection is used as the output of the global feature aggregator. The output of the last global feature aggregator passes through the batch normalization layer and the global average pooling layer to obtain the AU coding F t(u) 。
[0067] During the training process, in order to optimize the accuracy of AU recognition, a weighted binary cross-entropy loss function is used for supervision, and the loss function is expressed as:
[0068]
[0069] Where: represents the AU recognition probability, M (u) represents the total number of AUs, N represents the total number of sample images in the training set, represents the total number of sample images in the training set where the j-th AU appears, represents the appearance probability of the j-th AU in the training set, w j represents the weight to suppress the imbalance in the appearance frequencies between the j-th AU and other AUs, v j represents the weight to suppress the too low appearance frequency of the j-th AU (used for the first term of the cross-entropy, i.e., the appearance term), and represent the true appearance probability and the predicted appearance probability of the j-th AU in the t-th frame sample image, respectively.
[0070] The weight calculation is to address the issue of uneven AU appearance frequencies in AU recognition, and the form of the weight is:
[0071]
[0072] Where: represents the total number of sample images in the training set where the j-th AU appears; the weight w j is beneficial to suppressing the imbalance in appearance frequencies between different AUs. Low-frequency AUs will receive higher weights, making the model more sensitive to rare AUs and preventing high-frequency AUs from dominating the model training; the weight v j further adjusts the appearance frequency of each AU and is beneficial to suppressing the problem that the appearance frequency of each AU is often lower than the non-appearance frequency.
[0073] S03: Reshape and concatenate the AU-encoded time series, and perform multi-scale spatio-temporal feature fusion with the single-face video through a two-stream 3D convolutional network to generate fused expression features.
[0074] Reshape and concatenate the AU-encoded time series, convert it into a single-channel time series feature matrix I, and extract features from the single-face video to obtain a single-channel time series feature matrix II; the two-stream 3D convolutional network includes two inputs. The first input is the single-channel time series feature matrix I, and the second input is the single-channel time series feature matrix II. The first input extracts spatio-temporal feature I through an independent spatio-temporal feature extraction module, and the second input extracts spatio-temporal feature II through an independent spatio-temporal feature extraction module. After adding spatio-temporal feature I and spatio-temporal feature II, the fused expression feature F is obtained. (e) 。
[0075] The spatio-temporal feature extraction module consists of a cascaded 3D convolutional layer, a max pooling layer, a group of multi-scale 3D convolutional modules, and a global average pooling layer. As Figure 4 shown, the multi-scale 3D convolutional module contains a group of branches arranged in parallel. Each branch uses 3D convolutional kernels of different sizes (including 3×3×3 and 5×5×5 convolutional kernels) to capture information from different receptive fields respectively. The convolutional kernels of different sizes can extract features from different receptive fields, thus effectively capturing the spatio-temporal changes of local details and global structures. The outputs of each branch are concatenated at the channel level as the output of the multi-scale 3D convolutional module, and the output of the last multi-scale 3D convolutional module is dimension-reduced through the global average pooling layer to generate the corresponding spatio-temporal feature.
[0076] Through this multi-scale feature extraction method, the network can simultaneously focus on the local action features and global structure features in the video, thereby enhancing the sensitivity to facial expression changes.
[0077] S04: Perform multi-scale sampling on the fused expression feature F based on the sliding window strategy (e) for expression localization through a feature pyramid network combined with a dual localization mechanism of anchor-based and anchor-free, and filter the final expression localization through non-maximum suppression.
[0078] Perform multi-scale sampling on the fused expression feature F based on the sliding window strategy (e) It is necessary to set sliding windows of different lengths to cover the typical durations of macro-expressions and micro-expressions. There is partial overlap between two adjacent sliding windows in the time dimension. Uniformly sample T window frames from the single-face video segments within different sliding windows to obtain different window frame sequences.
[0079] Along the fused expression feature F (e)After applying a sliding window to the time dimension, first use a 1D convolutional layer and a max pooling layer to encode the window frame sequence, and then input it into a feature pyramid network with L levels. In each level of the feature pyramid network, a dual localization mechanism is executed in parallel (as Figure 5 shown), and a 1D convolutional layer is used to generate both anchor candidate intervals and anchor-free candidate intervals simultaneously.
[0080] (41) The expression localization of the anchor candidate intervals is achieved through intersection over union (IoU) threshold screening and dynamic regression parameter adjustment.
[0081] For an anchor candidate interval to be predicted, set the predefined anchor box as (c d , w d ). The predicted classification score is The IoU between the anchor candidate interval to be predicted and the closest true expression interval is p o . Denote the dynamic regression parameters of the anchor candidate interval to be predicted as {Δ c , Δ w}. Then the anchor candidate interval to be predicted is represented using the dynamic regression parameters as c = c d +αΔ c w d , w = w d exp(βΔ ω ).
[0082] During the training process, only the accurately predicted anchor candidate intervals are used to learn the IoU p o and the dynamic regression parameters {Δ c , Δ ω}; for the N p(ab) generated anchor candidate intervals, first calculate the IoU, then screen out M p(ab) positive candidate intervals (samples that match the true expression intervals) according to the IoU threshold, and finally calculate the IoU loss based on the positive candidate intervals At the same time, calculate the classification loss based on both positive and negative candidate intervals and the regression loss based only on positive candidate intervals in each level of the feature pyramid network and
[0083]
[0084] where: c d and w d represent the center point and width of the predefined anchor box respectively; c and w represent the predicted center point and predicted width of the anchor candidate interval respectively; Δ c represents the offset of c relative to c d , and Δ w represents the offset of w relative to w dOffset; Represents the score of an anchor candidate interval located as expression m, M (e) Is the total number of expression categories, and the prediction confidence is α and β are scale parameters. M p(ab) Is the number of positive samples corresponding to label assignment; Is the true label; l focal IoU(c, c′) of the focal loss represents the intersection over union between the predicted candidate interval and the true candidate interval; CSmoothL1(x) is the smooth L1 loss, which is used to alleviate the training instability problem when there are large deviations.
[0085] (42) Expression localization of the anchor-free candidate interval is achieved through foreground classification and dynamic regression parameters.
[0086] In the feature pyramid network, the input of the bottommost level is a sequence of window frames composed of T window frames. The output of the next level is used as the input of the upper level. Each level downsamples the input of this level (in the feature pyramid network, each time going up one level, the time dimension usually shrinks by half). Map the window frame positions in the window frame sequence to different mapping positions in the feature map: Let j ∈ {0, 1,..., T} represent the index of a certain window frame in the window frame sequence, and represent the mapping position corresponding to the index j in the i-th level as j i , Based on the mapping of the feature map, if j i The corresponding window frame j falls into a certain expression interval, then j i Is called the foreground, otherwise j i Is called the background (which means the window frame j does not belong to any expression interval); For an anchor-free candidate interval to be predicted, calculate the predicted classification score as The dynamic regression parameters are {Δ b , Δ e}.
[0087] During the training process, calculate the classification error of all N p(af) Anchor candidate intervals to obtain the classification loss Optimize the offsets of the M p(af) Anchor-free candidate intervals predicted as foreground to obtain the regression loss
[0088]
[0089] Where: Represents the score of an anchor-free candidate interval located as expression m, and the prediction confidence is max(s (af) ); Δ b = j - b, which is the starting frame offset and represents the predicted distance from window frame j to the starting frame b of the expression; Δ e= e - j, which is the end - frame offset, representing the predicted distance from window frame j to the end - expression frame b; M p(af) is the number of foreground points.
[0090] (43) During testing, the anchor - free localization module outputs the distances from the foreground points to the start - expression frame and to the end - expression frame, and the predicted confidence is the classification score s (af) the maximum value max(s (af) ) in; the anchor - based localization module outputs the interval position (c, w), the classification score s (ab) and the intersection - over - union ratio p o , and the predicted confidence is Each point in each layer of the Feature Pyramid Network has both anchor - free localization results and anchor - based localization results. After aggregating all the results of each layer, a non - maximum suppression operation is performed to obtain the final localization results of each expression.
[0091] (44) Calculate the total loss function of the double - anchored expression localization and recognition network: The loss function of the unified model includes a classification loss based on both positive and negative candidate intervals a regression loss based only on positive candidate intervals classification loss and regression loss Each loss term is assigned an appropriate weight. These loss terms are optimized by weighted summation during the training process of the unified framework:
[0092]
[0093] In each level of the Feature Pyramid Network, each window frame has both anchor - based and anchor - free expression localization results. Aggregate all the expression localization results of each level of the Feature Pyramid Network and perform a non - maximum suppression operation to obtain the final expression localization.
[0094] S05: Intercept the expression feature segment according to the start - frame and end - frame of the final expression localization, and input the expression feature segment into the expression localization and recognition network to predict the macro - expression category and micro - expression category.
[0095] According to the final expression localization, intercept the expression feature segment along the time dimension of the fused expression feature F (e) Input the intercepted expression feature segment into the expression localization and recognition network composed of a 2D convolutional layer, a temporal pooling layer, and a fully - connected classifier to predict the macro - expression category and micro - expression category, and use the multi - class cross - entropy loss L ce to calculate the expression recognition loss and achieve unified macro - expression and micro - expression recognition.
[0096] S06: Train the unified first - person fine - grained expression localization and recognition model composed of steps S01 - S05 by jointly optimizing the AU recognition loss, the dual - anchor localization loss, and the expression recognition loss.
[0097] In the loss function of the unified first - person fine - grained expression localization and recognition model, the AU recognition loss is the weighted binary cross - entropy loss L (u) , and the dual - anchor localization loss includes the classification loss Intersection over Union (IoU) loss Regression loss Classification loss and the distance regression loss The expression recognition loss is the multi - class cross - entropy loss L ce .
[0098] Train the unified first - person fine - grained expression localization and recognition model (as shown in Figure 6 ) composed of a frame - by - frame AU recognition network that aggregates local and global features, a spatio - temporal feature extraction network, and a dual - anchor expression localization and recognition network through an end - to - end method; First, train the frame - by - frame AU recognition network to extract accurate AU encodings as the input of the expression localization and recognition network. Then, the expression localization and recognition network learns the specific patterns and spatio - temporal correlation features of AUs, and uses the spatio - temporal correlation features between AUs to promote the localization and recognition of facial action units.
[0099] S07: Input a first - person perspective video with any number of frames into the trained unified first - person fine - grained expression localization and recognition model to identify macro - expression categories and micro - expression categories.
[0100] Directly output the prediction results of facial action units during prediction.
[0101] The method of the present invention can be fully implemented by a computer without manual assistance; this indicates that batch - processing and automation can be achieved in this case, which can greatly improve the processing efficiency and reduce the labor cost.
[0102] A joint expression localization and recognition device based on aggregated features and dual - anchor includes a single - face video extraction unit, a local - global feature aggregation unit, a spatio - temporal feature fusion unit, a dual - anchor localization unit, an expression recognition unit, and a joint training and optimization unit;
[0103] The single - face video extraction unit performs face tracking on the input first - person perspective video, extracts the trajectory of each face in the first - person perspective video, and forms individual temporally coherent single - face videos;
[0104] The local - global feature aggregation unit first inputs the single - face video into a 2D convolutional layer for pre - encoding to obtain a pre - encoded feature map, and then through Nlg) A cascaded local-global feature aggregation module is used to capture the local high-frequency action features and global low-frequency face features of a single face image. Then, through a batch normalization layer and a global average pooling layer, the AU code F is obtained. t(u) Finally, through a fully connected layer, the AU recognition probability is obtained.
[0105] For the spatio-temporal feature fusion unit, first, feature extraction is respectively performed on the time series of the single face video and the AU code F t(u) to obtain two single-channel time series feature matrices. Then, through a two-stream 3D convolutional network, multi-scale spatio-temporal feature fusion is performed on the two single-channel time series feature matrices to generate the fused expression feature F. (e) ;
[0106] For the dual-anchoring localization unit, based on the multi-layer time series features extracted by the feature pyramid network containing L levels, combined with the sliding window strategy, predefined anchor boxes and dynamic regression parameters are generated. Through non-maximum suppression, the expression localization interval corresponding to the maximum confidence value is selected, and at the same time, classification and regression with and without anchors are realized.
[0107] For the expression recognition unit, according to the final expression localization result of the dual-anchoring localization unit, along the time dimension of the fused expression feature F (e) expression feature segments are intercepted. Through the time series pooling layer, the feature dimension is compressed, and a fully connected classifier is used to distinguish macro-expression categories and micro-expression categories to achieve unified expression recognition.
[0108] For the joint training and optimization unit, the expression localization accuracy (intersection over union loss) and expression recognition performance (cross-entropy loss) are synchronously optimized through a multi-task loss function, and weight parameters are introduced to balance the gradients of the two types of losses; the parameters and loss values of the first-view fine-grained expression localization and recognition unified model composed of the local-global feature aggregation unit, spatio-temporal feature fusion unit, dual-anchoring localization unit, and expression recognition unit are calculated, and the parameters are updated based on the gradient optimization method.
[0109] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any form. Any technical solutions obtained by using equivalent replacements or equivalent transformations fall within the protection scope of the present invention.
Claims
1. A joint expression localization and recognition method based on aggregated features and dual anchoring, characterized in that: The steps are as follows: S01: Perform face tracking on the input first-person perspective video, extract the trajectory of each face in the first-person perspective video, and form individual single-face videos; S02: Perform local-global feature aggregation on each frame of the single-face image in the single-face video to generate AU codes; the local-global feature aggregation is implemented through a cascaded local feature aggregator and a global feature aggregator, where the local feature aggregator uses a multi-head convolution module and a multi-layer perceptron module to extract local high-frequency action features, and the global feature aggregator uses a multi-head self-attention module, a multi-head convolution module, and a multi-layer perceptron module to extract global low-frequency face features, and suppresses the gradient dispersion of the local feature aggregator and the global feature aggregator through residual connections; S03: Reshape and concatenate the time series encoded by AU, convert it into a single-channel time series feature matrix I, extract features from the single-face video to obtain a single-channel time series feature matrix II, and input the single-channel time series feature matrix I and the single-channel time series feature matrix II into a two-stream 3D convolutional network for multi-scale spatio-temporal feature fusion to generate fused expression features F (e) ; S04: Perform multi-scale sampling on the fused expression feature F based on the sliding window strategy (e) Perform facial expression localization by combining the feature pyramid network with dual localization mechanisms of anchor-based and anchor-free, and filter the final facial expression localization through non-maximum suppression S05: Intercept the expression feature segment according to the start frame and end frame of the final expression localization, and input the expression feature segment into the expression localization and recognition network to predict the macro-expression category and the micro-expression category; S06: Train the first-person perspective fine-grained expression localization and recognition unified model composed of steps S01 to S05 by jointly optimizing the AU recognition loss, the dual-anchor localization loss, and the expression recognition loss; S07: Input the first-person perspective video with any given number of frames into the trained first-person perspective fine-grained expression localization and recognition unified model to recognize the macro-expression category and the micro-expression category.
2. The combined expression localization and recognition method based on aggregated features and dual anchoring according to claim 1, wherein: In the step S02, first, the single-face video is input into a 2D convolutional layer for pre-encoding to obtain a pre-encoded feature map, and then, through N (lg) cascaded local-global feature aggregation modules to capture the local high-frequency action features and global low-frequency face features of the single-face image, and then, through a batch normalization layer and a global average pooling layer, the AU encoding F t (u) is obtained. Finally, through a fully connected layer, the AU recognition probability is obtained. The pre-encoded feature map serves as the input to the first local-global feature aggregation module, the output of the previous local-global feature aggregation module serves as the input to the next local-global feature aggregation module, and the output of the last local-global feature aggregation module; Each local-global feature aggregation module uses a sub-block embedding layer, N (l) cascaded local feature aggregators, and N (g) global feature aggregators for implementation; the sub-block embedding layer is used to embed and perform preliminary feature transformation on the input of the local-global feature aggregation module, and map it to a feature space suitable for subsequent processing. The local feature aggregator is composed of a multi-head convolution module and a multi-layer perceptron module, and is used to extract local high-frequency action features. The global feature aggregator is composed of a 2D pointwise convolution layer, a batch normalization layer, a multi-head self-attention module, a multi-head convolution module, and a multi-layer perceptron module, and is used to extract global low-frequency face features. The outputs of the local feature aggregator and the global feature aggregator both use residual connections; The sub-block embedding layer embeds and performs preliminary feature transformation on the input of the local-global feature aggregation module, maps it to a feature space suitable for subsequent processing, and uses the obtained feature map as the input of the first local feature aggregator; The implementation process of the local feature aggregator is as follows: in the multi-head convolution module, first divide the input feature map along the channel dimension into multiple sub-feature maps, and perform 2D convolution operations on each sub-feature map respectively. After concatenating the convolution results of each sub-feature map along the channel dimension, perform batch normalization, non-linear activation, and 2D pointwise convolution in sequence to obtain the output of the multi-head convolution module; in the multi-layer perceptron module, connect the output of the multi-head convolution module with the input through a residual connection as the input of the multi-layer perceptron module, perform 2D convolution, non-linear activation, and dropout in sequence to obtain the output of the multi-layer perceptron module, and enhance the expression ability of local features through the multi-layer perceptron module; use the feature map obtained by connecting the input and output of the multi-layer perceptron module through a residual connection as the output of the local feature aggregator, and use the output of the last local feature aggregator as the input of the first global feature aggregator; The implementation process of the global feature aggregator is as follows: First, the number of input channels is adjusted through a 2D pointwise convolutional layer, and then the global low-frequency face features are captured through a batch normalization layer and a multi-head self-attention module in sequence. The output of the multi-head self-attention module is connected to the input of the batch normalization layer through a residual connection and used as the input of the multi-head convolutional module. The output of the multi-head convolutional module is connected to the input through a residual connection and used as the input of the multi-layer perceptron module. The feature map obtained by connecting the input and output of the multi-layer perceptron module through a residual connection is used as the output of the global feature aggregator. After passing through the batch normalization layer and the global average pooling layer, the output of the last global feature aggregator obtains the AU code F t(u) .
3. The combined expression localization and recognition method based on aggregated features and dual anchoring according to claim 2, characterized in that: In the step S02, the AU recognition probability uses the weighted binary cross-entropy loss L (u) for supervision: Wherein: represents the AU recognition probability, M (u) represents the total number of AUs, N represents the total number of sample images in the training set, represents the total number of sample images with the j-th AU in the training set, represents the occurrence probability of the j-th AU in the training set, w j represents the weight for suppressing the frequency imbalance between the j-th AU and other AUs, v j represents the weight for suppressing the too low occurrence frequency of the j-th AU, and respectively represent the true occurrence probability and the predicted occurrence probability of the i-th AU in the t-th frame sample image.
4. The combined expression localization and recognition method based on aggregated features and dual anchoring according to claim 1, characterized in that: In the step S03, the dual-stream 3D convolutional network includes two inputs. The first input is a single-channel temporal feature matrix I, and the second input is a single-channel temporal feature matrix II. The first input extracts spatio-temporal feature I through an independent spatio-temporal feature extraction module, and the second input extracts spatio-temporal feature II through an independent spatio-temporal feature extraction module. After adding spatio-temporal feature I and spatio-temporal feature II, the fused expression feature F is obtained (e) ; The spatio-temporal feature extraction module is composed of a cascaded 3D convolutional layer, a max pooling layer, a group of multi-scale 3D convolutional modules, and a global average pooling layer; the multi-scale 3D convolutional module contains a group of branches arranged in parallel, and each branch uses 3D convolutional kernels of different sizes to capture information with different receptive fields respectively. After concatenating the outputs of each branch at the channel level, use it as the output of the multi-scale 3D convolutional module, and reduce the dimension of the output of the last multi-scale 3D convolutional module through the global average pooling layer to generate the corresponding spatio-temporal features.
5. The joint expression localization and recognition method based on aggregated features and dual anchoring according to claim 1, wherein: In the step S04, first, based on the sliding window strategy, the fused expression feature F (e) is sampled at multiple scales. Sliding windows with different lengths are set to cover the typical durations of macro-expressions and micro-expressions. There is partial overlap between two adjacent sliding windows in the time dimension. T window frames are uniformly sampled from the single-face video segments within different sliding windows to obtain different window frame sequences; Along the temporal dimension of the fused expression feature F (e) After applying a sliding window along the temporal dimension, first, a 1D convolutional layer and a max-pooling layer are used to encode the window frame sequence, and then it is input into a feature pyramid network with L levels. In each level of the feature pyramid network, a dual localization mechanism is executed in parallel, and a 1D convolutional layer is used to generate both anchor candidate intervals and anchor-free candidate intervals simultaneously; The expression localization of the anchored candidate intervals is achieved through intersection over union (IoU) threshold screening and dynamic regression parameter adjustment: For an anchored candidate interval to be predicted, a predefined anchor box is set as (c d , w d ). The predicted classification score is . The IoU between the anchored candidate interval to be predicted and the closest true expression interval is p o . The dynamic regression parameters of the anchored candidate interval to be predicted are denoted as {Δ c , Δ w}. Then the anchored candidate interval to be predicted is represented using the dynamic regression parameters as c = c d + αΔ c w d , w = w d exp(βΔ ω ). During the training process, only the accurately predicted anchored candidate intervals are used to learn the IoU p o and the dynamic regression parameters {Δ c , Δ ω}. For the generated N p(ab) anchored candidate intervals, first calculate the IoU, then screen out M p(ab) positive candidate intervals according to the IoU threshold, and finally calculate the IoU loss based on the positive candidate intervals At the same time, calculate the classification loss based on both positive and negative candidate intervals and the regression loss based only on positive candidate intervals at each level of the feature pyramid network where: c d and w d represent the center point and width of the predefined anchor box, respectively; c and w represent the predicted center point and predicted width of the anchor candidate interval; Δ c represents the offset of c relative to c d , and Δ w represents the offset of w relative to w d ; represents the score for the anchor candidate interval to be located as expression m, M (e) is the total number of expression types, and the predicted confidence is p o max(s (ab) ); α and β are scale parameters; The expression localization of the anchor-free candidate intervals is achieved through foreground classification and dynamic regression parameters: In the Feature Pyramid Network, the input of the bottommost level is a sequence of window frames consisting of T window frames. The output of the next level is used as the input of the previous level. Each level downsamples the input of this level and maps the window frame positions in the window frame sequence to different positions in the feature map. Let j ∈ {0, 1, …, T} represent the index of a certain window frame in the window frame sequence, and the corresponding mapping position of the index j in the i-th level is denoted as j i , based on the mapping of the feature map, if the window frame j corresponding to j i falls into a certain expression interval, then j i is called the foreground, otherwise j i is called the background; for an anchor-free candidate interval to be predicted, calculate the predicted classification score as The dynamic regression parameters are {Δ b , Δ e}; during the training process, calculate the classification errors of all N p(af) anchor candidate intervals to obtain the classification loss Optimize the offsets of the Mp (af) anchor-free candidate intervals predicted as foreground to obtain the regression loss Wherein: denotes the score of the anchor-free candidate interval located as expression m, and the prediction confidence is max(s (af) ); Δ b = j - b, which is the starting frame offset and represents the predicted distance from window frame j to the starting frame b of the expression; Δ e = e - j, which is the ending frame offset and represents the predicted distance from the window frame to the ending frame b of the expression; In each level of the Feature Pyramid Network (FPN), each window frame has the expression localization results with anchors and without anchors. Aggregate all the expression localization results at each level of the FPN and perform non-maximum suppression operation to obtain the final expression localization.
6. The combined expression localization and recognition method based on aggregated features and dual anchoring according to claim 1, characterized in that: In the step S05, according to the final expression localization, an expression feature segment is intercepted along the time dimension of the fused expression feature F (e) , and the intercepted expression feature segment is input into an expression localization and recognition network composed of a 2D convolutional layer, a temporal pooling layer, and a fully connected classifier to predict the macro-expression category and the micro-expression category, and a multi-class cross-entropy loss L ce is used to calculate the expression recognition loss.
7. The joint expression localization and recognition method based on aggregated features and dual anchoring according to claim 1, characterized in that: In the step S06, in the loss function of the first - perspective fine - grained expression localization and recognition unified model, the AU recognition loss is the weighted binary cross - entropy loss L (u) , the dual - anchor localization loss includes a classification loss IoU loss regression loss classification loss and distance regression loss The expression recognition loss is the multi - class cross - entropy loss L ce .
8. An apparatus for implementing the joint expression localization and recognition method based on aggregated features and dual anchoring according to any one of claims 1 to 7, characterized in that: It includes a single-face video extraction unit, a local-global feature aggregation unit, a spatio-temporal feature fusion unit, a dual-anchoring localization unit, an expression recognition unit, and a joint training and optimization unit. The single-face video extraction unit performs face tracking on the input first-person perspective video, extracts the trajectory of each face in the first-person perspective video, and forms individual temporally coherent single-face videos. The local-global feature aggregation unit first inputs the single-face video into a 2D convolutional layer for pre-encoding to obtain a pre-encoded feature map, and then passes it through N (lg) cascaded local-global feature aggregation modules to capture the local high-frequency action features and global low-frequency face features of the single-face image. Then, it passes through a batch normalization layer and a global average pooling layer to obtain the AU encoding F t(u) , and finally passes through a fully connected layer to obtain the AU recognition probability The spatio-temporal feature fusion unit first extracts features from the time series of the single-face video and the AU-encoded F t(u) respectively to obtain two single-channel temporal feature matrices, and then performs multi-scale spatio-temporal feature fusion on the two single-channel temporal feature matrices through a two-stream 3D convolutional network to generate the fused expression feature F (e) ; The dual-anchoring localization unit generates predefined anchor boxes and dynamic regression parameters based on the multi-level temporal features extracted by the feature pyramid network with L levels, combines with the sliding window strategy, selects the expression localization interval corresponding to the maximum confidence value through non-maximum suppression, and simultaneously realizes the classification and regression with and without anchors. The expression recognition unit intercepts expression feature segments along the time dimension of the fused expression feature F according to the final expression localization result of the dual-anchoring localization unit, compresses the feature dimension through a temporal pooling layer, and uses a fully connected classifier to distinguish macro-expression categories and micro-expression categories, thereby achieving unified expression recognition; (e) The expression recognition unit intercepts expression feature segments along the time dimension of the fused expression feature F according to the final expression localization result of the dual-anchoring localization unit, compresses the feature dimension through a temporal pooling layer, and uses a fully connected classifier to distinguish macro-expression categories and micro-expression categories, thereby achieving unified expression recognition; The joint training and optimization unit synchronously optimizes the expression localization accuracy and expression recognition performance through a multi-task loss function, introduces weight parameters to balance the gradients of the two types of losses; calculates the parameters and loss values of the first-person fine-grained expression localization and recognition unified model composed of the local-global feature aggregation unit, the spatio-temporal feature fusion unit, the dual-anchoring localization unit, and the expression recognition unit, and updates the parameters based on the gradient optimization method.
9. The device according to claim 8, characterized in that: The local-global feature aggregation module uses a sub-block embedding layer, N (l) cascaded local feature aggregators, and N (g) global feature aggregators. The sub-block embedding layer is used to perform preliminary feature transformation on the input and map it to a feature space suitable for subsequent processing. The local feature aggregators use a multi-head convolution module and a multi-layer perceptron module to extract local high-frequency action features. The global feature aggregators use a multi-head self-attention module, a multi-head convolution module, and a multi-layer perceptron module to extract global low-frequency face features. The gradient dispersion of the local feature aggregators and the global feature aggregators is suppressed through residual connections.
10. The device according to claim 8, characterized in that: For the dual-stream 3D convolutional network, independent spatio-temporal feature extraction modules are respectively used to extract spatio-temporal features from two single-channel temporal feature matrices, and the fused expression feature F is obtained after adding the two extracted spatio-temporal features together (e) ; The spatio-temporal feature extraction module is composed of cascaded 3D convolutional layers, max pooling layers, a group of multi-scale 3D convolutional modules, and a global average pooling layer; the multi-scale 3D convolutional module contains a group of branches arranged in parallel, each branch uses 3D convolutional kernels of different sizes to capture information of different receptive fields respectively, and the outputs of each branch are concatenated at the channel level as the output of the multi-scale 3D convolutional module. The output of the last multi-scale 3D convolutional module is dimension-reduced by the global average pooling layer to generate the corresponding spatio-temporal features.
Citation Information
Patent Citations
Establishment method of pig face facial expression recognition framework based on multi-task cascade
CN113065460A
Dynamic facial expression recognition method for modeling complex time-space relationship
CN119296157A
Face emotion recognition method based on dual-stream convolutional neural network
US20190311188A1
Cited By
Multi-object tracking method based on global-local feature joint modeling
CN121213616A