This invention discloses a joint expression localization and recognition method and apparatus based on aggregated features and dual anchoring. Addressing the need for fine-grained
expression analysis of single faces in first-person view videos, it constructs an end-to-end
deep learning framework to achieve unified localization and recognition of
macro and micro expressions. The method first processes the acquired single-face video data through a frame-by-frame AU recognition network that aggregates local and global features. Combining the advantages of convolutional neural networks and visual Transformers, multi-scale features are extracted and AU codes are generated. Then, a dual-
stream 3D convolutional network assisted by AU codes is used to fuse spatiotemporal features, combined with a multi-
scale sliding window strategy and a feature
pyramid network, and a dual anchoring mechanism is proposed to jointly locate expression intervals. Finally, video segments are extracted based on the localization results, and classification is completed through an expression localization and recognition network. This invention effectively solves the problems of multiple coexisting expressions, large duration differences, and difficulties in micro-expression detection in videos through local-global
feature aggregation, dual-
stream 3D convolutional networks, and dual anchoring strategies, significantly improving the accuracy and robustness of expression localization and recognition.