Multi-view facial expression recognition method based on hypergraph
Through the hypergraph structure and cross attention mechanism, the facial expression recognition method is integrated with the hypergraph structure and cross attention mechanism, the problems of information loss and feature differences under multi-view conditions are solved, and high-precision expression recognition effect is achieved.
Patent Information
- Application Number
- CN202510726140.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The existing facial expression recognition methods have information loss and large differences in feature expression under multi-view conditions, resulting in low recognition accuracy and lack of effective data structures to characterize the distribution characteristics of facial key points at different perspectives.
The hypergraph structure is used to characterize the topological structure of facial key points, and the facial features are extracted in combination with the dual-stream heterogeneous network, and the hypergraph convolution is used to capture higher-order correlation information, and image spatial features are extracted through the ResNet50 network, and feature fusion is performed using the cross attention mechanism.
The accuracy and adaptability of expression recognition under multi-view conditions are improved, the model's attention to local details and global context information is enhanced, and the accuracy and robustness of expression recognition are improved.
Smart Images

Figure CN120260101A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the problem of multi-view facial expression recognition in the field of deep learning, and particularly to a multi-view facial expression recognition method based on hypergraphs. Background Art
[0002] As an important part of affective computing, facial expression recognition technology can provide support for edge computing scenarios such as medical health, human-computer interaction, and intelligent driving by analyzing human emotional states through computer vision. In the past few decades, the field of facial expression recognition has developed from theoretical exploration to practical application, continuously promoting the progress of intelligent services. However, most of the existing facial expression recognition methods analyze frontal human faces, ignoring the problem of facial deflection that commonly exists in actual scenarios. The deflection of the facial angle will lead to the loss of expression information, seriously affecting the recognition accuracy. Therefore, exploring and researching multi-view human face expression recognition has important practical significance, which can not only make up for the limitations of existing methods in practical applications, but also further improve the adaptability and reliability of facial expression recognition technology.
[0003] Different from frontal facial expression recognition, the challenges of multi-view expression recognition mainly focus on two aspects: one is the information loss caused by facial deflection; the other is that the characteristic manifestation forms of the same expression are quite different under different views, increasing the intra-class difference. To solve the above problems, most researchers extract common features under different views or convert lateral human faces into frontal human faces for expression recognition. However, the existing methods ignore the distribution characteristics of facial key points under different views and lack an effective data structure to represent them. Therefore, this patent uses hypergraphs to represent and extract expression features under different views. First, a hypergraph is constructed according to the distribution characteristics of facial key points to capture complex human face expression patterns. Then, a two-stream heterogeneous network is used to extract facial features, that is, hypergraph convolution is used to mine high-order correlation information, and at the same time, a convolutional neural network is used to extract image spatial information. To achieve accurate feature localization, this patent uses the global geometric features obtained by hypergraph convolution to guide image feature localization, focusing on key facial regions. Finally, a cross-attention mechanism is used to deeply fuse these two heterogeneous features of images and geometry to achieve expression perception under different views. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi-view facial expression recognition method based on hypergraphs, which introduces a hypergraph structure to represent expression features under different views, uses it to guide image spatial feature localization, and finally completes the fusion of spatial features and geometric features through a cross-attention mechanism, improving the expression perception ability under different views.
[0005] For the convenience of description, the following concepts are introduced first: Hypergraph: A generalization in graph theory that can describe and analyze complex multi - variable relationships. Different from traditional graph structures, each edge in a hypergraph is called a hyper - edge, which allows connecting more than two nodes and can represent the relationships between entities more effectively.
[0006] ResNet50 Network (ResNet50): A convolutional neural network that deepens the network layers through residual connections while avoiding the problems of gradient vanishing and gradient explosion.
[0007] Cross - Attention Mechanism: It dynamically aggregates key information to achieve feature interaction by calculating the correlation weights between the query matrix and the key - value matrix.
[0008] The present invention specifically adopts the following technical solutions: A multi - perspective facial expression recognition method based on hypergraphs, characterized in that: a. Construct a hypergraph through facial key points to represent the topological structure information under different perspectives; b. Use a two - stream heterogeneous network to extract facial features. Extract geometric features in facial key points through hypergraph convolution, and extract image - space features through ResNet50; c. Use the geometric features of facial key points to guide image feature localization to achieve focused attention on key facial regions; d. Deeply fuse geometric features and spatial features through cross - attention to achieve the fusion of heterogeneous spaces; This method mainly includes the following steps: (1) Data pre - processing: Crop and scale the face images in the dataset to 224×224 pixels, and then perform normalization processing, and introduce the random erasing technique to enhance the model's robustness to local feature loss; (2) Facial key - point detection: Use the RetinaFace algorithm to detect the face image to obtain the spatial coordinates of facial key points; (3) Hypergraph construction: Utilize the key - point data obtained in step (2) to construct a hypergraph according to different facial regions, and connect different regions with hyper - edges; (4) Geometric feature extraction: Use hypergraph convolution to infer and update the hypergraph obtained in step (3) to effectively capture the global features of facial expressions; (5) Spatial feature extraction: Use the ResNet50 network to extract spatial features in the facial image to capture rich local emotional details in the facial image; (6) Guided feature localization: Use the geometric features obtained from hypergraph convolution in step (4) to guide the localization of spatial features in step (5), enabling the model to focus on facial regions that can better reflect emotions. (7) Feature fusion: There are semantic differences in the feature spaces constructed in steps (4) and (5), so a cross-attention mechanism is used for heterogeneous feature fusion to obtain a unified emotional description. (8) Model training: Based on an end-to-end learning framework, use the cross-entropy loss function for gradient backpropagation to optimize the model.
[0009] The beneficial effects of the present invention are as follows: (1) Use a hypergraph, a high-order association structure, to model facial key point information. The hypergraph can capture the complex relationships and high-order interactions between key points, providing a richer and more accurate representation of emotional information.
[0010] (2) Use a two-stream heterogeneous branch to extract facial emotion features. The feature encoder based on ResNet50 can capture the spatial information of the image and provide a fine-grained local feature description; the feature encoder based on hypergraph convolution is used to extract the geometric features of facial key points and perceive facial emotions from a global perspective.
[0011] (3) Use the geometric information extracted by hypergraph convolution to guide the learning process of ResNet50, thereby enhancing its feature perception ability and enabling it to focus on regions closely related to emotional expression.
[0012] (4) Deeply fuse the spatial features of the image with the geometric features in the facial key point topological structure in a cross-attention manner to integrate heterogeneous feature spaces and obtain complete emotion features. Brief Description of the Drawings
[0013] Figure 1 It is a schematic diagram of the hypergraph construction process.
[0014] Figure 2 It is a diagram of the overall model framework.
[0015] Figure 3 It is a structural diagram of the key point feature guidance module.
[0016] Figure 4 It is a structural diagram of the cross-fusion module. Detailed Embodiments
[0017] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is necessary to point out that the following embodiments are only used to further illustrate the present invention and should not be construed as limiting the protection scope of the present invention. Those skilled in the art can make some non-essential improvements and adjustments to the present invention according to the above-mentioned invention content and still fall within the protection scope of the present invention.
[0018] A multi-view facial expression recognition method based on a hypergraph specifically includes the following steps: (1) Data preprocessing To ensure the consistency and comparability of the experiments, the resolution of all input images is adjusted to 224×224 pixels. Then, normalization processing is performed, and a random erasing operation is added to enhance the model's perception ability of local feature loss; (2) Facial key point detection The RetinaFace algorithm is used to detect 68 facial key points of the face image, and these key points precisely outline the facial features of the contour, eyebrows, both eyes, nose, and mouth. To focus on the areas closely related to expression, this chapter selects 51 key points excluding the facial contour, as Figure 1 shown in (a).
[0019] (3) Hypergraph construction After obtaining the 51 key point features highly related to expression, the face is divided into six regions: left eyebrow, right eyebrow, left eye, right eye, nose, and mouth, as Figure 1 shown in (b). For each region, the key points in it are combined to form a hyperedge. To further capture the mutual correlation between different regions in emotional expression, two hyperedges are respectively constructed to connect regions 1 to 5 and regions 5 and 6, so as to enhance the information richness and expression ability of the model. The finally constructed hypergraph structure is as Figure 1 shown in (c). This structure can effectively capture the complex relationships between various facial regions in expression, and comprehensively reflect the feature information within a single region and the interaction between regions.
[0020] (4) Geometric feature extraction As Figure 2 shown, hypergraph convolution is used to extract geometric features from facial key points. Specifically, first, the information of the nodes associated with each hyperedge is aggregated to the hyperedge to complete hyperedge information collection. Then, the node features are updated using all the hyperedges connected to the nodes to complete node feature aggregation. Through 4 stages of hypergraph convolution, 2048-dimensional geometric features are finally obtained.
[0021] (5) Spatial feature extraction As Figure 2As shown, a convolutional neural network is used to extract spatial features from an image. Specifically, the ResNet50 network, which also consists of 4 stages, is used as the backbone network for feature extraction, and the final classification layer is adjusted to 2048 dimensions to obtain the spatial features of the facial image.
[0022] (6)Guiding Feature Localization As Figure 2 shown, the global geometric features obtained in each hypergraph convolution stage are used to guide the learning process of ResNet50. The structure of the key point feature guiding module is as Figure 3 shown. Specifically, for the obtained geometric features , first, the attention calculation method shown in Equation (1) is used to obtain multiple attention heads , and then different attention heads are concatenated through Equation (2) to focus on the characteristics of information from different dimensions and capture richer patterns and dependencies .
[0023] (1) (2) where Q 、 K and V are 's linear mappings, 、 、 and are trainable parameters, represents the length of the feature vector, represents the normalization exponential function, represents the concatenation operation, and what follows where represents 's calculation method. Then, the two-dimensional features are converted into three-dimensional features , and depthwise convolution and pointwise convolution are used to model the context information from the channel and feature map dimensions respectively to obtain the features , as shown in Equation (3). Then, the multi-layer perceptron shown in Equation (4) is used to obtain the key point mask . Finally, is passed to the ResNet50 network to guide it to focus on the regions more relevant to emotional expression.
[0024] (3) (4) where and represent depthwise convolution and pointwise convolution respectively, represents matrix multiplication.
[0025] (7)Feature fusion As Figure 4 shown, two heterogeneous features are fused through the cross-attention mechanism. Specifically, the features of the last stage of hypergraph convolution are used to generate the corresponding matrices , and . Similarly, the features of the last stage of ResNet50 are used to generate the corresponding matrices , and . Then, interactive perception is achieved through the calculation methods of equations (5) and (6) to obtain and , where and represent the vector lengths. Finally, the two are concatenated for facial expression classification. This fusion method not only strengthens the model's attention to local details but also improves the accuracy and robustness of emotion recognition by integrating global context information.
[0026] (5) (6) (8)Model training The model is trained in an end-to-end manner, and the cross-entropy loss function is used to optimize the model. The cross-entropy loss function is a commonly used loss function in classification problems and is used to measure the difference between the model's predicted probability distribution and the true label distribution. Its formula is shown in equation (7): (7) where represents the loss function,[[]] y represents the true label,[[]] represents the predicted value of the model,[[]] C represents the involved emotion categories.
Claims
1. A multi-view facial expression recognition method based on hypergraph, characterized in that: a. Construct a hypergraph through facial key points to represent the topological structure information under different views; b. Use a two-stream heterogeneous network to extract facial features, extract geometric features in facial key points through hypergraph convolution, and extract image space features through ResNet50; c. Use the geometric features of facial key points to guide image feature localization to achieve focused attention on key facial regions; d. Achieve the fusion of heterogeneous spaces by deeply fusing geometric features and space features through cross-attention; (1) Data preprocessing: Crop and scale the face images in the dataset to 224×224 pixels, and then perform normalization processing, and introduce the random erasing technique to enhance the model's robustness to local feature loss; (2) Facial key point detection: Use the RetinaFace algorithm to detect face images to obtain the spatial coordinates of facial key points; (3) Construct a hypergraph: Use the key point data obtained in step (2) to construct a hypergraph according to different facial regions, and use hyperedges to connect different regions; (4) Geometric feature extraction: Use hypergraph convolution to infer and update the hypergraph obtained in step (3) to effectively capture the global features of facial expressions; (5) Spatial feature extraction: Use the ResNet50 network to extract spatial features in facial images to capture rich local emotional details in facial images; (6) Guided feature localization: Use the geometric features obtained by hypergraph convolution in step (4) to guide the localization of spatial features in step (5) to make the model focus on facial regions that can better reflect emotions; (7) Feature fusion: There are semantic differences in the feature spaces constructed in steps (4) and (5), so use the cross-attention mechanism for heterogeneous feature fusion to obtain a unified emotional description; (8) Model training: Based on the end-to-end learning framework, use the cross-entropy loss function for gradient backpropagation to achieve model optimization.
2. The multi-view facial expression recognition method based on a hypergraph according to claim 1, characterized in that After the construction of the hypergraph is completed through steps (2) and (3); the hypergraph is a high-order association structure that can capture the complex relationships and high-order interactions between key points, thereby improving the accuracy of expression recognition.
3. The multi-view facial expression recognition method based on a hypergraph according to claim 1, characterized in that In steps (4) and (5), geometric features and spatial features are respectively extracted using two-stream heterogeneous branches; The topological graph structure constructed by facial key points can effectively capture the overall characteristics of the face and model expression information by combining global context information; at the same time, the feature encoder based on ResNet can effectively extract rich local emotional details in facial images.
4. The multi-view facial expression recognition method based on a hypergraph according to claim 1, characterized in that In step (6), geometric features are used to guide spatial feature localization; the feature encoder based on ResNet has an inductive bias and it is difficult to evaluate the importance of features from a global perspective; In contrast, the geometric features extracted by the hypergraph can effectively represent the overall characteristics of the face; Therefore, geometric features can effectively guide the localization of image features.
5. The multi-view facial expression recognition method based on a hypergraph according to claim 1, characterized in that Heterogeneous feature fusion is completed in step (7); geometric features and spatial features respectively describe global and local sentiment features, and cross-attention can be used to deeply fuse the two, realizing the integration of heterogeneous feature spaces and the extraction of final sentiment features.
Citation Information
Patent Citations
Expression recognition method based on key point spatial distribution and local and global feature reasoning
CN117831088A
AIGC video creation method for cultural relic explanation
CN119516059A
Multi-modal hypergraph emotion recognition method for video learning scene
CN119810887A
Multi-modal cognitive impairment evaluation system based on multi-dimensional cognitive function hypergraph
CN119818025A
Method and system for multimodal emotion recognition in conversation (ERC) based on graph neural network (GNN)
US20240355350A1