A multi-view facial expression recognition method based on hypergraphs

Through the hypergraph structure and cross attention mechanism, the facial expression recognition method is integrated with the hypergraph structure and cross attention mechanism, the information loss and feature difference of expression recognition in multi-view angles is solved, and the high-precision expression recognition effect is achieved.

CN120260101BActive Publication Date: 2025-07-29SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510726140.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-29
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The existing facial expression recognition method has low recognition accuracy when face deflection, and ignores the distribution characteristics of facial key points at different perspectives, resulting in large differences in information loss and feature expression, making it difficult to achieve accurate multi-view expression recognition.

Method used

The hypergraph structure is used to characterize the topological structure information of facial key points, and the facial features are extracted in combination with the dual-stream heterogeneous network. The hypergraph convolution is used to capture geometric features and guide image feature positioning. The heterogeneous features are fused through the cross attention mechanism to achieve expression recognition.

Benefits of technology

It improves the accuracy and robustness of facial expression recognition in multi-view angles, and enhances the model's adaptability and reliability to facial expression information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260101B_ABST
    Figure CN120260101B_ABST
Patent Text Reader

Abstract

The present invention proposes a multi-view facial expression recognition method based on hypergraphs, which mainly involves the problem of extracting emotional features in images for expression recognition in deep learning. First, a hypergraph structure is constructed according to the spatial distribution characteristics of facial key points to highlight the geometric features therein with a more reasonable spatial layout. Secondly, a key point feature guidance module is designed to use the high-order correlation information obtained by hypergraph convolution to guide the spatial feature localization of the ResNet50 network, enabling it to focus on the regions in the image that can better reflect emotions. In order to learn unified emotional features from two heterogeneous feature spaces, a cross-fusion module is further designed to integrate heterogeneous features with a cross-attention mechanism. The present invention fully considers the distribution characteristics of facial expressions from different perspectives, represents the geometric features in key points through hypergraphs, and uses them to guide the spatial feature localization of images, improving the adaptability of expression recognition technology in practical applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the problem of multi-view facial expression recognition in the field of deep learning, and particularly to a multi-view facial expression recognition method based on hypergraphs. Background Art

[0002] As an important part of affective computing, facial expression recognition technology can provide support for edge computing scenarios such as medical health, human-computer interaction, and intelligent driving by analyzing the emotional state of humans through computer vision. In the past few decades, the field of facial expression recognition has developed from theoretical exploration to practical application, continuously promoting the progress of intelligent services. However, most of the existing facial expression recognition methods analyze frontal human faces, ignoring the problem of facial deflection that commonly exists in actual scenarios. The deflection of the facial angle will cause the loss of expression information, seriously affecting the recognition accuracy. Therefore, exploring and researching multi-view human face expression recognition has important practical significance, which can not only make up for the limitations of existing methods in practical applications, but also further improve the adaptability and reliability of facial expression recognition technology.

[0003] Different from frontal facial expression recognition, the challenges of multi-view expression recognition mainly focus on two aspects: one is the information loss caused by facial deflection; the other is that the characteristic manifestation forms of the same expression are quite different under different views, increasing the intra-class difference. To solve the above problems, most researchers extract common features under different views or convert lateral human faces into frontal human faces for expression recognition. However, the existing methods ignore the distribution characteristics of facial key points under different views and lack an effective data structure to represent them. Therefore, this patent uses hypergraphs to represent and extract expression features under different views. First, a hypergraph is constructed according to the distribution characteristics of facial key points to capture complex facial expression patterns. Then, a two-stream heterogeneous network is used to extract facial features, that is, hypergraph convolution is used to mine high-order correlation information, and at the same time, a convolutional neural network is used to extract image spatial information. To achieve accurate feature localization, this patent uses the global geometric features obtained by hypergraph convolution to guide image feature localization, focusing on key facial regions. Finally, a cross-attention mechanism is used to deeply fuse these two heterogeneous features of images and geometry to achieve emotion perception under different views. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-view facial expression recognition method based on hypergraphs, which introduces a hypergraph structure to represent expression features under different views, uses it to guide image spatial feature localization, and finally completes the fusion of spatial features and geometric features through a cross-attention mechanism, improving the emotion perception ability under different views.

[0005] For convenience of description, the following concepts are first introduced:

[0006] Hypergraph: A generalization in graph theory that can describe and analyze complex multi - variable relationships. Different from traditional graph structures, each edge in a hypergraph is called a hyper - edge, which allows connecting more than two nodes and can more effectively represent the relationships between entities.

[0007] ResNet50 Network (ResNet50): A convolutional neural network that, through residual connections, deepens the network layers while avoiding the problems of gradient vanishing and gradient explosion.

[0008] Cross - Attention Mechanism: By calculating the correlation weights between the query matrix and the key - value matrix, it dynamically aggregates key information to achieve feature interaction.

[0009] The present invention specifically adopts the following technical solutions:

[0010] A multi - perspective facial expression recognition method based on hypergraph, characterized in that:

[0011] a. Construct a hypergraph through facial key points to represent the topological structure information from different perspectives;

[0012] b. Use a two - stream heterogeneous network to extract facial features. Extract geometric features in facial key points through hypergraph convolution, and extract image - space features through ResNet50;

[0013] c. Use the geometric features of facial key points to guide image feature localization to achieve focused attention on key facial regions;

[0014] d. Through cross - attention, deeply fuse geometric features and space features to achieve the fusion of heterogeneous spaces;

[0015] This method mainly includes the following steps:

[0016] (1) Data pre - processing: Crop and scale the face images in the dataset to 224×224 pixels, and then perform normalization processing, and introduce the random erasing technique to enhance the model's robustness to local feature loss;

[0017] (2) Facial key - point detection: Use the RetinaFace algorithm to detect the face image and obtain the spatial coordinates of facial key points;

[0018] (3) Construct a hypergraph: Utilize the key - point data obtained in step (2) to construct a hypergraph according to different facial regions, and connect different regions with hyper - edges;

[0019] (4) Geometric feature extraction: Use hypergraph convolution to infer and update the hypergraph obtained in step (3) to effectively capture the global features of facial expressions;

[0020] (5) Spatial feature extraction: Use the ResNet50 network to extract spatial features in the facial image and capture rich local emotional details in the facial image;

[0021] (6) Guided feature localization: Use the geometric features obtained by hypergraph convolution in step (4) to guide the localization of the spatial features in step (5), enabling the model to focus on facial regions that can better reflect emotions;

[0022] (7) Feature fusion: There are semantic differences in the feature spaces constructed in steps (4) and (5), so a cross-attention mechanism is used for heterogeneous feature fusion to obtain a unified emotional description;

[0023] (8) Model training: Based on the end-to-end learning framework, use the cross-entropy loss function for gradient backpropagation to optimize the model.

[0024] The beneficial effects of the present invention are:

[0025] (1) Adopt the hypergraph, a high-order association structure, to model the facial key point information. The hypergraph can capture the complex relationships and high-order interactions between key points, providing a richer and more accurate representation of emotional information.

[0026] (2) Use a two-stream heterogeneous branch to extract facial emotion features. The feature encoder based on ResNet50 can capture the spatial information of the image and provide a fine-grained local feature description; the feature encoder based on hypergraph convolution is used to extract the geometric features of facial key points and perceive facial emotions from a global perspective.

[0027] (3) Use the geometric information extracted by hypergraph convolution to guide the learning process of ResNet50, thereby enhancing its feature perception ability and enabling it to focus on regions closely related to emotional expression.

[0028] (4) Adopt the cross-attention method to deeply fuse the spatial features of the image with the geometric features in the facial key point topology, realize the integration of heterogeneous feature spaces, and obtain complete emotion features. Description of the Drawings

[0029] Figure 1 It is a schematic diagram of the hypergraph construction process.

[0030] Figure 2 It is a diagram of the overall model framework.

[0031] Figure 3 It is a structural diagram of the key point feature guidance module.

[0032] Figure 4 It is a structural diagram of the cross-fusion module. Detailed Embodiments

[0033] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is necessary to point out that the following embodiments are only used to further illustrate the present invention and cannot be construed as limiting the protection scope of the present invention. Those skilled in the art can make some non-essential improvements and adjustments to the present invention based on the above-mentioned invention content and implement it specifically, which should still fall within the protection scope of the present invention.

[0034] A multi-view facial expression recognition method based on hypergraphs specifically includes the following steps:

[0035] (1) Data preprocessing

[0036] To ensure the consistency and comparability of the experiment, the resolution of all input images is adjusted to 224×224 pixels. Then, standardization processing is performed, and a random erasing operation is added to enhance the model's perception ability of local feature loss;

[0037] (2) Facial key point detection

[0038] The RetinaFace algorithm is used to detect 68 facial key points of the face image, and these key points precisely outline the facial features of the contour, eyebrows, both eyes, nose, and mouth. To focus on the areas closely related to expression, this chapter selects 51 key points excluding the facial contour, as shown in Figure 1 (a).

[0039] (3) Hypergraph construction

[0040] After obtaining the 51 key point features highly related to expression, the face is divided into six regions: left eyebrow, right eyebrow, left eye, right eye, nose, and mouth, as shown in Figure 1 (b). For each region, the key points in it are combined to form a hyperedge. To further capture the mutual correlation between different regions in emotional expression, two hyperedges are respectively constructed to connect regions 1 to 5 and regions 5 and 6, so as to enhance the information richness and expression ability of the model. The finally constructed hypergraph structure is shown in Figure 1 (c). This structure can effectively capture the complex relationships between each facial region in expression, and comprehensively reflect the feature information within a single region and the interaction between regions.

[0041] (4) Geometric feature extraction

[0042] As shown in Figure 2As shown, hypergraph convolution is used to extract geometric features in facial keypoints. Specifically, first, the information of the nodes associated with each hyperedge is aggregated to the hyperedge to complete the collection of hyperedge information. Then, the features of the nodes are updated using all the hyperedges connected to the nodes to complete the aggregation of node features. Through four stages of hypergraph convolution, geometric features of 2048 dimensions are finally obtained.

[0043] (5)Spatial Feature Extraction

[0044] As Figure 2 shown, a convolutional neural network is used to extract spatial features in the image. Specifically, the ResNet50 network, which also consists of four stages, is used as the backbone network for feature extraction, and the final classification layer is adjusted to 2048 dimensions to obtain the spatial features of the facial image.

[0045] (6)Guided Feature Localization

[0046] As Figure 2 shown, the global geometric features obtained in each stage of hypergraph convolution are used to guide the learning process of ResNet50. The structure of the keypoint feature guidance module is as Figure 3 shown. Specifically, for the obtained geometric features , first, multiple attention heads are obtained using the attention calculation method shown in Equation (1) , and then different attention heads are concatenated through Equation (2) to focus on the characteristics of information from different dimensions and capture richer patterns and dependencies . .

[0047] (1)

[0048] (2)

[0049] Where Q , K and V are 's linear mappings, , , and are trainable parameters, represents the length of the feature vector, represents the normalization exponential function, represents the concatenation operation, and what is shown after where represents 's calculation method. Then, the two-dimensional features are converted into three-dimensional features , and channel-wise convolution and point-wise convolution are used to model the context information from the channel and feature map dimensions respectively to obtain the feature , as shown in Equation (3). Then, the multi-layer perceptron shown in Equation (4) is used Obtain the key point mask . Finally, is passed to the ResNet50 network to guide it to focus on regions more relevant to emotional expression.

[0050] (3)

[0051] (4)

[0052] Among them and represent channel-wise convolution and point-wise convolution respectively, represents matrix multiplication.

[0053] (7) Feature fusion

[0054] As Figure 4 shown, two heterogeneous features are fused through the cross-attention mechanism. Specifically, the features at the last stage of hypergraph convolution are used to generate the corresponding matrices , and . Similarly, the features at the last stage of ResNet50 are used to generate the corresponding matrices , and . Then, interactive perception is achieved through the calculation methods of Equation (5) and Equation (6) to obtain and , where and represent the vector lengths. Finally, the two are concatenated for facial expression classification. This fusion method not only strengthens the model's attention to local details but also improves the accuracy and robustness of emotion recognition by integrating global context information.

[0055] (5)

[0056] (6)

[0057] (8) Model training

[0058] The model is trained in an end-to-end manner, and the cross-entropy loss function is used to optimize the model. The cross-entropy loss function is a commonly used loss function in classification problems and is used to measure the difference between the model's predicted probability distribution and the true label distribution. Its formula is shown in Equation (7):

[0059] (7)

[0060] Among them represents the loss function, yIndicates the true label, Indicates the predicted value of the model, C Indicates the sentiment category involved.

Claims

1. A multi-view facial expression recognition method based on hypergraphs, characterized in that: a. Construct a hypergraph through facial key points to represent the topological structure information under different perspectives; b. Use a two-stream heterogeneous network to extract facial features, extract geometric features in facial key points through hypergraph convolution, and extract image space features through ResNet50; c. Use the geometric features of facial key points to guide image feature localization to achieve focused attention on key facial regions; d. Deeply fuse geometric features and spatial features through cross-attention to achieve the fusion of heterogeneous spaces; (1) Data preprocessing: Crop and scale the face images in the dataset to 224×224 pixels, and then perform standardization processing, and introduce the random erasing technique to enhance the model's robustness to local feature loss; (2) Facial key point detection: Use the RetinaFace algorithm to detect face images to obtain the spatial coordinates of facial key points; (3) Construct a hypergraph: Use the key point data obtained in step (2) to construct a hypergraph according to different facial regions, and use hyperedges to connect different regions; (4) Geometric feature extraction: Use hypergraph convolution to infer and update the hypergraph obtained in step (3) to effectively capture the global features of facial expressions; (5) Spatial feature extraction: Use the ResNet50 network to extract spatial features in facial images to capture rich local emotional details in facial images; (6) Guided feature localization: Use the geometric features obtained by hypergraph convolution in step (4) to guide the localization of spatial features in step (5), so that the model focuses on facial regions that can better reflect emotions; (7) Feature fusion: There are semantic differences in the feature spaces constructed in steps (4) and (5), so use the cross-attention mechanism for heterogeneous feature fusion to obtain a unified emotional description; (8) Model training: Based on an end-to-end learning framework, use the cross-entropy loss function for gradient backpropagation to optimize the model.

2. The multi-view facial expression recognition method based on a hypergraph according to claim 1, wherein After the hypergraph is constructed through steps (2) and (3); the hypergraph is a high-order association structure that can capture the complex relationships and high-order interactions between key points, thereby improving the accuracy of expression recognition.

3. The multi-view facial expression recognition method based on a hypergraph according to claim 1, wherein In steps (4) and (5), geometric features and spatial features are respectively extracted using two-stream heterogeneous branches; The topological graph structure constructed by facial key points can effectively capture the overall characteristics of the face and model expression information by combining global context information; at the same time, the feature encoder based on ResNet can effectively extract rich local emotional details in facial images.

4. The multi-view facial expression recognition method based on a hypergraph according to claim 1, wherein In step (6), geometric features are used to guide spatial feature localization; the feature encoder based on ResNet has an inductive bias and it is difficult to evaluate the importance of features from a global perspective; In contrast, the geometric features extracted by the hypergraph can effectively represent the overall characteristics of the face; Therefore, geometric features can effectively guide the localization of image features.

5. The multi-view facial expression recognition method based on a hypergraph according to claim 1, wherein Heterogeneous feature fusion is completed in step (7); geometric features and spatial features respectively describe global and local sentiment features, and cross-attention can deeply fuse the two to achieve the integration of heterogeneous feature spaces and the extraction of final sentiment features.

Citation Information

Patent Citations

  • Expression recognition method based on key point spatial distribution and local and global feature reasoning

    CN117831088A

  • AIGC video creation method for cultural relic explanation

    CN119516059A