Multi-view facial expression recognition method based on hypergraph

Through the hypergraph structure and cross attention mechanism, the facial expression recognition method is integrated with the hypergraph structure and cross attention mechanism, the problems of information loss and feature differences under multi-view conditions are solved, and high-precision expression recognition effect is achieved.

CN120260101AActive Publication Date: 2025-07-04SICHUAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510726140.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-04
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The existing facial expression recognition methods have information loss and large differences in feature expression under multi-view conditions, resulting in low recognition accuracy and lack of effective data structures to characterize the distribution characteristics of facial key points at different perspectives.

Method used

The hypergraph structure is used to characterize the topological structure of facial key points, and the facial features are extracted in combination with the dual-stream heterogeneous network, and the hypergraph convolution is used to capture higher-order correlation information, and image spatial features are extracted through the ResNet50 network, and feature fusion is performed using the cross attention mechanism.

Benefits of technology

The accuracy and adaptability of expression recognition under multi-view conditions are improved, the model's attention to local details and global context information is enhanced, and the accuracy and robustness of expression recognition are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260101A_ABST
    Figure CN120260101A_ABST
Patent Text Reader

Abstract

The invention provides a multi-view facial expression recognition method based on a hypergraph, and mainly relates to the problem of extracting emotional features in an image in deep learning to perform expression recognition. Firstly, a hypergraph structure is constructed according to spatial distribution characteristics of face key points, and geometric features in the hypergraph structure are highlighted through a more reasonable spatial layout; secondly, designing a key point feature guiding module, and guiding spatial feature positioning of the ResNet50 network by using high-order correlation information obtained by hypergraph convolution, so that the ResNet50 network pays attention to a region which can reflect emotion better in the image; in order to learn uniform emotional features from two heterogeneous feature spaces, a cross fusion module is further designed, and heterogeneous features are integrated by a cross attention mechanism. According to the method, the distribution characteristics of facial expressions at different visual angles are fully considered, geometric features in key points are represented through the hypergraph, and the geometric features are used for guiding image space feature positioning, so that the adaptability of an expression recognition technology in practical application is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the problem of multi-view facial expression recognition in the field of deep learning, and particularly to a multi-view facial expression recognition method based on hypergraphs. Background Art

[0002] As an important part of affective computing, facial expression recognition technology can provide support for edge computing scenarios such as medical health, human-computer interaction, and intelligent driving by analyzing human emotional states through computer vision. In the past few decades, the field of facial expression recognition has developed from theoretical exploration to practical application, continuously promoting the progress of intelligent services. However, most of the existing facial expression recognition methods analyze frontal human faces, ignoring the problem of facial deflection that commonly exists in actual scenarios. The deflection of the facial angle will lead to the loss of expression information, seriously affecting the recognition accuracy. Therefore, exploring and researching multi-view human face expression recognition has important practical significance, which can not only make up for the limitations of existing methods in practical applications, but also further improve the adaptability and reliability of facial expression recognition technology.

[0003] Different from frontal facial expression recognition, the challenges of multi-view expression recognition mainly focus on two aspects: one is the information loss caused by facial deflection; the other is that the characteristic manifestation forms of the same expression are quite different under different views, increasing the intra-class difference. To solve the above problems, most researchers extract common features under different views or convert lateral human faces into frontal human faces for expression recognition. However, the existing methods ignore the distribution characteristics of facial key points under different views and lack an effective data structure to represent them. Therefore, this patent uses hypergraphs to represent and extract expression features under different views. First, a hypergraph is constructed according to the distribution characteristics of facial key points to capture complex human face expression patterns. Then, a two-stream heterogeneous network is used to extract facial features, that is, hypergraph convolution is used to mine high-order correlation information, and at the same time, a convolutional neural network is used to extract image spatial information. To achieve accurate feature localization, this patent uses the global geometric features obtained by hypergraph convolution to guide image feature localization, focusing on key facial regions. Finally, a cross-attention mechanism is used to deeply fuse these two heterogeneous features of images and geometry to achieve expression perception under different views. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-view facial expression recognition method based on hypergraphs, which introduces a hypergraph structure to represent expression features under different views, uses it to guide image spatial feature localization, and finally completes the fusion of spatial features and geometric features through a cross-attention mechanism, improving the expression perception ability under different views.

[0005] For the convenience of description, the following concepts are introduced first: Hypergraph: A generalization in graph theory that can describe and analyze complex multi - variable relationships. Different from traditional graph structures, each edge in a hypergraph is called a hyper - edge, which allows connecting more than two nodes and can represent the relationships between entities more effectively.

[0006] ResNet50 Network (ResNet50): A convolutional neural network that deepens the network layers through residual connections while avoiding the problems of gradient vanishing and gradient explosion.

[0007] Cross - Attention Mechanism: It dynamically aggregates key information to achieve feature interaction by calculating the correlation weights between the query matrix and the key - value matrix.

[0008] The present invention specifically adopts the following technical solutions: A multi - perspective facial expression recognition method based on hypergraphs, characterized in that: a. Construct a hypergraph through facial key points to represent the topological structure information under different perspectives; b. Use a two - stream heterogeneous network to extract facial features. Extract geometric features in facial key points through hypergraph convolution, and extract image - space features through ResNet50; c. Use the geometric features of facial key points to guide image feature localization to achieve focused attention on key facial regions; d. Deeply fuse geometric features and spatial features through cross - attention to achieve the fusion of heterogeneous spaces; This method mainly includes the following steps: (1) Data pre - processing: Crop and scale the face images in the dataset to 224×224 pixels, and then perform normalization processing, and introduce the random erasing technique to enhance the model's robustness to local feature loss; (2) Facial key - point detection: Use the RetinaFace algorithm to detect the face image to obtain the spatial coordinates of facial key points; (3) Hypergraph construction: Utilize the key - point data obtained in step (2) to construct a hypergraph according to different facial regions, and connect different regions with hyper - edges; (4) Geometric feature extraction: Use hypergraph convolution to infer and update the hypergraph obtained in step (3) to effectively capture the global features of facial expressions; (5) Spatial feature extraction: Use the ResNet50 network to extract spatial features in the facial image to capture rich local emotional details in the facial image; (6) Guided feature localization: Use the geometric features obtained from hypergraph convolution in step (4) to guide the localization of spatial features in step (5), enabling the model to focus on facial regions that can better reflect emotions. (7) Feature fusion: There are semantic differences in the feature spaces constructed in steps (4) and (5), so a cross-attention mechanism is used for heterogeneous feature fusion to obtain a unified emotional description. (8) Model training: Based on an end-to-end learning framework, use the cross-entropy loss function for gradient backpropagation to optimize the model.

[0009] The beneficial effects of the present invention are as follows: (1) Use a hypergraph, a high-order association structure, to model facial key point information. The hypergraph can capture the complex relationships and high-order interactions between key points, providing a richer and more accurate representation of emotional information.

[0010] (2) Use a two-stream heterogeneous branch to extract facial emotion features. The feature encoder based on ResNet50 can capture the spatial information of the image and provide a fine-grained local feature description; the feature encoder based on hypergraph convolution is used to extract the geometric features of facial key points and perceive facial emotions from a global perspective.

[0011] (3) Use the geometric information extracted by hypergraph convolution to guide the learning process of ResNet50, thereby enhancing its feature perception ability and enabling it to focus on regions closely related to emotional expression.

[0012] (4) Deeply fuse the spatial features of the image with the geometric features in the facial key point topological structure in a cross-attention manner to integrate heterogeneous feature spaces and obtain complete emotion features. Brief Description of the Drawings

[0013] Figure 1 It is a schematic diagram of the hypergraph construction process.

[0014] Figure 2 It is a diagram of the overall model framework.

[0015] Figure 3 It is a structural diagram of the key point feature guidance module.

[0016] Figure 4 It is a structural diagram of the cross-fusion module. Detailed Embodiments

[0017] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is necessary to point out that the following embodiments are only used to further illustrate the present invention and should not be construed as limiting the protection scope of the present invention. Those skilled in the art can make some non-essential improvements and adjustments to the present invention according to the above-mentioned invention content and still fall within the protection scope of the present invention.

[0018] A multi-view facial expression recognition method based on a hypergraph specifically includes the following steps: (1) Data preprocessing To ensure the consistency and comparability of the experiments, the resolution of all input images is adjusted to 224×224 pixels. Then, normalization processing is performed, and a random erasing operation is added to enhance the model's perception ability of local feature loss; (2) Facial key point detection The RetinaFace algorithm is used to detect 68 facial key points of the face image, and these key points precisely outline the facial features of the contour, eyebrows, both eyes, nose, and mouth. To focus on the areas closely related to expression, this chapter selects 51 key points excluding the facial contour, as Figure 1 shown in (a).

[0019] (3) Hypergraph construction After obtaining the 51 key point features highly related to expression, the face is divided into six regions: left eyebrow, right eyebrow, left eye, right eye, nose, and mouth, as Figure 1 shown in (b). For each region, the key points in it are combined to form a hyperedge. To further capture the mutual correlation between different regions in emotional expression, two hyperedges are respectively constructed to connect regions 1 to 5 and regions 5 and 6, so as to enhance the information richness and expression ability of the model. The finally constructed hypergraph structure is as Figure 1 shown in (c). This structure can effectively capture the complex relationships between various facial regions in expression, and comprehensively reflect the feature information within a single region and the interaction between regions.

[0020] (4) Geometric feature extraction As Figure 2 shown, hypergraph convolution is used to extract geometric features from facial key points. Specifically, first, the information of the nodes associated with each hyperedge is aggregated to the hyperedge to complete hyperedge information collection. Then, the node features are updated using all the hyperedges connected to the nodes to complete node feature aggregation. Through 4 stages of hypergraph convolution, 2048-dimensional geometric features are finally obtained.

[0021] (5) Spatial feature extraction As Figure 2As shown, a convolutional neural network is used to extract spatial features from an image. Specifically, the ResNet50 network, which also consists of 4 stages, is used as the backbone network for feature extraction, and the final classification layer is adjusted to 2048 dimensions to obtain the spatial features of the facial image.

[0022] (6)Guiding Feature Localization As Figure 2 shown, the global geometric features obtained in each hypergraph convolution stage are used to guide the learning process of ResNet50. The structure of the key point feature guiding module is as Figure 3 shown. Specifically, for the obtained geometric features , first, the attention calculation method shown in Equation (1) is used to obtain multiple attention heads , and then different attention heads are concatenated through Equation (2) to focus on the characteristics of information from different dimensions and capture richer patterns and dependencies .

[0023] (1) (2) where Q 、 K and V are 's linear mappings, 、 、 and are trainable parameters, represents the length of the feature vector, represents the normalization exponential function, represents the concatenation operation, and what follows where represents 's calculation method. Then, the two-dimensional features are converted into three-dimensional features , and depthwise convolution and pointwise convolution are used to model the context information from the channel and feature map dimensions respectively to obtain the features , as shown in Equation (3). Then, the multi-layer perceptron shown in Equation (4) is used to obtain the key point mask . Finally, is passed to the ResNet50 network to guide it to focus on the regions more relevant to emotional expression.

[0024] (3) (4) where and represent depthwise convolution and pointwise convolution respectively, represents matrix multiplication.

[0025] (7)Feature fusion As Figure 4 shown, two heterogeneous features are fused through the cross-attention mechanism. Specifically, the features of the last stage of hypergraph convolution are used to generate the corresponding matrices , and . Similarly, the features of the last stage of ResNet50 are used to generate the corresponding matrices , and . Then, interactive perception is achieved through the calculation methods of equations (5) and (6) to obtain and , where and represent the vector lengths. Finally, the two are concatenated for facial expression classification. This fusion method not only strengthens the model's attention to local details but also improves the accuracy and robustness of emotion recognition by integrating global context information.

[0026] (5) (6) (8)Model training The model is trained in an end-to-end manner, and the cross-entropy loss function is used to optimize the model. The cross-entropy loss function is a commonly used loss function in classification problems and is used to measure the difference between the model's predicted probability distribution and the true label distribution. Its formula is shown in equation (7): (7) where represents the loss function,[[]] y represents the true label,[[]] represents the predicted value of the model,[[]] C represents the involved emotion categories.

Claims

1. A multi-view facial expression recognition method based on hypergraph, characterized in that: a. Construct a hypergraph through facial key points to represent the topological structure information under different views; b. Use a two-stream heterogeneous network to extract facial features, extract geometric features in facial key points through hypergraph convolution, and extract image space features through ResNet50; c. Use the geometric features of facial key points to guide image feature localization to achieve focused attention on key facial regions; d. Achieve the fusion of heterogeneous spaces by deeply fusing geometric features and space features through cross-attention; (1) Data preprocessing: Crop and scale the face images in the dataset to 224×224 pixels, and then perform normalization processing, and introduce the random erasing technique to enhance the model's robustness to local feature loss; (2) Facial key point detection: Use the RetinaFace algorithm to detect face images to obtain the spatial coordinates of facial key points; (3) Construct a hypergraph: Use the key point data obtained in step (2) to construct a hypergraph according to different facial regions, and use hyperedges to connect different regions; (4) Geometric feature extraction: Use hypergraph convolution to infer and update the hypergraph obtained in step (3) to effectively capture the global features of facial expressions; (5) Spatial feature extraction: Use the ResNet50 network to extract spatial features in facial images to capture rich local emotional details in facial images; (6) Guided feature localization: Use the geometric features obtained by hypergraph convolution in step (4) to guide the localization of spatial features in step (5) to make the model focus on facial regions that can better reflect emotions; (7) Feature fusion: There are semantic differences in the feature spaces constructed in steps (4) and (5), so use the cross-attention mechanism for heterogeneous feature fusion to obtain a unified emotional description; (8) Model training: Based on the end-to-end learning framework, use the cross-entropy loss function for gradient backpropagation to achieve model optimization.

2. The multi-view facial expression recognition method based on a hypergraph according to claim 1, characterized in that After the construction of the hypergraph is completed through steps (2) and (3); the hypergraph is a high-order association structure that can capture the complex relationships and high-order interactions between key points, thereby improving the accuracy of expression recognition.

3. The multi-view facial expression recognition method based on a hypergraph according to claim 1, characterized in that In steps (4) and (5), geometric features and spatial features are respectively extracted using two-stream heterogeneous branches; The topological graph structure constructed by facial key points can effectively capture the overall characteristics of the face and model expression information by combining global context information; at the same time, the feature encoder based on ResNet can effectively extract rich local emotional details in facial images.

4. The multi-view facial expression recognition method based on a hypergraph according to claim 1, characterized in that In step (6), geometric features are used to guide spatial feature localization; the feature encoder based on ResNet has an inductive bias and it is difficult to evaluate the importance of features from a global perspective; In contrast, the geometric features extracted by the hypergraph can effectively represent the overall characteristics of the face; Therefore, geometric features can effectively guide the localization of image features.

5. The multi-view facial expression recognition method based on a hypergraph according to claim 1, characterized in that Heterogeneous feature fusion is completed in step (7); geometric features and spatial features respectively describe global and local sentiment features, and cross-attention can be used to deeply fuse the two, realizing the integration of heterogeneous feature spaces and the extraction of final sentiment features.

Citation Information

Patent Citations

  • Expression recognition method based on key point spatial distribution and local and global feature reasoning

    CN117831088A

  • AIGC video creation method for cultural relic explanation

    CN119516059A

  • Multi-modal hypergraph emotion recognition method for video learning scene

    CN119810887A

  • Multi-modal cognitive impairment evaluation system based on multi-dimensional cognitive function hypergraph

    CN119818025A

  • Method and system for multimodal emotion recognition in conversation (ERC) based on graph neural network (GNN)

    US20240355350A1