Facial expression recognition system and method based on multi-dimensional mixed attention mechanism

Through the MobileNet V3 neural network based on feature pyramids and hybrid attention mechanisms, the problem of low recognition rate and high model complexity of facial expression recognition technology in dynamic environments is solved, real-time monitoring of emotional states in online education and optimization of teaching strategies, and learning efficiency and participation are improved.

CN120452042APending Publication Date: 2025-08-08SHANDONG MANAGEMENT UNIV

Patent Information

Application Number
CN202510591845.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing facial expression recognition technology has low recognition rate under natural conditions, and factors such as head posture offset, lighting changes, occlusion and motion blur are greatly affected, and the model is highly complex, making it difficult to effectively deploy on edge computing devices.

Method used

The MobileNet V3 neural network based on feature pyramids and hybrid attention mechanisms is adopted, combining a hybrid attention mechanism of frequency domain time domain fusion and spatial channel fusion, multi-scale features of facial images are extracted, and the accuracy and real-timeness of expression recognition are improved by analyzing the frequency domain energy distribution and spatial information.

Benefits of technology

In a dynamic learning environment, higher accuracy and real-time recognition of facial expressions are achieved, supporting real-time monitoring of emotional states and adjustment of teaching strategies in online education, and improving learning efficiency and participation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452042A_ABST
    Figure CN120452042A_ABST
Patent Text Reader

Abstract

The invention provides a facial expression recognition system and method based on a multi-dimensional mixed attention mechanism, and relates to the technical field of facial data processing, and the method comprises the steps: obtaining a human face image; inputting the preprocessed facial image into a facial expression recognition model, and outputting to obtain a facial expression category; wherein the facial expression recognition model is an enhanced MobileNet V3 neural network based on a feature pyramid and a mixed attention mechanism, after a facial image is subjected to depth separable convolution operation, multi-scale features of the facial image are extracted through the feature pyramid network, high-level semantic features and low-level detail features are obtained and fused, and the facial expression recognition model is obtained. And through a mixed attention mechanism of frequency domain and time domain fusion and space channel fusion, guiding a model to pay attention to effective features in an image by analyzing energy distribution in a frequency domain and extracting space information and channel information, obtaining a final feature vector, and realizing facial expression classification based on the final feature vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of facial data processing, and in particular to a facial expression recognition system and method based on a multi-dimensional hybrid attention mechanism. Background Art

[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.

[0003] With the widespread application of information technology in education, learning methods continue to expand and develop. From traditional face-to-face teaching to today's digital learning platforms, live streaming, and other diverse models, education increasingly relies on technology to improve teaching quality and effectiveness. At the same time, accurate monitoring of learners' learning status has become a key area of educational technology research. Within the educational implementation process, leveraging technologies such as computer vision to gain a deeper understanding of students' learning progress has become a promising approach. Facial expression recognition, as a key component, has garnered widespread attention.

[0004] Facial expression recognition is an important research direction in the field of computer vision and artificial intelligence, which involves multiple disciplines such as psychology, biology and computer science.

[0005] In online learning, the interaction between teachers and students is limited, making it difficult to evaluate learning outcomes using traditional methods. Therefore, it is crucial to use facial expression recognition technology to assess students' emotional states during learning and, therefore, to evaluate learning outcomes. Traditional facial expression recognition technology primarily relies on manual feature extraction, such as local binary patterns (LBP), Gabor feature extraction, and active shape models (ASM). These methods are typically used with small datasets and lack real-time and accuracy, especially in dynamically changing learning environments. With technological advancements, facial expression recognition technology has begun to move towards deep learning, leveraging techniques such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs) to extract the spatial and temporal features of facial expressions.

[0006] Although existing facial expression recognition technology has achieved certain results in experimental environments, it still faces many challenges in practical applications, especially in the field of education. The following technical issues exist: 1) Under natural conditions, such as in online learning, factors such as head posture deviation, lighting changes, occlusion, and motion blur have a significant impact on recognition rate.

[0007] 2) Most existing recognition technologies do not pay attention to the changes in facial expressions in the temporal and spatial dimensions, and have certain application limitations.

[0008] 3) The existing technical model is highly complex. In distributed application scenarios, it is limited by edge computing capabilities and has certain limitations on deployment and implementation. Summary of the Invention

[0009] In order to solve the above problems, the present disclosure proposes a facial expression recognition system and method based on a multi-dimensional hybrid attention mechanism, which uses an enhanced lightweight model MobileNet network and combines it with FPN to obtain feature information at different levels of facial data. Through a hybrid attention mechanism of frequency domain and time domain fusion and spatial channel fusion, feature vector information is obtained in different dimensions of frequency domain, time domain and spatial domain, so as to more effectively extract and analyze facial expression features to adapt to the dynamically changing learning environment, improve the accuracy and real-time performance of emotion recognition in the learning process, and thus better serve the field of education.

[0010] According to some embodiments, the present disclosure adopts the following technical solutions: Facial expression recognition method based on multi-dimensional hybrid attention mechanism, including: Acquire a facial image of a human face and preprocess the facial image; The preprocessed facial image is input into the facial expression recognition model, and the facial expression category is output; Among them, the facial expression recognition model is a MobileNetV3 neural network based on feature pyramid and hybrid attention mechanism. After the facial image is input, the facial image undergoes a depth-separable convolution operation, and then the feature pyramid network extracts the multi-scale features of the facial image, obtains high-level semantic features and low-level detail features and fuses them, and the fused features are then passed through a hybrid attention mechanism of frequency domain and time domain fusion and spatial channel fusion. By analyzing the energy distribution in the frequency domain and considering spatial information and channel information, the model is guided to focus on the effective features in the image to obtain the final feature vector, and facial expression classification is realized based on the final feature vector.

[0011] According to some embodiments, the present disclosure adopts the following technical solutions: Facial expression recognition system based on multi-dimensional hybrid attention mechanism, including: A data acquisition module is used to acquire facial images of human faces and pre-process the facial images; The expression recognition module is used to input the pre-processed facial image into the facial expression recognition model and output the facial expression category; Among them, the facial expression recognition model is a MobileNetV3 neural network based on feature pyramid and hybrid attention mechanism. After the facial image is input, the facial image undergoes a depth-separable convolution operation, and then the feature pyramid network extracts the multi-scale features of the facial image, obtains high-level semantic features and low-level detail features and fuses them, and the fused features are then passed through a hybrid attention mechanism of frequency domain and time domain fusion and spatial channel fusion. By analyzing the energy distribution in the frequency domain and considering spatial information and channel information, the model is guided to focus on the effective features in the image to obtain the final feature vector, and facial expression classification is realized based on the final feature vector.

[0012] According to some embodiments, the present disclosure adopts the following technical solutions: A computer program product includes a computer program, which, when executed by a processor, implements the facial expression recognition method based on a multi-dimensional hybrid attention mechanism.

[0013] According to some embodiments, the present disclosure adopts the following technical solutions: A non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the facial expression recognition method based on the multi-dimensional hybrid attention mechanism is implemented.

[0014] According to some embodiments, the present disclosure adopts the following technical solutions: An electronic device comprises: a processor, a memory, and a computer program; wherein the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the facial expression recognition method based on the multi-dimensional hybrid attention mechanism.

[0015] Compared with the prior art, the present invention has the following beneficial effects: The present invention discloses a facial expression recognition method based on a multi-dimensional hybrid attention mechanism. Facial images are processed through a MobileNet V3 neural network based on a feature pyramid and a hybrid attention mechanism. Feature information at different levels of facial data is acquired using an enhanced lightweight MobileNet network in combination with FPN. Feature vector information is acquired in different dimensions of the frequency domain, time domain, and spatial domain through a hybrid attention mechanism of frequency domain and time domain fusion and spatial channel fusion. Facial expression features are extracted and analyzed more effectively, facial expression classification is achieved, and students' emotional states are monitored in real time. Teachers can adjust teaching strategies in a timely manner, such as the speed of explanation and the difficulty of teaching content, to adapt to students' emotions and comprehension ability, thereby improving learning efficiency.

[0016] The present invention discloses a facial expression recognition method based on a multi-dimensional hybrid attention mechanism. Facial expression recognition technology can provide personalized learning suggestions for online learning platforms, adjust teaching content according to students' emotional feedback, make the learning process more in line with students' needs and interests, and enhance the learning experience. In online learning, teachers may not be able to directly observe students' emotional reactions. Facial expression recognition technology can help teachers understand students' emotional states, promote the implementation of emotional education, and improve students' participation and enthusiasm.

[0017] The present invention discloses a facial expression recognition method based on a multi-dimensional hybrid attention mechanism. By analyzing students' facial expressions, the education platform can better understand which teaching content or methods are more effective, thereby optimizing the allocation and utilization of teaching resources and improving the quality of teaching. It can also support remote teaching. In a remote teaching environment, facial expression recognition technology can help teachers overcome geographical barriers, better understand and respond to students' emotional needs, adapt to the dynamically changing learning environment, and improve the accuracy and real-time nature of emotion recognition during the learning process, thereby better serving the field of online education. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which constitute a part of the present disclosure, are used to provide a further understanding of the present disclosure. The exemplary embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation to the present disclosure.

[0019] Figure 1 This is a flow chart of a facial expression recognition method based on a multi-dimensional hybrid attention mechanism according to an embodiment of the present disclosure; Figure 2 Schematic diagram of the MobileNet V3 network structure of an embodiment of the present disclosure; Figure 3 Schematic diagram of the partial structure of the attention mechanism of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0020] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.

[0021] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present disclosure belongs.

[0022] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0023] Example 1 In one embodiment of the present disclosure, a facial expression recognition method based on a multi-dimensional hybrid attention mechanism is provided. Based on transfer learning, an enhanced lightweight model MobileNet network is used in combination with FPN to obtain feature information at different levels. A hybrid attention mechanism of frequency domain and time domain fusion and spatial channel fusion is used to obtain feature vector information in different dimensions such as frequency domain, time domain, and spatial domain. By testing the model, facial expression features can be more effectively extracted and analyzed. The steps include: Step 1: Obtain a facial image and preprocess the facial image; Step 2: Input the preprocessed facial image into the facial expression recognition model and output the facial expression category; Among them, the facial expression recognition model is a MobileNetV3 neural network based on feature pyramid and hybrid attention mechanism. When the facial image is input, the facial image undergoes a depth-separable convolution operation, and then passes through the feature pyramid network to extract the multi-scale features of the facial image, obtain high-level semantic features and low-level detail features and fuse them, and the fused features are then passed through a hybrid attention mechanism of frequency domain and time domain fusion and spatial channel fusion. By analyzing the energy distribution in the frequency domain, the spatial information and channel information are extracted, and the model is guided to focus on the effective features in the image to obtain the final feature vector, and facial expression classification is realized based on the final feature vector.

[0024] As an embodiment, the specific implementation process of the facial expression recognition method based on the multi-dimensional hybrid attention mechanism disclosed in the present invention is as follows: Step 1: Obtain a facial image and preprocess the facial image; Specifically, 1) obtaining facial images of human faces, including: collecting video streams, dividing the video streams into continuous image frames, and the acquisition frequency of the video images can be set according to actual conditions to ensure that sufficient expression change information can be obtained without affecting the learning experience, and obtaining facial images from the image frames for subsequent recognition processing.

[0025] 2) Preprocess the facial image, including: (1) Format normalization. The extracted facial image format is uniformly converted to a standard format suitable for subsequent processing (such as RGB format). The image resolution is also standardized to ensure the consistency of the image data quality and size input to subsequent modules.

[0026] (2) Posture Correction. 2D posture correction based on geometric transformation is used to detect facial key points and adjust the face to a standard posture that approximates the frontal face. Images with large posture deviations that cannot be accurately corrected can be marked as abnormal data and handled appropriately in subsequent processing (e.g., reminding students to adjust their posture during class).

[0027] (3) Illumination normalization. Perform illumination correction on facial images to reduce the impact of different lighting conditions on expression recognition. A histogram equalization method is used to adjust the pixel value distribution of the image to make the overall brightness and contrast of the image more uniform, thereby enhancing the recognizability of facial features.

[0028] (4) Image cropping and scaling. The image is cropped based on the facial region determined by the facial key points to remove unnecessary background information. The cropped facial image is then scaled to a fixed size suitable for input into the feature extraction network.

[0029] Step 2: Input the preprocessed facial image into the facial expression recognition model and output the facial expression category; Specifically, the facial expression recognition model is an enhanced MobileNet V3 neural network based on feature pyramid and hybrid attention mechanism. MobileNet V3 reduces the number of network parameters and computational complexity by introducing lightweight depth-wise separable convolutions while maintaining a high accuracy rate. The enhanced MobileNet V3 neural network based on feature pyramid and hybrid attention mechanism uses the MobileNet V3 network as the basic network architecture. It has efficient computing performance and good feature representation capabilities, and is suitable for use in scenarios such as online learning that have high real-time requirements. First, the parameter weights learned by the MobileNetV3 network on ImageNet are initialized using transfer learning. Then, a self-built facial expression dataset is used to train MobileNet V3, and the feature pyramid and hybrid attention mechanism are used to enhance the learning effect, thereby obtaining the common features of facial expressions during the model training process.

[0030] Furthermore, a feature pyramid is constructed based on the MobileNet V3 network. This feature pyramid effectively integrates feature information at different scales, addressing feature matching issues in facial expression recognition caused by factors such as varying facial size and expression amplitude. Feature maps are extracted at different layers of MobileNet V3 to obtain multi-scale features consisting of high-level semantic features and low-level detailed features. Then, through upsampling and downsampling operations and lateral connections, the multi-scale features of the feature pyramid are fused to form feature information rich in semantic information and detailed information.

[0031] Furthermore, the SE attention mechanism is used in the MobileNet V3 network. This attention mechanism focuses more on the dependency between channels and lacks the enhancement of frequency domain features and the processing of spatial position information. Therefore, in the present disclosure, the attention mechanism of the MobileNet V3 network is feature-enhanced and improved to a hybrid attention mechanism that integrates frequency domain and time domain, and spatial channel, to obtain feature vector information in different dimensions such as frequency domain, time domain, and spatial domain. The hybrid attention mechanism determines the importance of different frequency components to expression recognition by analyzing the energy distribution in the frequency domain. Higher weights are given to important frequency components to highlight expression-related features and suppress noise and irrelevant information. This approach can more effectively capture subtle changes in facial expressions, especially when the expression intensity is low or there is a certain amount of interference, thereby improving the accuracy of expression recognition. At the same time, the spatial information and channel information of the comprehensive input feature layer are also considered, so as to more effectively guide the network model to focus on the effective features in the image.

[0032] As an example, after a facial image is input into the MobileNet V3 neural network based on the feature pyramid and hybrid attention mechanism, Figure 2 As shown in the figure, after the facial image undergoes a depthwise separable convolution operation, it is passed through a feature pyramid network to extract multi-scale features of the facial image, obtain high-level semantic features and low-level detail features and fuse them. The fused features are then passed through a hybrid attention mechanism of frequency domain and time domain fusion and spatial channel fusion. By analyzing the energy distribution in the frequency domain and considering spatial information and channel information, the model is guided to focus on the effective features in the image to obtain the final feature vector, and facial expression classification is achieved based on the final feature vector.

[0033] Furthermore, after the depthwise separable convolution operation, the multi-scale features of the facial image are extracted through the feature pyramid network, and high-level semantic features and low-level detail features are obtained and fused.

[0034] Furthermore, the fused features are then processed through a hybrid attention mechanism combining frequency-domain and time-domain fusion and spatial-channel fusion. By analyzing the energy distribution in the frequency domain and taking into account spatial and channel information, the model is guided to focus on valid features in the image, resulting in the final feature vector. This includes acquiring feature vector information in different dimensions, such as the frequency, time, and spatial domains, through a hybrid attention mechanism combining frequency-domain and time-domain fusion and spatial-channel fusion. By analyzing the energy distribution in the frequency domain, the importance of different frequency components for expression recognition is determined. Higher weights are assigned to important frequency components, highlighting expression-related features and suppressing noise and irrelevant information. This approach can more effectively capture subtle changes in facial expressions, especially when the expression intensity is low or there is some interference, thereby improving the accuracy of expression recognition.

[0035] Specifically, if Figure 3 As shown in Figure 3, the hybrid attention mechanism mainly consists of three parts: (1) dual-pooled residual channel attention, which focuses more on channel information; (2) depthwise separable spatial attention, which focuses more on spatial information; and (3) block-wise frequency domain attention, which focuses more on frequency domain information.

[0036] (1) Double-pooled residual channel attention, performing a dual-channel pooling operation on the input feature information. The channel description vector obtained by global average pooling (GAP) captures the global statistical characteristics of the channel but is insensitive to spatial distribution; The channel description vector obtained by global maximum pooling (GMP) focuses on the local salient features of the channel and may ignore the overall distribution.

[0037]

[0038]

[0039] right and Splicing is performed along the channel direction, while modeling the global statistical characteristics and local significance of the channel to improve the robustness of feature expression.

[0040] (C is the number of channels) The channel dimension after splicing is 2C, and the dimension is reduced by the fully connected layer FC to reduce the target dimension: (r is the reduction rate, usually r=16), the weight matrix Dimensions: (Number of rows: target dimension , number of columns: input dimension 2C). Perform matrix multiplication:

[0041]

[0042] Activation function: δ (⋅) Use Leaky ReLU to enhance its nonlinearity: δ ( x ) ( =0.01) Dimensionality reduction can reduce the number of parameters and prevent overfitting. The second layer is fully connected Change the dimension from Restore to C , the formula is: , the final output is:

[0043] in, is the Sigmoid function, which constrains the weights between (0,1); is the residual term, which is used to preserve the distribution characteristics of the original features, where λ Is a trainable scalar parameter that can be automatically optimized during back propagation. The formula is:

[0044] The weight of the double-pooled residual channel attention after the entire processing: .

[0045] Among them, Diag( w ): weight vector Convert to a diagonal matrix to achieve independent scaling of each channel: ,and is an identity matrix multiplied by the learnable parameters ,Increase It can be used as a residual connection to keep the original feature information from being completely covered, and can be used To adjust the strength of the residual part and enhance the stability and expressiveness of the model.

[0046] (2) Deep Separable Spatial Attention, Spatial Feature Encoding ,in is average pooling, is the maximum pooling, Denotes splicing; use depthwise separable convolution (DWConv) to reduce the number of parameters, and the convolution kernel size is selected as 3×3 to balance the receptive field and the amount of computation. Perform GroupNorm (group normalization).

[0047] On mobile devices or lightweight models, BatchNorm is unstable when using small batches, so GroupNorm is used here to divide the channel dimension into G groups (currently using G=8), and calculate the mean and variance independently within each group. After normalization enhancement: Q , K =GroupNorm( Ms ) (divided into G The specific approach is: first divide the channel dimension C into G groups (G=8), with the number of channels in each group being C / G; then calculate the statistics of the channels within the group at each spatial position (h, w):

[0048]

[0049] Normalization and affine transformation:

[0050] in, , are learnable parameters.

[0051] The spatial attention weight after normalization enhancement is calculated as: (Q, K come from the input feature M s is the group normalization result of , and B is the learnable bias).

[0052] (3) Block frequency domain attention. In order to balance the frequency domain resolution and computational complexity, the block size is 8×8, and the DCT transform is performed on the block:

[0053] Among them, x and y are spatial coordinates, indicating the pixel position within the block; u and v are frequency domain coordinates, indicating the index of the frequency component.

[0054] In facial expression recognition, high-frequency components contain a lot of detailed information (such as wrinkles at the corners of the eyes, curvature of the mouth corners, etc.), and are sensitive to micro-expressions. However, high-frequency components are also prone to noise; low-frequency components represent the overall contour and brightness distribution, and are more critical to the overall category of expression (such as laughter and anger). At the same time, low frequencies may lose important details (such as subtle muscle movements). Therefore, frequency band energy weighting is used to measure the amount of information in each frequency band by calculating the energy of different frequency bands. The specific formula is: , MLP (Multi-layer Perceptron) can map the band energy to weights, thereby achieving adaptive band selection. Finally, it is subjected to inverse DCT transform, and its formula is:

[0055] Finally, the three parts of channel-space-frequency domain are jointly modeled to construct a composite attention mechanism and perform adaptive weight fusion:

[0056] The weight generation is:

[0057] Among them Represents the concatenation of global average pooling and global maximum pooling, weight matrix (3 corresponds to the channel, space and frequency domains), and the three dimensions are obtained by normalizing the Softmax function 、 and .

[0058] Composite loss function:

[0059] in, is the cross entropy loss, is the fusion weight The L1 regularization term, It is the attention module of each domain The gradient norm constraint of .

[0060] After processing through the feature pyramid and the improved attention mechanism of MobileNet V3, the final feature vector is obtained. This feature vector is input into the fully connected layer and classifier to classify facial expressions. To meet the needs of online learning, the expression categories include common expressions such as fatigue, disgust, confusion, excitement, concentration, and neutrality.

[0061] Furthermore, the present disclosure provides feedback and adjusts teaching strategies based on the recognized facial expression information, including: (1) Facial expression information statistics and analysis. Real-time collection and statistics of students’ facial expression information, including the frequency and duration of different expressions. By analyzing the data, the trend of students’ emotions throughout the learning process can be obtained. For example, it can be determined whether students show more confused expressions during the explanation of a certain knowledge point, or whether they show fatigue or boredom after a long period of study.

[0062] (2) Adjustment of teaching strategies. Based on the analysis results of facial expression information, the teaching strategies of the online learning platform are adjusted. If students frequently show confused expressions, additional explanatory materials for relevant knowledge points can be pushed to them, such as animated demonstrations, detailed text descriptions, etc. If students show signs of fatigue or boredom, some relaxing and interesting interactive links can be inserted in a timely manner, such as mini-games, fun quizzes, etc., to re-stimulate students' interest in learning. At the same time, teachers can also understand the emotional state of students in the entire class through background data and make macro adjustments to the teaching content and progress, such as re-explaining knowledge points that are generally difficult to understand or slowing down the teaching progress.

[0063] Example 2 In one embodiment of the present disclosure, a facial expression recognition system based on a multi-dimensional hybrid attention mechanism is provided, comprising: A data acquisition module is used to acquire facial images of human faces and pre-process the facial images; The expression recognition module is used to input the pre-processed facial image into the facial expression recognition model and output the facial expression category; Among them, the facial expression recognition model is an enhanced MobileNet V3 neural network based on feature pyramid and hybrid attention mechanism. When the facial image is input, the facial image undergoes a depth-separable convolution operation, and then passes through the feature pyramid network to extract the multi-scale features of the facial image, obtain high-level semantic features and low-level detail features and fuse them. The fused features are then passed through a hybrid attention mechanism of frequency domain and time domain fusion and spatial channel fusion. By analyzing the energy distribution in the frequency domain, spatial information and channel information are extracted, and the model is guided to focus on the effective features in the image to obtain the final feature vector, and facial expression classification is realized based on the final feature vector.

[0064] Example 3 In one embodiment of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the facial expression recognition method based on the multi-dimensional hybrid attention mechanism.

[0065] Example 4 In one embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, which is used to store computer instructions. When the computer instructions are executed by a processor, the facial expression recognition method based on the multi-dimensional hybrid attention mechanism is implemented.

[0066] Example 5 In one embodiment of the present disclosure, an electronic device is provided, comprising: a processor, a memory, and a computer program; wherein the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the facial expression recognition method based on the multi-dimensional hybrid attention mechanism.

[0067] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0068] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0069] Although the above describes the specific implementation methods of the present disclosure in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present disclosure. Those skilled in the art should understand that on the basis of the technical solution of the present disclosure, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present disclosure.

Claims

1. A facial expression recognition method based on a multi-dimensional hybrid attention mechanism, characterized in that: include: Acquire a facial image of a human face and preprocess the facial image; The preprocessed facial image is input into the facial expression recognition model, and the facial expression category is output; Among them, the facial expression recognition model is a MobileNet V3 neural network based on feature pyramid and hybrid attention mechanism. When the facial image is input, the facial image undergoes a depth-separable convolution operation, and then passes through the feature pyramid network to extract the multi-scale features of the facial image, obtain high-level semantic features and low-level detail features and fuse them. The fused features are then passed through a hybrid attention mechanism of frequency domain and time domain fusion and spatial channel fusion. By analyzing the energy distribution in the frequency domain, spatial information and channel information are extracted, and the model is guided to focus on the effective features in the image to obtain the final feature vector, and facial expression classification is realized based on the final feature vector.

2. The facial expression recognition method based on the multidimensional hybrid attention mechanism as claimed in claim 1, wherein Acquire facial images and preprocess them, including: capturing video streams and dividing them into continuous image frames. The video capture frequency is set according to specific needs, and facial images are extracted from the image frames. Preprocessing includes: converting the facial image format into a standard format and standardizing the image resolution. Then, illumination normalization, posture correction, and cropping and scaling are performed.

3. The facial expression recognition method based on the multidimensional hybrid attention mechanism as claimed in claim 1, wherein The facial expression recognition model is an enhanced MobileNet V3 neural network based on a feature pyramid and hybrid attention mechanism. The enhanced MobileNet V3 neural network based on a feature pyramid and hybrid attention mechanism uses the MobileNet V3 network as its basic network architecture. By introducing lightweight depthwise separable convolutions, the MobileNet V3 network reduces the number of network parameters and computational complexity, and learns the common features of facial images.

4. The facial expression recognition method based on a multidimensional hybrid attention mechanism as claimed in claim 1, wherein A feature pyramid network is built on the basis of the MobileNet V3 network. The feature pyramid network extracts feature maps at different layers of the MobileNet V3 network to obtain high-level semantic features and low-level detail features. It then fuses the high-level semantic features and low-level detail features through upsampling and downsampling operations as well as lateral connections.

5. The facial expression recognition method based on the multidimensional hybrid attention mechanism as claimed in claim 1, wherein The MobileNet V3 neural network introduces a hybrid attention mechanism, which is an improved attention mechanism that introduces frequency domain and time domain fusion and spatial channel fusion. Through the hybrid attention mechanism of frequency domain and time domain fusion and spatial channel fusion, feature vector information is obtained in different dimensions of frequency domain, time domain and spatial domain. By analyzing the energy distribution in the frequency domain, the importance of different frequency components to expression recognition is determined, and higher weights are given to important frequency components to highlight expression-related features.

6. The facial expression recognition method based on a multidimensional hybrid attention mechanism as claimed in claim 5, wherein The multi-dimensional hybrid attention mechanism mainly consists of three parts: (1) dual-pooled residual channel attention, which focuses more on channel information; (2) depthwise separable spatial attention, which focuses on spatial information; (3) block-wise frequency domain attention, which focuses on frequency domain information.

7. A facial expression recognition system based on a multi-dimensional hybrid attention mechanism, characterized by: include: A data acquisition module is used to acquire facial images of human faces and pre-process the facial images; The expression recognition module is used to input the pre-processed facial image into the facial expression recognition model and output the facial expression category; Among them, the facial expression recognition model is a MobileNet V3 neural network based on feature pyramid and hybrid attention mechanism. When the facial image is input, the facial image undergoes a depth-separable convolution operation, and then passes through the feature pyramid network to extract the multi-scale features of the facial image, obtain high-level semantic features and low-level detail features and fuse them. The fused features are then passed through a hybrid attention mechanism of frequency domain and time domain fusion and spatial channel fusion. By analyzing the energy distribution in the frequency domain, spatial information and channel information are extracted, and the model is guided to focus on the effective features in the image to obtain the final feature vector, and facial expression classification is realized based on the final feature vector.

8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the facial expression recognition method based on the multidimensional hybrid attention mechanism described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by the processor, the facial expression recognition method based on the multidimensional hybrid attention mechanism as described in any one of claims 1-6 is implemented.

10. An electronic device, characterized in that: include: A processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the facial expression recognition method based on the multidimensional hybrid attention mechanism as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Group image emotion recognition method based on attention mechanism and hybrid network

    CN110135251A

  • Three-dimensional expression face restoration method, system and device based on XR device in live scene

    CN116563506A

  • Smart campus safety management method and system based on face recognition

    CN118334559A

  • Facial expression and subjective emotion scale-based user emotion experience evaluation method

    CN119599738A

Cited By

  • Multi-branch pyramid expression recognition method based on hierarchical space-frequency fusion

    CN122290193A

  • A multi-branch pyramid facial expression recognition method based on hierarchical spatial-frequency fusion

    CN122290193B