Face expression recognition method and system based on face key point guidance, and medium
Through the face key point detector and multi-scale feature fusion technology, combined with self-attention and cross-scale cross-attention mechanism, the problems of weak global feature extraction ability and high computational complexity of expression recognition in the existing technology are solved, and the effect of efficiently identifying subtle expression changes is achieved.
Patent Information
- Application Number
- CN202510317459.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-18
AI Technical Summary
When existing facial expression recognition technology deals with complex practical application scenarios, it is difficult to effectively capture long-distance feature relationships, resulting in weak global feature extraction capabilities and high computational complexity, making it difficult to recognize tiny expression changes, and failing to fully utilize multi-scale information, resulting in loss of image details.
Multi-scale key point features and image features are extracted through the face key point detector, and the facial heat map is calculated for spatial attention weighting, combining self-attention and cross-scale attention mechanisms, dynamic pruning feature vectors are realized to achieve the fusion and optimization of multi-scale features.
It improves the accuracy of expression recognition, enhances the ability to represent subtle expressions, reduces the computational complexity and redundant information, and improves the recognition accuracy and efficiency of the model.
Smart Images

Figure CN120340086A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of computer vision, deep learning, and facial expression recognition, and particularly relates to a facial expression recognition method, system, and medium guided by facial key points. Background Art
[0002] Facial expressions are one of the most powerful, natural, and common signals for humans to express emotional states and intentions. Together with sounds, languages, postures, etc., they jointly construct the basic communication system of humans in a social environment and play a very important role in the process of interpersonal communication. With the rapid development of computer technology, especially the breakthrough achievements in the fields of artificial intelligence and computer vision, facial expression recognition technology has emerged and gradually become a research hotspot. The progress of computer vision technology enables computers to effectively detect and analyze human faces in images and videos; the emergence of deep learning algorithms has provided powerful technical support for facial expression recognition, significantly improving the accuracy of expression recognition. As a basic task in computer vision, facial expression recognition has shown great application potential in multiple fields, such as education, medical care, human-computer interaction, fatigue driving, intelligent security, etc.
[0003] Traditional facial expression recognition methods mainly rely on manually designed features, such as local binary pattern (LBP), histogram of oriented gradients (HOG), and scale-invariant feature transform (SIFT), etc.; these methods can achieve certain recognition effects in specific environments, but they are less robust to illumination changes, occlusions, and individual differences and are difficult to adapt to complex actual application scenarios. With the development of deep learning technology, facial expression recognition has gradually shifted from the traditional way of relying on manually designed features to deep neural networks. Some of the newer research works mostly use technologies such as convolutional neural networks, Transformers, generative adversarial networks, and contrastive learning to improve the model's ability to recognize facial expressions and have achieved high accuracy on some challenging datasets. There are also some works based on face-related tasks, such as face recognition and facial key point localization, and optimize the accuracy of facial expression recognition through a multi-task learning framework. Existing facial expression recognition technologies can be roughly divided into three types:
[0004] The first is an expression recognition method based on the Convolutional Neural Network (CNN) architecture, which classifies expressions by automatically extracting deep features of facial images. Compared with traditional methods, CNN eliminates the step of manually designing features, can learn more discriminative features from data, and improves the accuracy and robustness of recognition. Typical CNN architectures (such as VGG, ResNet, EfficientNet) have been applied to the expression recognition task. For example, in 2020, Wang et al. proposed a Region Attention Network (RAN) (Wang K, Peng X, Yang J, et al. Region attention networks for pose and occlusion robust facial expression recognition[J]. IEEE Transactions on Image Processing, 2020, 29: 4057-4069), which cropped the same face image to different degrees and sent the cropped images into the backbone CNN for feature extraction, and then obtained a compact face representation through the self-attention module and the relational attention module. However, RAN only relies on the attention focus of model adaptation, does not fully consider the relationship between facial structure changes and expressions, and it is difficult to accurately focus on the effective area when the pose changes too much. Moreover, extracting features from the same picture multiple times will cause excessive computational overhead. Hu et al. proposed a lightweight multi-scale network (Hu Z, Yan C. Lightweight multi-scale network with attention for facial expression recognition[C] / / 2021 4th International Conference on Advanced Electronic Materials, Computers and Software Engineering (AEMCSE). IEEE, 2021: 695-698), which combines the Xception model with the Convolutional Block Attention Module (CBAM) to learn key face features. However, the CBAM attention mechanism essentially relies on the automatic assignment of global feature weights and may perform poorly under severe occlusion, extreme lighting, or large pose changes.
[0005] The second is the facial expression recognition method based on the Transformer architecture. Leveraging the powerful feature modeling ability demonstrated by the Transformer architecture in computer vision tasks, especially its core self-attention mechanism, which can effectively capture long-range dependencies between elements, making it an important network architecture for facial expression recognition tasks. By capturing global features through the self-attention mechanism, it can better handle the long-range regional co-variation of different facial expression categories and improve the robustness of facial expression recognition. For example, Xue et al. proposed TransFER (Xue F, Wang Q, Guo G. Transfer: Learning relation-aware facial expression representations with transformers[C] / / Proceedings of the IEEE / CVF International conference on computer vision. 2021: 3601-3610). This method first uses a CNN as the backbone network to extract local features, constructs a preliminary facial expression feature representation, and then inputs it into the Transformer to calculate the global relationships between vectors. In this architecture, the CNN is responsible for local feature extraction, while the Transformer models long-range dependencies, thus obtaining a more abundant feature representation in facial expression recognition tasks. To further improve the efficiency of the model, subsequent researchers improved TransFER by adding a pooling module to discard facial expression-irrelevant information (Xue F, Wang Q, Tan Z, et al. Vision transformer with attentive pooling for robust facial expression recognition[J]. IEEE Transactions on Affective Computing, 2022, 14(4): 3244-3256), reducing interference while lowering the model's computational complexity. However, these methods only process the features of the higher layers of the CNN, which may lead to the loss of key local details at the bottom layer (such as micro-expressions), affecting the fine-grained recognition ability of facial expressions. In addition, the deep Transformer used in these methods still has a high computational complexity, a slow inference speed, and requires a large amount of data for training, otherwise it is prone to overfitting.
[0006] The third is an expression recognition method based on face-related tasks; in real-world scenarios, expression recognition is often affected by factors such as head pose, illumination changes, and individual identity. It may be difficult to obtain a robust recognition effect by relying solely on expression features for classification. Therefore, joint training of expression recognition with tasks such as face recognition and face keypoint detection can share information between different tasks, thereby improving the overall performance of expression recognition. For example, face keypoint detection can accurately locate key facial regions (such as eyes, eyebrows, mouth corners, etc.). The displacement and deformation of these regions are crucial for expression expression. Therefore, by jointly training expression recognition and keypoint detection, the model can use the structural information of the keypoints to assist in expression classification and improve the accuracy of recognition. As proposed in the existing technology paper "Poster: A Pyramid Cross-Fusion Transformer Network for Facial Expression Recognition", a POSTER model first extracts face keypoint features through a face keypoint detector, then extracts expression image features through a deep backbone network, and downsamples these two types of features respectively to obtain feature maps of different scales, forming a pyramid structure feature. Then, it calculates the cross-attention of the face keypoint features and expression image features at different scales, fully excavates the potential connection between the two, and finally outputs to the classifier for expression classification. However, the pyramid structure feature of POSTER is essentially obtained by upsampling and downsampling a single-scale feature, and does not truly retain multi-scale information. Moreover, its two-stream cross-attention mechanism significantly increases the computational complexity of the model and is not suitable for real-time expression recognition tasks.
[0007] Although the above technologies have made remarkable progress, they still face many challenges in practical applications. First, the facial expression recognition method based on the CNN architecture is limited by the convolutional receptive field and has difficulty in processing the relationships between long-distance features. Therefore, its global feature extraction ability is weak, and it is insufficient in adapting to the recognition challenges brought by complex expressions and pose changes. Moreover, as the convolutional layer deepens, the facial detail information extracted by the shallow network gradually gets lost, which is not conducive to the recognition of subtle expression changes. Second, for the facial expression recognition method based on the Transformer architecture, its computational complexity is high, it is prone to overfitting for small datasets, and the inference speed is slow. Moreover, the process of mapping features is usually implemented by a linear layer, which is relatively simple and difficult to achieve good extraction of local features. Although some studies have cascaded CNN and Transformer to integrate their advantages, usually only the high-level semantic information extracted by CNN is used as the input for attention calculation, resulting in the loss of underlying detail information and making it difficult to model the detailed changes of expressions. Finally, facial expression recognition based on face-related tasks is usually jointly learned through a multi-task framework. Therefore, the optimization objectives and learning rates of different tasks may conflict or vary. If the weight allocation between tasks is unreasonable, it may lead to the performance improvement of one task while the performance of other tasks deteriorates. Also, existing methods do not fully utilize multi-scale information, resulting in the loss of a large amount of image details (such as edge structures, facial textures, etc.), making it difficult to recognize micro-expression changes. In addition, existing methods have not fully modeled the correlation between facial structures and action changes, making it difficult to establish feature associations between cross-region expressions (such as the coordinated changes between eyes and mouth corners). As a result, when facing large-scale pose changes, it is unable to accurately focus on the key regions, and thus fails to effectively distinguish the importance of these regions during feature extraction and interaction, retaining too much redundant information, wasting the computational resources of the model, and reducing the recognition accuracy. Summary of the Invention
[0008] The main purpose of the present invention is to overcome the drawbacks and deficiencies of the prior art, and provide a method, system and medium for facial expression recognition guided by facial key points, aiming to make full use of the facial structure information contained in the facial key point features, improve the focus on key regions of expressions, relieve the influence brought by head pose changes, and achieve in-depth interaction between underlying texture information and high-level semantic information, enhance the detail expression ability, and thus improve the accuracy of facial expression recognition.
[0009] To achieve the above object, on the one hand, the present invention provides a method for facial expression recognition guided by facial key points, including the following steps:
[0010] Obtain a face image, and input it into a facial key point detector and a deep backbone network respectively to extract multi-scale key point features and image features; the multi-scale includes a bottom scale, a middle scale and a high scale;
[0011] Calculate the facial heatmap, and perform spatial attention weighting calculation with multi-scale image features to obtain multi-scale fused image features;
[0012] Linearly map the multi-scale fused image features and add position encoding to obtain multi-scale fused tokens;
[0013] Perform self-attention calculation and dynamic token pruning on the fused tokens of the bottom scale twice to obtain the first pruned feature and the second pruned feature;
[0014] Perform cross-scale cross-attention calculation and dynamic token pruning on the fused tokens of the middle scale and the first pruned feature to obtain the third pruned feature;
[0015] Perform cross-scale cross-attention calculation on the fused tokens of the high scale and the third pruned feature to obtain the high and middle layer attention features;
[0016] Concatenate the second pruned feature, the third pruned feature and the high and middle layer attention features and calculate the global multi-scale attention features, and input them into the classifier to obtain the expression classification result.
[0017] As a preferred technical solution, the face key point detector uses a pre-trained MobileFaceNet network;
[0018] The deep backbone network uses a pre-trained IR50 network.
[0019] As a preferred technical solution, the acquisition method of the facial heatmap is as follows:
[0020] Multi-scale key point features extracted by the face key point detector,
[0021] Or facial action unit features extracted by the facial action unit detector,
[0022] Or face segmentation features extracted based on the face segmenter;
[0023] Perform max pooling operation on the multi-scale key point features, facial action unit features or face segmentation features along the channel dimension respectively, and then calculate to obtain the facial heatmap.
[0024] As a preferred technical solution, the obtaining of the multi-scale fused image features is specifically as follows:
[0025] Perform dimension expansion operation on the facial heatmap along the channel dimension respectively to match the dimension of the multi-scale image features, and obtain the expanded heatmap;
[0026] Perform dot product operation on the expanded heatmap and the multi-scale image features on the same channel dimension to obtain the multi-scale fused image features.
[0027] As a preferred technical solution, obtaining the first pruning feature and the second pruning feature specifically includes:
[0028] Performing a first self-attention calculation on the bottom-scale fused tokens to obtain a first attention weight and a first attention feature;
[0029] Performing a first dynamic token pruning on the first attention feature according to the first attention weight, calculating the importance score of each bottom-scale fused token in the first attention feature, and retaining the first set number of bottom-scale fused tokens with the highest importance scores to obtain the first pruning feature; the first set number is the number of middle-scale fused tokens;
[0030] Performing a second self-attention calculation on the first pruning feature to obtain a second attention weight and a second attention feature;
[0031] Performing a second dynamic token pruning on the second attention feature according to the second attention weight, calculating the importance score of each bottom-scale fused token in the second attention feature, and retaining the second set number of bottom-scale fused tokens with the highest importance scores to obtain the second pruning feature; the second set number is the number of high-scale fused tokens.
[0032] As a preferred technical solution, obtaining the third pruning feature specifically includes:
[0033] Calculating cross-scale cross-attention based on the middle-scale fused tokens and the first pruning feature to obtain a middle-bottom attention weight and a middle-bottom attention feature;
[0034] Performing dynamic token pruning on the middle-bottom attention feature using the middle-bottom attention weight, calculating the importance score of each middle-scale fused token in the middle-bottom attention feature, and retaining the third set number of middle-scale fused tokens with the highest importance scores to obtain the third pruning feature; the third set number is the number of high-scale fused tokens.
[0035] As a preferred technical solution, obtaining the expression classification result specifically includes:
[0036] Concatenating the second pruning feature, the third pruning feature, and the high-middle attention feature to obtain a global feature;
[0037] Performing a self-attention calculation on the global feature to obtain a global multi-scale self-attention feature;
[0038] Inputting the global multi-scale self-attention feature into a classifier including a fully connected layer and a softmax layer to obtain an expression classification result.
[0039] As a preferred technical solution, the total loss function of the method is expressed as:
[0040]
[0041] where N is the number of face images, y i is the one-hot encoding of the true label of the i-th face image, and p i is the classification prediction probability of the i-th face image.
[0042] On the other hand, the present invention also provides a face expression recognition system based on face key point guidance, including a multi-scale feature extraction module, a multi-scale feature fusion module, a multi-scale feature mapping module, a bottom-scale feature processing module, a middle-scale feature processing module, a high-scale feature processing module, and an expression classification prediction module;
[0043] The multi-scale feature extraction module is used to obtain a face image and input it into a face key point detector and a depth backbone network respectively to extract multi-scale key point features and image features; the multi-scale includes bottom scale, middle scale, and high scale;
[0044] The multi-scale feature fusion module is used to perform max pooling operations on the multi-scale key point features along the channel dimension respectively to calculate the facial heat map, and perform spatial attention weighting calculation with the multi-scale image features to obtain multi-scale fused image features;
[0045] The multi-scale feature mapping module is used to linearly map the multi-scale fused image features and add position encoding to obtain multi-scale fused tokens;
[0046] The bottom-scale feature processing module is used to perform self-attention calculations and dynamic token pruning on the bottom-scale fused tokens twice to obtain a first pruned feature and a second pruned feature;
[0047] The middle-scale feature processing module is used to perform cross-scale cross-attention calculations and dynamic token pruning on the middle-scale fused tokens and the first pruned feature to obtain a third pruned feature;
[0048] The high-scale feature processing module is used to perform cross-scale cross-attention calculations on the high-scale fused tokens and the third pruned feature to obtain high-middle layer attention features;
[0049] The expression classification prediction module is used to splice the second pruned feature, the third pruned feature, and the high-middle layer attention features and calculate global multi-scale attention features, and input them into a classifier to obtain an expression classification result.
[0050] On the other hand, the present invention also provides a computer-readable storage medium storing a program, which when executed by a processor, implements the above-mentioned face expression recognition method guided by face key points.
[0051] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0052] 1. Explicit guidance of facial key points to achieve feature focusing on key areas
[0053] The present invention explicitly extracts geometric structure information through a face key point detector, calculates a facial heat map to weight image features, thereby assigning higher weights to key facial areas (such as eyes, mouth, etc.) that reflect expression changes; while assigning lower weights to areas with weaker relevance to expression, such as hair and background, enabling the model to focus on mining the core features of expressions, effectively focusing on the core information of expressions, excluding the interference of redundant information, and improving the accuracy of expression recognition.
[0054] 2. Multi-scale feature modeling to enhance detail capture and feature fusion
[0055] The present invention adopts a multi-scale feature framework to retain underlying detail information and can recognize subtle expression changes; the proposed cross-scale cross-attention mechanism uses the global semantic information encoded at a high level as a query to actively filter the underlying detail information, providing more targeted retrieval conditions for the attention mechanism, selecting the most important details that match the global context from the low-level features, optimizing the global feature representation, enhancing the feature expression ability of subtle expressions, thereby improving the accuracy of expression classification and achieving higher recognition accuracy.
[0056] 3. Dynamic pruning mechanism to improve inference efficiency
[0057] The dynamic token pruning strategy designed by the present invention only retains several feature vectors with the greatest importance for expression recognition, and prunes the remaining features, alleviating the problems of feature redundancy and high computational complexity caused by multi-scale feature calculation, and effectively improving the inference speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0059] Figure 1 It is the overall flowchart of the face expression recognition method guided by face key points in the embodiments of the present invention.
[0060] Figure 2It is the flowchart framework of the face expression recognition method guided by face key points in the embodiment of the present invention.
[0061] Figure 3 It is the operation schematic diagram of self-attention calculation and dynamic token pruning in the embodiment of the present invention.
[0062] Figure 4 It is the operation schematic diagram of cross-scale cross-attention calculation in the embodiment of the present invention.
[0063] Figure 5 It is the structural schematic diagram of the face expression recognition system guided by face key points in the embodiment of the present invention.
[0064] Figure 6 It is the structural schematic diagram of the computer-readable storage medium in the embodiment of the present invention. Detailed implementation manners
[0065] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0066] Referring to "embodiment" in the present application means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in the present application may be combined with other embodiments.
[0067] As Figure 1 、 2 shown, the face expression recognition method guided by face key points in this embodiment mainly includes four steps:
[0068] Step 1, multi-scale face key point-guided feature fusion: aiming to extract multi-scale face key point features and image features through a convolutional neural network, obtain a facial heat map through the key point features, and perform spatial attention weighting on the image features to obtain fused image features.
[0069] S1. Obtain a face image, and input it into a face key point detector and a deep backbone network respectively to extract multi-scale key point features and image features; among them, the multi-scale includes a bottom scale, a middle scale and a high scale.
[0070] Specifically, given a face image First, input them into the face key point detector F D and the deep backbone network F B In the above example, key point features at different levels are extracted. and image features
[0071]
[0072] Among them, L, M, and H are the labels of the bottom, middle, and high scale feature maps respectively. They are face key point detector F D The extracted feature maps of different scales: bottom, middle, and high. They are the deep backbone network F B The extracted feature maps are of different scales: low, medium, and high. Low-scale features mainly contain fine-grained information such as edges and textures, which help capture local expression changes (such as the corners of the mouth turning up); medium-scale features integrate local information and can enhance the model's understanding of facial structure; high-scale features integrate global semantic information (such as facial posture and expression category) to improve the robustness and generalization ability of the model. In an embodiment of the present invention, a pre-trained MobileFaceNet network (Chen S, Liu Y, Gao X, et al. Mobilefacenets: Efficient cnns for accurate real-time faceverification on mobiledevices[C] / / Chinese conference on biometricrecognition. Cham: Springer International Publishing, 2018: 428-438) is used as a face key point detector, and a pre-trained IR50 network (Deng J, Guo J, Xue N, et al. Arcface: Additiveangular margin loss for deep face recognition[C] / / Proceedings of the IEEE / CVFconference on computer vision and pattern recognition. 2019: 4690-4699) is used as a multi-scale deep backbone network.
[0073] S2. Calculate the facial heat map and perform spatial attention weighted calculation with multi-scale image features to obtain multi-scale fused image features.
[0074] After obtaining the multi-scale face key-point features, it is necessary to calculate the facial heat map, and the acquisition methods include:
[0075] First, the multi-scale key-point features extracted by the face key-point detector,
[0076] or the facial action unit features extracted by the facial action unit detector (Action Unit),
[0077] or the face segmentation features extracted based on the face parcer (Face Parcer);
[0078] Then, perform max-pooling operations on the multi-scale key-point features, facial action unit features, or face segmentation features along the channel dimension respectively to calculate the facial heat map H = [H L , H M , H H . The max-pooling operation can highlight the significant features in the feature map and suppress the secondary features, so that the heat map can more clearly reflect the feature distribution of the facial key areas. The above process can be expressed by the following formula:
[0079]
[0080] Among them, are the facial heat maps of different scales at the bottom, middle, and high levels respectively, and MaxPooling(·) is the max-pooling operation along the channel dimension.
[0081] Subsequently, use the heat maps [H L , H M , H H to perform spatial attention weighting calculations on the extracted image features respectively. The purpose is to focus the attention on the facial areas that are more critical for expression by leveraging the characteristics of the weight distribution in each area of the heat map, thereby enhancing the ability to capture effective expression features. Before the calculation, since there are differences in dimensions between the heat map and the image features, it is necessary to first perform a dimension expansion operation on the heat map. Specifically, along the channel dimension, duplicate the facial heat maps [H L , H M , H H c L , c M , c H times (c L , c M , c H are the channel dimensions of the image features respectively) to make their dimensions match the image features respectively. The expanded heat map is marked as [H L ′, H′M , H′ H . Then, perform a dot product operation on the expanded heatmap and the image features in the same channel dimension to re-weight the image features, and then obtain multi-scale fused image features This process can be expressed by the following formula:
[0082] [H′ L , H′ M , H′ H = expand([H L , H M , H H ),
[0083]
[0084] The fused image features can effectively integrate the spatial attention information carried by the heatmap and the original image features, providing a more representative feature representation for subsequent facial expression recognition tasks.
[0085] S3. Perform a linear mapping on the multi-scale fused image features and add positional encoding to obtain multi-scale fused tokens.
[0086] To calculate the subsequent feature attention, the fused image features need to be linearly mapped and converted into a token form similar to that in Transformer, denoted as t X , and add positional encoding:
[0087]
[0088] Among them, represents the fused image features of different scales, Linear(·) is the linear mapping function, is the positional encoding, N X is the number of tokens after mapping, d represents the token dimension, is the calculated fused tokens of different scales.
[0089] Step 2. Bottom-layer self-attention calculation and dynamic token pruning: Aims to perform self-attention calculation and dynamic token pruning on the extracted bottom-layer features to achieve the interactive fusion and compression of detailed information in different bottom-layer regions.
[0090] S4. Perform self-attention calculation and dynamic token pruning on the fused tokens of the bottom scale twice to obtain the first pruned feature and the second pruned feature.
[0091] Furthermore, as Figure 3 shown, the specific steps to obtain the first pruned feature and the second pruned feature are as follows:
[0092] S4.1. First, perform self-attention calculation on the fused tokens t at the bottom scale L to obtain the first attention weights and the first attention features.
[0093] Specifically, the self-attention calculation process for t L is as follows:
[0094] First, map the fused tokens t at the bottom scale L to three matrices Q L , K L , and V L :
[0095] Q L = t L W Q , K L = t L W K , V L = t L W V ,
[0096] where W Q , W K , W V are the mapping matrices respectively.
[0097] Then, use Q L and K L to calculate the first attention weights A L , and then apply these weights to V L to achieve long-range and in-depth interaction of features in different regions:
[0098]
[0099] Attention(Q L , K L , V L ) = A L V L ,
[0100] where softmax(·) is the softmax activation function, is the normalization scaling factor, and is the first attention weight matrix.
[0101] Finally, calculate the first attention feature
[0102] t' L = Attention(Q L , K L , VL ) + t L ,
[0103]
[0104] The entire self-attention calculation process can be simplified and expressed by the formula .
[0105] S4.2. Perform the first dynamic token pruning on the first attention feature according to the first attention weight, calculate the importance score of each bottom-scale fusion token in the first attention feature, and retain the first set number of bottom-scale fusion tokens with the highest importance scores to obtain the first pruned feature; where the first set number is the number of medium-scale fusion tokens.
[0106] Furthermore, in order to remove redundant and irrelevant information, this method is based on the first attention weight A calculated in S4.1 L to prune the first attention feature . The pruning methods can be: 1) Map the attention weight through an MLP layer, use the output as the importance scores of different tokens, and prune the tokens with low importance; 2) Calculate the maximum value of each row based on the attention weight as the token importance score for pruning; 3) Calculate the similarity between different tokens (such as cosine similarity), merge the tokens with high similarity to reduce the number of tokens; 4) Learn the importance of different tokens through learnable weight parameters and prune the tokens with low importance; 5) Introduce an adaptive pooling layer to automatically determine which tokens can be pooled and merged to reduce the number of tokens.
[0107] In this embodiment, the method 1) is adopted for token pruning. Specifically, first, use the MLP to linearly map the attention weight A L to learn the importance score of each token:
[0108] s L = MLP(A L ),
[0109] where represents the importance score of each token. Then, perform the pruning operation based on the gating mechanism, and only retain the top k L most important tokens:
[0110]
[0111] For the first pruning of the bottom-scale fusion tokens, the number of tokens k retained after pruning LSet to be the same as the number of meso-scale fusion tokens (k L = N M ), where N M is the number of tokens in the meso-scale fusion tokens, that is which will be used for cross-layer cross-attention in meso-scale calculations later.
[0112] S4.3. Perform a second self-attention calculation on the first pruned feature to obtain the second attention weight and the second attention feature.
[0113] S4.4. Perform a second dynamic token pruning on the second attention feature according to the second attention weight, calculate the importance score of each bottom-scale fusion token in the second attention feature, and retain the second set number of bottom-scale fusion tokens with the highest importance score to obtain the second pruned feature; where the second set number is the number of high-scale fusion tokens.
[0114] Perform the same operation again after the first self-attention calculation and dynamic token pruning, aiming to further fuse and compress the feature information, which will be used for global feature fusion and classification later.
[0115] Step 3. Middle and high-level cross-scale cross-attention calculation and dynamic token pruning: aiming to implement a cross-layer cross-attention mechanism to promote the fusion of features at different levels and further enhance the ability of features to represent expression details.
[0116] Among multi-scale features, high-level features have strong global semantic information but lack details, while low-level features contain rich details but lack global information. Therefore, the present invention designs a cross-scale cross-attention mechanism to fully fuse the detail information (low-level features) and global information (high-level features). Specifically, since high-level features focus on the overall expression semantics, they can selectively focus on key regions (such as eyebrows, eyes, mouth corners, etc.) when calculating attention, while ignoring irrelevant low-level noise, which enables high-level features to focus on local information that is most helpful for expressing emotions. In contrast, low-level features contain a large amount of local edge and texture information and are suitable as a more comprehensive memory bank. Therefore, the present invention uses high-level features as the Query in attention calculation to dominate information acquisition, and low-level information as the Key and Value to provide rich reference information; this design allows high-level features to actively query low-level information that matches themselves to enhance their own detail expression ability without being interfered by redundant information, making the attention mechanism more robust.
[0117] S5. Cross-scale cross-attention calculation and dynamic token pruning are performed on the medium-scale fused tokens and the first pruned feature to obtain the third pruned feature.
[0118] In the present invention, the underlying features capture the long-range dependencies of the features in each region by calculating self-attention, and for the middle-level features, cross-attention with the underlying features is calculated to achieve cross-layer feature cross-fusion. Specifically, the steps for obtaining the third pruned feature are as follows:
[0119] S5.1. Based on the medium-scale fused tokens and the first pruned feature, calculate the cross-scale cross-attention to obtain the middle-bottom attention weights and the middle-bottom attention features.
[0120] Specifically, for the medium-scale fused tokens t M , calculate the cross-attention with the first pruned feature . First, map t M to the Query matrix Q M , map to the Key matrix and the Value matrix . Then calculate the middle-bottom cross-attention weights and the middle-bottom attention features:
[0121]
[0122] Among them, is the middle-bottom cross-attention weight, is the middle-bottom cross-attention feature. The above cross-attention calculation process can be simplified as being represented by the formula .
[0123] S5.2. Use the middle-bottom attention weights to perform dynamic token pruning on the middle-bottom attention features, calculate the importance scores of each medium-scale fused token in the middle-bottom attention features, and retain the third set number of medium-scale fused tokens with the highest importance scores to obtain the third pruned feature; where the third set number is the number of high-scale fused tokens.
[0124] For the middle-bottom attention features , then use the middle-bottom attention weight A LM to perform dynamic token pruning operation on it to obtain the third pruned feature
[0125] s LM = MLP(A LM ),
[0126]
[0127] Among them, s LM is the importance score of each token calculated based on the cross-attention weights in the middle and lower layers, and k M is the number of tokens retained after pruning, which is set to be the same as the number of tokens in the high-scale fusion (k M = N H ), where N H is the number of tokens in the high-scale fusion tokens. Similarly, the method of dynamic token pruning is the same as above and will not be elaborated here.
[0128] S6. Calculate the cross-scale cross-attention between the high-scale fusion tokens and the third pruned feature to obtain the middle and high-level attention features.
[0129] Similarly, the third pruned feature is used to calculate the middle and high-level cross-attention features with the high-scale fusion tokens
[0130]
[0131] For the output Since it contains high-level semantic information and has a large feature density, no pruning operation is performed on it, but it is directly used for global feature fusion and classification.
[0132] Step 4: Global Feature Fusion and Classification: The purpose is to fuse the expression features of different scales, ensuring the effective combination of local detail information (low-level features) and overall structural information (high-level features), thereby improving the sensitivity to micro-expression changes and enhancing the understanding of complex expression patterns.
[0133] S7. Concatenate the second pruned feature, the third pruned feature, and the middle and high-level attention features to calculate the global multi-scale attention feature, and input it into the classifier to obtain the expression classification result.
[0134] Furthermore, the steps to obtain the expression classification result are as follows:
[0135] S7.1. Concatenate the bottom second pruned feature the middle third pruned feature and the high-level middle and high-level attention feature together to obtain the global feature t global :
[0136]
[0137] S7.2. Then perform self-attention calculation on the global feature to obtain the global multi-scale self-attention feature:
[0138]
[0139] S7.3. Finally, the global multi-scale self-attention features are input into a classifier containing a fully-connected layer and a softmax layer to calculate the final facial expression classification prediction P:
[0140]
[0141] where Linear() is a linear mapping function.
[0142] Furthermore, the method uses a cross-entropy loss function for optimization, and the total loss function is expressed as follows:
[0143]
[0144] where N is the number of face images, y i is the one-hot encoding of the true label of the i-th face image, and p i is the classification prediction probability of the i-th face image.
[0145] In summary, on the one hand, the present invention directly extracts multi-scale features through a face key point detector and a deep backbone network, and performs spatial attention weighting on the image features by calculating the facial heat map, so that these two kinds of information are fully fused, improving the model's recognition ability for facial structures and subtle changes. On the other hand, the present invention innovatively designs a cross-scale feature interaction mechanism, which can fully fuse global semantic information and local detail features, enabling high-level features to actively query low-level information that matches them to enhance their own detail expression ability without being interfered by redundant information; at the same time, a feature adaptive compression mechanism is also proposed based on the attention weight and the dynamic token pruning strategy, aiming to enhance the feature expression ability in the attention mechanism, reduce redundant information and optimize the model calculation efficiency. On the other hand, the present invention directly calculates the global self-attention of expression features at different scales through global feature fusion and classification, ensuring the cross-scale deep fusion of local detail information and overall structure information, thereby improving the accuracy of facial expression recognition.
[0146] To verify the superiority of the method of the present invention, comparative experiments were conducted on three datasets, namely RAF-DB, FERPlus, and AffectNet. The above datasets have a wide range of sample sources and rich expression changes, and contain a large number of natural face images with different illuminations, poses, races, and age distributions. They are highly challenging and suitable for studying the facial expression recognition task for real-world scenarios. During the experiment, all input images were scaled to 112×112 pixels; for multi-scale image feature extraction, IR50 was used as the deep backbone network and pre-trained on the Ms-Celeb-1M dataset; for multi-scale facial keypoint feature extraction, MobileFaceNet was used as the facial keypoint detector and pre-trained on the 300W dataset. During the training process of the overall model, the parameters of MobileFaceNet were kept fixed to ensure the accuracy of keypoint feature localization. The sizes of the bottom, middle, and high-scale feature maps of the facial keypoint features and image features extracted by this method are as follows: the size of the bottom-scale feature map is 28×28, the size of the middle-scale feature map is 14×14, and the size of the high-scale feature map is 7×7. For the bottom-scale self-attention and cross-scale cross-attention modules, the number of attention layers was set to 1, and for the global multi-scale self-attention module, the number of attention layers was set to 2. For the first and second pruning operations on the bottom-scale features, the number of retained tokens was 196 and 49 respectively, and for the pruning operation on the middle scale, the number of retained tokens was 49. The Adam optimizer was used to update the network parameters (β1 = 0.9, β2 = 0.999), and the learning rate decay strategy was adopted for iterative training for 200 rounds, with a batch size of 64. For the RAF-DB and FERPlus datasets, the initial learning rate was set to 3.5e -5 , and for the AffectNet dataset, the initial learning rate was set to 1e -6 . In this paper, the model was built using PyTorch and trained with 1 NVIDIA RTX 3090 GPU. Several existing methods used in the comparative experiment were compared with the method of the present invention. The existing methods include:
[0147] SCN (Wang K, Peng X, Yang J, et al. Suppressing uncertainties for large-scale facial expression recognition[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2020: 6897-6906),
[0148] DACL (Farzaneh AH, Qi X. Facial expression recognition in the wild via deep attentive center loss [C] / / Proceedings of the IEEE / CVF winter conference on applications of computer vision. 2021:2402-2411),
[0149] KTN (Li H, Wang N, Ding X, et al. Adaptively learning facial expression representation via cf labels and distillation [J]. IEEE Transactions on Image Processing, 2021, 30:2016-2028),
[0150] Meta-Face2Exp (Zeng D, Lin Z, Yan X, et al. Face2exp: Combating data biases for facial expression recognition [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022:20291-20300),
[0151] EAC (Zhang Y, Wang C, Ling X, et al. Learn from all: Erasing attention consistency for noisy label facial expression recognition [C] / / European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2022:418-434),
[0152] At the same time, the accuracy rate was used as an indicator for comparison, and the results are shown in Table 1 below:
[0153] Table 1 Comparison of accuracy rates on the RAF-DB, FERPlus, and AffectNet datasets (%)
[0154]
[0155] As can be seen from Table 1, compared with the existing methods, the method of the present invention has achieved the highest classification accuracy on the RAF-DB, FERPlus and AffectNet datasets, which are 91.85%, 66.37% and 90.52% respectively, indicating that the method of the present invention can effectively focus on the core information of the expression, while retaining key details, optimizing the global feature expression, thereby improving the recognition accuracy. On the RAF-DB dataset, compared with SCN (87.03%) and KTN (88.07%), the method of the present invention has improved by 4.82% and 3.78% respectively. SCN mainly uses uncertainty modeling to improve the robustness of the model, but lacks an effective feature fusion strategy. The method of this paper effectively optimizes the fusion of features of different scales through a cross-scale cross-attention mechanism, so that the model can more accurately identify fine-grained expression changes. On the AffectNet dataset, the method of the present invention is 1.05% and 2.40% higher than EAC and KTN, respectively. EAC prevents the model from memorizing samples of noisy labels by erasing consistency loss, and KTN adopts a coarse-to-fine strategy to determine the category of the expression from easy to difficult. However, these methods mainly focus on overall modeling and still have certain limitations in the depiction of local details, resulting in low accuracy on the AffectNet dataset. The present invention explicitly models the facial structure to guide the model to focus on key areas and fully explore local details. In addition, the method of the present invention combines global and local features to effectively reduce noise interference and improve the robustness of the model. In large-scale, multi-category expression recognition tasks, the present invention demonstrates excellent generalization capabilities, especially in complex environments, it can still maintain high-precision recognition performance.
[0156] It should be noted that, for the sake of convenience, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited to the described order of actions, because according to the present invention, certain steps can be performed in other orders or simultaneously.
[0157] Based on the same idea as the facial expression recognition method based on facial key point guidance in the above embodiment, the present invention also provides a facial expression recognition system based on facial key point guidance, which can be used to execute the above facial expression recognition method based on facial key point guidance. For ease of explanation, the structural schematic diagram of the embodiment of the facial expression recognition system based on facial key point guidance only shows the parts related to the embodiment of the present invention. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and may include more or fewer components than the illustrated structure, or combine certain components, or arrange the components differently.
[0158] likeFigure 5 As shown in Figure 5 , another embodiment of the present invention provides a face expression recognition system guided by face key points, including a multi-scale feature extraction module, a multi-scale feature fusion module, a multi-scale feature mapping module, a bottom-scale feature processing module, a middle-scale feature processing module, a high-scale feature processing module, and an expression classification and prediction module;
[0159] Among them, the multi-scale feature extraction module is used to obtain a face image and input it into a face key point detector and a deep backbone network respectively to extract multi-scale key point features and image features; the multi-scale includes bottom scale, middle scale, and high scale;
[0160] The multi-scale feature fusion module is used to perform maximum pooling operations on the multi-scale key point features along the channel dimension respectively to calculate the facial heat map, and perform spatial attention weighting calculation with the multi-scale image features to obtain multi-scale fused image features;
[0161] The multi-scale feature mapping module is used to linearly map the multi-scale fused image features and add position encoding to obtain multi-scale fused tokens;
[0162] The bottom-scale feature processing module is used to perform self-attention calculation and dynamic token pruning on the bottom-scale fused tokens twice to obtain a first pruned feature and a second pruned feature;
[0163] The middle-scale feature processing module is used to perform cross-scale cross-attention calculation and dynamic token pruning on the middle-scale fused tokens and the first pruned feature to obtain a third pruned feature;
[0164] The high-scale feature processing module is used to perform cross-scale cross-attention calculation on the high-scale fused tokens and the third pruned feature to obtain high-middle layer attention features;
[0165] The expression classification and prediction module is used to splice the second pruned feature, the third pruned feature, and the high-middle layer attention features and calculate the global multi-scale attention features, and input them into a classifier to obtain the expression classification result.
[0166] It should be noted that the face expression recognition system guided by face key points of the present invention corresponds one-to-one with the face expression recognition method guided by face key points of the present invention. The technical features and their beneficial effects described in the embodiments of the above face expression recognition method guided by face key points are all applicable to the embodiments of the face expression recognition system guided by face key points. For specific content, reference can be made to the description in the method embodiments of the present invention, which will not be repeated here. This is hereby declared.
[0167] In addition, in the implementation manner of the face expression recognition system guided by face key points in the above embodiments, the logical division of each program module is only an example. In practical applications, according to needs, for example, considering the configuration requirements of the corresponding hardware or the convenience of software implementation, the above functions can be assigned to different program modules to complete, that is, the internal structure of the face expression recognition system guided by face key points is divided into different program modules to complete all or part of the functions described above.
[0168] As Figure 6 shown, in one embodiment, a computer-readable storage medium is provided, storing a program in a memory. When the program is executed by a processor, the face expression recognition method guided by face key points is implemented, specifically as follows:
[0169] Obtain a face image and input it into a face key point detector and a depth backbone network respectively to extract multi-scale key point features and image features; among them, the multi-scale includes a bottom scale, a middle scale, and a high scale;
[0170] Perform max-pooling operations on the multi-scale key point features along the channel dimension respectively to calculate the facial heat map, and perform spatial attention weighting calculation with the multi-scale image features to obtain multi-scale fused image features;
[0171] Linearly map the multi-scale fused image features and add position encoding to obtain multi-scale fused tokens;
[0172] Perform self-attention calculations and dynamic token pruning on the fused tokens of the bottom scale twice to obtain a first pruned feature and a second pruned feature;
[0173] Perform cross-scale cross-attention calculations and dynamic token pruning on the fused tokens of the middle scale and the first pruned feature to obtain a third pruned feature;
[0174] Perform cross-scale cross-attention calculations on the fused tokens of the high scale and the third pruned feature to obtain high-middle layer attention features;
[0175] Concatenate the second pruned feature, the third pruned feature, and the high-middle layer attention features and calculate the global multi-scale attention features, and input them into a classifier to obtain the expression classification result.
[0176] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0177] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0178] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention should be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A face expression recognition method guided by face key points, characterized in that It includes the following steps: Obtain a face image and input it into a face key point detector and a depth backbone network respectively to extract multi-scale key point features and image features; the multi-scale includes a bottom scale, a middle scale, and a high scale; Calculate a facial heat map and perform spatial attention weighted calculation with the multi-scale image features to obtain multi-scale fused image features; Linearly map the multi-scale fused image features and add position encoding to obtain multi-scale fused tokens; Perform self-attention calculation and dynamic token pruning on the fused tokens of the bottom scale twice to obtain a first pruned feature and a second pruned feature; Perform cross-scale cross-attention calculation and dynamic token pruning on the fused tokens of the middle scale and the first pruned feature to obtain a third pruned feature; Perform cross-scale cross-attention calculation on the fused tokens of the high scale and the third pruned feature to obtain high-middle layer attention features; Concatenate the second pruned feature, the third pruned feature, and the high-middle layer attention features and calculate global multi-scale attention features, then input them into a classifier to obtain an expression classification result.
2. The method for face expression recognition guided by face key points according to claim 1, wherein, The face key point detector uses a pre-trained MobileFaceNet network; The depth backbone network uses a pre-trained IR50 network.
3. The method for facial expression recognition guided by facial key points according to claim 1, wherein The method for obtaining the facial heat map is as follows: Multi-scale key point features extracted by a face key point detector, Or facial action unit features extracted by a facial action unit detector, Or face segmentation features extracted based on a face segmenter; Perform max pooling operations on the multi-scale key point features, facial action unit features, or face segmentation features along the channel dimension respectively, and then calculate to obtain the facial heat map.
4. The method for face expression recognition guided by face key points according to claim 1, wherein The specific method for obtaining the multi-scale fused image features is as follows: Perform dimension expansion operations on the facial heat map along the channel dimension respectively to match the dimensions of the multi-scale image features, and obtain the expanded heat map; Perform dot product operations on the expanded heat map and the multi-scale image features on the same channel dimension to obtain multi-scale fused image features.
5. The method for facial expression recognition guided by facial key points according to claim 1, wherein The specific method for obtaining the first pruned feature and the second pruned feature is as follows: Perform the first self-attention calculation on the fused tokens of the bottom scale to obtain a first attention weight and a first attention feature; Perform the first dynamic token pruning on the first attention feature according to the first attention weight, calculate the importance scores of each bottom-scale fused token in the first attention feature, and retain the first set number of bottom-scale fused tokens with the highest importance scores to obtain the first pruned feature; the first set number is the number of fused tokens of the middle scale; Perform the second self-attention calculation on the first pruned feature to obtain a second attention weight and a second attention feature; Perform the second dynamic token pruning on the second attention feature according to the second attention weight, calculate the importance scores of each bottom-scale fused token in the second attention feature, and retain the second set number of bottom-scale fused tokens with the highest importance scores to obtain the second pruned feature; the second set number is the number of fused tokens of the high scale.
6. The face expression recognition method based on face key point guidance according to claim 1, characterized in that The obtaining of the third pruning feature is specifically as follows: Based on the medium-scale fused tokens and the first pruning feature, calculate the cross-scale cross-attention to obtain the middle-bottom attention weights and the middle-bottom attention features; Use the middle-bottom attention weights to perform dynamic token pruning on the middle-bottom attention features, calculate the importance scores of each medium-scale fused token in the middle-bottom attention features, and retain the third set number of medium-scale fused tokens with the highest importance scores to obtain the third pruning feature; the third set number is the number of high-scale fused tokens.
7. The method for facial expression recognition guided by facial key points according to claim 1, wherein The obtaining of the expression classification result is specifically as follows: Concatenate the second pruning feature, the third pruning feature, and the high-middle layer attention features to obtain the global feature; Perform self-attention calculation on the global feature to obtain the global multi-scale self-attention feature; Input the global multi-scale self-attention feature into a classifier including a fully connected layer and a softmax layer to obtain the expression classification result.
8. The method for facial expression recognition guided by facial key points according to claim 1, wherein The total loss function of the method is expressed as: where N is the number of face images, and y i is the one-hot encoding of the true label of the i-th face image, and p i is the classification prediction probability of the i-th face image.
9. A face expression recognition system guided by face key points, characterized in that, Applied to the face expression recognition method guided by face key points described in any one of claims 1-8, including a multi-scale feature extraction module, a multi-scale feature fusion module, a multi-scale feature mapping module, a bottom-scale feature processing module, a middle-scale feature processing module, a high-scale feature processing module, and an expression classification prediction module; The multi-scale feature extraction module is used to obtain a face image and input it into a face key point detector and a deep backbone network respectively to extract multi-scale key point features and image features; the multi-scale includes bottom scale, middle scale, and high scale; The multi-scale feature fusion module is used to perform maximum pooling operations on the multi-scale key point features along the channel dimension respectively to calculate the facial heat map, and perform spatial attention weighted calculation with the multi-scale image features to obtain the multi-scale fused image features; The multi-scale feature mapping module is used to linearly map the multi-scale fused image features and add position encoding to obtain the multi-scale fused tokens; The bottom-scale feature processing module is used to perform two self-attention calculations and dynamic token pruning on the bottom-scale fused tokens successively to obtain the first pruning feature and the second pruning feature; The middle-scale feature processing module is used to perform cross-scale cross-attention calculation and dynamic token pruning on the middle-scale fused tokens and the first pruning feature to obtain the third pruning feature; The high-scale feature processing module is used to perform cross-scale cross-attention calculation on the high-scale fused tokens and the third pruning feature to obtain the high-middle layer attention features; The expression classification prediction module is used to concatenate the second pruning feature, the third pruning feature, and the high-middle layer attention features and calculate the global multi-scale attention features, and input them into the classifier to obtain the expression classification result.
10. A computer-readable storage medium storing a program, characterized in that, When the program is executed by a processor, it implements the face expression recognition method guided by face key points described in any one of claims 1-8.
Citation Information
Cited By
Expression recognition Transform compression method and system based on key feature region protection
CN121214528A
An expression recognition transformer compression method and system based on key feature region protection
CN121214528B
Edge end power transmission tower corrosion diagnosis method and system based on corrosion significance perception
CN121582233A
Corrosion diagnosis method and system for edge power transmission tower based on corrosion significance perception
CN121582233B