A method for facial expression recognition based on population characteristic factors

CN117558049BActive Publication Date: 2026-09-08SHANGHAI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311517820.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2026-09-08
Estimated Expiration
2043-11-15

AI Technical Summary

Technical Problem

医疗领域的智能陪护场合下同时存在性别,年龄,种族等人口统计学因素的多样性影响表情表达,然而出于保护隐私信息等多方面的考虑,人口因素特征难以完整收集的问题,导致表情识别精度低于预期,影响了智能陪护场景下的效果

Benefits of technology

[0041] (1) This invention uses population factor multi-feature extraction technology, utilizes the influence and relationship between population factor feature information (gender, age, race) on each other's recognition tasks, establishes a population factor restrictive enhancement feature framework, obtains the corresponding population factor characteristic information space on the facial expression object image data, expands the range of facial information space, and improves the accuracy of facial expression recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117558049B_ABST
    Figure CN117558049B_ABST
Patent Text Reader

Abstract

The application relates to a face expression recognition method based on population characteristic factors, which comprises the following steps: face recognition, obtaining a face expression object; analyzing the relationship between population factor characteristic information, training a population factor restrictive reinforcement characteristic framework; inputting the face object into the population factor restrictive reinforcement characteristic framework, obtaining population factor characteristic information space; obtaining a multi-type space characteristic information region; learning the multi-type space characteristic information region by using a backbone network, obtaining face multi-block characteristics; obtaining face multi-block characteristic attention confidence; fusing the face multi-block characteristic attention confidence and corresponding face multi-block characteristics, obtaining block adaptive characteristic expression of a face image, and judging the face expression by fusing the block adaptive characteristic expression. Compared with the prior art, the application has the advantages of effectively expanding the object face information range, improving the face region characteristic effective utilization rate and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of facial expression recognition technology, and in particular to a method for facial expression recognition based on demographic factors. Background Technology

[0002] Analyzing and identifying the gender, age, ethnicity, and emotions of individuals using facial features is a crucial topic in computer vision with a wide range of applications. In particular, variations in facial features and textures pose challenges to various facial recognition tasks. As a primary means of conveying social information, the face is a major carrier of human emotional expression. Facial expressions contain complex information, such as temperament and personality, emotional state, and psychopathological manifestations. Approximately 55% of emotional expression is conveyed through the face. Over the past decade, with the development of artificial intelligence, various applications of facial expression recognition have seen broad and significant growth prospects, including security, lie detection, predicting fatigued driving, intelligent tutoring systems, healthcare, and intelligent companionship.

[0003] Existing research on facial expression recognition mainly focuses on improving the effective utilization of expression features, thereby enhancing the accuracy of facial recognition tasks, through feature extraction, feature parameter sharing, and regional feature measurement. Currently, common facial recognition methods for feature extraction are mainly divided into manual feature extraction and deep feature learning models. Manual feature extraction includes geometric feature extraction methods and texture appearance feature extraction methods. These two methods extract facial image features through the geometric relationships between facial components and extract micro-texture information through local coding schemes. Manual feature extraction suffers from problems such as high noise and shallow information; therefore, improving the efficiency of expression feature extraction has become the main direction for the development of expression recognition.

[0004] With rising labor costs and increasing demand for companionship services, intelligent companionship has become a crucial development direction for various companionship scenarios, including healthcare and elderly care. Facial expression recognition is essential for understanding the state and needs of those being accompanied. In intelligent companionship settings within the medical field, multiple demographic factors are present. Psychological and physiological studies show that facial expressions vary with age, gender, and ethnic background. The diversity of demographic factors such as gender, age, and ethnicity influencing facial expressions in intelligent companionship settings is significant. However, due to privacy concerns and other considerations, it is difficult to fully collect all demographic features, leading to lower-than-expected accuracy in facial expression recognition and impacting the effectiveness of intelligent companionship. Current technologies lack methods to integrate demographic features into facial expressions to improve recognition accuracy; therefore, methods based on demographic factors for facial expression recognition are needed. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of existing technologies by providing a facial expression recognition method based on demographic factors. This method analyzes the relationships between demographic feature information to obtain a multi-level, multi-category demographic characteristic information space of the facial object, including gender, age, and ethnicity. By connecting this multi-level, multi-category facial feature information space with the facial object, multi-block facial features are obtained, achieving complementarity of information at different levels and categories. An attention mechanism is used to learn the key areas of interest for different types of feature information, inferring the interrelationships between different areas of interest, and obtaining adaptive block features. This effectively expands the range of facial information and improves the effective utilization rate of facial region features.

[0006] The objective of this invention can be achieved through the following technical solutions:

[0007] This invention discloses a method for facial expression recognition based on demographic factors, comprising the following steps:

[0008] Step 1: Preprocess the sample image to be judged to obtain the facial expression object;

[0009] Step 2: Analyze the relationships between demographic features in the facial expression objects and train a demographic-restricted enhancement feature framework;

[0010] Step 3: Input the facial expression object into the population factor-restricted enhancement feature framework to obtain the population factor characteristic information space;

[0011] Step 4: Connect the facial expression object with the population factor characteristic information space to expand the facial information space range and obtain multi-type spatial feature information regions;

[0012] Step 5: Use the backbone network of the deep learning algorithm to learn the multi-type spatial feature information regions to obtain multi-block facial features;

[0013] Step 6: Analyze the correlation between the multi-block facial features and facial expression recognition, and use the self-attention mechanism to obtain the self-attention confidence of the multi-block facial features.

[0014] Step 7: Analyze the interrelationship between facial expression objects and the spatial regions of the multi-category feature information, and obtain the facial multi-block feature attention confidence by utilizing the self-attention confidence and multi-feature attention mechanism of the facial multi-block features;

[0015] Step 8: Fuse the attention confidence of the multi-block facial features with the corresponding multi-block facial features to obtain the block adaptive feature representation of the facial image, and judge facial expressions by fusing the block adaptive feature representation.

[0016] Further, step 1 specifically includes: performing face recognition on the image, aligning and cutting it to obtain the cut facial expression object.

[0017] Furthermore, step 2 specifically includes: using a feature constraint reinforcement mechanism to measure the confidence level of the interaction between gender, age, and race characteristics; the process of calculating the confidence levels of gender, age, and race characteristics is shown in the formula:

[0018]

[0019]

[0020]

[0021] Among them, U i gender For the confidence level of gender characteristics, U i ethnicity For confidence levels of racial characteristics, U i age Let λ1, λ2, and λ3 be the confidence scores for the age feature, λ1, λ2, and λ3 be the confidence scores for the population feature, and ||·|| represent the absolute value function.

[0022] By using a feature constraint reinforcement mechanism, the confidence levels of the mutual influence of gender, age, and race are calculated to obtain the three category reinforcement features of object i: gender, age, and race, as shown in the formula:

[0023]

[0024]

[0025]

[0026] in, These are the restricted reinforcement features for the object's age, gender, and race, respectively. gender C ethnicity C age This represents the result of the activation function for gender, race, and age classification. K represents the length of the gender class, M represents the length of the race class, Z represents the length of the age class, and U... i gender For the confidence level of gender characteristics, U i ethnicity For confidence levels of racial characteristics, U i age The confidence level for the age characteristic;

[0027] By combining the restrictive features of age, race, and gender, a comprehensive demographic feature is obtained, as shown in the following formula:

[0028]

[0029] Among them, A i As a comprehensive characteristic of demographics, These are the restricted reinforcement features based on the object's age, gender, and race, respectively.

[0030] Furthermore, the parameters of the trained population factor-restricted enhanced feature framework are fixed, and the facial expression object is input into the framework. Feature heatmaps, comprehensive feature heatmaps, and inverse gradient heatmaps of the comprehensive features for each category of gender, age, and race are extracted from the framework. The feature heatmaps, comprehensive feature heatmaps, and inverse gradient heatmaps of the comprehensive features for each category of gender, age, and race are used together as the population factor characteristic information space.

[0031] Furthermore, step 4 specifically includes: vertically connecting the facial expression object with the demographic characteristic information space.

[0032] Furthermore, step 5 specifically includes: extracting high semantic features from multiple types of spatial feature information regions using the ResNet18 backbone network of the deep learning algorithm, thereby obtaining deep features of different categories of spatial feature information regions.

[0033] Furthermore, step 6 specifically includes: calculating the depth features of different categories of spatial feature information regions using a self-attention mechanism, analyzing the contribution of the depth features of different categories of spatial feature information regions to facial expression recognition, and obtaining the self-attention confidence of facial expressions for multiple types of spatial features.

[0034] Furthermore, step 7 specifically includes: setting up a multi-feature attention mechanism, analyzing the relationship between population factor spatial features in the multi-type spatial features, as well as the complementarity and correlation between population factor spatial features and facial expression object features, through the self-attention confidence of the facial expressions of the multi-type spatial features, and obtaining the facial multi-block feature attention confidence.

[0035] Furthermore, the contribution of different regions to facial expression judgment is measured by the attention confidence score of multi-region facial features. The specific process is expressed by the following formula:

[0036]

[0037]

[0038] in: The confidence level of the adaptive features for all feature regions. The confidence level of the adaptive features for facial expression objects. Adaptive feature confidence for multi-type feature spatial blocks of population factors.

[0039] Further, step 8 specifically includes: non-uniformly fusing the attention confidence of the multi-block facial features with the corresponding regional features, fusing the expression task sharing ratio according to the regional features, obtaining the block adaptive features of the facial image, and judging facial expressions through the block adaptive features of the facial image.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] (1) This invention uses population factor multi-feature extraction technology, utilizes the influence and relationship between population factor feature information (gender, age, race) on each other's recognition tasks, establishes a population factor restrictive enhancement feature framework, obtains the corresponding population factor characteristic information space on the facial expression object image data, expands the range of facial information space, and improves the accuracy of facial expression recognition.

[0042] (2) The facial multi-block features obtained by deep feature extraction of multi-type spatial feature information regions are analyzed. The correlation between individual facial block features and target tasks is analyzed. Through the self-attention mechanism, the self-attention confidence of facial expressions of multi-type spatial features is extracted, so that the model can learn the proportion of different facial spatial features in task learning and optimize the range of attention features for expression recognition.

[0043] (3) Reasoning the relationship between population factor spatial features in multi-type spatial features, as well as the complementarity and association between population factor spatial features and facial expression object features. Using a multi-feature attention mechanism, the facial multi-block features (population factor features and facial expression object features) are associated within sub-regions and between sub-regions. The relationship between population factor features and the complementary relationship between population factor features and facial expression object features are inferred, which improves the efficiency of facial expression feature extraction.

[0044] (4) The confidence level of facial multi-block feature attention is obtained through a multi-feature attention mechanism. The range of attention features for facial expressions is optimized, and the accuracy of facial expression recognition is improved. Attached Figure Description

[0045] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0046] Figure 2 This is a schematic diagram of the overall structure of the present invention;

[0047] Figure 3 This is a schematic diagram of the population factor-restrictive enhanced feature framework of the present invention;

[0048] Figure 4This is a schematic diagram illustrating the learning rate change during the training process on the CK+ dataset according to the present invention.

[0049] Figure 5 This is a schematic diagram of the (TSNE) distribution of the validation set for the ablation experiments of this invention on the CK+ dataset. Detailed Implementation

[0050] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0051] Example

[0052] This embodiment is based on psychological and physiological research, as well as the relationship between facial expressions and demographic characteristic regions, and addresses the differences in facial expressions with age, gender, ethnic background, and other demographic factors. Compared to children, the tendency for vertical movement in adult facial expressions is increased due to muscle tissue development. Men have a more pronounced vertical upward component in posed smiles and lips, while women have a more pronounced horizontal component in posed smiles. Older adults use more facial muscles when expressing emotions than younger people, and the intensity of activity in different AU regions differs from that of younger people.

[0053] Different genders, ages, races, and other demographic factors result in variations in the way facial expressions are expressed, as well as differences in muscle direction and intensity. This affects the accuracy of facial expression recognition. To address the challenge of comprehensively collecting these demographic features, a demographic-restricted enhanced feature framework is used to obtain a spatial information space of the subject's demographic characteristics. An attention mechanism is employed to learn the key areas of interest for different types of feature information, and the contribution of each area of ​​interest to facial expression recognition is measured.

[0054] like Figure 1 As shown, this embodiment provides a method for facial expression recognition based on demographic factors, which includes the following specific steps:

[0055] Step 1: Perform face recognition on the sample image object i to be judged using a third-party tool, and perform preprocessing for alignment and segmentation to obtain the segmented facial expression object x. i .

[0056] Step 2: Analyze the relationships between demographic characteristics and train the Demographic Restrictive Enhanced Feature Framework (MFIS).

[0057] Use the aligned and cropped facial expression object x i Feature extraction is performed using a deep learning network to obtain the basic object features T(x) of the facial expression object.i ,w i And through different classification activation functions δ gender (·), δ ethnicity (·), δ age (·) Obtain the basic demographic characteristics (gender, race, age) of the facial expression subject, as shown in Formula 1:

[0058]

[0059] C gender C ethnicity C age This represents the result of the classification activation function.

[0060] Use softmax as the activation function. This represents the probability that the i-th object is of the j-th gender. This represents the probability that the i-th object belongs to the j-th race group. Let K represent the probability that the i-th object belongs to the j-th age group, K represent the length of the gender category, M represent the length of the race category, and Z represent the length of the age category. The predicted probabilities of the gender, race, and age of the i-th object are shown in Formula 2:

[0061]

[0062] There are mutual influences among demographic characteristics. A feature constraint reinforcement mechanism is used to measure the confidence level of the mutual influence between different types of demographic characteristics. To assess the reliability of different characteristics, this embodiment calculates the confidence level of different types of characteristics, and uses the highest prediction score among each type of characteristic for object i. The confidence scores (U) for different types of features are calculated by taking the highest score for the gender of object i. i gender For the confidence level of gender characteristics, U i ethnicity For confidence levels of racial characteristics, U i age (The confidence level of the age feature) is calculated as shown in Formula 3:

[0063]

[0064] Where λ1, λ2, and λ3 are the confidence parameters of population characteristics, and ||·|| represents the absolute value function. Through a feature constraint reinforcement mechanism, the confidence of the influence of multi-category population factor characteristics on each other is calculated, obtaining the reinforced features for each category of object i. These are enhancement features based on the subject's age, race, and gender.

[0065]

[0066] Step 3: Transfer the facial expression object x i Input the aforementioned population factor restrictive enhancement feature framework to obtain the population factor characteristic information space.

[0067] The setup process for obtaining population factor characteristic information space is as follows: Heatmaps corresponding to different categories of enhanced features are extracted using the population factor restrictive enhancement feature framework in step 2. This study examines the information space of facial attention areas categorized by age, race, and gender. From a comprehensive perspective, it analyzes the general characteristics of demographics, superimposing multi-category features to reinforce the information space, thus obtaining a comprehensive demographic feature A. i As shown in Formula 5:

[0068]

[0069] Extract A i The corresponding heatmap A i (Heatmap) yields spatial blocks of interest representing general demographic features. To further study the spatial block variation trends of comprehensive demographic features, comprehensive feature A... i Reverse gradient thermogram A i (backforward heatmap) is used to map the heatmaps corresponding to multi-class restricted reinforcement features. Comprehensive Feature Heatmap A i (heatmap) together serve as a multi-category feature information space for demographic factors such as age, race, and gender.

[0070] Step 4: Connect the facial expression object with the population factor characteristic information space to obtain multi-type spatial feature information regions and expand the range of facial information space.

[0071] Specifically, by obtaining the multi-category information space of facial demographic factors in step 3, and combining it with the facial expression object x i Vertically splicing them together to form a multi-block emoji object E i . in Let m be the multi-category information space of facial demographic factors corresponding to object i, and m be the number of multi-category facial space blocks. Let E be the multi-block expression object. i The specific definition is shown in Formula 6:

[0072]

[0073] Step 5: Use the ResNet18 backbone network in the deep learning algorithm to learn the multi-type spatial feature information regions to obtain multi-block facial features.

[0074] Specifically, the multi-block emoji object E obtained in step 4... i The input is fed into the feature extraction backbone network to extract multi-block facial expression features F. Let... in The extracted features are the j-th multi-class information space corresponding to the facial image of object i. For the feature weights of facial expression objects, Feature weights are assigned to the spatial blocks of multi-category facial feature information. By fusing regions, the scope of comprehensive information feature extraction for facial expression objects is expanded, providing a wider range of information features for expression recognition.

[0075] Step 6: Analyze the correlation between the facial multi-block features and the expression task, and use the self-attention mechanism to obtain the self-attention confidence of the facial multi-block features.

[0076] Specifically, the steps for obtaining the self-attention confidence of facial multi-block features in this embodiment are as follows: Considering that facial expression objects and demographic attributes correspond to different contributions of multi-class feature information spaces to facial expression recognition, this embodiment uses a self-attention module to observe the association between high semantic features of facial expression objects and multi-class feature information spaces. Since facial expression object features and multi-class feature information space features focus on different aspects, considering the mutual influence of multi-class feature information spaces based on demographic factors and their supplementary information to facial expression features, a multi-class space attention learning mechanism is introduced to infer the interrelationship between facial expression objects and multi-class feature information spaces, as well as their contribution to expression recognition.

[0077] The facial expression object comprehensive information features obtained in step 5 are used as the input to the multi-feature attention network, which... Specifically, the self-attention mechanism for measuring the importance of multi-block features consists of a pooling operation, a fully connected (FC) layer, and an activation function, which can be expressed by the formula:

[0078]

[0079] Where fc(·) represents a pooling operation and a fully connected (FC) layer, σ(·) represents the activation function, which is the sigmoid function used here. After attention weighting by the sigmoid function, the weights of different regions show the correlation between different types of features and facial expression task judgment. The larger the value, the greater the contribution of the feature value of that region to the correct judgment of the task, and vice versa.

[0080] Because the multi-category information features based on population factors have certain internal connections, all multi-category feature information spaces are merged into a comprehensive category information space. make Extract comprehensive category information features from the backbone network. The self-attention mechanism, which measures the contribution of comprehensive category information features, consists of a representative pooling operation, a fully connected linear module (FC) layer, and an activation function, as expressed by the following formula:

[0081]

[0082] Step 7: Analyze the interrelationship between facial expression objects and multi-category feature information space, and obtain the facial multi-block feature attention confidence by utilizing the self-attention confidence of facial multi-block features and the multi-feature attention mechanism.

[0083] Specifically, to further understand the complementary relationship between demographic features and facial features, this embodiment infers the connection between different category feature information space blocks and comprehensive category information space blocks, and measures the contribution of different blocks to facial expression task judgment by using the facial multi-block feature attention confidence. The specific process can be described using a formula...

[0084] Equations 9 and 10 can be used to represent:

[0085]

[0086]

[0087] in The confidence level of the adaptive features for all feature regions. The confidence level of the adaptive features for facial expression objects. Adaptive feature confidence for multi-type feature spatial blocks of population factors.

[0088] Step 8: Fuse the attention confidence of the multi-block facial features with the corresponding multi-block facial features to obtain the block adaptive feature representation of the facial image, and judge facial expressions by fusing the block adaptive feature representation.

[0089] Specifically, the process of obtaining the adaptive feature representation is as follows: by weighting the adaptive feature confidence with the corresponding feature information, the adaptive feature vector is obtained. The weighting calculation is shown in Formula 11:

[0090]

[0091] The adaptive feature vectors are fused by connecting them in a simple way as shown in Equation 12, and then passed through a fully connected (FC) layer and a softmax activation function to obtain the final discrimination result.

[0092]

[0093] To verify the effectiveness of the multi-feature attention mechanism and adaptive expression proposed in this embodiment, the MDFAM algorithm based on the facial expression recognition method using demographic factors proposed in this embodiment was run on a GPU RTX 3080ti. The deep learning framework PyTorch was used as the experimental platform. The SGD algorithm was selected for gradient descent and backpropagation calculations. The demographic factor information space was obtained using a trained demographic factor restricted enhanced feature framework (MFIS) pre-trained on the UTKFace dataset. The ResNet18 network was used as the backbone network to extract the feature attributes of image objects. The effectiveness of this embodiment was tested on the public dataset CK+.

[0094] like Figure 4 As shown, Figure 4 This is a schematic diagram illustrating the learning rate change during the training process on the CK+ dataset in this embodiment. Figure 4 The diagram illustrates the changes and determines the initial learning rate for MDFAM training.

[0095] To verify the advancement of the population factor multi-feature attention strategy proposed in this invention, this embodiment compares the performance of several representative facial expression recognition algorithms on the CK+ dataset, as shown in Table 1.

[0096] CMCNN SE 98.33 RGAAM SE 99.59 DSAN SE 98.92 WS-LGRN SE 98.37 SEIIL SE 98.77 CFNET SE 99.07 ViT+SE SE 99.8 DeepEmotion SE 98.0 Self-adaptive net SE 98.5 NEDP SE 98.68 MDFAM SE 100

[0097] Table 1 Comparison on the CK+ dataset

[0098] As shown in Table 1, on the CK+ dataset, the proposed population factor multi-feature attention strategy outperforms other algorithms in accuracy, achieving an average recognition accuracy of 100% on CK+. In methods utilizing only static images for facial expression recognition, the experimental results reach an advanced level.

[0099] Among the comparative methods, CMCNN, DSAN and other methods use multi-task shared parameters and multi-task collaborative attention. This method effectively extracts facial expression features and thus achieves good results. Compared with these methods, the MDFAM proposed in this embodiment effectively improves the extraction of effective feature information by sharing multi-type feature information of population factors and combining it with reasoning on the relationship between multi-type features and facial expression features, thus achieving a good judgment accuracy.

[0100] It is worth noting that both RGAAM and CFNET methods achieved excellent recognition accuracy exceeding 99%. Both methods focus on different facial regions and employ multi-level, multi-granular feature extraction. Compared to RGAAM and CFNET, MDFAM obtains a demographic feature information space through a two-stage network. This provides effective complementary information for facial expression regions and lays the foundation for subsequent inference and effective extraction between multiple feature types, thus achieving higher accuracy.

[0101] Self-adaptive Net, as a facial expression recognition method based on prior knowledge, can achieve complementary information about facial expression objects by decoupling the original space and the expression subspace, while reducing the influence of personal facial identity on facial expression recognition. In contrast, the MDFAM proposed in this embodiment uses a population-constrained enhanced feature framework to obtain population feature information, and combines physiological research to effectively supplement facial expression features, thus achieving better recognition accuracy than Self-adaptive Net.

[0102] To describe in more detail the role of demographic features and adaptive learning mechanisms in the multi-feature attention mechanism in this embodiment, extensive ablation evaluation experiments were conducted on these two modules. Figure 5 This diagram illustrates the distribution of the validation set (TSNE) for the ablation experiments of this invention on the CK+ dataset. Each colored dot represents a facial expression; the greater the distance between the centers of different expressions, the better the algorithm's ability to distinguish between categories and the better its recognition performance. The diagram shows that the distance between the centers of each category is greater for algorithms using demographic features and adaptive learning mechanisms than for algorithms that do not use demographic features or only use demographic features without adaptive learning mechanisms.

[0103] in, Figure 5 (a) For comparison, the distribution of facial expression recognition labels without demographic features represents the distribution of 7 facial expression predictions using ResNet18 baseline (here, points of the same color represent the same facial expression, and the farther apart the centers of different facial expression types are, the more different the facial expression characteristics are, and therefore the better the facial expression recognition effect).

[0104] Figure 5 (b) is the distribution of facial expression recognition labels using only demographic features, representing the distribution of 7 facial expression predictions using a multi-factor demographic feature strategy for facial expression recognition.

[0105] Figure 5(c) is the distribution of facial expression recognition using demographic features and adaptive learning mechanisms, representing the distribution of 7 facial expression predictions using a multi-feature attention mechanism based on demographic factors.

[0106] This embodiment addresses the issue that facial expressions vary with demographic factors such as age, gender, and ethnicity. Based on the influence and relationships between demographic features (gender, age, ethnicity) on the recognition task, a demographic-constrained enhanced feature framework is established and trained. This framework is used to obtain the demographic characteristic information space corresponding to the facial expression image object. Combining the demographic characteristic information space and the facial expression image object expands the information space of the facial expression object. By combining a multi-type spatial feature self-attention mechanism and a multi-block facial feature attention mechanism, the connections between demographic spatial features within the multi-type spatial features, as well as the complementarity and correlation between demographic spatial features and facial expression object features, and their contribution to the facial expression task, are obtained. This embodiment expands the feature space to a wider range while optimizing feature extraction efficiency, thus improving the accuracy of facial expression recognition.

[0107] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A method for facial expression recognition based on demographic factors, characterized in that, Includes the following steps: Step 1: Preprocess the sample image to be judged to obtain the facial expression object; Step 2: Analyze the relationships between demographic features in the facial expression objects and train a demographic-restricted enhancement feature framework; Step 3: Input the facial expression object into the population factor-restricted enhancement feature framework to obtain the population factor characteristic information space; Step 4: Connect the facial expression object with the population factor characteristic information space to expand the facial information space range and obtain multi-type spatial feature information regions; Step 5: Use the backbone network of the deep learning algorithm to learn the multi-type spatial feature information regions to obtain multi-block facial features; Step 6: Analyze the correlation between the multi-block facial features and facial expression recognition, and use the self-attention mechanism to obtain the self-attention confidence of the multi-block facial features. Step 7: Analyze the interrelationship between facial expression objects and the spatial regions of the multi-category feature information, and obtain the facial multi-block feature attention confidence by utilizing the self-attention confidence and multi-feature attention mechanism of the facial multi-block features; Step 8: Fuse the attention confidence of the multi-block facial features with the corresponding multi-block facial features to obtain the block adaptive feature representation of the facial image, and judge facial expressions by fusing the block adaptive feature representation; Step 2 specifically includes: using a feature constraint reinforcement mechanism to measure the confidence level of the interaction between gender, age, and race characteristics; the process of calculating the confidence level of gender, age, and race characteristics is shown in the formula: in, The confidence level for gender characteristics. Confidence level for racial characteristics The confidence level for the age characteristic. , , For confidence parameters of population characteristics, Represents the absolute value function; By employing a feature constraint reinforcement mechanism, the confidence levels of the influence of gender, age, and race on each other are calculated to obtain the object. The three categories of enhanced features—gender, age, and race—are shown in the formula: in, These are the restricted reinforcement features based on the subject's age, gender, and race. , , This represents the result of the activation function for classifying gender, race, and age. K Indicates the length of the gender class. M The length representing the racial category, Z The length representing the age category, The confidence level for gender characteristics. Confidence level for racial characteristics The confidence level for the age characteristic; By combining the restrictive features of age, race, and gender, a comprehensive demographic feature is obtained, as shown in the following formula: in, As a comprehensive characteristic of demographics, These are the restricted reinforcement features based on the object's age, gender, and race, respectively. The parameters of the trained population factor-restricted enhanced feature framework are fixed, and the facial expression object is input into the framework. The feature heatmaps of gender, age and race for individual categories, the comprehensive feature heatmap and the inverse gradient heatmap of the comprehensive feature are extracted from the framework. The feature heatmaps of gender, age and race for individual categories, the comprehensive feature heatmap and the inverse gradient heatmap of the comprehensive feature are used together as the population factor characteristic information space.

2. The method for facial expression recognition based on demographic factors according to claim 1, characterized in that, Step 1 specifically includes: performing face recognition on the image, aligning and cutting it to obtain the cut facial expression object.

3. The method for facial expression recognition based on demographic factors according to claim 1, characterized in that, Step 4 specifically includes: vertically connecting the facial expression object with the population factor characteristic information space.

4. The method for facial expression recognition based on demographic factors according to claim 1, characterized in that, Step 5 specifically includes: extracting high semantic features from multiple types of spatial feature information regions using the ResNet18 backbone network of the deep learning algorithm, thereby obtaining deep features of different categories of spatial feature information regions.

5. The facial expression recognition method based on demographic factors according to claim 4, characterized in that, Step 6 specifically includes: calculating the depth features of different categories of spatial feature information regions using a self-attention mechanism, analyzing the contribution of the depth features of different categories of spatial feature information regions to facial expression recognition, and obtaining the self-attention confidence of facial expressions for multiple types of spatial features.

6. The method for facial expression recognition based on demographic factors according to claim 1, characterized in that, Step 7 specifically includes: setting up a multi-feature attention mechanism, analyzing the relationship between population factor spatial features in the multi-type spatial features and the complementarity and correlation between population factor spatial features and facial expression object features through the self-attention confidence of the multi-type spatial features, and obtaining the facial multi-block feature attention confidence.

7. The method for facial expression recognition based on demographic factors according to claim 6, characterized in that, The contribution of different regions to facial expression judgment is measured by the attention confidence score of multi-region facial features. The specific process is expressed by the following formula: in: The confidence level of the adaptive features for all feature regions. The confidence level of the adaptive features for facial expression objects. Adaptive feature confidence for multi-type feature spatial blocks of population factors.

8. The method for facial expression recognition based on demographic factors according to claim 1, characterized in that, Step 8 specifically includes: non-uniformly fusing the attention confidence of the multi-block facial features with the corresponding regional features, fusing the expression task sharing ratio according to the regional features, obtaining the block adaptive features of the facial image, and judging facial expressions through the block adaptive features of the facial image.

Citation Information

Patent Citations

  • Anordning vid dörrar

    SE1000097A1

  • Attribute-level multi-modal emotion classification method based on cross-modal translation

    CN115186683A

  • Age identification method based on 3D face image multi-feature fusion

    CN115512403A