Cascade classification facial expression recognition method based on D-GFK network

Through a cascading classification method based on D-GFK network, combined with a generative adversarial network, an improved dense connection network and a graph convolutional neural network, the accuracy of face expression recognition in different lighting and occlusion environments is solved, especially in pessimistic expression recognition, which achieves higher recognition accuracy and model adaptability.

CN120088828APending Publication Date: 2025-06-03JIANGSU MOORE ACOUSTIC TECH RES INST CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510109320.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing facial expression recognition technology has low recognition accuracy under different lighting conditions and under occlusion environments, especially pessimistic expressions such as disgust and sadness are easily confused.

Method used

The cascading classified facial expression recognition method based on D-GFK network is adopted, and the occlusion area is modified by generating an adversarial network, and the improved dense connection network and graph convolution neural network are used to divide the face expressions in a coarse and fine granularity, and the multi-model cascading emotion recognition model is integrated to optimize the model to improve the recognition accuracy.

Benefits of technology

Improve the accuracy of facial expression recognition under different lighting conditions and under occlusion environments, especially in pessimistic expression recognition, which significantly improves the prediction accuracy and enables the model to quickly adapt to new environments and new tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088828A_ABST
    Figure CN120088828A_ABST
Patent Text Reader

Abstract

The invention discloses a cascade classification facial expression recognition method based on a D-GFK network, and relates to the technical field of image processing, the cascade classification facial expression recognition method comprises the following steps: collecting and preprocessing a facial image, extracting a shielding image from the preprocessed facial image, and inputting the shielding image into a generative adversarial network for correction; inputting the corrected face image into an improved dense connection network for coarse-grained division of face expressions; respectively inputting the corrected face image into an image convolutional neural network model and a face key point recognition model to carry out fine-grained division of face expressions; and fusing output results of the image convolutional neural network model and the face key point recognition model to obtain a multi-model cascade emotion recognition model. According to the method, the problem of facial expression recognition under different illumination conditions and in a shielding environment is solved, and meanwhile, the accuracy of facial expression recognition under pessimistic expression recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing. Specifically, it relates to a cascaded classification face expression recognition method based on a D-GFK network. Background Art

[0002] A person's expression usually reflects their inner emotions. Face expression recognition is a technology that infers a person's emotional state by analyzing human facial expressions, using artificial intelligence technology to recognize and interpret the expressions on a person's face. Well-known psychologists proposed the concept of six basic human emotions, and later a neutral expression was added, constituting seven basic expressions for face expression recognition, including anger, fear, disgust, happiness, sadness, surprise, and calmness, which are represented by labels 1 to 7 respectively. Face expression recognition has wide applications in fields such as social media analysis, psychological research, and user experience design. However, due to the changes in the subject's environment and the diversity of facial appearances, face expression recognition technology still faces challenges. When there is occlusion in the subject's environment or the light is complex, the recognition accuracy will decrease, and pessimistic expressions such as disgust and sadness are easily confused, which pose major challenges to face expression recognition.

[0003] Numerous researchers have tried to use different methods to improve the accuracy of face recognition. For example, facial images are divided into high-weight and low-weight ones, and the expression distributions are obtained through a convolutional neural network (CNN) and a graph convolutional (GCN) network respectively, and then the emotion labels are fused for prediction. This method well solves the problem of difficult recognition of face expressions under complex light and with partial occlusion, but has certain requirements for performance. Researchers proposed a geometric perception algorithm framework based on geometric and appearance knowledge, using a convolutional neural network to observe the entire face, and at the same time using an image convolutional network to mine the facial structure information behind different expressions, and inferring facial expressions from the visual and structural aspects. However, since facial geometric information is easily occluded, it does not have good stability in the case of blur or occlusion; researchers proposed a hierarchical attention network with progressive feature fusion, designed multiple feature extraction modules based on multiple feature aggregation blocks, fused different gradient features, improved the stability of the network under different lighting environments, and at the same time enhanced the feature discrimination of the image in the key parts of the face, improving the accuracy of face recognition. However, this network model is not lightweight enough and has limitations in complex environments; researchers proposed a facial expression recognition method based on an association graph (HASs), formulated the face recognition task as a vertex prediction problem, used vertex confidence to find high-order neighbors, used graph convolution to infer and find high-order neighbors, and determined the vertex category. It can achieve good accuracy with relatively low computing power, but there are problems such as too large number of parameters and excessive dependence on the graph structure, and a neural network needs to be added to improve the generalization ability.

[0004] Although facial expression recognition has broad application prospects in many fields, there are still some limitations. Human expressions are complex and variable. The same expression may have different interpretations, and there are also differences in expression expressions among different individuals. Secondly, facial expression recognition is sensitive to factors such as light, occlusion, and clarity. Minor changes may also affect the accuracy of facial expression recognition. Although single-model facial expression recognition can complete tasks, the overall accuracy is low. It has good recognition ability for expressions with large differences, such as happiness and anger, but has very weak recognition ability for expressions with small differences and is extremely prone to confusing expressions such as disgust and sadness. At the same time, the single model also has low recognition of expressions under different lighting conditions and in the presence of occlusion. Changes in lighting conditions may cause changes in features such as the contrast and brightness of the image, thus affecting the extraction and recognition of expression features. At the same time, in the case of occlusion, some facial features may be hidden or blurred, leading to algorithm recognition deviation.

[0005] In response to the problems in the related art, no effective solution has been proposed yet. Summary of the Invention

[0006] In response to the problems in the related art, the present invention proposes a cascaded classification facial expression recognition method based on the D-GFK network to overcome the above-mentioned technical problems existing in the existing related art.

[0007] Therefore, the specific technical solution adopted by the present invention is as follows:

[0008] A cascaded classification facial expression recognition method based on the D-GFK network, the cascaded classification facial expression recognition method includes the following steps:

[0009] Collect facial images and perform preprocessing, extract occluded images from the preprocessed facial images and input them into the generative adversarial network for correction;

[0010] Input the corrected facial images into the improved dense connection network for coarse-grained division of facial expressions;

[0011] Input the corrected facial images into the graph convolutional neural network model and the facial key point recognition model respectively for fine-grained division of facial expressions;

[0012] Fuse the output results of the graph convolutional neural network model and the facial key point recognition model to obtain a multi-model cascaded emotion recognition model, and optimize the multi-model cascaded emotion recognition model.

[0013] Preferably, collecting facial images and performing preprocessing, extracting occluded images from the preprocessed facial images and inputting them into the generative adversarial network for correction includes the following steps:

[0014] Collect a face image using an image capture device, and identify the face region in the face image using a Haar cascade classifier;

[0015] Crop the face region from the face image; convert the cropped face image into a face grayscale image, and perform image enhancement processing on the face grayscale image;

[0016] Determine whether there is occlusion in the face grayscale image after image enhancement processing. If there is occlusion in the face grayscale image, use a generative adversarial network to correct the occluded area of the image.

[0017] Preferably, identifying the face region in the face image using a Haar cascade classifier includes the following steps:

[0018] Calculate the integral image of the face image, and use the integral image to obtain the sum of pixel values of any rectangular region in the face image;

[0019] Calculate the Haar feature value based on the sum of pixel values of the rectangular region, and measure the pixel intensity difference of the rectangular region in the face image through the Haar feature value;

[0020] Train a number of weak classifiers in sequence to form a strong classifier, and identify the face region in the face image in combination with the Haar feature value.

[0021] Preferably, inputting the corrected face image into an improved dense connection network for coarse-grained division of facial expressions includes the following steps:

[0022] Improve two dense connection blocks of the dense connection network into a 2D selective scanning module based on VMamba;

[0023] Set and adjust the configuration parameters of each layer of neural network in the improved dense connection network, and initialize the input weights of each layer of neural network;

[0024] Introduce a global attention mechanism into the improved dense connection network, and use the corrected face image as a training set to train the improved dense connection network;

[0025] Continuously repeat the training process of the improved dense connection network until the recognition accuracy of facial expressions output by the improved dense connection network meets the predetermined value.

[0026] Preferably, inputting the corrected face image into a graph convolutional neural network model and a face key point recognition model respectively for fine-grained division of facial expressions includes the following steps:

[0027] Establish a graph convolutional neural network model based on the improved dense connection network, and input the corrected face image into the graph convolutional neural network model to obtain a feature matrix based on global features;

[0028] Input the corrected face image into the face key point recognition model to generate face key point coordinates, calculate the face key point coordinates, and obtain a feature matrix based on local features;

[0029] Fuse the feature matrix based on global features and the feature matrix based on local features to obtain a fused feature;

[0030] Generate an association graph from the fused feature through dimensionality reduction technology, and obtain a first prediction result through the association graph; input the face key point coordinates into the graph convolutional neural network model to output a second prediction result; fuse and output the first prediction result and the second prediction result.

[0031] Preferably, inputting the corrected face image into the face key point recognition model to generate face key point coordinates, calculating the face key point coordinates, and obtaining a feature matrix based on local features includes the following steps:

[0032] Use an edge detection algorithm to extract face feature points in the corrected face image, and calculate face geometric information based on the face feature points;

[0033] Extract features from the face geometric information, and establish a face key point recognition model based on the feature extraction results;

[0034] Use the corrected face image as the input of the face key point recognition model, and output the face key point coordinates through the face key point recognition model;

[0035] Calculate the feature vector of the face key points based on the face key point coordinates, and use the feature vector of the face key points as the feature matrix based on local features.

[0036] Preferably, generating an association graph from the fused feature through dimensionality reduction technology and obtaining a first prediction result through the association graph includes the following steps:

[0037] Obtain the high-dimensional features corresponding to each face image in the existing dataset, and reduce the high-dimensional space where the high-dimensional features are located to a two-dimensional space;

[0038] Correspond each face image to the two-dimensional coordinates, mark the two-dimensional coordinates in the two-dimensional coordinate system in the form of points, and output the trained association graph;

[0039] Extract the fused feature and input it into the trained association graph, obtain the node information associated with the fused feature, and calculate the Euclidean distance between the new node and the remaining associated nodes to generate a first prediction result of the facial expression.

[0040] Preferably, obtaining the high-dimensional features corresponding to each face image in the existing dataset includes the following steps:

[0041] Calculate the similarity of high-dimensional data, calculate the joint conditional probability of each pair of data points in the high-dimensional space based on the similarity of high-dimensional data, and ensure the symmetry of the joint probability of each pair of data points;

[0042] Calculate the similarity of low-dimensional data, calculate the asymmetric metric based on the similarity of low-dimensional data, and realize that the data point distribution in the low-dimensional space is close to the data distribution in the high-dimensional space.

[0043] Preferably, fuse the output results of the graph convolutional neural network model and the facial key point recognition model to obtain a multi-model cascaded emotion recognition model, and optimize the multi-model cascaded emotion recognition model, including the following steps:

[0044] Record the accuracy rates of the classification results output by the graph convolutional neural network model and the facial key point recognition model respectively;

[0045] Fuse the output results of the graph convolutional neural network model and the facial key point recognition model to generate a multi-model cascaded emotion recognition model, and input the facial images with correct classification results into the multi-model cascaded emotion recognition model to learn new image data;

[0046] Perform weighted fusion on the output results of the multi-model cascaded emotion recognition model to obtain the predicted probability values of the new facial expression under different labels;

[0047] Extract the maximum predicted probability value of the facial expression under different labels, and output the label corresponding to the maximum predicted probability value as the predicted result of the new facial expression.

[0048] Preferably, the expression of the integral image is:

[0049] I sum (x,y) = Σ x'≤x,y'≤y I(x',y');

[0050] In the formula, I sum (x,y) represents the pixel value of the pixel point (x,y) in the integral image;

[0051] I(x',y') represents the pixel value of the original pixel point (x',y') in the original image;

[0052] x represents the abscissa of the pixel point;

[0053] y represents the ordinate of the pixel point;

[0054] x′ represents the abscissa of the original pixel point;

[0055] y′ represents the ordinate of the original pixel point.

[0056] The beneficial effects of the present invention are:

[0057] 1. A cascade classification face expression recognition method based on the D-GFK network provided by the present invention is used for face expression recognition, which solves the problems of face expression recognition under different lighting conditions and in occluded environments, and at the same time improves the accuracy of face expression recognition under pessimistic expression recognition.

[0058] 2. The present invention uses an improved cascade classification neural network model to conduct preliminary classification of face expressions, then uses the neural network model to extract the features of the training set, and at the same time constructs an association graph to reclassify each vertex, solving the problem that traditional neural network classification only focuses on global features and lacks attention to local features, resulting in insufficient prediction accuracy for pessimistic expressions. Furthermore, it effectively improves the prediction accuracy of pessimistic expressions.

[0059] 3. The present invention fuses the results of the two models to obtain a multi-model cascade emotion recognition model based on the dense connection network and the graph convolutional network, and adds the image data with correct predictions to the existing graph convolutional neural network model and face key point model to continuously learn new data and gradually improve the performance. Furthermore, the model can quickly adapt to new environments and new tasks and maintain high-efficient emotion recognition capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0061] Figure 1 is a flowchart of a cascade classification face expression recognition method based on the D-GFK network according to an embodiment of the present invention;

[0062] Figure 2 is a specific implementation schematic diagram of training an improved dense connection network in a cascade classification face expression recognition method based on the D-GFK network according to an embodiment of the present invention;

[0063] Figure 3 is a specific implementation schematic diagram of a graph convolutional neural network in a cascade classification face expression recognition method based on the D-GFK network according to an embodiment of the present invention;

[0064] Figure 4 is a specific implementation schematic diagram of a face key point recognition model in a cascade classification face expression recognition method based on the D-GFK network according to an embodiment of the present invention;

[0065] Figure 5It is a specific implementation schematic diagram of a cascaded emotion recognition model in a cascaded classification face expression recognition method based on the D-GFK network according to an embodiment of the present invention. Specific implementation manner

[0066] To further illustrate each embodiment, the present invention provides accompanying drawings. These drawings are part of the disclosure of the present invention, mainly used to illustrate the embodiments, and can be used in conjunction with the relevant descriptions in the specification to explain the operating principle of the embodiments. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention.

[0067] Since traditional algorithms only focus on global features and ignore local features; traditional algorithm models are relatively large and have high requirements for devices; traditional algorithms have insufficient prediction accuracy for pessimistic expressions; traditional algorithms have insufficient prediction accuracy under occluded environments and lighting changes; traditional algorithms usually only focus on global features and do not pay attention to the geometric features of the human face; traditional algorithms cannot update the model in a timely manner. In view of the above defects of traditional algorithms, the present invention uses a cascaded classification network based on the D-GFK network (DenseNet-Graph Convolutional Network and Face Key Point Network) for face expression prediction, complements the occluded area through a generative adversarial network, and predicts optimistic and calm expressions through an improved convolutional neural network, obtains the results of re-prediction in the input graph convolutional network and face key point recognition, and finally predicts the expressions separately through a hierarchical cascaded network, thereby improving the accuracy of the prediction results.

[0068] According to an embodiment of the present invention, a cascaded classification face expression recognition method based on the D-GFK network is provided.

[0069] Now, the present invention will be further described in conjunction with the accompanying drawings and specific implementation manners, as Figure 1 shown, the cascaded classification face expression recognition method based on the D-GFK network according to an embodiment of the present invention includes the following steps:

[0070] S1. Collect a face image and perform preprocessing, and extract an occluded image from the preprocessed face image and input it into a generative adversarial network for correction.

[0071] Among them, collecting a face image and performing preprocessing, and extracting an occluded image from the preprocessed face image and inputting it into a generative adversarial network for correction includes the following steps:

[0072] Use an image capture device to collect a face image, and use a Haar cascade classifier to identify the face area in the face image.

[0073] Among them, identifying the face region in the face image using the Haar cascade classifier includes the following steps:

[0074] Calculate the integral image of the face image, and use the integral image to obtain the sum of pixel values of any rectangular region in the face image.

[0075] Calculate the Haar feature value based on the sum of pixel values of the rectangular region (the Haar feature is a convolution template, and the Haar feature includes edge features, linear features, central features, and diagonal features, which are combined into a feature template), and measure the pixel intensity difference of the rectangular region in the face image through the Haar feature value.

[0076] Train a number of weak classifiers in sequence to form a strong classifier, and combine the Haar feature value to identify the face region in the face image.

[0077] Crop the face region from the face image; convert the cropped face image into a face grayscale image, and perform image enhancement processing on the face grayscale image.

[0078] Judge whether there is occlusion in the face grayscale image after image enhancement processing. If there is occlusion in the face grayscale image, use the generative adversarial network to correct the occluded area of the image.

[0079] It should be noted that as Figure 2 shown, first take the original face image as the input. Since there is a lot of interference information in the original picture that is irrelevant to the extraction of expression features, it is necessary to preprocess the original face image; use the Haar cascade classifier to translate the window from left to right and from top to bottom to detect the face region in the image, return the coordinates of the rectangular box containing the face region, and exclude the interference information; use the obtained coordinate information to frame a rectangle in the image and crop it, and then recalculate and allocate the pixel values of the image to unify the size of all images to 100×100 pixels; perform grayscale processing on the normalized image, replace each pixel point with a grayscale value, and divide the dataset into two parts: occluded and non-occluded using the existing dataset. Use the generative adversarial network to train a pre-trained model for judging the occlusion situation of the picture, and input the picture after grayscale processing into the generative adversarial network to complete the occlusion part.

[0080] The principle of the Haar cascade classifier is to efficiently detect the target object in the image by quickly calculating simple Haar features, using the AdaBoost algorithm (Adaptive Boosting) to select the most effective features, and arranging multiple weak classifiers hierarchically. Its specific methods include:

[0081] When extracting Haar features, calculate the integral image:

[0082] I sum (x,y) = Σ x'≤x,y'≤y I(x',y');

[0083] In the formula, I sum (x,y) represents the pixel value of the pixel point (x,y) in the integral image; I(x',y') represents the pixel value of the original pixel point (x',y') in the original image; x represents the abscissa of the pixel point; y represents the ordinate of the pixel point; x′ represents the abscissa of the original pixel point; y′ represents the ordinate of the original pixel point.

[0084] Calculate the sum of pixel values in the rectangular region:

[0085] S = I sum (x 2 ,y 2 ) - I sum (x 1 - 1,y 2 ) - I sum (x 2 ,y 1 - 1) + I sum (x 1 - 1,y 1 - 1);

[0086] In the formula, (x 1 ,y 1 ) and (x 2 ,y 2 ) represent the coordinates of the upper left corner and the lower right corner of the rectangular region respectively.

[0087] Calculate the Haar feature value:

[0088] feature = S white - S black ;

[0089] In the formula, S white and S black are the sums of pixels in the white and black rectangular regions respectively, where the judgment and selection of the white and black regions are implemented through the AdaBoost algorithm.

[0090] It should be noted that the AdaBoost algorithm includes:

[0091] Calculate the initial sample weights:

[0092]

[0093] In the formula, represents the weight of the i-th sample in the first round of iteration, and N represents the total number of samples.

[0094] Training weak classifiers:

[0095]

[0096] Where \(t\) represents the number of iterations, represents the weight of the \(i\)-th sample in the \(t\)-th iteration, \(N\) represents the total number of samples, and \(\varepsilon\) t represents the classification error of the \(t\)-th weak classifier, \(I\) represents the indicator function, and \(h\) t (\(x\) i ) represents the weak classifier, which refers to the prediction result of the \(i\)-th sample \(x\) i , and \(y\) i represents the true label of the \(i\)-th sample.

[0097] Calculating the weight of the weak classifier:

[0098]

[0099] Where \(\alpha\) t represents the weight of the \(t\)-th weak classifier, and \(\varepsilon\) t represents the classification error of the \(t\)-th weak classifier.

[0100] The final strong classifier \(H(x)\) is the weighted sum of each weak classifier:

[0101]

[0102] Where \(T\) represents the number of iterations, \(\alpha\) t represents the weights of \(t\) weak classifiers, \(h\) t (\(x\) i ) represents the weak classifier, and \(\text{sign}(\cdot)\) represents the sign function, specifically returning the sign of the input.

[0103] It should be noted that the gray value of each pixel is calculated as follows:

[0104] Gray = 0.299×R + 0.587×G + 0.144×B;

[0105] Where Gray represents the gray value, R represents the pixel value of the red channel, G represents the pixel value of the green channel, and B represents the pixel value of the blue channel.

[0106] S2. Input the corrected face image into the improved dense connection network for coarse-grained division of facial expressions.

[0107] Among them, inputting the corrected face image into the improved dense connection network for coarse-grained division of facial expressions includes the following steps:

[0108] Improve two dense connection blocks of the dense connection network into a 2D Selective Scan Module (SS2D) based on VMamba;

[0109] Set and adjust the configuration parameters of each layer of the neural network in the improved dense connection network, and initialize the input weights of each layer of the neural network.

[0110] Introduce a global attention mechanism into the improved dense connection network, and use the corrected face images as the training set to train the improved dense connection network.

[0111] Continuously repeat the training process of the improved dense connection network until the recognition accuracy of the face expression output by the improved dense connection network meets the predetermined value.

[0112] It should be noted that, as Figure 3 shown, input the preprocessed training set into the improved dense connection network, set the number and position distribution of the convolutional layer, pooling layer, and fully connected layer of the improved dense convolutional neural network, and determine the overall model structure.

[0113] Adjust various parameters of the convolutional layer, pooling layer, and fully connected layer of each layer according to the above structure, such as the convolutional kernel size, internal parameters of the convolutional kernel, types of pooling matrices, sizes of pooling matrices and moving strides, types of activation functions of the fully connected layer, etc., initialize the input weights of each layer of the neural network, and then use the grayscale face expression images for model training.

[0114] It should be noted that the improved dense connection network (the dense connection network is a dense convolutional neural network) consists of alternately connected dense blocks and transition layers. In the dense blocks, each layer is directly connected to all subsequent layers to enhance feature transmission. Therefore, each subsequent layer will receive the feature maps from all previous layers. Among them, denote X r as the output layer of the r-th layer, and its expression:

[0115] X r = H r ([X 0 , X 1 ,..., X r-1 );

[0116] In the formula, [X 0 , X 1 ,..., X r-1 represents the concatenation of the feature maps generated in layers 0,..., r - 1, and H r represents the r-th dense block.

[0117] Compared with the traditional dense connection network, the improvement of the present invention lies in that first, the first two dense connection blocks of the dense connection network are improved into a 2D selective scanning module based on VMamba, and its expression is: the input image is X∈H×W×C, where S represents the image, H represents the image height, W represents the image width, and C represents the number of image channels. Then, the input two-dimensional feature map is flattened into a one-dimensional vector along four different directions, where the four directions are from top left to bottom right, from bottom left to top right, from bottom right to top left, and from top right to bottom left. Then, feature extraction is performed on each scanning path through an independent selective scanning spatial state sequence model, and its state update formula is:

[0118]

[0119] In the formula, h a+1 represents the hidden state at time step a + 1, represents the exponential matrix of the state transition matrix A at time step Δ a , h a represents the hidden state at time step a, B a represents the input matrix at time step a, u a represents the input vector at time step a, represents the negative exponential matrix of the state transition matrix A at time step Δ a , Δ a represents the time step from time step a to a + 1.

[0120] The output formula is:

[0121]

[0122] In the formula, w T represents the weight vector at time T, ⊙ represents the Hadamard product, K i represents the key matrix at time i, V i represents the value matrix at time i, represents the transpose matrix of K i , w i represents the weight vector at time i.

[0123] Secondly, the traditional classifier is modified into a classifier with a complete binary tree structure, and its expression is: Tree = {V, ε}, where Tree represents the complete binary tree, V = {ν 1 ,..., ν n}, V represents the nodes, {ν 1 ,..., ν n} represents the 1st to the nth nodes, n represents the total number of nodes, ε = {e 1 ,..., e k}, ε represents the edges, {e 1,...,e k} represents the 1st to the kth edges, where k represents the total number of edges, and each leaf node represents a classification result.

[0124] S3. Input the corrected face images into the graph convolutional neural network model and the face key point recognition model respectively for fine-grained division of facial expressions.

[0125] Among them, inputting the corrected face images into the graph convolutional neural network model and the face key point recognition model respectively for fine-grained division of facial expressions includes the following steps:

[0126] Build a graph convolutional neural network model based on the improved dense connection network, and input the corrected face images into the graph convolutional neural network model to obtain a feature matrix based on global features.

[0127] Input the corrected face images into the face key point recognition model to generate face key point coordinates, and calculate the face key point coordinates to obtain a feature matrix based on local features.

[0128] Among them, inputting the corrected face images into the face key point recognition model to generate face key point coordinates, and calculating the face key point coordinates to obtain a feature matrix based on local features includes the following steps:

[0129] Use the edge detection algorithm to extract face feature points in the corrected face images, and calculate the face geometric information based on the face feature points.

[0130] Extract features from the face geometric information, and build a face key point recognition model based on the feature extraction results.

[0131] Take the corrected face images as the input of the face key point recognition model, and output the face key point coordinates through the face key point recognition model.

[0132] Calculate the feature vectors of the face key points based on the face key point coordinates, and use the feature vectors of the face key points as the feature matrix based on local features.

[0133] Fuse the feature matrix based on global features and the feature matrix based on local features to obtain the fused features.

[0134] Generate an association graph from the fused features through dimensionality reduction technology, and obtain the first prediction result through the association graph; input the face key point coordinates into the graph convolutional neural network model to output the second prediction result; fuse and output the first prediction result and the second prediction result.

[0135] Among them, generating an association graph from the fused features through dimensionality reduction technology, and obtaining the first prediction result through the association graph includes the following steps:

[0136] Obtain the high-dimensional features corresponding to each face image in the existing dataset, and reduce the high-dimensional space where the high-dimensional features are located to a two-dimensional space.

[0137] Among them, obtaining the high-dimensional features corresponding to each face image in the existing dataset includes the following steps:

[0138] Calculate the similarity of high-dimensional data, and calculate the joint conditional probability of each pair of data points in the high-dimensional space based on the similarity of high-dimensional data, while ensuring the symmetry of the joint probability of each pair of data points.

[0139] Calculate the similarity of low-dimensional data, and calculate the asymmetric metric based on the similarity of low-dimensional data to make the data point distribution in the low-dimensional space close to the data distribution in the high-dimensional space.

[0140] Correspond each face image to two-dimensional coordinates, mark the two-dimensional coordinates in the two-dimensional coordinate system in the form of points, and output the completed training association graph.

[0141] Extract the fusion features and input them into the completed training association graph, obtain the node information associated with the fusion features, and calculate the Euclidean distance between the new node and the remaining associated nodes to generate the first prediction result of the facial expression.

[0142] It should be noted that as Figure 4 shown, first use the existing dataset to conduct a large number of trainings in an improved dense connection network to obtain a primary model in pth format, and then input the preprocessed facial expression map to be predicted into the already trained primary model to obtain the global facial expression features; use the existing dataset again, and conduct a large number of trainings using the regression tree model to obtain a facial key point recognition model in dat format.

[0143] Input the preprocessed facial expression image to be predicted into the facial key point recognition model to obtain the facial key point coordinates, perform operations on the obtained facial key point coordinates to obtain local feature data, and fuse the global features output by the convolutional model with the local features obtained from the facial key points to obtain the fused features.

[0144] Use the t-SNE (t-distributed Stochastic Neighbor Embedding) dimensionality reduction technology to draw an association graph for the fused features, obtain the prediction result through the association graph, input the facial key points into the convolutional neural network to output the prediction result, and finally fuse the two prediction results for output.

[0145] It should be noted that a pre-trained model is used to extract features to construct an association graph. By extracting the features of different face images and combining them in the form of a graph, the relationships between the training set images are captured. Using a graph convolutional neural network for label prediction can make full use of the topological information in the graph structure, further improving the performance and stability of the model.

[0146] It should be noted that the feature formula of the image is extracted using the trained model:

[0147] v = reshape(ConvOut, H×W×C);

[0148] In the formula, v represents the feature vector, reshape represents a function, ConvOut represents the convolutional output, and H, W, and C respectively represent the height, width, and number of channels of the convolutional layer output.

[0149] Output the feature vector V based on the output of the fully connected layer of the convolutional neural network 1 , use the coordinates of the output face key points, and through the calculation of the key point coordinate data, obtain data such as eyebrow height, eyebrow width, eye height, eye width, mouth height, and mouth width, and output the feature vector V based on the recognition of face key points 2 , the feature vector V output by the fully connected layer of the convolutional neural network 1 and the feature vector V output by the face key point recognition 2 are concatenated to obtain a richer and more representative feature representation, and its expression is:

[0150] V con = [V 1 , V 2 ;

[0151] Use the t-SNE algorithm to reduce the dimension of the extracted features, and reduce the high-dimensional feature V con to a two-dimensional space. Among them, the specific steps of the t-SNE algorithm include:

[0152] Calculate the similarity of high-dimensional data using the Gaussian distribution, and its calculation formula is:

[0153]

[0154] In the formula, p j|i is the conditional probability that x i and x j are neighbors, x i , x j are data points in the high-dimensional space, σ i is the standard deviation of the Gaussian kernel related to the point x i , ‖x i -x j ‖ 2is the square of the Euclidean distance between point x i and point x j ; exp is the exponential function, and k is other data points in the high-dimensional space.

[0155] Calculate the joint probability of each pair of data points in the high-dimensional space to ensure symmetry. Its symmetrization formula is as follows:

[0156]

[0157] In the formula, p ij is the symmetric joint probability between point x i and point x j ; p j|i is the conditional probability that x i selects point x j as a neighbor; p i|j is the conditional probability that x j selects point x i as a neighbor, and N is the total number of data points.

[0158] Calculate the similarity of low-dimensional data. Its calculation formula is:

[0159]

[0160] In the formula, q ij is the similarity between point y i and point y j in the low-dimensional space; y i and y j are data points in the low-dimensional space; ‖y i - y j ‖ 2 is the square of the Euclidean distance between point y i and point y j ; k and l are other data points in the low-dimensional space; (1 + ||y i - y j || 2 ) -1 is the similarity measure in the t-distribution.

[0161] Calculate the KL divergence that minimizes the difference. Its calculation formula is:

[0162]

[0163] In the formula, C represents the KL divergence, specifically the objective function to be optimized; P i and Q i represent the probability distributions in the high-dimensional space and the low-dimensional space respectively; p i j and q i j represent the samples x i and xj The conditional probability distribution between, ∑ i and ∑ j respectively represent the samples x in the high-dimensional space i and x j for summation, represents the relative distance between samples in two spaces. i and j represent indices on different dimensions and are used to describe the two-dimensional probability distribution.

[0164] Based on the above t-SNE algorithm, the high-dimensional features corresponding to each picture in the existing dataset are obtained, and then reduced to a two-dimensional space. Each picture corresponds to a set of two-dimensional coordinates, and the two-dimensional coordinates are marked in the two-dimensional coordinate system in the form of points, and a trained association graph is output.

[0165] When a face image is input, its features are extracted and put into the already trained association graph to obtain the associated node information. By calculating the Euclidean distance between the new node and other associated nodes, the number of labels corresponding to the associated nodes is comprehensively used to predict the facial expression. The calculation formula of the Euclidean distance is as follows:

[0166]

[0167] In the formula, d represents the Euclidean distance, (x 1 , y 1 ) and (x 2 , y 2 ) respectively represent the coordinates of two data points.

[0168] The facial expression recognition method based on facial key points is a technology that uses the key point information on the human face to infer facial expressions. In this method, first, the face is located through a face detection algorithm, and then the key points with semantic meanings on the human face are identified using a key point detection algorithm.

[0169] It should be noted that the Canny edge detection algorithm is often used to locate the edge of the face in facial key point detection, and then assist in extracting the feature points of the face. Its calculation formula is as follows:

[0170]

[0171] In the formula, G x and G y respectively represent the gradients of the image in the horizontal and vertical directions. By calculating the gradient magnitude, the edge information in the image can be found, and then the geometric information of the face can be obtained.

[0172] Extract the geometric features, use graph convolution to obtain a model based on facial key points, extract the facial key point features by inputting the face image, and input it into the graph neural network to obtain the predicted facial expression. Its calculation formula is as follows:

[0173]

[0174] In the formula, A represents the adjacency matrix,  represents the processed adjacency matrix, D represents the transition matrix, H (1) represents the input feature, and W (1) represents the network parameter, represents the activation function. Among them, the relationship formula between A and  is:

[0175]

[0176] In facial expression recognition, the convolutional neural network pays more attention to global information, while facial key point recognition pays more attention to local information. Only focusing on limited information cannot accurately recognize expressions. Therefore, the present invention uses a method of feature fusion based on the D-GFK network to splice the feature vector output from the fully connected layer of the convolutional neural network and the feature vector output from facial key point recognition to obtain richer and more representative features.

[0177] S4. Combine the output results of the graph convolutional neural network model and the facial key point recognition model to obtain a multi-model cascaded emotion recognition model, and optimize the multi-model cascaded emotion recognition model.

[0178] Among them, combining the output results of the graph convolutional neural network model and the facial key point recognition model to obtain a multi-model cascaded emotion recognition model, and optimizing the multi-model cascaded emotion recognition model includes the following steps:

[0179] Record the accuracy rates of the classification results output by the graph convolutional neural network model and the facial key point recognition model respectively.

[0180] Combine the output results of the graph convolutional neural network model and the facial key point recognition model to generate a multi-model cascaded emotion recognition model, and input the facial images with correct classification results into the multi-model cascaded emotion recognition model to learn new image data.

[0181] Perform weighted fusion on the output results of the multi-model cascaded emotion recognition model to obtain the predicted probability values of the new facial expression under different labels.

[0182] Extract the maximum predicted probability value of the facial expression under different labels, and output the label corresponding to the maximum predicted probability value as the predicted result of the new facial expression.

[0183] It should be noted that, such as Figure 5As shown in the figure, the present invention corrects the prediction result of the previous network by combining the recognition advantages of multiple models. When the improved dense connection network outputs an optimistic expression prediction, the prediction result is directly output; when a pessimistic expression is output, it enters the subsequent model for more accurate prediction. Furthermore, through this cascade classification mechanism, the advantages of multiple models can be more accurately integrated, and the accuracy of the recognition result can be improved.

[0184] Therefore, a cascade classification mechanism is added to multiple models, and it is determined whether to enter the secondary prediction by judging the primary prediction result. The dense connection network is used to make the first prediction on the input image, and its calculation formula is as follows:

[0185]

[0186] In the formula, y represents the prediction result, and R represents the recognition result.

[0187] According to the first prediction result, when the predicted value is not an "optimistic" or "calm" expression, it is input into the secondary network, and finally the average of the probability distributions is used to predict the final result, and its calculation formula is as follows:

[0188]

[0189] In the formula, Average represents the comprehensive probability distribution, P GCN represents the probability distribution of the graph convolutional neural network model, and P KeyPoint represents the probability distribution of the facial key point recognition model.

[0190] When a new picture is input and is manually marked as a correct prediction, its features are recorded and added as new data to the graph convolutional neural network model and the facial key point recognition model, thereby improving the performance and accuracy of the model.

[0191] In summary, by means of the above technical solutions of the present invention, a cascaded classification facial expression recognition method based on the D-GFK network provided by the present invention is used for facial expression recognition, which solves the problems of facial expression recognition under different lighting conditions and in an occluded environment, and at the same time improves the accuracy of facial expression recognition under pessimistic expression recognition. The present invention uses an improved cascaded classification neural network model to perform preliminary classification of facial expressions, then uses the neural network model to extract the features of the training set, and at the same time constructs an association graph to classify each vertex again, solving the problem that traditional neural network classification only focuses on global features and lacks attention to local features, resulting in insufficient prediction accuracy for pessimistic expressions, and thus effectively improving the prediction accuracy of pessimistic expressions. The present invention fuses the results of the two models to obtain a multi-model cascaded emotion recognition model based on the dense connection network and the graph convolutional network, and adds the image data with correct predictions to the existing graph convolutional neural network model and the facial key point model to continuously learn new data and gradually improve the performance, so that the model can quickly adapt to new environments and new tasks and maintain high-efficient emotion recognition ability.

[0192] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A cascade classification facial expression recognition method based on D-GFK network, characterized in that: The cascade classification facial expression recognition method comprises the following steps: Collect and preprocess the face image, extract the occluded image from the preprocessed face image and input it into the generative adversarial network for correction; Input the corrected face image into the improved densely connected network to perform coarse-grained division of facial expressions; The corrected face image is input into the graph convolutional neural network model and the face key point recognition model respectively to perform fine-grained division of facial expressions; The output results of the graph convolutional neural network model and the facial key point recognition model are integrated to obtain a multi-model cascade emotion recognition model, and the multi-model cascade emotion recognition model is optimized.

2. A cascade classification facial expression recognition method based on D-GFK network according to claim 1, characterized in that: The method of collecting and preprocessing a face image, extracting an occluded image from the preprocessed face image and inputting it into a generative adversarial network for correction comprises the following steps: Using an image capture device to collect a face image, and using a cascade classifier to identify a face region in the face image; Cropping a face region from a face image; converting the cropped face image into a face grayscale image, and performing image enhancement processing on the face grayscale image; Determine whether there is occlusion in the grayscale image of the face after image enhancement processing. If there is occlusion in the grayscale image of the face, use the generative adversarial network to correct the occluded area of ​​the image.

3. A cascade classification facial expression recognition method based on D-GFK network according to claim 2, characterized in that: The method of using a cascade classifier to identify a face region in a face image comprises the following steps: Calculate the integral image of the face image, and use the integral image to obtain the sum of pixel values ​​of any rectangular area in the face image; The eigenvalue is calculated based on the sum of the pixel values ​​of the rectangular area, and the eigenvalue is used to measure the difference in pixel intensity of the rectangular area in the face image; Several weak classifiers are trained in sequence to form a strong classifier, and the facial area in the face image is recognized by combining the feature values.

4. A cascade classification facial expression recognition method based on D-GFK network according to claim 3, characterized in that: The step of inputting the corrected face image into the improved densely connected network to perform coarse-grained segmentation of facial expressions comprises the following steps: Improve the two densely connected blocks of the densely connected network into a 2D selective scanning module based on VMamba; Set and adjust the configuration parameters of each layer of the neural network in the improved densely connected network, and initialize the input weights of each layer of the neural network; A global attention mechanism is introduced into the improved densely connected network, and the modified face images are used as training sets to train the improved densely connected network; The training process of the improved densely connected network is continuously repeated until the recognition accuracy of the facial expression output by the improved densely connected network meets a predetermined value.

5. A cascade classification facial expression recognition method based on D-GFK network according to claim 4, characterized in that: The step of inputting the corrected face image into the graph convolutional neural network model and the face key point recognition model for fine-grained division of facial expressions comprises the following steps: A graph convolutional neural network model is established based on the improved densely connected network, and the corrected face image is input into the graph convolutional neural network model to obtain a feature matrix based on global features; Input the corrected face image into the face key point recognition model to generate the face key point coordinates, and calculate the face key point coordinates to obtain a feature matrix based on local features; The feature matrix based on global features and the feature matrix based on local features are fused to obtain fused features; The fusion features are used to generate a correlation graph through dimensionality reduction technology, and the first prediction result is obtained through the correlation graph; the coordinates of the facial key points are input into the graph convolutional neural network model to output the second prediction result; the first prediction result and the second prediction result are fused and output.

6. A cascade classification facial expression recognition method based on D-GFK network according to claim 5, characterized in that: The step of inputting the corrected face image into the face key point recognition model to generate the face key point coordinates, and calculating the face key point coordinates to obtain a feature matrix based on local features comprises the following steps: Extracting facial feature points from the corrected face image using an edge detection algorithm, and calculating facial geometric information based on the facial feature points; Extract features from facial geometry information and establish a facial key point recognition model based on the feature extraction results; The corrected face image is used as the input of the face key point recognition model, and the face key point recognition model outputs the coordinates of the face key points; The feature vectors of the facial key points are calculated based on the coordinates of the facial key points, and the feature vectors of the facial key points are used as feature matrices based on local features.

7. A cascade classification facial expression recognition method based on D-GFK network according to claim 6, characterized in that: The step of generating a correlation graph by using a dimensionality reduction technique for the fusion features and obtaining a first prediction result by using the correlation graph comprises the following steps: Obtain the high-dimensional features corresponding to each face image in the existing data set, and reduce the high-dimensional space where the high-dimensional features are located to a two-dimensional space; Correspond each face image to the two-dimensional coordinates, mark the two-dimensional coordinates in the two-dimensional coordinate system in the form of points, and output the trained association graph; The fusion feature is extracted and input into the trained association graph to obtain the node information associated with the fusion feature, and the Euclidean distance between the new node and the remaining associated nodes is calculated to generate the first prediction result of the facial expression.

8. A cascade classification facial expression recognition method based on D-GFK network according to claim 7, characterized in that: The method of obtaining the high-dimensional features corresponding to each face image in the existing data set comprises the following steps: Calculate the similarity of high-dimensional data, and calculate the joint conditional probability of each pair of data points in the high-dimensional space based on the similarity of high-dimensional data, while ensuring the symmetry of the joint probability of each pair of data points; Calculate the similarity of low-dimensional data, and calculate asymmetric metrics based on the similarity of low-dimensional data to make the distribution of data points in low-dimensional space close to the data distribution in high-dimensional space.

9. A cascade classification facial expression recognition method based on D-GFK network according to claim 8, characterized in that: The output results of the fused graph convolutional neural network model and the face key point recognition model are used to obtain a multi-model cascade emotion recognition model, and optimizing the multi-model cascade emotion recognition model includes the following steps: Record the accuracy of the classification results output by the graph convolutional neural network model and the face key point recognition model respectively; The output results of the graph convolutional neural network model and the facial key point recognition model are integrated to generate a multi-model cascade emotion recognition model, and the facial images with correct classification results are input into the multi-model cascade emotion recognition model to learn new image data; Perform weighted fusion on the output results of the multi-model cascade emotion recognition model to obtain the predicted probability values ​​of new facial expressions under different labels; The maximum predicted probability value of facial expressions under different labels is extracted, and the label corresponding to the maximum predicted probability value is output as the new facial expression prediction result.

10. A cascade classification facial expression recognition method based on D-GFK network according to claim 9, characterized in that: The expression of the integral graph is: I sum (x,y)=∑ x'≤x,y'≤y I(x',y'); In the formula, I sum (x, y) represents the pixel value of the pixel point (x, y) in the integral image; I(x',y') represents the pixel value of the original pixel point (x',y') in the original image; x represents the horizontal coordinate of the pixel; y represents the vertical coordinate of the pixel; x′ represents the horizontal coordinate of the original pixel; y′ represents the vertical coordinate of the original pixel.