Training Method and Recognition Method of Facial Expression Recognition Model Based on Feature Fusion
By using feature fusion and contrast learning methods in the expression recognition model, and using attention weights for feature fusion and uncertainty description, the problem of poor expression recognition performance in the prior art is solved, and higher expression recognition accuracy is achieved.
Patent Information
- Application Number
- CN202211201514.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-09-29
AI Technical Summary
Existing expression recognition models have problems with poor recognition performance when dealing with occlusion and subjective expression images, especially due to insufficient performance caused by a single loss value optimization and uncertainty description method.
A training method based on feature fusion is adopted to extract features through comparative learning methods, and an initial expression recognition model is constructed, including pre-trained feature extraction network, flip feature fusion network, classification network and out-of-order fusion network, and feature fusion and uncertainty description is used using attention weights.
The model's recognition performance of expression images is improved. Through feature clustering and the use of attention weights, the model's attention to areas with high uncertainty is enhanced, thereby improving the accuracy of expression recognition.
Smart Images

Figure CN115690872B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a training method and an identification method for an expression recognition model based on feature fusion. Background Art
[0002] Expression image classification is an image processing method that separates different categories of expression images according to different features reflected by different categories of targets in image information. Expression image classification can be divided into a classification method based on traditional features and a classification method based on machine learning. The classification method based on traditional features classifies the expression of an image target based on image features such as color, texture, shape, and spatial relationship of key points.
[0003] With the wide application of machine learning in various fields, there have also emerged various classification methods based on deep learning for expression image classification, such as Convolutional Neural Networks (CNN), Residual Neural Network (ResNet), Generative Adversarial Nets (GAN), Deep Belief Network (DBN), etc. The expression image classification method based on machine learning does not require manual feature extraction and excessive professional knowledge, and can automatically and robustly extract image features through the training of a neural network model, thus obtaining good classification results and being widely used by researchers.
[0004] In the field of computer vision, expression image classification is slightly different from general picture classification. The dataset used for expression image classification is generally collected from the Internet. Some of these pictures are occluded, and for some human expressions, they are not obvious, with a large degree of subjectivity, which will lead to mislabeling during data annotation. These problems will deepen the training of the network and affect the performance of the network.
[0005] In related technologies, some training methods send the original picture and the flipped original picture into a convolutional neural network respectively to obtain their features and attention maps, and optimize the parameters of the network by comparing the attention maps of the two, so that the network can pay more attention to the overall part of the picture, thereby improving the network's recognition ability for expression pictures; other methods send the picture into a convolutional neural network to obtain the features of the picture and a value describing the uncertainty. For pictures with larger occlusion or mislabeled image tags, the uncertainty is greater. The network pays more attention to pictures with greater uncertainty, so that the network can improve its recognition ability for pictures that are difficult to recognize.
[0006] However, the method of optimizing the network parameters by comparing the attention maps of the two makes the loss value used to optimize the network relatively single, so that the performance of the trained network is poor; in the method of training the network by making the network pay more attention to pictures with greater uncertainty, only one value is used to describe the uncertainty of the entire picture, making the uncertainty low, so that the performance of the trained network is poor. Summary of the invention
[0007] In order to solve the above problems existing in the related art, the present invention provides a training method and a recognition method of an expression recognition model based on feature fusion. The technical problem to be solved by the present invention is achieved by the following technical solutions:
[0008] The present invention provides a training method for an expression recognition model based on feature fusion, comprising:
[0009] Preprocessing the images in the pre-acquired training set to obtain preprocessed images;
[0010] Using a plurality of pre-processed images in the training set, a feature extraction network is trained by a contrast learning method to obtain a pre-trained feature extraction network; the pre-trained feature extraction network is used to shorten the distance between feature maps of images of the same category in the feature space and to increase the distance between feature maps of images of different categories in the feature space;
[0011] Constructing an initial expression recognition model according to the pre-trained feature extraction network; the initial expression recognition model comprises: the pre-trained feature extraction network, an attention-based flip feature fusion network, a classification network and a random fusion network;
[0012] Each time, a plurality of pre-processed images are selected as original images, each original image is flipped to obtain a flipped image of each original image, and the original image and the flipped image are input into the initial expression recognition model;
[0013] Obtaining a first feature map of each image input each time through the pre-trained feature extraction network;
[0014] The second feature map and attention weight of the first feature map, the first global fusion feature including the first fusion features of each original image, and the flipping attention loss are obtained through the flipping feature fusion network; the first fusion feature of each original image is determined according to the second feature map and attention weight of the original image and the second feature map and attention weight of the flipped image of the original image; the flipping attention loss is determined according to the attention weights of each original image and the attention weights of the flipped image of each original image; the second feature map and the attention weight have the same dimension;
[0015] The order of different first fusion features in the first global fusion feature is adjusted to a preset order through the scrambled feature fusion network to obtain a first scrambled global fusion feature, and a second global fusion feature including each second fusion feature is obtained according to the first global fusion feature, the attention weight corresponding to the first global fusion feature, the first scrambled global fusion feature, and the attention weight corresponding to the first scrambled global fusion feature;
[0016] The classification loss corresponding to the first global fusion feature and the scrambled fusion loss corresponding to the second global fusion feature are respectively obtained through the classification network;
[0017] According to the obtained flipping attention loss, classification loss, and scrambled fusion loss each time, the flipping feature fusion network and / or the classification network are adjusted until the number of training times reaches a first preset number, and a pre-trained flipping feature fusion network and a pre-trained classification network are obtained;
[0018] The model composed of the pre-trained feature extraction network, the pre-trained flipping feature fusion network, and the pre-trained classification network is used as a pre-trained facial expression recognition model.
[0019] The present invention also provides a facial expression recognition method based on feature fusion, and the method includes:
[0020] Obtain at least one image to be recognized;
[0021] Preprocess the at least one image to be recognized to correspondingly obtain at least one preprocessed image;
[0022] Extract the first feature map of each preprocessed image in the at least one preprocessed image through the pre-trained feature extraction network in the above-mentioned pre-trained facial expression recognition model;
[0023] Generate a first global fusion feature according to the first feature map of each preprocessed image through the pre-trained flipping feature fusion network in the pre-trained facial expression recognition model;
[0024] Through the pre-trained classification network, according to the first global fusion feature, predict the facial expression category corresponding to each first fusion feature in the first global fusion feature, and obtain the facial expression category of the preprocessed image corresponding to each first fusion feature.
[0025] The present invention has the following beneficial technical effects:
[0026] By using the method of contrastive learning, the feature extraction network clusters the extracted features in the feature space, making the distance between features of the same class closer and the distance between features of different classes farther apart. The features obtained by the feature extraction network are further used to obtain features and attention maps, and the attention maps are used as the weights for feature fusion to obtain fused features. This can improve the discriminability of the fused features obtained by the trained model, thereby improving the performance of the trained model; and, attention weights with the same dimension as the features are used to describe the uncertainty of the features, that is, each value in the features has an attention value to describe its uncertainty. In this way, the parts with large uncertainty in the two pictures can be fused separately during the disordered fusion process, improving the model's attention to the parts with large uncertainty in the fused features, thereby improving the performance of the trained model.
[0027] The performance of the model trained by the above method is relatively high. Therefore, the recognition result of the expression category of the image to be recognized by the model trained by the above method is more accurate, thereby improving the accuracy of the expression recognition of the image.
[0028] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. Description of the Drawings
[0029] Figure 1 It is an optional flowchart of the training method of the expression recognition model based on feature fusion provided by the embodiment of the present invention;
[0030] Figure 2 It is the network structure diagram of the exemplary initial expression recognition model provided by the embodiment of the present invention;
[0031] Figure 3 It is an optional flowchart of the expression recognition method based on feature fusion provided by the embodiment of the present invention. Detailed Embodiments
[0032] The present invention will be further described in detail below in conjunction with specific embodiments, but the embodiments of the present invention are not limited thereto.
[0033] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0034] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0035] Although the present invention has been described in connection with various embodiments herein, however, in the process of implementing the claimed invention, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality of cases. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0036] Figure 1 is an optional flowchart of a method for training an expression recognition model based on feature fusion provided by an embodiment of the present invention, as Figure 1 shown, the method includes the following steps:
[0037] S101. Preprocess the images in the pre-acquired training set to obtain preprocessed images.
[0038] Exemplarily, the expression images in the training files in the RAF-DB dataset can be selected as the training set, and the expression images in the test files can be selected as the test set; wherein, the test set is used to test the performance of the pre-trained expression recognition model.
[0039] Exemplarily, each training image in the training set can be preprocessed, so as to enlarge the resolution of the image from 48*48 to 224*224, and horizontally flip each obtained preprocessed image, and randomly occlude a part of the image (occlude the image with a black rectangle of a random size). Similarly, for the test samples, only the resolution of their images is enlarged from 48*48 to 224*224.
[0040] S102. Using multiple pre-processed images in the training set, a feature extraction network is trained by a comparative learning method to obtain a pre-trained feature extraction network; the pre-trained feature extraction network is used to shorten the distance between feature maps of images of the same category in the feature space, and to increase the distance between feature maps of images of different categories in the feature space.
[0041] In some embodiments, a contrastive learning model including a feature extraction network and a fully connected layer feature classification head can be constructed first, and at least two preprocessed images are selected each time to input the contrastive learning model, and the initial feature map of each input image is extracted respectively through the feature extraction network; the feature vector of the image corresponding to the initial feature map in the projection space is obtained through the fully connected layer feature classification head; then, based on the feature vector of each image in the projection space and the real expression category corresponding to each image, a contrastive learning loss corresponding to the at least two preprocessed images is calculated, and then the network parameters of the feature extraction network are adjusted according to the contrastive learning loss obtained this time, thereby completing one training, and this is done until the number of training times reaches a second preset number of times, and a pre-trained feature extraction network is obtained.
[0042] Exemplarily, the feature extraction network can be the remaining network part after deleting the classification network in the residual network ResNet18, and keeping the parameter settings of ResNet18 unchanged; the fully connected layer feature classification head consists of a fully connected layer, a ReLU activation function layer and a fully connected layer connected in sequence.
[0043] Exemplarily, the number of iterations of contrastive learning s1, the maximum number of iterations S1 (for example, S1=25), and the contrastive learning model of the s1th iteration can be initialized as The weight parameter is And let s1 = 1; randomly select M training samples X from the preprocessed training sample set R with replacement as the comparative learning model The input and The output of is sent to a fully connected layer feature classification head (projection classification head), and the output set of the projection classification head is Among them, M ≥ 2, Indicates that the mth training sample is learned by contrast model The output of (the feature vector of the mth training sample in the projection space); using the InfoNCE loss function, according to the feature vector of the training sample to which each initial feature map belongs in the projection space, and the real expression category of the training sample to which the initial feature map belongs, the comparative learning model is calculated The contrastive learning loss L infoNCE , then, find L infoNCE Weight parameter The partial derivative of Then, the gradient descent method is adopted, and by performing backpropagation in the weight parameters are updated, and the updated weight parameters are obtained. According to the updated weight parameters the updated contrastive learning model is obtained After that, it is judged whether s1≥S1 holds. If so, the updated contrastive learning model is used as the obtained trained contrastive learning model Otherwise, let s1 = s1 + 1, and continue to train the obtained updated contrastive learning model using the above training principle. Through such cyclic iteration, until s1≥S1, the last obtained updated contrastive learning model is used as the trained contrastive learning model Among them, the trained feature extraction network in the trained contrastive learning model is the obtained pre-trained feature extraction network.
[0044] S103. Construct an initial facial expression recognition model according to the pre-trained feature extraction network; the initial facial expression recognition model includes: a pre-trained feature extraction network, an attention-based flipped feature fusion network, a classification network, and a disordered fusion network.
[0045] In the embodiment of the present invention, the output end of the pre-trained feature extraction network is connected to the input end of the attention-based flipped feature fusion network, the output end of the attention-based flipped feature fusion network is respectively connected to the input end of the classification network and the input end of the disordered fusion network, and the output end of the disordered fusion network is connected to the input end of the classification network.
[0046] S104. Each time, multiple preprocessed images are selected as the original images, and each original image is flipped to obtain the flipped image of each original image, and the original image and the flipped image are input into the initial facial expression recognition model.
[0047] In the embodiment of the present invention, each time the initial facial expression recognition model is trained, at least two preprocessed images can be selected as the original images, and the images selected this time are input into the initial facial expression recognition model at the same time.
[0048] S105. The first feature map of each image input each time is obtained through the pre-trained feature extraction network.
[0049] S106. Obtain the second feature map and attention weights of the first feature map, the first global fusion feature including the first fusion features of each original image, and the flipped attention loss through the flipped feature fusion network; the first fusion feature of each original image is determined according to the second feature map of this original image, the attention weights, and the second feature map and attention weights of the flipped image of this original image; the flipped attention loss is determined according to the attention weights of each original image and the attention weights of the flipped images of each original image; the dimensions of the second feature map and the corresponding attention weights are the same.
[0050] In some embodiments, the flipped feature fusion network can be used to obtain, according to the first feature maps of each original image in multiple original images input each time, the second feature map corresponding to the first feature map and the attention weights of the second feature map; determine the first fusion feature corresponding to each original image according to the second feature map and attention weights of each original image, and the second feature map and attention weights of the flipped image of this original image, so as to obtain the first global fusion feature including the first fusion features of each original image in multiple original images input each time; perform a flipping process on the attention weights of the flipped images of each original image in multiple original images input each time to obtain the flipped attention weights of each flipped image; determine the flipped loss value of each original image according to the attention weights of each image in multiple original images input each time and the flipped attention weights of the flipped image of this original image; obtain the flipped attention loss corresponding to multiple original images input each time according to the flipped loss value of each original image.
[0051] Exemplarily, the training iteration number s2 of the initial facial expression recognition model, the maximum iteration number S2 (for example, S2 = 60), and the weight parameters of the initial facial expression recognition model trained in the s2-th iteration can be initialized first as for and let S2 = 1; take M images (M training samples) randomly selected with replacement from the preprocessed images in the training sample set R as x, use x as the input of the model and flip x to get x flip , and at the same time use x flip also as the input of to obtain the output features F = {f1, f2, …, f m , …, f M}, the corresponding attention weights α att = {α1, α2, …, α m , …, α M}, the features of the corresponding M flipped images (flipped features) the corresponding attention weights β of the M flipped images att = {β1, β2, …, β m , …, βM}, M≥2, where, f m represents the output feature of the m-th original image in the model α m represents the attention weight of the m-th original image, represents the output feature of the flipped image (the m-th flipped image) of the m-th original image in the model β m represents the attention weight of the m-th flipped image. The feature F is put into one-to-one correspondence with the attention weight α att and the flipped feature F flip is put into one-to-one correspondence with the attention weight β att and they are fused to obtain the first global fusion feature where, represents the first fusion feature corresponding to the m-th original image, and,
[0052] Exemplarily, the above-mentioned flipped attention loss corresponding to multiple original images input each time can be obtained according to the flipped loss value of each original image, and can be realized by the following formula:
[0053]
[0054] where, L att represents the flipped attention loss, N represents the number of multiple original images input each time, i represents the i-th original image among the N original images, α i represents the attention weight of the i-th original image, β i represents the attention weight of the flipped image of the i-th original image, represents the flipped attention weight of the flipped image of the i-th original image.
[0055] Here, the flipping direction used to obtain the flipped attention weight of each flipped image is the same as the flipping direction used to obtain this flipped image. For example, it can be horizontal flipping or vertical flipping, etc.
[0056] S107. Adjust the order of different first fusion features in the first global fusion feature to a preset order through a disordered feature fusion network to obtain a first disordered global fusion feature, and obtain a second global fusion feature including each second fusion feature according to the first global fusion feature, the attention weight corresponding to the first global fusion feature, the first disordered global fusion feature, and the attention weight corresponding to the first disordered global fusion feature.
[0057] In the embodiments of the present invention, the preset order can be set according to actual needs, and the present invention does not limit this.
[0058] Exemplarily, when the first global fusion feature is When, the first scrambled global fusion feature can be
[0059] In some embodiments, for each first fusion feature in the first global fusion feature, a first fusion sub-value can be calculated according to the first fusion feature and the attention weight of the original image to which the first fusion feature belongs; according to the first fusion feature, a first fusion feature with the same position sorting as the first fusion feature is determined from the first scrambled global fusion feature as the target fusion feature; according to the target fusion feature and the attention weight of the original image to which the target fusion feature belongs, a second fusion sub-value is calculated; according to the first fusion sub-value and the second fusion sub-value, a second fusion feature jointly corresponding to the original image to which the first fusion feature belongs and the original image to which the target fusion feature belongs is obtained, so as to obtain a second global fusion feature including each second fusion feature.
[0060] Exemplarily, for f1 fuse1 , the corresponding second fusion feature f1 fuse1 can be calculated using the following formula fuse2 :
[0061]
[0062] Where, α1 is the attention weight corresponding to the original image to which f1 fuse1 belongs, is the above-mentioned first scrambled global fusion feature in which, the first fusion feature with the same position sorting as f1 fuse1 , α7 is the attention weight corresponding to the original image to which belongs, α1·f1 fuse1 is the first fusion sub-value, is the second fusion sub-value. Thus, for and , the obtained second global fusion feature is
[0063] S108. Respectively obtain the classification loss corresponding to the first global fusion feature and the scrambled fusion loss corresponding to the second global fusion feature through the classification network.
[0064] In some embodiments, the classification loss corresponding to the first global fusion feature can be obtained through the classification network, which can be achieved by the following method: for each first fusion feature in the first global fusion feature, through the classification network, obtain the first output value of the first fusion feature; the first output value is the probability value of the first fusion feature belonging to each preset expression category; according to the probability values of the first fusion feature belonging to each preset expression category and the probability value corresponding to the true expression category of the original image corresponding to the first fusion feature in the first output value, calculate the classification sub-loss corresponding to the first fusion feature; according to the classification sub-losses of each first fusion feature in the first global fusion feature and the number of multiple original images input each time, determine the classification loss.
[0065] Here, according to the classification sub-losses of each first fusion feature in the first global fusion feature and the number of multiple original images input each time, the classification loss can be determined, which can be achieved by the following formula:
[0066]
[0067] where, L cls represents the classification loss, N represents the number of multiple original images input each time, i represents the i-th original image among the N original images, represents the classification sub-loss of the first fusion feature corresponding to the i-th original image, represents the probability value corresponding to the true expression category of the i-th original image among the probability values of the first fusion feature corresponding to the i-th original image belonging to each preset expression category, C represents the number of preset expression categories, j represents the j-th category among the C preset expression categories, represents the probability value predicted by the classification network that the first fusion feature corresponding to the i-th original image belongs to the j-th category, and exp(.) represents the exponential function with the natural constant e as the base.
[0068] In some embodiments, the scrambled fusion loss corresponding to the second global fusion feature can be obtained through the classification network, which can be achieved by the following method: for each second fusion feature in the second global fusion feature, through the classification network, obtain the second output value of the second fusion feature; the second output value is the probability value of the second fusion feature belonging to each preset expression category; according to the probability values of the second fusion feature belonging to each preset expression category and the probability values corresponding to the true expression categories of the two original images corresponding to the second fusion feature in the second output value, calculate the scrambled sub-loss corresponding to the second fusion feature; according to the scrambled sub-losses of each second fusion feature in the second global fusion feature and the number of multiple original images input each time, obtain the scrambled fusion loss.
[0069] Here, according to the scrambled sub-loss of each second fusion feature in the second global fusion feature and the number of multiple original images input each time, a scrambled fusion loss can be obtained, which can be achieved through the following formula:
[0070]
[0071] Among them, L s represents the scrambled fusion loss, L z represents the scrambled sub-loss of the z-th second fusion feature, N represents the number of multiple original images input each time, i and j respectively represent the i-th original image and the j-th original image among the N original images, the i-th original image and the j-th original image are the two original images corresponding to the z-th second fusion feature, represents the probability value corresponding to the true expression category of the i-th original image among the probability values of the second fusion feature corresponding to the i-th original image belonging to each preset expression category, represents the probability value corresponding to the true expression category of the j-th original image among the probability values of the second fusion feature corresponding to the j-th original image belonging to each preset expression category, C represents the number of preset expression categories, c represents the c-th category among the C preset expression categories, represents the probability value predicted by the classification network that the z-th second fusion feature belongs to the c-th category, and exp(.) represents the exponential function with the natural constant e as the base.
[0072] S109. According to the obtained flipping attention loss, classification loss, and scrambled fusion loss each time, adjust the flipping feature fusion network and / or the classification network until the number of training times reaches the first preset number of times, and obtain the pre-trained flipping feature fusion network and the pre-trained classification network.
[0073] Here, the first preset number of times can be set according to actual needs, and the embodiments of the present invention do not limit this.
[0074] Exemplarily, when obtaining the flipping attention loss, classification loss, and scrambled fusion loss of this time each time, these three losses can be added according to a certain weight coefficient to obtain the total loss L total = λ1L cls + λ2L att + λ3L s , and then obtain the partial derivative of L total with respect to the weight parameter Then, through the gradient descent method, by backpropagating in to update the weight parameter , and obtain the updated weight parameter Thus, the updated After that, it is judged whether s2≥S2 holds. If so, the updated one obtained is used as the trained model Otherwise, let s2 = s2 + 1, and continue to perform iterative training on the updated one obtained by using the above training principle until s2≥S2 holds, and then use the last updated one obtained as the trained model used as the trained model
[0075] Exemplarily, λ1 can be 1, λ2 can be 0.5, and λ3 can be 5.
[0076] Here, the weight parameter can be the weight parameter of the flipping feature fusion network and / or the weight parameter of the classification network, or can be only the weight parameter of the classification network or the weight parameter of the flipping feature fusion network.
[0077] S110. Use the model composed of the pre-trained feature extraction network, the pre-trained flipping feature fusion network, and the pre-trained classification network as the pre-trained facial expression recognition model.
[0078] In the embodiment of the present invention, when the trained model is obtained, the pre-trained feature extraction network, the pre-trained flipping feature fusion network, and the pre-trained classification network in can be used as the pre-trained facial expression recognition model, wherein the output end of the pre-trained feature extraction network is connected to the input end of the pre-trained flipping feature fusion network, and the output end of the pre-trained flipping feature fusion network is connected to the input end of the pre-trained classification network.
[0079] Exemplarily, Figure 2 is the network structure diagram of the initial facial expression recognition model, as shown in Figure 2As shown, during each model training, four original images and four flipped images obtained by horizontally flipping the four original images are simultaneously input into a pre-trained feature extraction network. The pre-trained feature extraction network outputs the first feature map of each of the four original images and the first feature map of each of the four flipped images. The obtained first feature maps are all input into a flipped feature fusion network. Through the flipped feature fusion network, the second feature map of the first feature map of each original image, the attention weight corresponding to the second feature map, the second feature map of the first feature map of each flipped image, and the attention weight corresponding to the second feature map are obtained. According to the second feature map and attention weight of each original image, and the second feature map and attention weight of the flipped image of this original image, the first fusion feature corresponding to this original image is determined, thereby obtaining a first global fusion feature containing four first fusion features. And according to the second feature map and attention weight of the flipped image, the flipped attention weight is obtained. According to the attention weight corresponding to the second feature map of the original image and the flipped attention weight, the flipped attention loss for this time is calculated. Then, the first global fusion feature is respectively input into a classification network and a scrambled fusion network. Through the classification network, an output value is obtained, and according to the output value, the classification loss for this time is obtained. And through the scrambled fusion network, a first scrambled global fusion feature of the first global fusion feature is obtained. According to the first global fusion feature, the attention weight corresponding to the first global fusion feature, the first scrambled global fusion feature, and the attention weight corresponding to the first scrambled global fusion feature, a second global fusion feature containing four second fusion features is obtained. Then, the second global fusion feature is input into the classification network. Through the classification network, an output value is obtained, and according to the output value, the scrambled fusion loss for this time is obtained. Finally, according to the flipped attention loss, classification loss, and scrambled fusion loss obtained this time, the total loss for this time is obtained. According to the total loss, the weight parameters of the flipped feature fusion network and the classification network are adjusted until, when the number of training times reaches a first preset number of times, a pre-trained flipped feature fusion network and a pre-trained classification network are obtained.
[0080] In the embodiments of the present invention, by using the method of contrastive learning, the feature extraction network clusters the extracted features in the feature space, making the distance between features of the same class closer and the distance between features of different classes farther apart. The features obtained by the feature extraction network are further used to obtain features and attention maps, and the attention maps are used as the weights for feature fusion to obtain fused features. This can improve the discriminability of the fused features obtained by the trained model, thereby improving the performance of the trained model. Moreover, attention weights with the same dimension as the features are used to describe the uncertainty of the features, that is, each value in the features has an attention value to describe its uncertainty. In this way, the parts with large uncertainty in the two images can be fused separately during the disordered fusion process, improving the model's attention to the parts with large uncertainty in the fused features, thereby improving the performance of the trained model.
[0081] The embodiments of the present invention also provide an expression recognition method based on feature fusion, as Figure 3 shown. This method includes:
[0082] S201. Obtain at least one image to be recognized.
[0083] Here, one or more images to be recognized can be obtained.
[0084] S202. Preprocess at least one image to be recognized to correspondingly obtain at least one preprocessed image.
[0085] Here, the resolution of each image to be recognized can be enlarged from 48*48 to 224*224 to obtain the preprocessed image of each image to be recognized.
[0086] S203. Through the pre-trained feature extraction network in the above-mentioned pre-trained expression recognition model, extract the first feature maps of each preprocessed image in at least one preprocessed image.
[0087] S204. Through the pre-trained flipping feature fusion network in the pre-trained expression recognition model, generate the first global fusion feature according to the first feature maps of each preprocessed image.
[0088] Here, the principle of S204 is the same as the principle of generating the first global fusion feature in S106 above.
[0089] S205. Through the pre-trained classification network, predict the expression categories corresponding to each first fusion feature in the first global fusion feature to obtain the expression categories of the preprocessed images corresponding to each first fusion feature.
[0090] In the embodiments of the present invention, since the performance of the model trained by the above method is relatively high, the recognition result of the expression category of the image by the trained model is more accurate, thereby improving the accuracy of the expression recognition of the image.
[0091] The training effect of the above training method in the embodiments of the present invention is further described below through experimental data.
[0092] The hardware platform used in the simulation experiment is a CPU Intel(R)Xeon(R)Gold 6240CPU with a main frequency of 2.60GHz and 16G RAM. The software platform is Python3.8. The operating system is Ubuntu 18.04.3LTS x64. The expression image dataset used in the simulation experiment is the RAF-DB dataset. The expression images in this dataset are collected from the Internet and have the same image size. The RAF-DB dataset contains 15,539 expression images of 7 categories. In the simulation experiment, 12,271 images are selected to form the training set R, and the remaining 3,068 images form the test set E.
[0093] The present invention is compared and simulated with the classification accuracy of the existing expression image classification method based on feature fusion network (hereinafter referred to as RUL). The results are shown in Table 1. Referring to Table 1, the classification accuracy of the present invention on the test sample set E is 89.37%, and the classification accuracy of RUL on the test sample set E is 88.90%. Compared with RUL in the prior art, the classification accuracy of the present invention is increased by 0.47%.
[0094]
[0095] Table 1
[0096] Based on the result analysis in the above simulation experiment, the method proposed by the present invention can effectively solve the problem that the traditional expression recognition network cannot correctly classify complex data sets, and further solve the problem that the deep convolutional neural network has a low classification accuracy for expression images.
[0097] The embodiments of the present invention also provide a computer-readable storage medium. A computer program is stored in the computer-readable storage medium, and the computer program is used to execute some or all of the steps in the above control method for secretly approaching a target aircraft. The computer-readable storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0098] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as falling within the protection scope of the present invention.
Claims
1. A training method for an expression recognition model based on feature fusion, characterized in that, include: Preprocessing the images in the pre-acquired training set to obtain preprocessed images; Using a plurality of pre-processed images in the training set, a feature extraction network is trained by a contrast learning method to obtain a pre-trained feature extraction network; the pre-trained feature extraction network is used to shorten the distance between feature maps of images of the same category in the feature space and to increase the distance between feature maps of images of different categories in the feature space; Constructing an initial expression recognition model according to the pre-trained feature extraction network; the initial expression recognition model comprises: the pre-trained feature extraction network, an attention-based flip feature fusion network, a classification network and a random fusion network; Each time, a plurality of pre-processed images are selected as original images, each original image is flipped to obtain a flipped image of each original image, and the original image and the flipped image are input into the initial expression recognition model; Obtaining a first feature map of each image input each time through the pre-trained feature extraction network; The second feature map and attention weight of the first feature map, the first global fusion feature including the first fusion features of each original image, and the flipping attention loss are obtained through the flipping feature fusion network; the first fusion feature of each original image is determined according to the second feature map and attention weight of the original image and the second feature map and attention weight of the flipped image of the original image; the flipping attention loss is determined according to the attention weights of each original image and the attention weights of the flipped image of each original image; the second feature map and the attention weight have the same dimension; The order of different first fusion features in the first global fusion feature is adjusted in a preset order through the random feature fusion network to obtain a first random global fusion feature, and a second global fusion feature including each second fusion feature is obtained according to the first global fusion feature, the attention weight corresponding to the first global fusion feature, the first random global fusion feature, and the attention weight corresponding to the first random global fusion feature; Obtaining, through the classification network, a classification loss corresponding to the first global fusion feature and a disordered fusion loss corresponding to the second global fusion feature; According to the flipped attention loss, the classification loss and the out-of-order fusion loss obtained each time, adjusting the flipped feature fusion network and / or the classification network until the number of training times reaches a first preset number of times, thereby obtaining a pre-trained flipped feature fusion network and a pre-trained classification network; A model consisting of the pre-trained feature extraction network, the pre-trained flipped feature fusion network and the pre-trained classification network is used as a pre-trained expression recognition model.
2. The training method of the expression recognition model based on feature fusion according to claim 1, wherein The method of using a plurality of pre-processed images in the training set to train a feature extraction network by a contrast learning method to obtain a pre-trained feature extraction network includes: Constructing a contrastive learning model including the feature extraction network and the fully connected layer feature classification head; Selecting at least two pre-processed images each time and inputting them into the contrastive learning model; Extracting an initial feature map of each input image through the feature extraction network; Obtain the feature vector of the image corresponding to the initial feature map in the projection space through the fully connected layer feature classification head; Based on the feature vectors of each image in the projection space and the true expression category corresponding to each image, calculate a contrastive learning loss corresponding to the at least two preprocessed images; According to the contrastive learning loss obtained each time, adjust the network parameters of the feature extraction network until the number of training times reaches the second preset number of times, and obtain the pre-trained feature extraction network.
3. The training method of the expression recognition model based on feature fusion according to claim 1, wherein, The obtaining of the second feature map and the attention weight of the first feature map, the first global fusion feature including the first fusion features of each original image, and the flipped attention loss through the flipped feature fusion network includes: Through the flipped feature fusion network, according to the first feature maps of each original image in the multiple original images input each time, obtain the second feature map corresponding to the first feature map and the attention weight of the second feature map; According to the second feature map and the attention weight of each original image, and the second feature map and the attention weight of the flipped image of this original image, determine the first fusion feature corresponding to this original image, so as to obtain the first global fusion feature including the first fusion features of each original image in the multiple original images input each time; Perform the flipping process on the attention weights of the flipped images of each original image in the multiple original images input each time to obtain the flipped attention weights of each flipped image; According to the attention weight of each image in the multiple original images input each time and the flipped attention weight of the flipped image of this original image, determine the flipped loss value of this original image; According to the flipped loss value of each original image, obtain the flipped attention loss corresponding to the multiple original images input each time.
4. The training method of the expression recognition model based on feature fusion according to claim 1, characterized in that The obtaining of the second global fusion feature including each second fusion feature according to the first global fusion feature, the attention weight corresponding to the first global fusion feature, the first shuffled global fusion feature, and the attention weight corresponding to the first shuffled global fusion feature includes: For each first fusion feature in the first global fusion feature, calculate a first fusion sub-value according to the first fusion feature and the attention weight of the original image to which the first fusion feature belongs; According to the first fusion feature, determine the first fusion feature with the same position sorting as the first fusion feature in the first shuffled global fusion feature as the target fusion feature; According to the target fusion feature and the attention weight of the original image to which the target fusion feature belongs, calculate a second fusion sub-value; According to the first fusion sub-value and the second fusion sub-value, obtain the second fusion feature jointly corresponding to the original image to which the first fusion feature belongs and the original image to which the target fusion feature belongs, so as to obtain the second global fusion feature including each second fusion feature.
5. The training method of the expression recognition model based on feature fusion according to claim 1, wherein Obtain the classification loss corresponding to the first global fusion feature through the classification network, including: For each first fusion feature in the first global fusion feature, through the classification network, obtain a first output value of this first fusion feature; the first output value is the probability value of this first fusion feature belonging to each preset expression category; According to the probability value of this first fusion feature belonging to each preset expression category, and the probability value corresponding to the true expression category of the original image corresponding to this first fusion feature in the first output value, calculate the classification sub-loss corresponding to this first fusion feature; According to the classification sub-loss of each first fusion feature in the first global fusion feature, and the number of multiple original images input each time, determine the classification loss.
6. The training method of the expression recognition model based on feature fusion according to claim 5, wherein, The classification loss is calculated by the following formula: Among them, L cls represents the classification loss, N represents the number of multiple original images input each time, i represents the i-th original image among the N original images, represents the classification sub-loss of the first fusion feature corresponding to the i-th original image, represents the probability value corresponding to the true expression category of the i-th original image among the probability values of the first fusion feature corresponding to the i-th original image belonging to each preset expression category, C represents the number of preset expression categories, and j represents the j-th category among the C preset expression categories, represents the probability value that the first fusion feature corresponding to the i-th original image predicted by the classification network belongs to the j-th category, and exp(.) represents the exponential function with the natural constant e as the base.
7. The training method of the expression recognition model based on feature fusion according to claim 1, characterized in that Obtaining the disordered fusion loss corresponding to the second global fusion feature through the classification network includes: For each second fusion feature in the second global fusion feature, through the classification network, obtain a second output value of this second fusion feature; the second output value is the probability value of this second fusion feature belonging to each preset expression category; According to the probability value of this second fusion feature belonging to each preset expression category, and the probability values corresponding to the true expression categories of the two original images corresponding to this second fusion feature in the second output value, calculate the disordered sub-loss corresponding to this second fusion feature; According to the disordered sub-loss of each second fusion feature in the second global fusion feature, and the number of multiple original images input each time, obtain the disordered fusion loss.
8. The training method of the expression recognition model based on feature fusion according to claim 7, characterized in that The disordered fusion loss is calculated by the following formula: Among them, L s represents the scrambled fusion loss, and L z represents the scrambled sub-loss of the z-th second fusion feature, N represents the number of multiple original images in each input, i and j respectively represent the i-th original image and the j-th original image among the N original images, the i-th original image and the j-th original image are two original images corresponding to the z-th second fusion feature, represents the probability value corresponding to the true expression category of the i-th original image among the probability values of the second fusion feature corresponding to the i-th original image belonging to each preset expression category, represents the probability value corresponding to the true expression category of the j-th original image among the probability values of the second fusion feature corresponding to the j-th original image belonging to each preset expression category, C represents the number of preset expression categories, c represents the c-th category among the C preset expression categories, and μ yc represents the probability value predicted by the classification network that the z-th second fusion feature belongs to the c-th category, and exp(.) represents the exponential function with the natural constant e as the base.
9. The training method of the expression recognition model based on feature fusion according to claim 3, characterized in that The flipped attention loss is implemented by the following formula: Among them, L att represents the flipped attention loss, N represents the number of multiple original images input each time, i represents the i-th original image among the N original images, and α i represents the attention weight of the i-th original image, and β i represents the attention weight of the flipped image of the i-th original image, represents the flipped attention weight of the flipped image of the i-th original image.
10. An expression recognition method based on feature fusion, characterized in that The method includes: Obtain at least one image to be recognized; Preprocess the at least one image to be recognized to correspondingly obtain at least one preprocessed image; Through the pre-trained feature extraction network in the pre-trained expression recognition model according to any one of claims 1 to 9 above, extract the first feature maps of each preprocessed image in the at least one preprocessed image; Through the pre-trained flipped feature fusion network in the pre-trained expression recognition model, generate a first global fusion feature according to the first feature maps of each preprocessed image; Through the pre-trained classification network, according to the first global fusion feature, predict the expression category corresponding to each first fusion feature in the first global fusion feature, and obtain the expression category of the preprocessed image corresponding to each first fusion feature.
Citation Information
Patent Citations
Video pedestrian re-identification method based on generative adversarial network and attention mechanism
CN113221641A
Multi-head posture facial expression recognition method and application thereof
CN113221799A