A training method for a micro-expression discrimination model based on a generative adversarial network
Through the micro-expression discrimination model training method based on the generative adversarial network, the problem of difficult micro-expression recognition in the prior art is solved, and higher discrimination accuracy and performance are achieved.
Patent Information
- Application Number
- CN202210562150.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-05-23
AI Technical Summary
The existing micro-expression recognition technology has the problems of short detection window, lack of large-scale data sets, and low efficiency of manual feature extraction methods, which makes it difficult to automatically recognize micro-expression.
The micro-expression discriminant model training method based on the generative adversarial network is used to pre-train the generator by generating the L1 norm of the image and the target image as the first loss function of pixel reconstruction; the discriminator is pre-trained using the cross-entropy loss of the classification result and the sum of the cross-entropy loss of the enhanced label as the second loss function; then the pre-trained generator and discriminator are connected into the adversarial network model and jointly trained to obtain a micro-expression discriminant model with better performance.
The micro-expression discrimination model trained through this method has better performance and higher discrimination accuracy, and can more effectively identify micro-expression, overcoming the problem of difficulty in recognition in the prior art.
Smart Images

Figure CN115019124B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of micro-expression recognition, and in particular relates to a micro-expression recognition model training method based on a generative adversarial network. Background Art
[0002] Facial micro-expressions are facial movements that occur in a spontaneous and involuntary manner when a person is experiencing a particular emotion. These micro-expressions can often be found in high-stakes situations, such as criminal interrogations, interviews, political debates, and poker games. Therefore, multi-task computational analysis and automated recognition of micro-expressions has become an emerging field in face research. In the past five years, researchers have shown increasing interest in micro-expression research. At present, some commonly used datasets have been published, including the SMIC dataset, the FACS-encoded Chinese Academy of Sciences micro-expression database CASMEII, and the spontaneous micro-facial movement dataset SAMM, which have promoted the further development of this research field. Since the launch of the Micro-Expression Competition MEGC, the field of micro-expression research has begun to focus on two main micro-expression tasks: localization and recognition. For the micro-expression localization task, two long video benchmark datasets were introduced in the MEGC competition, and the performance indicators were standardized. As for the micro-expression recognition task, a comprehensive composite database was introduced in MEGC, which includes samples collected from different environments and a wide range of subjects as a standard for the micro-expression recognition task.
[0003] Micro-expressions have the characteristics of low intensity, weak amplitude and short duration, which makes most spontaneous micro-expressions cannot be observed by the naked eye, and is a challenging pattern recognition task. Due to the particularity of micro-expressions, it has attracted more and more researchers to enter the research of micro-expression recognition. Due to the short detection window, there is currently no large-scale dataset. Existing micro-expression recognition technology mainly focuses on manual feature extraction methods, such as optical flow (OF) and local binary pattern (LBP). Optical flow is a method of observing and analyzing the motion changes of objects in image sequences. It obtains the motion information of objects by calculating the pixel changes in adjacent frames. The local binary pattern method obtains texture features by calculating the relationship between pixels, and then classifies them according to the extracted features. There are also related technologies that use LBP-TOP to extract the maximum value of spatiotemporal features.
[0004] Since convolutional neural networks have been widely used in the field of computer vision, many researchers in the field of micro-expressions have proposed many new recognition technologies. Some researchers have published the dual time-scale convolutional neural network (DTSCNN) method, which is far superior to other manual feature extraction methods. Some researchers have also published the STSTNet method, which combines the spatial and temporal information embedded in the micro-expression video clips to prevent overfitting problems. Some researchers have published technologies that enhance the vertex frame information to optimize micro-expression classification.
[0005] In recent years, image enhancement technology has also turned to deep learning methods, including information enhancement technology and technology that increases the number of images through generative methods. Generative Adversarial Networks (GANs) and their variants have been widely studied as generative methods. GANs-based related technologies are applied to various visual application fields, such as facial pose, style transfer, biomedical image synthesis, and expression generation. Summary of the invention
[0006] In view of the above problems existing in the prior art, the present invention provides a micro-expression discrimination model training method based on a generative adversarial network. The micro-expression discrimination model finally trained has better performance and higher discrimination accuracy.
[0007] The present invention adopts the following technical solutions:
[0008] A micro-expression discrimination model training method based on a generative adversarial network comprises the following steps:
[0009] S1, using the L1 norm of the generated image and the target image as the first loss function for pixel reconstruction, and performing the first pre-training on the generator;
[0010] S2, using the sum of the cross entropy loss of the classification result and the cross entropy loss of the enhanced label as the second loss function, and performing a second pre-training on the discriminator;
[0011] S3, connecting the first pre-trained generator and the second pre-trained discriminator to generate an adversarial network model;
[0012] S4. Based on the first loss function, the second loss function, and the introduction of the adversarial loss function, the adversarial network model is jointly trained to obtain a trained micro-expression discrimination model. In the adversarial loss function, the generator aims to minimize the adversarial loss function, and the discriminator aims to maximize the adversarial loss function.
[0013] As a preferred solution, in the first pre-training process of the generator and the joint training process of the adversarial network model, the original facial micro-expression image and the facial micro-expression image near the vertex frame generated by the target are derived from the video frame sequence of the same subject.
[0014] As a preferred solution, the generator includes an encoder, a decoder, and a controller;
[0015] The encoder is used to convert the input video frame sequence into a coding vector;
[0016] A controller, used for adding the corresponding dimension noise and the corresponding dimension enhancement label as control items to the encoding vector as control items to obtain a controlled encoding vector;
[0017] A decoder is used to convert the control encoded vector into an output image sequence.
[0018] As a preferred solution, the encoder includes 5 sequentially connected encoding layers, each encoding layer includes a connected conv_block unit layer and a maximum pooling layer, and the conv_block unit layer includes 2 convolutional layers, 2 batch normalization layers and 2 times ReLU functions.
[0019] As a preferred solution, the decoder comprises five sequentially connected decoding layers corresponding to the five encoding layers in the encoder.
[0020] As a preferred solution, the generator also includes 5 skip connection layers, which connect each encoding layer to the corresponding decoding layer.
[0021] As a preferred solution, in step S1, the first loss function when the generator is first pre-trained is:
[0022] L pixel (G) = || G(x weak )-x apex || 1 ,
[0023] Among them, x weak is the original weak facial micro-expression image, x apex is the strong facial micro-expression image near the original vertex frame, G(x weak ) is composed of the original weak facial micro-expression image x weak The generated corresponding strong facial micro-expression image near the pseudo vertex frame.
[0024] As a preferred solution, in step S2, the second loss function when performing the second pre-training on the discriminator is:
[0025] L cls+el (D) = L 1 (y precls,y cls )+λL 2 (y preel ,y el ),
[0026] Among them, L 1 (y precls ,y cls ) represents the actual micro-expression type y cls and the predicted micro-expression type y precls The cross entropy loss of L 2 (y preel ,y el ) is the enhanced label y el and predict the augmented label y preel The cross entropy loss of , λ represents the first weighted value.
[0027] As a preferred solution, in step S4, the optimization goal of the joint training of the adversarial network model is:
[0028] L MEEGAN =L adv +τL pixel (G)+μL cls+el (D)
[0029] Among them, L adv represents the adversarial loss function, τ represents the second weighted value, and μ represents the third weighted value.
[0030] As a preferred solution, in step S4, the adversarial loss function introduced is specifically:
[0031]
[0032] Among them, G represents the generator, D represents the discriminator, logD(x apex ) is [1 0] T and [D(x apex ) 1-D(x apex )] T The cross entropy between Represents a strong facial micro-expression image x apex According to the distribution probability p(x apex ) is sampled when logD(x apex )'s expected value; log(1-D(G(x weak )))Yes[1 0] T and [D(G(x weak )) 1-D(G(x weak ))] T The cross entropy between Indicates that when the input weak facial micro-expression image x weak According to the distribution probability p(xweak ) is sampled when log(1-D(G(x weak )))'s expected value.
[0033] The beneficial effects of the present invention are:
[0034] A micro-expression discrimination model is generated by pre-training and joint training. The micro-expression discrimination model has better performance and higher discrimination accuracy.
[0035] The encoder uses conv_block unit layer and maxpooling layer to downsample weak micro-expression images. After conv_block and maxpooling layer structures are repeated 5 times, the input image is encoded into a 1024-dimensional encoding vector. Among them, conv_block contains 2 convolutional layers, 2 batch normalization layers, and 2 Relu functions. Designing such a complex network structure improves the acquisition of more details in the generation process.
[0036] The generator also includes skip-connections, which connect each encoding layer to the corresponding decoding layer and pass the compressed feature information in the model in a jump-like manner. Through the skip connection layer, spatial information can be passed directly from the encoder to the decoder. Therefore, the generator can obtain more detailed information during the reconstruction of the image. In addition, the skip connection layer can provide better training results. Because through the skip connection layer, the gradient can flow in the deep data of the entire model structure. The addition of the skip connection layer is also because the encoder and decoder of the generator have the same underlying structure.
[0037] In the first pre-training process of the generator and the joint training process of the adversarial network model, the original facial micro-expression image and the facial micro-expression image near the vertex frame generated by the target are derived from the video frame sequence of the same subject. Such a training mode can eliminate the excessive influence of facial information. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0039] Figure 1 It is a flow chart of a micro-expression discrimination model training method based on a generative adversarial network according to the present invention;
[0040] Figure 2 It is a schematic diagram of the customized enhanced label marking method of the present invention;
[0041] Figure 3 It is a structural diagram of the generator part;
[0042] Figure 4 It is a schematic diagram of the structure of the discriminator part;
[0043] Figure 5 It is the pre-training scheme diagram of the generator;
[0044] Figure 6 It is the pre-training scheme diagram of the discriminator;
[0045] Figure 7 It is the overall structure diagram of the micro-expression discrimination model;
[0046] Figure 8 It is a diagram of the joint training scheme for the adversarial network model. DETAILED DESCRIPTION
[0047] The following describes the embodiments of the present invention through specific embodiments, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0048] Reference Figure 1 This embodiment provides a micro-expression discrimination model training method based on a generative adversarial network, comprising the steps of:
[0049] S1, using the L1 norm of the generated image and the target image as the first loss function for pixel reconstruction, and performing the first pre-training on the generator;
[0050] S2, using the sum of the cross entropy loss of the classification result and the cross entropy loss of the enhanced label as the second loss function, and performing a second pre-training on the discriminator;
[0051] S3, connecting the first pre-trained generator and the second pre-trained discriminator to generate an adversarial network model;
[0052] S4. Based on the first loss function, the second loss function, and the introduction of the adversarial loss function, the adversarial network model is jointly trained to obtain a trained micro-expression discrimination model. In the adversarial loss function, the generator aims to minimize the adversarial loss function, and the discriminator aims to maximize the adversarial loss function.
[0053] Specific:
[0054] Reference Figure 2 , the figure shows the customized enhanced labeling method of the present invention. The present invention defines 10% of the video frame images near the vertex frame as strong expressions, and the other 90% of the video frame images as weak expressions. The micro-expression image of the vertex frame has the most obvious muscle movement and is close to the macroscopic facial expression image. From the starting frame to the ending frame, the remaining 90% of the video frame images are marked as weak expression images, which are not obvious but can still be classified.
[0055] Reference Figure 3 , the figure is a schematic diagram of the structure of the generator part, refer to Figure 5 , the figure shows the pre-training scheme diagram of the generator.
[0056] The generator uses an "encoder-decoder" structure network, namely the U-Net model, to perform image enhancement generation of identity teams so that identity information can be stably preserved during the generation process. Identity teams are defined as the original facial micro-expression image and the target generated vertex frame near the strong facial micro-expression image in the generation training process are derived from the same subject's video frame sequence. In the generator, the encoder converts the input video frame sequence into an encoded vector; the decoder converts the encoded vector generated by the encoder into an output image sequence.
[0057] One of the main tasks of the present invention is micro-expression enhancement, that is, generating a strong facial micro-expression image near the corresponding pseudo vertex frame from the original weak facial micro-expression image. In the case of the model framework of the present invention, the U-Net model will be used as a generator to perform the generation task of micro-expression enhancement.
[0058] In the encoder, the weaker micro-expression images are downsampled using the conv_block unit layer and the max pooling layer. After the conv_block and max pooling layer structures are repeated 5 times, the input image is encoded into a 1024-dimensional encoding vector. Among them, conv_block contains 2 convolutional layers (Conv layer), 2 batch normalization layers (Batchnormal layer) and 2 Relu functions. Such a complex network structure is designed to obtain more details in the generation process. Before decoding, 100-dimensional noise z and 1-dimensional enhanced label c are added to the 1024-dimensional encoding vector as control items. Then this 1024-dimensional encoding vector becomes a new 1125-dimensional encoding vector.
[0059] The decoder uses an upsampling process that is exactly the same as the encoder. The network structures of the encoder and decoder in this model are basically symmetrical. In addition, the end-to-end structure of the U-Net model also plays an important role in the generator, so that the output result can have the same size as the input image.
[0060] In addition, the U-Net model also adds skip-connections. The skip connection layer connects each encoding layer to the corresponding decoding layer and passes the compressed feature information in the model in a jump manner. Through the skip connection layer, spatial information can be passed directly from the encoder to the decoder. Therefore, the generator can obtain more detailed information during the image reconstruction process. In addition, the skip connection layer can provide better training results. Because through the skip connection layer, the gradient can flow in the deep data of the entire model structure. The addition of the skip connection layer is also because the encoder and decoder of the U-Net model have the same underlying structure.
[0061] The original facial micro-expression images and the strong facial micro-expression images near the vertex frames generated by the target in the pre-training process of the generator are derived from the video frame sequence of the same subject. This training mode can eliminate the excessive influence of facial information. In order to improve the image quality, the L1 norm of the generated image and the target image is used as the target loss function for pixel reconstruction, denoted as L pixel . That is, the optimization goal of the generator G is:
[0062] L pixel (G) = || G(x weak )-x apex || 1 ,
[0063] Among them, x weak is the original weak facial micro-expression image, x apex is the strong facial micro-expression image near the original vertex frame, G(x weak ) is composed of the original weak facial micro-expression image x weak The generated corresponding pseudo vertex frame near the strong facial micro-expression image, here it should be noted that: x apex That is, the target image described in step S1, G(x weak ) is the generated image described in step S1.
[0064] Reference Figure 4 , the figure is a schematic diagram of the structure of the discriminator part, refer to Figure 6 , the figure shows the pre-training scheme of the discriminator.
[0065] The model of the discriminator D can use the best deep learning classifier network, such as VGG, ResNet, etc.
[0066] The training process of the discriminator D includes:
[0067] (1) Encode facial information through a pre-trained main model to obtain all features;
[0068] (2) The extracted feature information will be input into two linear layers for micro-expression classification and enhanced detection respectively.
[0069] (3) Optimize the loss function in the pre-training process, that is, the cross entropy loss L of the classification result cls and the cross entropy loss L of the enhanced labels el The sum of L cls+el , recorded as:
[0070] In step S2, the second loss function when performing the second pre-training on the discriminator is:
[0071] L cls+el (D) = L 1 (y precls ,y cls )+λL 2 (y preel ,y el ),
[0072] Among them, L 1 (y precls ,y cls ) represents the actual micro-expression type y cls and the predicted micro-expression type y precls The cross entropy loss of L 2 (y preel ,y el ) is the enhanced label y el and predict the augmented label y preel In addition, a weight value λ is added to control the balance between the two losses.
[0073] Therefore, in addition to classifying micro-expressions, the model can also evaluate the function of the generator, that is, the enhancement effect of the generator. The discriminator designed by the present invention needs to classify weaker micro-expression frames and micro-expression frames near the vertex frames to achieve the function of evaluating the enhancement effect of the generator, which is similar to the true and false judgment function in the discriminator of the original generative adversarial network model.
[0074] Figure 7 represents the overall structure of the micro-expression discrimination model, and Figure 8 A diagram showing the joint training scheme of the adversarial network model. The former mainly introduces the specific structure of the model, while the latter mainly introduces the training scheme of the overall model. The training scheme diagram mainly explains the design of input and output and the key points of identity corresponding training.
[0075] Since the convergence speeds of the generator and discriminator mentioned above may be different. In most cases, the discriminator converges faster than the generator. For example, the loss function of the discriminator steadily decreases and becomes very small, or steadily increases and becomes very large. These are all signs that the network is not balanced. In rare cases, the situation is just the opposite.
[0076] In the above pre-training process, the generator and the discriminator are pre-trained separately. The generator will use the original weaker facial micro-expression frames in the database and the original stronger facial micro-expression frame images near the vertex frame for pre-training, so that the weak micro-expression frame images are generated to the vertex micro-expression frames. The key is to use identity-paired data to exclude the influence of identity information. At the same time, the generator will be controlled using the pixel reconstruction loss function so that the generator can achieve enhanced generation while maintaining the quality of the generated image. When the discriminator is pre-trained, the existing micro-expression classification labels in the database and the manually annotated weak and strong labels will be used for classification. This ensures that the discriminator has a certain ability to classify micro-expressions and distinguish between weak expressions and vertex strong expressions.
[0077] After pre-training, the generator G and the discriminator D will be connected together as a complete adversarial network model. During the joint training process, none of the components will be frozen.
[0078] The basic optimization goal of the overall structure joint training is a minimax objective function, that is, the generator goal is to minimize the loss function, and the discriminator goal is to maximize the loss function. That is, the adversarial loss function introduced in step S4 is:
[0079]
[0080] Among them, G represents the generator, D represents the discriminator, log D(x apex ) is [1 0] T and [D(x apex ) 1-D(x apex )] T The cross entropy between Represents a strong facial micro-expression image x apex According to the distribution probability p(x apex ) when sampling log D(x apex ); similarly, log(1-D(G(x weak )))Yes[1 0] T and [D(G(x weak )) 1-D(G(x weak ))] T The cross entropy between Indicates that when the input weak facial micro-expression image x weakAccording to the distribution probability p(x weak ) is sampled when log(1-D(G(x weak )))'s expected value.
[0081] Combining the optimization objectives of the generator and discriminator mentioned above, two weighted values τ and μ are used to control the balance between the generator and the discriminator, and the optimization objective L for joint training of the adversarial network model is MEEGAN for:
[0082] L MEEGAN =L adv +τL pixel (G)+μL cls+el (D)
[0083] Among them, L adv is the above-mentioned adversarial loss function, that is
[0084] After the micro-expression discrimination model of the present invention is obtained through the above training steps, the recognition method includes the following steps:
[0085] A. Preprocess the input image, including capturing the facial expression image, face alignment and scaling;
[0086] B. Input the input image into the generator part, and enhance the micro-expressions in the input image, wherein the micro-expression enhancement amplitude is kept within the micro-expression domain, and the identity information remains unchanged;
[0087] C. Input the enhanced image output from step B into the discriminator to obtain the classification result and enhanced label prediction result.
[0088] Specifically, the preprocessing process uses a more mainstream toolkit to extract facial key points, and uses a face calibration method to correct the face. Image scaling ensures that the image color is not distorted and the main features are not lost. Images that exceed a certain size are reduced to meet the input size of the model in step B.
[0089] Specifically, the generator model is a symmetrical encoder-decoder structure whose input image size is the same as the output image size.
[0090] Specifically, the classification result obtained by the discriminator is a three-category result of micro-expressions, including positive, negative, and surprised. The enhanced label is a self-defined strength judgment label of the present invention, defining the video frame image near the vertex frame as a strong expression and the other video frame images as a weak expression.
[0091] Experimental results:
[0092] The following will illustrate the effect of the present invention in accordance with the above implementation scheme combined with the experimental results. A control experiment is presented in the form of U-net+backbone network for comparison.
[0093]
[0094] Table 1 Comparison of accuracy and F1 value results of different backbone network construction models
[0095] Table 1 above lists the accuracy and F1 value results after joint training on Casme II and SAMM. Compared with other groups, U-Net combined with ResNet101 has the best results, with an accuracy of 88.6% when trained on the Casme II database and tested on Casme II, 79.1% when trained on the Casme II database and tested on SMIC, 84.5% when trained on the SAMM database and tested on it; and 77.8% when trained on the SAMM database and tested on SMIC. As can be seen from the above table, U-Net has similar performance when combined with different ResNets, but there are still differences in different data sets. The model combined with U-Net and ResNet performs better than other groups.
[0096] The method of the present invention is compared with other existing methods. The accuracy and F1 score are listed in Table 2 below, which proves that the present invention has better performance, especially in Case II and cross-dataset experiments. This shows that the enhanced generation of micro-expressions is feasible. The training scheme with paired identity information can reduce the deformation caused by the influence of identity information during the generation process, thereby making the classification in the discrimination process more reasonable.
[0097] The experimental results are compared with the related methods of the prior art. The references of the prior art are as follows:
[0098] Literature [1] Liong ST, See J, Phan CW, et al. Less is more: microexpressionrecognition from video using apex frame [C]. Signal Process.: Image Commun., 2018, 62: 82-92.
[0099] Literature [2] Liong ST, Gan YS, See J, et al. Shallow triple stream three-dimensional cnn (ststnet) for micro-expression recognition [C]. 2019 14th IEEEInternational Conference on Automatic Face&Gesture Recognition, 2019: 1-5.
[0100] Literature[3]Gan YS, Liong ST, Yau WC, et al.OFF-ApexNet on micro-expression recognition system[J].Signal Processing:Image Communication, 2019,74:129-139.
[0101]
[0102] Table 2 Comparison of the accuracy and F1 score of the method of the present invention and the prior art method
[0103] The embodiments described above are merely descriptions of preferred implementations of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope of the present invention.
Claims
1. A training method for a micro-expression discrimination model based on a generative adversarial network, characterized in that, it includes the steps: S1. Use the L1 norm of the generated image and the target image as the first loss function for pixel reconstruction, and perform the first pre-training on the generator; S2. Use the sum of the cross-entropy loss of the classification result and the cross-entropy loss of the enhanced label as the second loss function, and perform the second pre-training on the discriminator; S3. Connect the generator after the first pre-training and the discriminator after the second pre-training to generate a generative adversarial network model; S4. Based on the first loss function and the second loss function, and introduce an adversarial loss function, and perform joint training on the generative adversarial network model to obtain a trained micro-expression discrimination model. In the adversarial loss function, the generator aims to minimize the adversarial loss function, and the discriminator aims to maximize the adversarial loss function; In step S1, the first loss function during the first pre-training of the generator is: L pixel (G) = ||G(x weak ) - x apex || 1 , Among them, x weak is the original weak facial micro-expression image, and x apex is the strong facial micro-expression image near the original vertex frame. G(x weak ) is the corresponding strong facial micro-expression image near the pseudo vertex frame generated from the original weak facial micro-expression image x weak ; In step S2, the second loss function during the second pre-training of the discriminator is: L cls+el (D) = L 1 (y precls , y cls ) + λL 2 (y preel , y el ), Among them, L 1 (y precls , y cls ) represents the cross-entropy loss between the actual micro-expression category y c l s and the predicted micro-expression category y precls ; L 2 (y preel , y el ) is the cross-entropy loss between the enhanced label y el and the predicted enhanced label y preel , and λ represents the first weighting value; In step S4, the specifically introduced adversarial loss function is: Among them, G represents the generator and D represents the discriminator. logD(x apex ) is [1 0] T and [D(x apex ) 1 - D(x apex )] T the cross - entropy between denotes the expected value of logD(x apex ) when the strong face micro - expression image x apex is sampled according to the distribution probability p(x apex ); log(1 - D(G(x weak ))) is [1 0] T and [D(G(x weak )) 1 - D(G(x weak ))] T the cross - entropy between denotes the expected value of log(1 - D(G(x weak ))) when the input weak face micro - expression image x weak is sampled according to the distribution probability p(x weak ).
2. The training method for a micro-expression discrimination model based on a generative adversarial network according to claim 1, characterized in that, during the first pre-training process of the generator and the joint training process of the generative adversarial network model, the original human face micro-expression image and the human face micro-expression image near the target-generated vertex frame are from the video frame sequence of the same subject.
3. The training method for a micro-expression discrimination model based on a generative adversarial network according to claim 1, characterized in that, the generator includes an encoder, a decoder, and a controller; The encoder is used to convert the input video frame sequence into an encoded vector; The controller is used to add the corresponding-dimensional noise and the corresponding-dimensional enhanced label as control items to the encoded vector as control items to obtain a controlled encoded vector; The decoder is used to convert the controlled encoded vector into an output image sequence.
4. The training method for a micro-expression discrimination model based on a generative adversarial network according to claim 3, characterized in that, the encoder includes 5 sequentially connected encoding layers, and each encoding layer includes a connected conv_block unit layer and a max pooling layer. The conv_block unit layer includes 2 convolutional layers, 2 batch normalization layers, and 2 Relu functions.
5. The training method for a micro-expression discrimination model based on a generative adversarial network according to claim 4, characterized in that, the decoder includes 5 sequentially connected decoding layers corresponding to the 5 encoding layers in the encoder.
6. The training method for a micro-expression discrimination model based on a generative adversarial network according to claim 5, characterized in that, the generator further includes 5 skip connection layers, and the skip connection layers connect each encoding layer to the corresponding decoding layer.
7. The training method for a micro-expression discrimination model based on a generative adversarial network according to claim 1, characterized in that, in step S4, the optimization objective for the joint training of the generative adversarial network model is: L MEEGAN = L adv + τL pixel (G) + μL cls+el (D), Among them, L adv represents the adversarial loss function, τ represents the second weighting value, and μ represents the third weighting value.
Citation Information
Patent Citations
Image inpainting method and system based on antagonistic generation neural network
CN109191402A
Micro-expression type discrimination method based on transfer learning and auto-encoder data enhancement
CN111767842A