Cross-dataset micro-expression recognition method based on contribution of facial regions of interest

Through a group sparse model based on the contribution of the area of ​​interest of faces, the problem of difference in feature distribution between different data sets in micro-expression recognition is solved, and the recognition accuracy and classification stability are improved.

CN113971825BActive Publication Date: 2025-05-09SHANDONG FOREIGN TRADE VOCATIONAL COLLEGE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110903686.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-06
Publication Date
2025-05-09
Estimated Expiration
2041-08-06

AI Technical Summary

Technical Problem

The existing micro-expression recognition methods have large differences in feature distribution when the training samples and the samples to be identified come from different micro-expression data sets, resulting in a significant reduction in the recognition effect.

Method used

A cross-dataset micro-expression recognition method based on the contribution of the area of ​​interest of faces is adopted. A group sparse model is established for the MDMO features of the source facial image sequence, and a micro-expression type recognition is performed on the target facial image sequence.

Benefits of technology

The difference in MDMO feature distribution of training samples and test samples from different micro-expression data sets is narrowed, the recognition accuracy is improved, and the classification stability of different target data sets and micro-expression categories is better.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113971825B_ABST
    Figure CN113971825B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-dataset micro-expression recognition method based on the contribution of a region of interest of a human face, comprising: S1, a micro-expression sample preprocessing step; S2, a main direction average optical flow feature extraction step, calculating the optical flow field of each facial image sequence, and extracting MDMO features; S3, constraining the feature structure of the target sample according to the feature distribution characteristics of the source facial image sequence; S4, establishing a group sparse model for the MDMO features of the source facial image sequence, and quantifying the contribution of each region of interest; S5, using the group sparse model to identify the type of micro-expression of the target facial image sequence, and outputting the recognition result. The method of the present invention has a higher recognition accuracy, and has better classification stability for different target data sets and different micro-expression categories, and shows strong adaptability to test samples with different characteristics, and can greatly improve the performance of cross-dataset micro-expression recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and in particular, relates to a cross-dataset micro-expression recognition method based on the contribution of a facial region of interest. Background Art

[0002] Facial expressions are a direct reflection of human emotional states, and can usually be divided into macro-expressions and micro-expressions. In the past few years, academic research on facial expression recognition has mainly focused on macro-expressions. Unlike conventional macro-expressions, micro-expressions are quick and unconscious tiny facial movements that humans involuntarily reveal when they experience emotional fluctuations and try to conceal their inner emotions. Its special feature is that it can neither be disguised nor forcibly suppressed. Therefore, micro-expressions can be used as a reliable basis for analyzing and judging people's true emotions and psychological emotions, and have strong practical value and application prospects in clinical diagnosis, negotiation, teaching evaluation, polygraph detection and interrogation.

[0003] Micro-expressions last very short and fleeting, only less than half a second. The amplitude of facial muscle movement caused by micro-expressions is very small, only appears in a few small local facial areas, and usually does not appear in the upper and lower halves of the face at the same time. This makes micro-expressions difficult to be observed by the human eye, and the accuracy of manual recognition is not high. In addition, manual recognition of micro-expressions requires professional training and rich classification experience, which is time-consuming and laborious, and difficult to promote and apply on a large scale in real-life scenarios. Based on the large social needs and technological advances, in recent years, the use of computer vision and pattern recognition technology to achieve automatic identification of micro-expressions has been increasingly attracting the attention of scientific researchers.

[0004] At present, there are relatively few studies on micro-expression recognition using image processing technology, and the technology is still in its infancy. Due to the differences between micro-expressions and macro-expressions in duration, action intensity, and facial area, the current more mature macro-expression recognition methods are not suitable for micro-expression recognition.

[0005] The process of automatic micro-expression recognition can be divided into two stages: first, extracting micro-expression features, that is, extracting useful feature information from facial video clips to describe the micro-expressions contained in the video clips; then, classifying micro-expressions, using classifiers to classify the emotion categories to which the extracted features belong. In these two stages, feature selection is particularly important for micro-expression recognition. Therefore, most micro-expression recognition research focuses on feature extraction, aiming to effectively describe the subtle changes of micro-expressions by designing reliable micro-expression features in order to complete the micro-expression recognition task.

[0006] It should be pointed out that the development of micro-expression recognition research depends largely on a complete facial micro-expression dataset. By reviewing previous research work, it can be found that most of the existing micro-expression recognition methods are developed and evaluated when the training samples and test samples come from the same dataset. In this case, it can be considered that the training samples and test samples follow the same or similar feature distribution. However, it is obvious that in real applications, the training samples and the samples to be recognized often come from two completely different micro-expression datasets (referred to as the source dataset and the target dataset, respectively). The video clips in these two datasets will differ in terms of lighting conditions, shooting equipment, parameter settings, and background environment. Therefore, in this case, due to the heterogeneous video quality, the training samples and the samples to be recognized will be very different, resulting in a large difference in their feature distribution states, which greatly reduces the recognition effect of the existing micro-expression recognition methods. Summary of the invention

[0007] The present invention aims to solve the technical problem in the prior art that training samples and samples to be recognized in micro-expression recognition often come from two completely different micro-expression data sets, and there are also large differences in feature distribution states, which leads to a significant reduction in the recognition effect of the existing micro-expression recognition methods. A cross-dataset micro-expression recognition method based on the contribution of facial regions of interest is proposed to solve the above problem.

[0008] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solutions:

[0009] A cross-dataset micro-expression recognition method based on the contribution of facial regions of interest, comprising:

[0010] S1. Micro-expression sample preprocessing steps include:

[0011] S11, sampling the source micro-expression dataset and the target micro-expression dataset respectively, capturing video frames, and arranging them in sequence to obtain a source image sequence and a target image sequence respectively;

[0012] S12, downsampling the source image sequence and the target image sequence to adjust the size of the image;

[0013] S13, locating the face region in the image sequence, and performing facial image cropping on each image sequence to obtain a source facial image sequence and a target facial image sequence;

[0014] S14, performing facial landmark detection on the first frame image in each facial image sequence to obtain Q feature points describing key positions of the face;

[0015] S15, dividing the facial image into N specific non-overlapping but closely adjacent regions of interest using the coordinates of the feature points, wherein N<Q, and Q and N are both positive integers;

[0016] S16, graying each facial image sequence;

[0017] S2, extracting the main direction average optical flow feature step, calculating the optical flow field of each facial image sequence, and extracting the MDMO feature, wherein the MDMO feature is the main direction average optical flow feature based on the optical flow;

[0018] S3, constraining the feature structure of the target sample according to the feature distribution characteristics of the source facial image sequence, where the target sample is a test sample in the target micro-expression dataset;

[0019] S4, establishing a group sparse model for the MDMO features of the source facial image sequence, and quantifying the contribution of each region of interest;

[0020] S5. Use the group of sparse models to identify micro-expression types of the target facial image sequence, and output a recognition result.

[0021] Further, step S13 includes:

[0022] Perform face detection on the first frame of each image sequence to locate the facial area. Using the center point of the original rectangular bounding box as a reference, expand the front face selection box of the image to the surrounding areas in the same proportion to obtain the facial area.

[0023] According to the position and size of the detected facial region, other images in the image sequence are cropped to obtain a source facial image sequence and a target facial image sequence.

[0024] Furthermore, in step S15, the regions of interest are divided according to the facial action units in the facial action coding system, and each region of interest corresponds to a facial action unit.

[0025] Furthermore, after step S16, the following steps are further included:

[0026] S17, normalizing the number of frames of each facial image sequence, and using a time interpolation model to normalize the number of frames of each facial image sequence.

[0027] Furthermore, the method for calculating the optical flow field of each facial image sequence in step S2 is:

[0028] Calculate each frame f in the facial image sequence except the first frame i (i>1) and the optical flow vector [V x ,V y] and converted to polar coordinates (ρ, θ), where V x and V y are the x-component and y-component of the optical flow velocity, ρ and θ are the amplitude and angle of the optical flow velocity, respectively.

[0029] Furthermore, the method for extracting MDMO features in step S2 is:

[0030] In each frame f i (i>1), each region of interest All optical flow vectors in (k=1,2,…,N) are classified into 8 direction bins according to their angles, and the bin with the largest number of optical flow vectors is selected as the main direction, denoted as Bmax;

[0031] Calculate the average value of all optical flow vectors belonging to Bmax and define it as The main direction optical flow of is the average amplitude of the optical flow velocity, is the average angle of the optical flow speed;

[0032] Through an atomic optical flow feature Ψ i To represent each frame f i (i>1):

[0033]

[0034] Ψ i The dimension of is 2N, and an m-frame micro-expression video clip Γ can be represented as a set of atomic optical flow features:

[0035] Γ=(Ψ2,Ψ3,…,Ψ m ) (2)

[0036] For all i (i>1) (k=1,2,…,N) take the average, that is:

[0037]

[0038] is the average optical flow vector in the main direction of the kth region of interest;

[0039]

[0040] Pair Vector The amplitude in is normalized:

[0041]

[0042] In formula (5), Substitute into formula (4) and replace Get a new 2N-dimensional row vector As the MDMO features describing the video segment Γ:

[0043]

[0044] Furthermore, the method for constraining the characteristic structure of the target facial image sequence in step S3 is:

[0045] The MDMO features of the source facial image sequence are The MDMO features of the target facial image sequence are: Where d is the dimension of the feature vector, n s and n t are the number of source samples and the number of target samples, respectively. The source samples are the training samples in the source micro-expression dataset. The feature transformation of the target samples meets the following two requirements:

[0046] S31. The characteristics of the source sample should remain unchanged during this process, that is, the following conditions must be met:

[0047]

[0048] Where G is the target sample feature transformation operator;

[0049] S32, using function f G (X s ,X t ) as the regularization term of formula (7), and the objective function is obtained:

[0050]

[0051] Where λ is the weight coefficient, which is used to adjust the balance of the two terms in the objective function;

[0052] The target sample feature transformation operator G is determined by kernel mapping and linear projection operations.

[0053] Furthermore, the method for determining the target sample feature transformation operator G is as follows:

[0054] The source samples are projected from the original feature space to the Hilbert space through a kernel mapping operator φ;

[0055] Through a projection matrix φ(C)∈R ∞×d Transform the source sample from the Hilbert space back to the original feature space, G can be expressed as G(·) = φ(C) T The form of φ(·);

[0056] The objective function in formula (8) is rewritten as:

[0057]

[0058] Minimize the maximum mean difference distance MMD of the objective function in the Hilbert space; treat MMD as a regular term f G (X s ,X t ):

[0059]

[0060] Where H represents the Hilbert space, 1 s and 1 t They are of length n s and n t A column vector whose elements are all 1;

[0061] The MMD in formula (10) is transformed into the following form, as f G (X s ,X t ):

[0062]

[0063] Replace f in formula (11) G (X s ,X t ) into formula (9), the objective function becomes:

[0064]

[0065] The optimization problem shown in formula (12) can be converted into a solvable form by calculating the kernel function instead of the inner product operation in the kernel space, including: Let φ(C) = [φ(X s ),φ(X t )]P, where the linear coefficient matrix Then formula (12) is rewritten as the following form as the final objective function:

[0066]

[0067] in The calculation formula of the four kernel matrices is K ss =φ(X s ) T φ(X s ), K st =φ(X s ) T φ(X t ), K ts =φ(X t ) T φ(Xs ) and K tt =φ(X t ) T φ(X t );

[0068] In formula (13), an L1 norm of P is added as a constraint term of the objective function, that is, where p i is the i-th column of P, and the sparsity of P is adjusted by the weighting coefficient μ.

[0069] Furthermore, in step S4, groups are used as sparse representation units, and each group is composed of an MDMO feature matrix of a face region of interest, so as to quantify the contribution of each face region of interest, including:

[0070] The MDMO feature matrix corresponding to the M micro-expression training samples is X = [x1,…,x M ]∈R d×M , where d is the dimension of the feature vector, d = 2N;

[0071] The label vector is used to represent the categories of micro-expressions, including:

[0072] Let L = [l1,…,l M ]∈R c×M represents the label matrix corresponding to the feature matrix X, where c is the number of micro-expression types; the kth column l k =[l k,1 ,…,l k,c ] T (1≤k≤M) is a column vector whose elements are either 0 or 1 according to the following rules:

[0073]

[0074] The above label vectors are a set of orthogonal bases, which are expanded into a vector space containing label information. A projection matrix U is introduced to establish the connection between the feature space and the label space of the sample. The projection matrix U is obtained by solving the objective function:

[0075]

[0076] U in formula (15) T X can be rewritten by matrix decomposition as Where N is the number of face regions of interest, N = 36; X i is the MDMO feature matrix of the i-th region of interest; U i Yes X i The corresponding sub-projection matrix; use Replace U in formula (15)T X, we can get the equivalent formula:

[0077]

[0078] In formula (16), a weighting coefficient β is introduced for each region of interest i , and add a beta i The non-negative L1 norm of As a regular term, a linear group sparse model is formed:

[0079]

[0080] Where μ is a trade-off coefficient that determines the number of non-zero elements in the learned weight vector β;

[0081] The linear kernel of the sparse model of the group is expanded into a nonlinear kernel, using the nonlinear mapping φ:R d →F put X i and U i Mapped to the kernel space F, that is, using and Replace X in formula (17) respectively i and U i :

[0082]

[0083] By replacing the inner product operation in the kernel space with the kernel function, in the kernel space F, Each column of It can be expressed as Right now A linear combination of j is a linear coefficient vector; so can be Represents, where P = [p1,…,p c ];

[0084] Will Substitute into formula (18) and add an L1 norm about P As a constraint, to ensure that p j The sparsity of and avoid overfitting when optimizing the objective function, the final form of the group sparse model is obtained:

[0085]

[0086] in is the Gram matrix; λ is the weighting coefficient used to adjust the sparsity of P;

[0087] The alternating direction method is used to solve the optimization problem of formula (19), that is, to update the parameters P and β alternately and iteratively.i until the objective function converges.

[0088] Further, step S5 includes:

[0089] For the training samples in the source micro-expression dataset, the optimal parameter value is learned through iteration and Finally, the group sparse model is used as a classifier to predict the label vector of the test samples in the target dataset, that is, to identify the types of micro-expressions;

[0090] For a test sample, let its feature vector be x t ∈R 72×1 , we can predict the label vector l of the sample by solving the following optimization problem: t :

[0091]

[0092] in It can be calculated by the kernel function selected when learning the group sparse model;

[0093] Assume that the obtained label vector is Then the micro-expression type of the test sample is in express The kth element of .

[0094] Compared with the prior art, the advantages and positive effects of the present invention are:

[0095] The cross-dataset micro-expression recognition method based on the contribution of the face region of interest of the present invention constrains the feature structure of the target sample according to the feature distribution characteristics of the source facial image sequence, reduces the difference in MDMO feature distribution between training samples and test samples from different micro-expression data sets, has higher recognition accuracy, and has better classification stability for different target data sets and different micro-expression categories. It shows strong adaptability to test samples with different characteristics, and can greatly improve the performance of cross-dataset micro-expression recognition.

[0096] After reading the specific embodiments of the present invention in conjunction with the accompanying drawings, other features and advantages of the present invention will become more clear. BRIEF DESCRIPTION OF THE DRAWINGS

[0097] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0098] Figure 1 It is a principle block diagram of an embodiment of a method for cross-dataset micro-expression recognition based on contribution of a facial region of interest proposed by the present invention;

[0099] Figure 2 is a schematic diagram of dividing the region of interest of a face in the first embodiment;

[0100] Figure 3 This is a recognition result diagram of the method in Example 1 using the CASMEII->CASME dataset;

[0101] Figure 4 This is a recognition result diagram of the method in Example 1 using the CASMEII->SMIC-HS data set;

[0102] Figure 5 This is a recognition result diagram of the method in Example 1 using the SMIC-HS->CASME data set;

[0103] Figure 6 This is a recognition result diagram of the method in Example 1 using the SMIC-HS->CASMEII data set;

[0104] Figure 7 This is a diagram of the recognition results of the method in Example 1 using the SMIC-HS->SAMM data set. DETAILED DESCRIPTION

[0105] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0106] It should be noted that in the description of the present invention, the terms "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicating directions or positional relationships are based on the directions or positional relationships shown in the drawings, which are only for the convenience of description, and do not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first" and "second" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.

[0107] Embodiment 1

[0108] This embodiment proposes a cross-dataset micro-expression recognition method based on the contribution of the face interest region, such as Figure 1 As shown, including:

[0109] S1. Micro-expression sample preprocessing steps include:

[0110] S11, sampling the source micro-expression dataset and the target micro-expression dataset respectively, capturing video frames, and arranging them in sequence to obtain a source image sequence and a target image sequence respectively;

[0111] S12, downsampling the source image sequence and the target image sequence to adjust the size of the image;

[0112] S13, locating the face region in the image sequence, and performing facial image cropping on each image sequence to obtain a source facial image sequence and a target facial image sequence;

[0113] S14, performing facial landmark detection on the first frame image in each facial image sequence to obtain Q feature points describing key positions of the face;

[0114] S15, dividing the facial image into N specific non-overlapping but closely adjacent regions of interest using the coordinates of the feature points, wherein N<Q, and Q and N are both positive integers;

[0115] S16, graying each facial image sequence;

[0116] S2, extracting the main direction average optical flow feature step, calculating the optical flow field of each facial image sequence, and extracting the MDMO feature, wherein the MDMO feature is the main direction average optical flow feature based on the optical flow;

[0117] S3, constraining the feature structure of the target sample according to the feature distribution characteristics of the source facial image sequence, where the target sample is a test sample in the target micro-expression dataset;

[0118] S4, establishing a group sparse model for the MDMO features of the source facial image sequence, and quantifying the contribution of each region of interest;

[0119] S5. Use the group of sparse models to identify micro-expression types of the target facial image sequence, and output a recognition result.

[0120] The sample in the micro-expression dataset refers to a complete video clip of a micro-expression with a certain emotion, including three important video frames: onset frame, apex frame, and offset frame. Onset refers to the moment when the micro-expression begins to appear; apex refers to the moment when the micro-expression has the largest amplitude; and offset refers to the moment when the micro-expression disappears.

[0121] In step S11, the facial micro-expression video clip is first converted into an image sequence. For a micro-expression video clip Γ, a plurality of continuous static images, ie, video frames, are captured from the video clip by setting a sampling interval.

[0122] Then, the number of redundant frames is reduced by interval sampling. Assuming the original frame rate is m frames / second and the video duration is tvid seconds, the video has a total of m×tvid frames. Assuming the sampling period is tsam seconds, the corresponding number of frames is m×tsam, which means that one frame is extracted every m×tsam frames. In this way, the resulting image sequence contains only [tvid / tsam] frames, where [] is a rounding function.

[0123] In step S12, downsampling processing based on bicubic interpolation may be performed on the video frames in all the image sequences, and the width of the video frames may be uniformly adjusted to 500 pixels, while keeping the aspect ratio unchanged.

[0124] Since the duration of micro-expression video clips is short, the amplitude of the position movement (including translation and rotation) of the head in multiple consecutive frames of each image sequence is very small and has been roughly aligned; at the same time, in order to improve the efficiency of the algorithm, this embodiment only uses the face detector proposed by Masayuki Tanaka to perform face detection on the first frame of each image sequence to locate the facial area. This detection algorithm can not only detect multiple frontal faces appearing in the same image with high precision, but also simultaneously detect their corresponding left eyes, right eyes, mouths and noses. In particular, when the input image is rotated or the head of the person in the image is tilted, the detection effect is still excellent.

[0125] It should be pointed out that the algorithm can only detect images with three color channels, while the SAMM data set provides a grayscale image sequence with a single color channel. Therefore, in actual operation, all of them are converted into three-channel forms. In addition, after a large number of experiments, it was found that there are differences in objective conditions such as lighting conditions, background complexity, and skin color and face shape of subjects in different data sets, which will make the face detection algorithm under the same parameter setting unable to accurately locate the front face area in all sample image sequences. For example, in the face area determined by the detection of some subjects, a part of the chin is missing. To solve this problem, this embodiment uses the center point of the original rectangular bounding box as a reference, and appropriately expands the front face selection box in all sample image sequences in the same proportion to ensure that the facial area of ​​the appropriate size is obtained.

[0126] For each image sequence, a region cropping operation is performed on all frames according to the position and size of the facial region detected in the first frame to form a new facial image sequence.

[0127] Step S13 includes:

[0128] The first frame image in each image sequence is subjected to face detection to locate the facial region, and the front face selection frame of the image is expanded outwards in the same proportion to obtain the facial region, taking the center point of the original rectangular bounding box as a reference. In this embodiment, the Masayuki Tanaka detector can be used to perform face detection to locate the facial region.

[0129] According to the position and size of the detected facial region, other images in the image sequence are cropped to obtain a source facial image sequence and a target facial image sequence.

[0130] like Figure 2 As shown in FIG. 1 , an example of using the DFRM algorithm to detect facial landmark points on a micro-expression data set, where the “+” marker represents the detected key feature points. In step S14, the DFRM algorithm (Discriminative Fitting of Response Map) that relies on a texture model is used to perform facial landmark point detection on the first frame image in each facial image sequence to obtain Q feature points that describe the key positions of the face. In this embodiment, 66 feature points are obtained as an example for explanation.

[0131] In step S15, the regions of interest are divided according to the facial action units in the facial action coding system, and each region of interest corresponds to a facial action unit.

[0132] There are many strategies for dividing the face ROI, but the general principle is that it should be neither too dense nor too sparse. If it is divided too densely, redundant information may be introduced; if it is divided too sparsely, useful information may be missed.

[0133] Since micro-expressions only involve contraction or relaxation of local facial muscles, this embodiment further divides the facial area into 36 specific non-overlapping but closely adjacent regions of interest by using the coordinates of key feature points obtained by the DFRM algorithm, while excluding some irrelevant regions.

[0134] By using the coordinates of the key feature points obtained by the DFRM algorithm, the facial area is further divided into N specific non-overlapping but closely adjacent regions of interest, while excluding some irrelevant regions. In this embodiment, N is 36, such as Figure 2As shown in the figure, the positions and sizes of these regions of interest are uniquely determined by 66 feature points, and the division is based on the facial action unit (AU) in the Facial Action Coding System (FACS). Each region of interest corresponds to a part of the facial action unit, which can better reflect the apparent changes caused by facial muscle movements. The combination of all regions of interest can express almost all types of micro-expressions.

[0135] Step S16 converts all color sample image sequences into grayscale image sequences to prevent the color information therein from being affected by light.

[0136] After step S16, the following steps are also included:

[0137] S17, normalizing the number of frames of each facial image sequence, and using a time interpolation model to normalize the number of frames of each facial image sequence.

[0138] In this embodiment, the temporal interpolation model (TIM) proposed by Zhou et al. can be used to normalize the number of frames of each sample, and the required number of frames can be interpolated from the low-dimensional manifold structure established by the face image sequence, thereby avoiding too few or too many frames.

[0139] The method for calculating the optical flow field of each facial image sequence in step S2 is:

[0140] Calculate each frame f in the facial image sequence except the first frame i (i>1) and the optical flow vector [V x ,V y ] and converted to polar coordinates (ρ, θ), where V x and V y are the x-component and y-component of the optical flow velocity, ρ and θ are the amplitude and angle of the optical flow velocity, respectively.

[0141] In step S2, an improved main direction mean optical flow (MDMO) feature based on optical flow is extracted.

[0142] As a preferred embodiment, the method for extracting MDMO features is:

[0143] The optical flow field of the grayscale image sequence is calculated using the Robust Local Optical Flow (RLOF) algorithm based on the Hampel estimator to quantitatively estimate the movement of the subject's facial muscles.

[0144] For a micro-expression image sequence with m frames (f1, f2, ..., f m ), considering that the change between two adjacent frames is very small, the optical flow change presented is not obvious, therefore, this embodiment calculates each frame except the first frame f i (i>1) and the optical flow vector [V x ,V y ](where V x and V y The x and y components of the optical flow velocity are respectively, and the Cartesian coordinates are converted into polar coordinates (ρ, θ) (where ρ and θ are amplitude and angle, respectively) to facilitate subsequent feature extraction.

[0145] In each frame f i (i>1), each region of interest All optical flow vectors in (k=1,2,…,N) are classified into 8 direction bins according to their angles, and the bin with the largest number of optical flow vectors is selected as the main direction, denoted as Bmax;

[0146] Calculate the average value of all optical flow vectors belonging to Bmax and define it as The main direction optical flow of is the average amplitude of the optical flow velocity, is the average angle of the optical flow speed;

[0147] Through an atomic optical flow feature Ψ i To represent each frame f i (i>1):

[0148]

[0149] Ψ i The dimension of is 2N, and an m-frame micro-expression video clip Γ can be represented as a set of atomic optical flow features:

[0150] Γ=(Ψ2,Ψ3,…,Ψ m ) (2)

[0151] For all i (i>1) (k=1,2,…,N) take the average, that is:

[0152]

[0153] is the average optical flow vector in the main direction of the kth ROI. The above formula means that the average optical flow vector in the main direction of the ROI (all kth ROI) at the same position in all frames (starting from the 2nd frame) in the current video clip is taken to obtain the average optical flow vector in the main direction of the kth ROI.

[0154]

[0155] Considering that the magnitude of the main direction in different video clips may vary greatly, the vector The amplitude in is normalized:

[0156]

[0157] In formula (5), Substitute into formula (4) and replace Get a new 2N-dimensional row vector As the MDMO features describing the video segment Γ:

[0158]

[0159] Step S3 uses a transfer learning method to narrow the difference in MDMO feature distribution between training samples and test samples from different micro-expression datasets.

[0160] Assuming that the label information of the target sample is completely unknown, the feature structure of the target sample needs to be transformed according to the feature distribution characteristics of the source sample.

[0161] The method for constraining the characteristic structure of the target facial image sequence in step S3 is:

[0162] The MDMO features of the source facial image sequence are The MDMO features of the target facial image sequence are: Where d is the dimension of the feature vector, n s and n t are the number of source samples and the number of target samples, respectively. The source samples are the training samples in the source micro-expression dataset. The feature transformation of the target samples meets the following two requirements:

[0163] S31. The characteristics of the source sample should remain unchanged during this process, that is, the following conditions must be met:

[0164]

[0165] Where G is the target sample feature transformation operator;

[0166] After S32 and G transform the target sample features, the newly reconstructed target sample features and the source sample features should have the same or similar distribution characteristics. To this end, the function f is used. G (X s ,X t ) as the regularization term of formula (7), and the objective function is obtained:

[0167]

[0168] Where λ is the weight coefficient, which is used to adjust the balance of the two terms in the objective function;

[0169] The target sample feature transformation operator G is determined by kernel mapping and linear projection operations.

[0170] Preferably, the method for determining the target sample feature transformation operator G is:

[0171] First, the source sample is projected from the original feature space to the Hilbert space through a kernel mapping operator φ; then a projection matrix φ(C)∈R ∞×d Transform the source sample from the Hilbert space back to the original feature space. Based on this, G can be expressed as G(·) = φ(C) T The form of φ(·);

[0172] The objective function in formula (8) is rewritten as:

[0173]

[0174] In order to eliminate the difference in feature distribution between source samples and target samples, the maximum mean difference distance MMD of the objective function in the Hilbert space can be minimized; MMD is used as the regularization term f G (X s ,X t ):

[0175]

[0176] Where H represents the Hilbert space, 1 s and 1 t They are of length n s and n t and the elements are all 1 column vector; however, directly using MMD as f G (X s ,X t ), it is necessary to learn the optimal kernel mapping operator φ, which is obviously very difficult. Therefore, the MMD in formula (10) is transformed into the following form, as f G (X s ,X t ):

[0177]

[0178] It can be shown that minimizing the MMD in formula (10) is equivalent to minimizing f in formula (11) G (X s ,X t ). In this way, f G (X s ,X t ) only needs to learn the optimal φ(C), and φ(C) is also the variable that needs to be learned in formula (9).

[0179] Replace f in formula (11) G (X s ,X t ) into formula (9), the objective function becomes:

[0180]

[0181] The optimization problem shown in formula (12) can be converted into a solvable form by calculating the kernel function instead of the inner product operation in the kernel space, including: Let φ(C) = [φ(X s ),φ(X t )]P, where the linear coefficient matrix Then formula (12) is rewritten as the following form as the final objective function:

[0182]

[0183] in The calculation formula of the four kernel matrices is K ss =φ(X s ) T φ(X s ), K st =φ(X s ) T φ(X t ), K ts =φ(X t ) T φ(X s ) and K tt =φ(X t ) T φ(X t );

[0184] To prevent overfitting when optimizing the objective function, an L1 norm of P is added to formula (13) as a constraint term of the objective function, namely: where p i is the i-th column of P, and the sparsity of P is adjusted by the weighting coefficient μ.

[0185] Step S4 establishes a group sparse model based on the 72-dimensional MDMO features and micro-expression label information from the 36 facial regions of interest, using groups as sparse representation units. Each group is composed of an MDMO feature matrix of a facial region of interest, and quantifies the contribution of each facial region of interest, including:

[0186] The MDMO feature matrix corresponding to the M micro-expression training samples is X = [x1,…,x M ]∈R d×M , where d is the dimension of the feature vector, d = 2N;

[0187] The label vector is used to represent the categories of micro-expressions, including:

[0188] Let L = [l1,…,l M ]∈R c×M represents the label matrix corresponding to the feature matrix X, where c is the number of micro-expression types; the kth column l k =[l k,1 ,…,l k,c ] T (1≤k≤M) is a column vector whose elements are either 0 or 1 according to the following rules:

[0189]

[0190] The above label vectors are a set of orthogonal bases, which are expanded into a vector space containing label information. A projection matrix U is introduced to establish the connection between the feature space and the label space of the sample. The projection matrix U is obtained by solving the objective function:

[0191]

[0192] U in formula (15) T X can be rewritten by matrix decomposition as Where N is the number of face regions of interest, N = 36; X i is the MDMO feature matrix of the i-th region of interest; U i Yes X i The corresponding sub-projection matrix; use Replace U in formula (15) T X, we can get the equivalent formula:

[0193]

[0194] In order to numerically measure the specific contribution of each facial region of interest to the occurrence of micro-expressions, a weighting coefficient β is introduced for each region of interest in formula (16): i, and add a beta i The non-negative L1 norm of As a regular term, a linear group sparse model is formed:

[0195]

[0196] Where μ is a trade-off coefficient that determines the number of non-zero elements in the learned weight vector β;

[0197] The regularization term in formula (17) has two benefits. First, during the learning of the model, the regions of interest that have little contribution to micro-expression recognition will be discarded (the corresponding coefficient β i is 0); secondly, each selected region of interest will be assigned a positive rational number weight to measure its contribution.

[0198] In order to improve the classification performance of the group sparse model, the linear kernel of the group sparse model is further expanded to a nonlinear kernel, and the nonlinear mapping φ:R d →F put X i and U i Mapped to the kernel space F, that is, using and Replace X in formula (17) respectively i and U i :

[0199]

[0200] By replacing the inner product operation in the kernel space with the kernel function, in the kernel space F, Each column of It can be expressed as Right now A linear combination of j is a linear coefficient vector; so can be Represents, where P = [p1,…,p c ];

[0201] Will Substitute into formula (18) and add an L1 norm about P As a constraint, to ensure that p j The sparsity of and avoid overfitting when optimizing the objective function, the final form of the group sparse model is obtained:

[0202]

[0203] in is the Gram matrix; λ is the weighting coefficient used to adjust the sparsity of P;

[0204] The alternating direction method is used to solve the optimization problem of formula (19), that is, to update the parameters P and β alternately and iteratively. i until the objective function converges.

[0205] Step S5 includes:

[0206] For the training samples in the source micro-expression dataset, the optimal parameter value is learned through iteration and Finally, the group sparse model is used as a classifier to predict the label vector of the test samples in the target dataset, that is, to identify the types of micro-expressions;

[0207] For a test sample, let its feature vector be x t ∈R 2N×1 , we can predict the label vector l of the sample by solving the following optimization problem: t :

[0208]

[0209] in It can be calculated by the kernel function selected when learning the group sparse model;

[0210] Assume that the obtained label vector is Then the micro-expression type of the test sample is in express The kth element of .

[0211] In order to verify the effectiveness of the cross-dataset micro-expression recognition algorithm based on facial ROIs Contribution Quantification (FRCQ) proposed in the present invention, a large number of cross-dataset micro-expression recognition experiments in pairs were conducted on four micro-expression datasets: CASME, CASMEII, SMIC-HS, and SAMM. One of them acts as a source dataset to provide training samples; the other acts as a target dataset to provide test samples.

[0212] This example compares the FRCQ algorithm with three more advanced micro-expression recognition algorithms. The experimental comparison results are as follows: Figures 3 to 7As shown. None of the three comparison methods made any changes to the extracted features, and all used the widely used support vector machine with polynomial kernel as the classifier. Among them, comparison method 1 extracts the LBP-TOP features of the entire facial area (referred to as LBP-TOP-Whole+SVM); comparison method 2 extracts the LBP-TOP features of each face ROI (36 in total) and connects them into a combined feature (referred to as LBP-TOP-ROIs+SVM); comparison method 3 extracts the original MDMO features of the face (referred to as MDMO+SVM).

[0213] Due to the limitation of length and space, only a part of the experimental results are shown here. In the following description, the symbol "A->B" is used to represent the micro-expression recognition experiment from source dataset A to target dataset B.

[0214] A.CASMEII->CASME, such as Figure 3 The figure shows the comparison results of different methods in the cross-dataset micro-expression recognition experiment from the source dataset CASME II to the target dataset CASME. From top to bottom are the confusion matrix and F1-Measure bar chart, and the recognition accuracy from left to right is 50%, 20.31%, 53.13% and 67.19% respectively.

[0215] B. CASMEII->SMIC-HS, such as Figure 4 The figure shows the comparison results of different methods in the cross-dataset micro-expression recognition experiment from the source dataset CASME II to the target dataset SMIC-HS. From top to bottom are the confusion matrix and F1-Measure bar chart, and the recognition accuracy from left to right is 35.48%, 26.45%, 46.45% and 49.03% respectively.

[0216] C.SMIC-HS->CASME, e.g. Figure 5 The figure shows the comparison results of different methods in the cross-dataset micro-expression recognition experiment from the source dataset SMIC-HS to the target dataset CASME. From top to bottom are the confusion matrix and F1-Measure bar chart, and the recognition accuracy from left to right is 50.00%, 46.88%, 57.81% and 62.50% respectively.

[0217] D.SMIC-HS->CASMEII, such as Figure 6 The figure shows the comparison results of different methods in the cross-dataset micro-expression recognition experiment from the source dataset SMIC-HS to the target dataset CASME II. From top to bottom are the confusion matrix and F1-Measure bar chart, and the recognition accuracy from left to right is 22.12%, 27.43%, 63.72% and 71.68% respectively.

[0218] E.SMIC-HS->SAMM, such as Figure 7 The figure shows the comparison results of different methods in the cross-dataset micro-expression recognition experiment from the source dataset SMIC-HS to the target dataset SAMM. From top to bottom are the confusion matrix and F1-Measure bar chart, and the recognition accuracy from left to right is 32.33%, 43.61%, 45.11% and 51.13% respectively.

[0219] In the above five groups of comparative experiments, in order to quantitatively compare and analyze the recognition effect and overall recognition effect of each method for three types of micro-expressions, namely, positive, negative and surprise, this embodiment draws a confusion matrix and an F1-Measure bar chart respectively, and gives the overall recognition accuracy of each method.

[0220] By observing the various confusion matrices given, it is not difficult to find that compared with the currently more advanced combination method of "LBP-TOP or MDMO features + support vector machine", the FRCQ method proposed in the present invention always maintains a high level of recognition accuracy for the three types of micro-expressions, and the numerical fluctuation between classes is very small. In particular, in the two groups of experiments CASME II->CASME and SMIC-HS->CASMEII, the recognition accuracy of the FRCQ method for the three types of micro-expressions exceeded 60%. In terms of overall recognition accuracy, the FRCQ method achieved the highest value in all five groups of comparative experiments.

[0221] exist Figures 3 to 7 In the F1-Measure bar chart shown, the four recognition methods have their own strengths, and each method has its own specific micro-expression categories that it is good at classifying. However, except for the FRCQ method, the F1-Measure values ​​of the other three methods are not stable enough and fluctuate to varying degrees. This shows that they are not adaptable to the micro-expression image sequences in the target dataset, and the quality of classification is somewhat accidental, and they are not suitable for classifying target samples that are significantly different from the source samples. Obviously, the F1-Measure value of the FRCQ method is higher than that of other methods as a whole, and it always maintains a high value. This shows that its classification quality is higher, and it has better discrimination of small differences in facial detail features; at the same time, the classification performance is more stable and more robust, and it can successfully complete the cross-dataset classification task.

[0222] In this embodiment, a large number of cross-dataset micro-expression recognition experiments were carried out in pairs on four spontaneous micro-expression datasets, namely CASME, CASME II, SMIC-HS and SAMM. The experimental results show that the recognition strategy proposed in the present invention is effective. Compared with several existing advanced recognition methods, the recognition effect is better: not only is the recognition accuracy higher, but the classification stability for different target datasets and different micro-expression categories is better, and it shows strong adaptability to test samples with different characteristics, which can greatly improve the performance of cross-dataset micro-expression recognition.

[0223] The micro-expression recognition scheme proposed in the present invention makes it possible to automatically analyze large-scale micro-expression video clips in real time and even apply them in natural scenes. It has important scientific value and broad application prospects in many fields such as clinical diagnosis, social interaction and national security.

[0224] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for a person skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to replace some of the technical features therein by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions claimed to be protected by the present invention.

Claims

1. A cross-dataset micro-expression recognition method based on the contribution of facial interest regions, characterized in that: include: S1. Micro-expression sample preprocessing steps include: S11, sampling the source micro-expression dataset and the target micro-expression dataset respectively, capturing video frames, and arranging them in sequence to obtain a source image sequence and a target image sequence respectively; S12, downsampling the source image sequence and the target image sequence to adjust the size of the image; S13, locating the face region in the image sequence, and performing facial image cropping on each image sequence to obtain a source facial image sequence and a target facial image sequence; S14, performing facial landmark detection on the first frame image in each facial image sequence to obtain Q feature points describing key positions of the face; S15, using the coordinates of the feature points to divide the facial image into N non-overlapping but closely adjacent regions of interest, where N<Q, and Q and N are both positive integers; S16, graying each facial image sequence; S2, extracting the main direction average optical flow feature step, calculating the optical flow field of each facial image sequence, and extracting the MDMO feature, wherein the MDMO feature is the main direction average optical flow feature based on the optical flow; S3, constraining the feature structure of the target sample according to the feature distribution characteristics of the source facial image sequence, where the target sample is a test sample in the target micro-expression dataset; S4, establishing a group sparse model for the MDMO features of the source facial image sequence, and quantifying the contribution of each region of interest; S5, using the group of sparse models to identify the types of micro-expressions of the target facial image sequence, and outputting the identification results; In step S4, groups are used as sparse representation units, and each group is composed of an MDMO feature matrix of a face region of interest, so as to quantify the contribution of each face region of interest, including: The MDMO feature matrix corresponding to the M micro-expression training samples is X = [x1,…,x M ]∈R d×M , where d is the dimension of the feature vector, d = 2N; The label vector is used to represent the categories of micro-expressions, including: Let L = [l1,…,l M ]∈R c×M represents the label matrix corresponding to the feature matrix X, where c is the number of micro-expression types; the kth column l k =[l k,1 ,…,l k,c ] T (1≤k≤M) is a column vector whose elements are either 0 or 1 according to the following rules: The label vector is a set of orthogonal bases, which are expanded into a vector space containing label information. A projection matrix U is introduced to establish the connection between the feature space and the label space of the sample. The projection matrix U is obtained by solving the objective function: U in formula (15) T X can be rewritten by matrix decomposition as Where N is the number of face regions of interest, N = 36; X i is the MDMO feature matrix of the i-th region of interest; U i Yes X i The corresponding sub-projection matrix; use Replace U in formula (15) T X, we get the equivalent formula: In formula (16), a weighting coefficient β is introduced for each region of interest i , and add a beta i The non-negative L1 norm of As a regular term, a linear group sparse model is formed: Where μ is a trade-off coefficient that determines the number of non-zero elements in the learned weight vector β; The linear kernel of the group sparse model is expanded into a nonlinear kernel, using the nonlinear mapping φ:R d →F put X i and U i Mapped to the kernel space F, that is, using and Replace X in formula (17) respectively i and U i : By replacing the inner product operation in the kernel space with the kernel function, in the kernel space F, Each column of Expressed as Right now A linear combination of j is a linear coefficient vector; so can be Represents, where P = [p1,…,p c ]; Will Substitute into formula (18) and add an L1 norm about P As a constraint, to ensure that p j The sparsity of and avoid overfitting when optimizing the objective function, the final form of the group sparse model is obtained: in is the Gram matrix; λ is the weighting coefficient used to adjust the sparsity of P; The alternating direction method is used to solve the optimization problem of formula (19), that is, to update the parameters P and β alternately and iteratively. i until the objective function converges.

2. The micro-expression recognition method according to claim 1, characterized in that: Step S13 includes: Perform face detection on the first frame of each image sequence to locate the facial area. Using the center point of the original rectangular bounding box as a reference, expand the front face selection box of the image to the surrounding areas in the same proportion to obtain the facial area. According to the position and size of the detected facial region, other images in the image sequence are cropped to obtain a source facial image sequence and a target facial image sequence.

3. The micro-expression recognition method according to claim 1, characterized in that: In step S15, the regions of interest are divided according to the facial action units in the facial action coding system, and each region of interest corresponds to a facial action unit.

4. The micro-expression recognition method according to claim 1, characterized in that: After step S16, the following steps are also included: S17, normalizing the number of frames of each facial image sequence, and using a time interpolation model to normalize the number of frames of each facial image sequence.

5. The micro-expression recognition method according to claim 1, characterized in that: The method for calculating the optical flow field of each facial image sequence in step S2 is: Calculate each frame f in the facial image sequence except the first frame i (i>1) and the optical flow vector [V x ,V y ] and converted to polar coordinates (ρ, θ), where V x and V y are the x-component and y-component of the optical flow velocity, ρ and θ are the amplitude and angle of the optical flow velocity, respectively.

6. The micro-expression recognition method according to claim 5, characterized in that: The method for extracting MDMO features in step S2 is: In each frame f i (i>1), each region of interest All optical flow vectors in are classified into 8 direction bins according to their angles, and the bin with the largest number of optical flow vectors is selected as the main direction, denoted as Bmax; Calculate the average value of all optical flow vectors belonging to Bmax and define it as The main direction optical flow of is the average amplitude of the optical flow velocity, is the average angle of the optical flow speed; Through an atomic optical flow feature Ψ i To represent each frame f i (i>1): Ψ i The dimension of is 2N, and an m-frame micro-expression video clip Γ can be represented as a set of atomic optical flow features: C=(Ψ2,Ψ3,…,Ψ m ) (2) For all i (i>1) Take the average, that is: is the average optical flow vector in the main direction of the kth region of interest; Pair Vector The amplitude in is normalized: In formula (5), Substitute into formula (4) and replace Get a new 2N-dimensional row vector As the MDMO features describing the video segment Γ:

7. The micro-expression recognition method according to claim 1, characterized in that: The method for constraining the characteristic structure of the target facial image sequence in step S3 is: The MDMO features of the source facial image sequence are The MDMO features of the target facial image sequence are: Where d is the dimension of the feature vector, n s and n t are the number of source samples and the number of target samples, respectively. The source samples are the training samples in the source micro-expression dataset. The feature transformation of the target samples meets the following two requirements: S31. The characteristics of the source sample should remain unchanged during this process, that is, the following conditions must be met: Where G is the target sample feature transformation operator; S32, using function f G (X s ,X t ) as the regularization term of formula (7), and the objective function is obtained: Where λ is the weight coefficient, which is used to adjust the balance of the two terms in the objective function; The target sample feature transformation operator G is determined by kernel mapping and linear projection operations.

8. The micro-expression recognition method according to claim 7, characterized in that: The method for determining the target sample feature transformation operator G is: The source samples are projected from the original feature space to the Hilbert space through a kernel mapping operator φ; Through a projection matrix φ(C)∈R ∞×d Transform the source sample from the Hilbert space back to the original feature space, G is expressed as G(·) = φ(C) T The form of φ(·); The objective function in formula (8) is rewritten as: Minimize the maximum mean difference distance MMD of the objective function in the Hilbert space; treat MMD as a regular term f G (X s ,X t ): Where H represents the Hilbert space, 1 s and 1 t They are of length n s and n t A column vector whose elements are all 1; The MMD in formula (10) is transformed into the following form, as f G (X s ,X t ): Replace f in formula (11) G (X s ,X t ) into formula (9), the objective function becomes: Formula (12) is converted into a solvable form by calculating the kernel function instead of the inner product operation in the kernel space. Including: Let φ(C)=[φ(X s ),φ(X t )]P, where the linear coefficient matrix Then formula (12) is rewritten as the following form as the final objective function: in The calculation formula of the four kernel matrices is K ss =φ(X s ) T φ(X s ), K st =φ(X s ) T φ(X t ), K ts =φ(X t ) T φ(X s ) and K tt =φ(X t ) T φ(X t ); In formula (13), an L1 norm of P is added as a constraint term of the objective function, that is, where p i is the i-th column of P, and the sparsity of P is adjusted by the weighting coefficient μ.

9. The micro-expression recognition method according to claim 1, characterized in that: Step S5 includes: For the training samples in the source micro-expression dataset, the optimal parameter value is learned through iteration and Finally, the group sparse model is used as a classifier to predict the label vector of the test samples in the target dataset, that is, to identify the types of micro-expressions; The feature vector of the test sample is x t ∈R 72×1 , predict the label vector l of the test sample t : in Obtained through calculation of the kernel function selected when learning the group sparse model; is the label vector, then the micro-expression type of the test sample is in express The kth element of .

Citation Information

Patent Citations

  • Micro-expression identification method based on informative and representative active learning

    CN108830222A