Contrastive micro-expression recognition method based on text position attention
By introducing text position attention and contrast loss in micro-expression recognition technology, combined with optical flow characteristics, the problem of underutilizing text information and facial motion unit feature weight allocation in the prior art is solved, and more efficient micro-expression recognition is achieved.
Patent Information
- Application Number
- CN202411124661.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-08-14
AI Technical Summary
The existing micro-expression recognition technology mainly relies on optical flow characteristics, does not fully utilize text information, and does not consider the weight allocation problem of facial motor unit features during the fusion process.
A contrasting micro-expression recognition method based on text position attention is adopted. By extracting optical flow position attention, face light flow characteristics and text position characteristics, the micro-expression comparison model is trained, and the positive and negative example contrast loss of text position attention is analyzed to effectively distinguish micro-expression categories.
By combining optical flow attention and text position characteristics, mutual confusion in micro-expression recognition is effectively avoided, and the accuracy and efficiency of micro-expression recognition are improved.
Smart Images

Figure CN119028001B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of information technology and artificial intelligence, and in particular to a comparative micro-expression recognition method based on text position attention. Background Art
[0002] Micro-expressions refer to brief and subtle changes in human facial expressions, which usually last from 1 / 25 to 1 / 5 second. Due to their brevity and subtlety, micro-expressions are usually difficult to detect with the naked eye. However, micro-expressions have important application value in psychology, crime investigation, psychotherapy and other fields. For example, in interpersonal communication, micro-expressions can provide key clues to the emotional changes of the other party. In recent years, with the development of computer vision and deep learning technology, automated micro-expression recognition methods have gradually become a research hotspot. These deep learning-based methods can accurately capture and analyze subtle changes in facial expressions, thereby improving the accuracy and efficiency of recognition.
[0003] China's public patent number CN 111461021 B "A method for detecting micro-expressions based on optical flow" includes the following steps: using the open source toolkit dlib to perform face recognition on each frame of a video sample, and marking the region of interest ROI and the non-deformable region ROI'; calculating the dense optical flow of two adjacent frames in the video sample, extracting the optical flow in the region of interest and the optical flow in the non-deformable region, and performing subtraction processing on the two to remove the influence of head shaking; defining the angle area under polar coordinates, and calculating the main optical flow of the region of interest ROI of each frame of the video sample in the defined polar coordinate angle area, and sequentially representing the main optical flow trajectory of all frames in the video sample by amplitude and angle; capturing the micro-expression occurrence frame in the video sample according to the amplitude and angle trajectory, and making micro-expression annotations. Thus, the movement of micro-expressions in each region of interest ROI is displayed in real time.
[0004] China's public patent number CN 118366202 A "A method for micro-expression recognition based on Transformer motion feature fusion" includes the following steps: A. Preprocessing micro-expression videos; B. Dividing the obtained horizontal optical flow and vertical optical flow feature maps into test sets and training sets; C. Constructing a Transformer-based motion feature fusion recognition model, including a shared Transformer feature encoder, a feature selection fusion module, and a global cross-attention module; D. Constructing spatial consistency constraint loss, cross entropy loss, and contrast loss, training a Transformer-based motion feature fusion recognition model to obtain a Transformer recognition model with strong discriminative ability; E. Classification and recognition, classifying and recognizing the test set according to the trained Transformer recognition model. Effectively obtain and integrate more discriminative micro-expression facial motion unit change features.
[0005] China's patent publication number CN 113723287 B "Micro-expression recognition method, device and medium based on bidirectional recurrent neural network" includes the following steps: pre-processing the original micro-expression data to obtain a facial behavior coding sequence; performing facial behavior coding on the facial behavior coding sequence to obtain a micro-expression coding feature vector; inputting the micro-expression coding feature vector into a trained bidirectional recurrent neural network; extracting micro-expression temporal features from the feature vector output by the bidirectional recurrent neural network based on a temporal attention mechanism; using the micro-expression temporal features to identify the expression category of the micro-expression through a Softmax function. Fully identify the features in micro-expressions and improve the recognition accuracy of micro-expressions.
[0006] Existing patents have some shortcomings when performing micro-expression recognition: first, they mainly rely on optical flow features for micro-expression detection, but do not fully utilize information from other modalities such as text; second, although more discriminative micro-expression facial motion unit change features are obtained and fused, the weight distribution problem of different facial motion unit features in the fusion process is not fully considered; third, micro-expression features are extracted through facial behavior coding sequences, but the use of optical flow features to capture and combine the motion information of facial micro-expressions is not considered. Summary of the invention
[0007] The present invention aims to remedy the deficiencies of the prior art and provide a comparative micro-expression recognition method based on text position attention.
[0008] The present invention is achieved through the following technical solutions:
[0009] A comparative micro-expression recognition method based on text position attention specifically comprises the following steps:
[0010] S1: extract optical flow position attention;
[0011] S2: Extract facial optical flow features;
[0012] S3: extract text position features;
[0013] S4: training micro-expression comparison model;
[0014] S5: test model;
[0015] The step S1 of extracting the optical flow position attention specifically includes the following steps:
[0016] S1-1: Input micro-expression video dataset;
[0017] S1-2: Calculate the amplitude of micro-expression movements;
[0018] S1-3: Get {framei The first frame frame1 in} is used as the starting frame frame of the micro-expression movement begin ∈R 3 ×H×W , get the optical flow sampling image frame begin 、frame peak ;
[0019] S1-4: Use optical flow to sample image frame begin 、frame peak Extracting micro-expression optical flow features[u e ,v e ], where u e Represents frame begin The horizontal component of the optical flow feature at the pixel point (a, b), v e Represents frame begin The vertical component of the optical flow feature at the middle pixel (a, b);
[0020] S1-5: Extracting micro-expression optical flow change rate features OF e ;
[0021] S1-6: Extracting optical flow position attention A o .
[0022] The input micro-expression video dataset described in step S1-1 specifically includes the following steps:
[0023] S1-1-1: Micro-expression video dataset Data =<X,Text,Y,> , where X represents the micro-expression video set, Text represents the micro-expression category text set; Y represents the label set of the micro-expression category;
[0024] S1-1-2: x∈X, x represents a micro-expression video in the dataset, x={frame i}, frame i represents the i-th frame of the video, i=1,…,L, L represents the total number of frames of the video;
[0025] S1-1-3: Text = {happy, disgusted, fearful, angry, sad, surprised, other expressions};
[0026] S1-1-4: y∈Y, y represents the category label of a micro-expression video. The value of y is {1,…,7}, which corresponds to the text one by one. 1 represents happiness, 2 represents disgust, 3 represents fear, 4 represents anger, 5 represents sadness, 6 represents surprise, and 7 represents other expressions.
[0027] The calculation of the micro-expression movement amplitude described in step S1-2 specifically includes the following steps:
[0028] S1-2-1: Calculate optical flow features OF;
[0029] S1-2-2: Calculate the movement distance D;
[0030] S1-2-3: Calculate D i The number of moving pixels num greater than the threshold 0.2 i ;
[0031] S1-2-4: For movement distance D = {D i Each of them is represented by the optical flow feature OF i The generated movement distance D i , call step S1-2-3, get the number of moving pixels num = {num i}, i = 1, 2, ..., L-1;
[0032] S1-2-5: Get the frame with the maximum amplitude of micro-expression peak ∈R 3×H×W .
[0033] The calculation of the optical flow feature OF described in step S1-2-1 specifically includes the following steps:
[0034] S1-2-1-1: Given a video frame i A pixel point (a, b) in the video frame is calculated i and frame i+1 The pixel optical flow features OF i (a,b), is:
[0035] OF i (a,b)=(u i (a,b),v i (a,b))
[0036] Where a=1,2,…,W,b=1,2,…,H,i=1,2,…,L-1,W represents the width of the image, H represents the height of the image, L represents the total number of frames of the video, u i (a,b) represents frame i The optical flow feature OF at the pixel point (a, b) i The horizontal component of (a,b), v i (a,b) represents frame i The optical flow feature OF at the pixel point (a, b) i The vertical component of (a, b);
[0037] S1-2-1-2: video frame iFor each pixel point (a, b) in the image, the optical flow feature OF is calculated. i =OF i (a,b)}, a=1,2,…,W, b=1,2,…,H, i=1,2,…,L-1,OF i ∈R 2×H×W ;
[0038] S1-2-1-3: i} each of which consists of a video frame frame i and frame i+1 Generated optical flow features OF i , call S1-2-1-2, calculate the optical flow feature OF = {OF i};
[0039] The calculation of the movement distance D described in step S1-2-2 specifically includes the following steps:
[0040] S1-2-2-1: Given the optical flow feature OF i A pixel point (a, b) in , calculate its pixel point movement distance D i (a,b), is:
[0041]
[0042] S1-2-2-2: Optical flow features OF i For each pixel point (a, b), call S1-2-2-1 to calculate the movement distance D i ={D i (a,b)}, a=1,2,…,W, b=1,2,…,H, i=1,2,…,L-1;
[0043] S1-2-2-3: Yes {D i Each of them is composed of optical flow features OF i The generated movement distance D i , call S1-2-2-2, calculate the movement distance D = {D i}, i = 1, 2, ..., L-1;
[0044] The calculation D described in step S1-2-3 i The number of moving pixels num greater than the threshold 0.2 i , specifically including the following steps:
[0045] S1-2-3-1: Set the number of moving pixels num i The initial value is set to 0;
[0046] S1-2-3-2: If D iThe movement distance D at the pixel point (a, b) i (a,b) is greater than the threshold 0.2, the number of moving pixels num i Add 1, otherwise the number of moving pixels num i The value remains unchanged;
[0047] S1-2-3-3: Move the pixel point to a distance D i For all the pixels (a, b) in , a=1,2,…,W, b=1,2,…,H, call S1-2-3-2 to calculate num i ;
[0048] Step S1-2-5 of obtaining the micro-expression maximum amplitude frame peak ∈R 3×H×W , specifically including the following steps:
[0049] S1-2-5-1: Get the number of moving pixels num = {num i}, the largest num i , use the value of i+1 as the subscript index peak;
[0050] S1-2-5-2: Get {frame i The maximum amplitude frame of micro-expression in} peak .
[0051] The feature of the optical flow change rate of micro-expression extracted in step S1-5 e , specifically including the following steps:
[0052] S1-5-1: Using micro-expression optical flow features [u e ,v e ], calculate the optical flow change rate z e , the calculation formula is as follows:
[0053]
[0054] in, is the micro-expression optical flow feature [u e ,v e ];
[0055] S1-5-2: Generate optical flow change rate features, OF e =[u e ,v e ,ze],OF e ∈R 3×H×W ;
[0056] Step S1-6 extracts the optical flow position attention A o, specifically including the following steps:
[0057] S1-6-1: Optical flow change rate feature OF e , input into a convolutional layer, and get the feature F e ∈R 3×h×w , where the convolution kernel size is [k,k] and the step size is k.
[0058] S1-6-2: For feature F e ∈R 3×h×w , perform maximum pooling in the first dimension to obtain the optical flow position attention A o ∈R 1×h×w .
[0059] The extraction of facial optical flow features described in step S2 specifically includes the following steps:
[0060] S2-1: Get the frame with the maximum micro-expression movement amplitude in S1-2-5 peak ∈R 3×H×W ;
[0061] S2-2: The frame with the largest micro-expression movement amplitude peak ∈R 3×H×W , divided into image blocks [k,k] with a height of k and a width of k, and the local image block frame is obtained peak = {LI s}, LI s ∈R 3×k×k , s represents the subscript index of the local block, s = 1, ..., hw,
[0062] S2-3: Extracting facial visual features F V ;
[0063] S2-4: The facial visual features F V With optical flow position attention A o Point product to generate the face optical flow feature F VO ∈R D ×h×w ,for:
[0064] F VO =F V ⊙A o
[0065] Here, ⊙ represents the dot product.
[0066] Step S2-3 described in extracting facial visual features F V , specifically including the following steps:
[0067] S2-3-1: Given a local image block LIs ∈R 3×k×k , flatten the sth block to obtain a local block vector
[0068] S2-3-2: For each local image block LI s , call S2-3-1 to get the local block vector {LV s}, s=1,…,hw;
[0069] S2-3-3: The local block vector {LV s}, merge along the first dimension to generate the feature with the maximum amplitude of micro-expression movement
[0070] S2-3-4: The feature F with the maximum micro-expression movement amplitude peak , input a linear layer, and obtain the features Where D is the embedding dimension;
[0071] S2-3-5: In the characteristics On the top, add position feature F p ∈R hw×D , get the features
[0072] S2-3-6: In the characteristics Add a token vector V along the first dimension token ∈R 1×D , get the features
[0073] S2-3-7: Features Input into the Vision Transformer with CLIP pre-trained weights to obtain the Transformer feature F T ∈R (hw+1)×D ;
[0074] S2-3-8: Along the Transformer feature F T The first dimension of F is to take the first hw features and obtain F T ′∈R hw×D ;
[0075] S2-3-9: For F T ′∈R hw×D , perform matrix dimension flipping to obtain F T ″∈R D×hw ;
[0076] S2-3-10: For F T ″∈R D×hw, perform matrix dimension transformation to obtain the face visual feature F V ∈R D×h×w .
[0077] The extraction of text position features described in step S3 specifically includes the following steps:
[0078] S3-1: Get the micro-expression category text set Text;
[0079] S3-2: Generate text prompts f ;
[0080] S3-3: Text prompt t f Input a fully connected layer to generate text location prompt t p ∈R 7×hw ;
[0081] S3-4: Prompt text position p , perform matrix dimension transformation to generate text position attention A T ∈R 7×h×w ;
[0082] S3-5: Extract text position feature F TP .
[0083] Step S3-2 generates a text prompt t f , specifically including the following steps:
[0084] S3-2-1: Input the micro-expression category text set Text, and tokenize it with CLIP pre-trained weights to get the word tag w t ∈R 7×n , 7 is the number of micro-expression categories, and n is the number of tokens;
[0085] S3-2-2: Calculate text tag weight t t ;
[0086] S3-2-3: Input word mark w t , into encode_taken with CLIP pre-trained weights, and get the word embedding w e ∈R 7×n×D , D is the embedding dimension;
[0087] S3-2-4: Calculate text embedding weight t e ,
[0088] S3-2-5: Input text label weight t t and text embedding weight t e , get the text prompt t from encode_text with CLIP pre-trained weightsf ∈R 7×D ;
[0089] The calculation text tag weight t described in step S3-2-2 t , specifically including the following steps:
[0090] S3-2-2-1: Initialize text tag weight t t ∈R 7×n , the initial value of the weight is all zero;
[0091] S3-2-2-2: Mark the word w t ∈R 7×n , extract w t The subscript ind of the maximum value in the j-th row, j = 1, ..., 7, ind = 1, ..., n;
[0092] S3-2-2-3: Set p r To indicate the prefix subscript, p r =10;
[0093] S3-2-2-4: Set p o To indicate the suffix subscript, p o =10;
[0094] S3-2-2-5: Mark the word w t ∈R 7×n , extract the value of the jth row and the indth column, and assign it to the text tag weight t t ∈R 7×n In the jth row, the (p r +p o +ind) column, is:
[0095] t t [j,p r +p o +ind]=w t [j,ind]
[0096] S3-2-2-6: For word tag w t ∈R 7×n For each line in , call steps S3-2-2-2, S3-2-2-3, S3-2-2-4, S3-2-2-5 to get the text tag weight t t ∈R 7×n ;
[0097] The calculated text embedding weight t described in step S3-2-4 e , specifically including the following steps:
[0098] S3-2-4-1: Initialize text embedding weight t e∈R 7×n×D , the weight is initialized to a random value that follows a standard normal distribution;
[0099] S3-2-4-2: Call step S3-2-2-2 to obtain word tag w t The subscript ind of the maximum value in the j-th row, j = 1, ..., 7, ind = 1, ..., n;
[0100] S3-2-4-3: embedding word w e ∈R 7×n×D , extract the vector of the jth row and the first column, and assign it to the text embedding weight t e ∈R 7×n×D In , the vector in the jth row and the first column is:
[0101] t e [j,1,1:D]=w e [j,1,1:D]
[0102] S3-2-4-4: embedding word w e , extract the vector from row j, column 1 to row j, column ind, and assign it to the text embedding weight t e In the jth row, the (p r +1) column to the jth row (p r +ind) columns, which is:
[0103] t e [j,p r +1:p r +ind,1:D]=w e [j,1:ind,1:D]
[0104] S3-2-4-5: embedding word w e , extract the vector of the jth row and indth column and assign it to the text embedding weight t e In the jth row, the (p r +p o +ind) columns, which is:
[0105] t e [j,p r +p o +ind,1:D]=w e [j,ind,1:D]
[0106] S3-2-4-6: For word tag w t ∈R 7×n For each row in , call steps S3-2-4-2, S3-2-4-3, S3-2-4-4, S3-2-4-5 to get the text embedding weight t e ∈R7×n×D ;
[0107] The extracted text position feature F in step S3-5 TP , specifically including the following steps:
[0108] S3-5-1: Focus on the text position A T ∈R 7×h×w , divided into text position local attention
[0109] A T ={LAT j}, LAT j ∈R h×w , j represents the subscript index of the micro-expression category, j = 1,…,7;
[0110] S3-5-2: Local attention LAT given text position j , the face optical flow feature F obtained by step S2-4 VO ∈R D×h×w , with LAT j Perform dot multiplication to obtain the local feature LFT of the text position j ∈R D×h×w ;
[0111] LFT j =F VO ⊙LAT j
[0112] Among them, ⊙ represents the dot product;
[0113] S3-5-3: Local attention LAT for each text position j , call step S3-5-2, and obtain the text position local feature LFT = {LFT j}, LFT j ∈R D×h×w , j = 1,…,7;
[0114] S3-5-4: Given a text position local feature LFT j , perform dimension expansion and obtain LFT j ∈R 1×D×h×w ;
[0115] S3-5-5: For each text position local feature LFT j , call step S3-5-4, get {LFT j},
[0116] LFT j ∈R 1×D×h×w , j = 1,…,7;
[0117] S3-5-6: Local feature of text position {LFT j}, merge along the first dimension to obtain the text position feature F TP ∈R 7×D×h×w
[0118] S3-5-7: Text position feature F TP , input a fully connected layer to transform the dimension to F TP ∈R D×h×w .
[0119] The training micro-expression comparison model described in step S4 specifically includes the following steps:
[0120] S4-1: Calculate micro-expression contrast loss co n;
[0121] S4-2: Calculate normal / abnormal position loss ab ;
[0122] S4-3: Calculate micro-expression prediction score S pre ;
[0123] S4-4: Calculate micro-expression feature loss fea ;
[0124] S4-5: Calculate the total loss of the model;
[0125] S4-6: Use the total loss of the model to train and obtain the optimal model parameters;
[0126] Calculate the micro-expression contrast loss loss described in step S4-1 con , specifically including the following steps:
[0127] S4-1-1: Text prompt t f ∈R 7×D , divided into text prompt vector t f ={tv j}, tv j ∈R 1×D , j represents the subscript index of the micro-expression category, j = 1,…,7;
[0128] S4-1-2: Calculate micro-expression contrast loss con ,for:
[0129]
[0130] Calculation of normal / abnormal position loss loss in step S4-2 ab , specifically including the following steps:
[0131] S4-2-1: Attention to the text position A T ∈R 7×h×w , transform the matrix dimension to obtain the text position attention matrix A TM ∈R 7×hw ;
[0132] S4-2-2: Text position attention matrix A TM , input a fully connected layer to get the text position attention vector A TV ∈R 1×hw ;
[0133] S4-2-3: Calculate the average value S of normal / abnormal micro-expression scores ab ;
[0134] S4-2-4: Obtain the micro-expression category label y corresponding to the micro-expression video x, and calculate the binary label y ab ,for:
[0135]
[0136] S4-2-5: Calculate micro-expression position loss ab ,for:
[0137] loss ab =-[y ab log(S ab )+(1-y ab )log(1-S ab )]
[0138] The calculation of the micro-expression prediction score S in step S4-3 pre , specifically including the following steps:
[0139] S4-3-1: The text position feature F obtained by S3-5 TP ∈R D×h×w , transform the matrix dimension to F TP ′∈R D×hw ;
[0140] S4-3-2: Text position feature F TP ′ Input a fully connected layer to obtain the micro-expression prediction feature F pre ∈R 7×hw , 7 is the number of micro-expression categories;
[0141] S4-3-3: Micro-expression prediction feature F pre , divided into micro-expression prediction local features F pre ={LFpre j}, LFpre j ∈R 1×hw, j represents the subscript index of the micro-expression category, j = 1,…,7;
[0142] S4-3-4: Given micro-expression prediction local features LFpre j ∈R 1×hw , take LFpre j The first K blocks of the numerical size in the , and merge the first K blocks to generate the micro-expression prediction K value feature LFKpre j ∈R 1×K ;
[0143] S4-3-5: Predict local features LFpre for each micro-expression j , call step S4-3-4 to generate micro-expression prediction K value features {LFKpre j}, LFKpre j ∈R 1×K , j = 1,…,7;
[0144] S4-3-6: Predicting micro-expression K value features {LFKpre j}, merge along the first dimension to generate the micro-expression prediction score feature F s , F s ∈R 7×K ;
[0145] S4-3-7: Micro-expression prediction score feature F s , average pooling is performed along the second dimension, and the softmax function is used for normalization to obtain the micro-expression prediction score S pre ∈R 7×1 ;
[0146] Calculate the micro-expression feature loss loss described in step S4-4 fea , specifically including the following steps:
[0147] S4-4-1: Convert the micro-expression category label y corresponding to the micro-expression video x into a one-hot vector y h ;
[0148] S4-4-2: Score the micro-expression prediction pre ∈P 7×1 , traverse each category one by one, and get S pre = {Spre j}, Spre j ∈R 1×1 , j represents the subscript index of the micro-expression category, j = 1,…,7;
[0149] S4-4-3: Calculate micro-expression feature loss fea ,for:
[0150]
[0151] The micro-expression category label y corresponding to the micro-expression video x described in step S4-4-1 is converted into a one-hot vector y h , specifically including the following steps:
[0152] S4-4-1-1: Given a yh j ∈R 1×1 , j represents the subscript index of the micro-expression category, j = 1, ..., 7, and its onehot value is calculated as
[0153]
[0154] S4-4-1-2: Yes j Each yh in j , j = 1, ..., 7, call step S4-4-1-1, calculate its onehot value, and obtain the onehot vector y h ={yh j},y h ∈R 7×1 , yh j ∈R 1×1 .
[0155] The test model described in step S5 specifically includes the following steps:
[0156] S5-1: Input micro-expression video frame;
[0157] S5-2: According to the optimal parameters of the model, call step S1 to extract the optical flow position attention A o ;
[0158] S5-3: Based on the optimal model parameters, call step S2 to extract the face optical flow feature F VO ;
[0159] S5-4: Based on the optimal model parameters, call step S3 to extract the text position feature F TP ;
[0160] S5-5: Based on the optimal model parameters, call step S4-3 to calculate the micro-expression prediction score S pre ;
[0161] S5-6: Identify micro-expression categories based on micro-expression prediction scores.
[0162] The advantages of the present invention are: the present invention uses video optical flow features to calculate optical flow attention, which is used to analyze the visual features of Vision Transformer and extract micro-expression visual features of optical flow attention; for the difficulty of finding the position of micro-expressions, the token of the micro-expression text is extracted, the text position attention is learned, and the micro-expression features of the text position attention are extracted. In order to effectively distinguish the significant positions of different micro-expressions, the positive and negative example comparison loss of the text position attention is analyzed to effectively separate the easily confused micro-expression categories. The present invention effectively avoids mutual confusion in micro-expression recognition by considering the optical flow attention features, text position attention, and text position comparison relationship. BRIEF DESCRIPTION OF THE DRAWINGS
[0163] Figure 1 It is a flow chart of the comparative micro-expression recognition method based on text position attention of the present invention;
[0164] Figure 2 The optical flow position attention process of the present invention is extracted;
[0165] Figure 3 The process of extracting the optical flow features of the face of the present invention;
[0166] Figure 4 The process of extracting text position features of the present invention;
[0167] Figure 5 The process of training the micro-expression comparison model of the present invention;
[0168] Figure 6 This is the test model process of the present invention. DETAILED DESCRIPTION
[0169] In order to make the purpose, technical solution and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments. The present invention is a method for comparing micro-expression recognition based on text position attention. The specific process is as follows Figure 1 As shown, the implementation scheme of the present invention is divided into the following steps:
[0170] S1: Extract optical flow position attention, such as Figure 2 As shown;
[0171] S1-1: Input micro-expression video dataset. The specific steps are as follows:
[0172] S1-1-1: Micro-expression video dataset Data =<X,Text,Y,> , where X represents the micro-expression video set, Text represents the micro-expression category text set; Y represents the label set of the micro-expression category;
[0173] S1-1-2: x∈X, x represents a micro-expression video in the dataset, x={frame i}, frame i represents the i-th frame of the video, i=1,…,L, L represents the total number of frames of the video;
[0174] S1-1-3: Text = {happy, disgusted, fearful, angry, sad, surprised, other expressions};
[0175] S1-1-4: y∈Y, y represents the category label of a micro-expression video, and the value of y is {1,…,7}, which corresponds to the text one by one (1 represents happiness, 2 represents disgust, 3 represents fear, 4 represents anger, 5 represents sadness, 6 represents surprise, and 7 represents other expressions);
[0176] S1-2: Calculate the micro-expression movement amplitude. The specific steps are as follows:
[0177] S1-2-1: Calculate the optical flow feature OF. The specific steps are as follows:
[0178] S1-2-1-1: Given a video frame i A pixel point (a, b) in the video frame is calculated i and frame i+1 The pixel optical flow features OF i (a,b), is:
[0179] OF i (a,b)=(u i (a,b),v i (a,b))
[0180] Where a=1,2,…,W,b=1,2,…,H,i=1,2,…,L-1,W represents the width of the image, H represents the height of the image, L represents the total number of frames of the video, u i (a,b) represents frame i At the pixel point (a, b), the optical flow feature OF i The horizontal component of (a,b), v i (a,b) represents frame i At the pixel point (a, b), the optical flow feature OF i The vertical component of (a, b);
[0181] S1-2-1-2: video frame i For each pixel point (a, b) in , call S1-2-1-1 to calculate the optical flow feature OF i =OF i(a,b)}, a=1,2,…,W, b=1,2,…,H, i=1,2,…,L-1,OF i ∈R 2×H×W ;
[0182] S1-2-1-3: i}, each of which is composed of a video frame i and frame i+1 Generated optical flow features OF i , call S1-2-1-2, calculate the optical flow feature OF = {OF i};
[0183] S1-2-2: Calculate the movement distance D. The specific steps are as follows:
[0184] S1-2-2-1: Given the optical flow feature OF i A pixel point (a, b) in , calculate its pixel point movement distance D i (a,b), is:
[0185]
[0186] S1-2-2-2: Optical flow features OF i For each pixel point (a, b), call S1-2-2-1 to calculate the movement distance D i ={D i (a,b)}, a=1,2,…,W, b=1,2,…,H, i=1,2,…,L-1;
[0187] S1-2-2-3: Yes {D i}, each of which is represented by the optical flow feature OF i The generated movement distance D i , call S1-2-2-2, calculate the movement distance D = {D i}, i = 1, 2, ..., L-1;
[0188] S1-2-3: Calculate D i The number of moving pixels num greater than the threshold 0.2 i , the specific steps are as follows:
[0189] S1-2-3-1: Set num i The initial value is set to 0;
[0190] S1-2-3-2: If D i The movement distance D at the pixel point (a, b) i (a,b) is greater than the threshold 0.2, num i Add 1, otherwise numi The value remains unchanged;
[0191] S1-2-3-3: Move the pixel point to a distance D i For all the pixels (a, b) in , a=1,2,…,W, b=1,2,…,H, call S1-2-3-2 to calculate num i ;
[0192] S1-2-4: For movement distance D = {D i Each of} is represented by the optical flow feature OF i The generated movement distance D i , call S1-2-3, get the number of moving pixels num = {num i}, i = 1, 2, ..., L-1;
[0193] S1-2-5: Get the frame with the maximum amplitude of micro-expression peak ∈R 3×H×W , the specific steps are as follows:
[0194] S1-2-5-1: Get the number of moving pixels num = {num i}, the largest num i , use the value of i+1 as the subscript index peak;
[0195] S1-2-5-2: Get {frame i The maximum amplitude frame of micro-expression in} peak ;
[0196] S1-3: Get {frame i The first frame frame1 in} is used as the starting frame frame of the micro-expression movement begin ∈R 3 ×H×W , get the optical flow sampling image frame begin 、frame peak ;
[0197] S1-4: Use optical flow to sample image frame begin 、frame peak Call S1-2-1-1, S1-2-1-2 to extract micro-expression optical flow features [u e ,v e ], where u e Represents frame begin At the middle pixel (a, b), the horizontal component of the optical flow feature, v e Represents frame begin The vertical component of the optical flow feature at the middle pixel (a, b);
[0198] S1-5: Extracting micro-expression optical flow change rate features OF e , the specific steps are as follows:
[0199] S1-5-1: Using micro-expression optical flow features [u e ,v e ], calculate the optical flow change rate z e , the calculation formula is as follows:
[0200]
[0201] in, is the micro-expression optical flow feature [u e ,v e ];
[0202] S1-5-2: Generate optical flow change rate features, OF e =[u e ,v e ,ze],OF e ∈R 3×H×W ;
[0203] S1-6: Extracting optical flow position attention A o , the specific steps are as follows:
[0204] S1-6-1: Optical flow change rate feature OF e , input into a convolutional layer, and get the feature F e ∈R 3×h×w , where the convolution kernel size is [k,k] and the step size is k.
[0205] S1-6-2: For feature F e ∈R 3×h×w , perform maximum pooling in the first dimension to obtain the optical flow position attention A o ∈R 1×h×w ;
[0206] S2: Extract facial optical flow features, such as Figure 3 As shown;
[0207] S2-1: Get the frame with the maximum micro-expression movement amplitude in S1-2-5 peak ∈R 3×H×W ;
[0208] S2-2: The frame with the largest micro-expression movement amplitude peak ∈R 3×H×W , divided into image blocks [k,k] with a height of k and a width of k, and the local image block frame is obtained peak = {LIs}, LI s ∈R 3×k×k , s represents the subscript index of the local block, s = 1, ..., hw,
[0209] S2-3: Extracting facial visual features F V , the specific steps are as follows:
[0210] S2-3-1: Given a local image block LI s ∈R 3×k×k , flatten the sth block to obtain a local block vector
[0211] S2-3-2: For each local image block LI s , call S2-3-1 to get the local block vector {LV s}, s=1,…,hw;
[0212] S2-3-3: The local block vector {LV s}, merge along the first dimension to generate the feature with the maximum amplitude of micro-expression movement
[0213] S2-3-4: The feature F with the maximum micro-expression movement amplitude peak , input a linear layer, and obtain the features Where D is the embedding dimension;
[0214] S2-3-5: In the characteristics On the top, add position feature F p ∈R hw×D , get the features
[0215] S2-3-6: In the characteristics Add a token vector V along the first dimension token ∈R 1×D , get the features
[0216] S2-3-7: Features Input into the Vision Transformer with CLIP pre-trained weights to obtain the Transformer feature F T ∈R (hw+1)×D ;
[0217] S2-3-8: Along the Transformer feature F T The first dimension of F is to take the first hw features and obtain F T ′∈Rhw×D ;
[0218] S2-3-9: For F T ′∈R hw×D , perform matrix dimension flipping to obtain F T ″∈R D×hw ;
[0219] S2-3-10: For F T ″∈R D×hw , perform matrix dimension transformation to obtain the face visual feature F V ∈R D×h×w ;
[0220] S2-4: The facial visual features F V With optical flow position attention A o Point product to generate the face optical flow feature F VO ∈R D ×h×w ,for:
[0221] F VO =F V ⊙A o
[0222] Among them, ⊙ represents the dot product;
[0223] S3: Extract text position features, such as Figure 4 As shown:
[0224] S3-1: Get the micro-expression category text set Text;
[0225] S3-2: Generate text prompts f , the specific steps are as follows:
[0226] S3-2-1: Input the micro-expression category text set Text, and tokenize it with CLIP pre-trained weights to get the word tag w t ∈R 7×n , 7 is the number of micro-expression categories, and n is the number of tokens;
[0227] S3-2-2: Calculate text tag weight t t , the specific steps are as follows:
[0228] S3-2-2-1: Initialize text tag weight t t ∈R 7×n , the initial value of the weight is all zero;
[0229] S3-2-2-2: Mark the word w t ∈R 7×n , extract w tThe subscript ind of the maximum value in the j-th row, j = 1, ..., 7, ind = 1, ..., n;
[0230] S3-2-2-3: Set p r To indicate the prefix subscript, p r =10;
[0231] S3-2-2-4: Set p o To indicate the suffix subscript, p o =10;
[0232] S3-2-2-5: Mark the word w t ∈R 7×n , extract the value of the jth row and the indth column, and assign it to the text tag weight t t ∈R 7×n In the jth row, the (p r +p o +ind) column, is:
[0233] t t [j,p r +p o +ind]=w t [j,ind]
[0234] S3-2-2-6: For word tag w t ∈R 7×n For each line in, call S3-2-2-2, S3-2-2-3, S3-2-2-4, S3-2-2-5 to get the text tag weight t t ∈R 7×n ;
[0235] S3-2-3: Input word mark w t , into encode_taken with CLIP pre-trained weights, and get the word embedding w e ∈R 7×n×D , D is the embedding dimension;
[0236] S3-2-4: Calculate text embedding weight t e , the specific steps are as follows:
[0237] S3-2-4-1: Initialize text embedding weight t e ∈R 7×n×D , the weight is initialized to a random value that follows a standard normal distribution;
[0238] S3-2-4-2: Call S3-2-2-2 to get the word token w t The subscript ind of the maximum value in the j-th row, j = 1, ..., 7, ind = 1, ..., n;
[0239] S3-2-4-3: embedding word w e ∈R 7×n×D , extract the vector of the jth row and the first column, and assign it to the text embedding weight t e ∈R 7×n×D In , the vector in the jth row and the first column is:
[0240] t e [j,1,1:D]=w e [j,1,1:D]
[0241] S3-2-4-4: embedding word w e , extract the vector from row j, column 1 to row j, column ind, and assign it to the text embedding weight t e In the jth row, the (p r +1) column to the jth row (p r +ind) columns, which is:
[0242] t e [j,p r +1:p r +ind,1:D]=W e [j,1:ind,1:D]
[0243] S3-2-4-5: embedding word w e , extract the vector of the jth row and indth column and assign it to the text embedding weight t e In the jth row, the (p r +p o +ind) columns, which is:
[0244] t e [j,p r +p o +ind,1:D]=w e [j,ind,1:D]
[0245] S3-2-4-6: For word tag w t ∈R 7×n For each line in , call S3-2-4-2, S3-2-4-3, S3-2-4-4, S3-2-4-5 to get the text embedding weight t e ∈R 7×n×D ;
[0246] S3-2-5: Input text label weight t t and text embedding weight t e , get the text prompt t from encode_text with CLIP pre-trained weights f ∈R 7×D ;
[0247] S3-3: Text prompt t f Input a fully connected layer to generate text location prompt t p ∈R 7×hw ;
[0248] S3-4: Prompt text position p , perform matrix dimension transformation to generate text position attention A T ∈R 7×h×w ;
[0249] S3-5: Extract text position feature F TP , the specific steps are as follows:
[0250] S3-5-1: Focus on the text position A T ∈R 7×h×w , divided into text position local attention A T ={LAT j}, LAT j ∈R h×w , j represents the subscript index of the micro-expression category, j = 1,…,7;
[0251] S3-5-2: Local attention LAT given text position j , the face optical flow feature F obtained by S2-4 VO ∈R D ×h×w , with LAT j Perform dot multiplication to obtain the local feature LFT of the text position j ∈R D×h×w ;
[0252] LFT j =F VO ⊙LAT j
[0253] Among them, ⊙ represents the dot product;
[0254] S3-5-3: Local attention LAT for each text position j , call S3-5-2, get the text position local feature LFT = {LFT j}, LFT j ∈R D×h×w , j = 1,…,7;
[0255] S3-5-4: Given a text position local feature LFT j , perform dimension expansion and obtain LFT j ∈R 1×D×h×w ;
[0256] S3-5-5: For each text position local feature LFT j , call S3-5-4, get {LFT j}, LFT j ∈R 1 ×D×j×w , j = 1,…,7;
[0257] S3-5-6: Local feature of text position {LFT j}, merge along the first dimension to obtain the text position feature F TP ∈R 7×D×h×w ;
[0258] S3-5-7: Text position feature F TP , input a fully connected layer to transform the dimension to F TP ∈R D×h×w ;
[0259] S4: Train a micro-expression comparison model, such as Figure 5 As shown:
[0260] S4-1: Calculate micro-expression contrast loss con , the specific steps are as follows:
[0261] S4-1-1: Text prompt t f ∈R 7×D , divided into text prompt vector t f ={tv j}, tv j ∈R 1×D , j represents the subscript index of the micro-expression category, j = 1,…,7;
[0262] S4-1-2: Calculate micro-expression contrast loss con ,for:
[0263]
[0264] S4-2: Calculate normal / abnormal position loss ab , the specific steps are as follows:
[0265] S4-2-1: Attention to the text position A T ∈R 7×j×w , transform the matrix dimension to obtain the text position attention matrix A TM ∈R 7×hw ;
[0266] S4-2-2: Text position attention matrix A TM , input a fully connected layer to get the text position attention vector A TV ∈R1×hw ;
[0267] S4-2-3: Calculate the average value S of normal / abnormal micro-expression scores ab , the specific steps are as follows:
[0268] S4-2-3-1: Text position attention vector A TV ∈R 1×hw , divided into text position attention local blocks, A TV ={LATV s}, LATV s ∈R 1×1 , s represents the subscript index of the local block, s = 1,…,hw;
[0269] S4-2-3-2: Text position attention vector A TV ∈R 1×hw , according to each text position attention local block LATV s The numerical values are sorted in descending order to obtain the descending sequence Ades={Ldes s}, Ades∈R 1×hw , Ldes s ∈R 1×1 , s=1,…,hw;
[0270] S4-2-3-3: Calculate the average value S of the normal / abnormal micro-expression scores of the first K blocks ab ,for:
[0271]
[0272] in,
[0273] S4-2-4: Obtain the micro-expression category label y corresponding to the micro-expression video x, and calculate the binary label y ab ,for:
[0274]
[0275] S4-2-5: Calculate micro-expression position loss ab ,for:
[0276] loss ab =-[y ab log(S ab )+(1-y ab )log(1-S ab )]
[0277] S4-3: Calculate micro-expression prediction score S pre , the specific steps are as follows:
[0278] S4-3-1: The text position feature F obtained by S3-5 TP ∈R D×h×w , transform the matrix dimension to F TP ′∈R D×hw ;
[0279] S4-3-2: Text position feature F TP ′ Input a fully connected layer to obtain the micro-expression prediction feature F pre ∈R 7×hw , 7 is the number of micro-expression categories;
[0280] S4-3-3: Micro-expression prediction feature F pre , divided into micro-expression prediction local features F pre ={LFpre j}, LFpre j ∈R 1×hw , j represents the subscript index of the micro-expression category, j = 1,…,7;
[0281] S4-3-4: Given micro-expression prediction local features LFpre j ∈R 1×hw , take LFpre j The first K blocks of the numerical size in the , and merge the first K blocks to generate the micro-expression prediction K value feature LFKpre j ∈R 1×K ;
[0282] S4-3-5: Predict local features LFpre for each micro-expression j , call S4-3-4 to generate micro-expression prediction K value features {LFKpre j}, LFKpre j ∈R 1×K , j = 1,…,7;
[0283] S4-3-6: Predicting micro-expression K value features {LFKpre j}, merge along the first dimension to generate the micro-expression prediction score feature F s , F s ∈R 7×K ;
[0284] S4-3-7: Micro-expression prediction score feature F s , average pooling is performed along the second dimension, and the softmax function is used for normalization to obtain the micro-expression prediction score S pre ∈R 7×1 ;
[0285] S4-4: Calculate micro-expression feature lossfea , the specific steps are as follows:
[0286] S4-4-1: Convert the micro-expression category label y corresponding to the micro-expression video x into a one-hot vector y h , the specific steps are as follows:
[0287] S4-4-1-1: Given a yh j ∈R 1×1 , j represents the subscript index of the micro-expression category, j = 1, ..., 7, and its onehot value is calculated as
[0288]
[0289] S4-4-1-2: Yes j Each yh in j , j = 1, ..., 7, call S4-4-1-1, calculate its onehot value, and get the onehot vector y h ={yh j},y h ∈R 7×1 , yh j ∈R 1×1 ;
[0290] S4-4-2: Score the micro-expression prediction pre ∈R 7×1 , traverse each category one by one, and get S pre = {Spre j}, Spre j ∈R 1×1 , j represents the subscript index of the micro-expression category, j = 1,…,7;
[0291] S4-4-3: Calculate micro-expression feature loss fea ,for:
[0292]
[0293] S4-5: Calculate the total loss of the model, which is:
[0294] loss=loss ab +loss fea +λloss con
[0295] Where λ = 1 × 10 -4 ;
[0296] S4-6: Use the total loss of the model to train and obtain the optimal model parameters;
[0297] S5: Test model, e.g. Figure 6 As shown:
[0298] S5-1: Input micro-expression video frame;
[0299] S5-2: According to the optimal parameters of the model, call S1 to extract the optical flow position attention A o ;
[0300] S5-3: Based on the optimal model parameters, call S2 to extract the face optical flow feature F VO ;
[0301] S5-4: Call S3 based on the optimal model parameters to extract text position feature F TP ;
[0302] S5-5: Based on the optimal model parameters, call S4-3 to calculate the micro-expression prediction score S pre ;
[0303] S5-6: Identify micro-expression categories based on micro-expression prediction scores.
Claims
1. A comparative micro-expression recognition method based on text position attention, characterized by: The specific steps include: S1: extract optical flow position attention; S2: Extract facial optical flow features; S3: extract text position features; S4: training micro-expression comparison model; S5: test model; The step S1 of extracting the optical flow position attention specifically includes the following steps: S1-1: Input micro-expression video dataset; Micro-expression video dataset Data =<X,Text,Y,> , where X represents the micro-expression video set, Text represents the micro-expression category text set; Y represents the label set of the micro-expression category; x∈X, x represents a micro-expression video in the dataset, x={frame i }, frame i represents the i-th frame of the video, i=1,…,L, L represents the total number of frames of the video; Text = {happy, disgusted, fearful, angry, sad, surprised, other expressions}; y∈Y, y represents the category label of a micro-expression video. The value of y is {1,…,7}, which corresponds to the text one by one. 1 represents happiness, 2 represents disgust, 3 represents fear, 4 represents anger, 5 represents sadness, 6 represents surprise, and 7 represents other expressions. S1-2: Calculate the micro-expression movement amplitude and obtain the frame with the maximum micro-expression amplitude peak ∈R 3×H×W ; S1-3: Get {frame i The first frame frame1 in} is used as the starting frame frame of the micro-expression movement begin ∈R 3×H×W , get the optical flow sampling image frame begin 、frame peak ; S1-4: Using optical flow to sample image frame begin 、frame peak Extracting micro-expression optical flow features[u e ,v e ], where u e Represents frame begin The horizontal component of the optical flow feature at the pixel point (a, b), v e Represents frame begin The vertical component of the optical flow feature at the middle pixel (a, b); S1-5: Extracting micro-expression optical flow change rate features OF e ; S1-6: Extracting optical flow position attention A o ; The extraction of text position features described in step S3 specifically includes the following steps: S3-1: Get the micro-expression category text set Text; S3-2: Generate text prompts f ; S3-3: Text prompt t f Input a fully connected layer to generate text location prompt t p ∈R 7×hw ; S3-4: Prompt the text position p , perform matrix dimension transformation to generate text position attention A T ∈R 7×h×w ; S3-5: Extract text position feature F TP .
2. The method for comparing micro-expression based on text position attention according to claim 1, characterized in that: The input micro-expression video dataset described in step S1-1 specifically includes the following steps: S1-1-1: Micro-expression video dataset Data =<X,Text,Y,> , where X represents the micro-expression video set, Text represents the micro-expression category text set; Y represents the label set of the micro-expression category; The calculation of the micro-expression movement amplitude described in step S1-2 specifically includes the following steps: S1-2-1: Calculate optical flow features OF; S1-2-2: Calculate the movement distance D; S1-2-3: Calculate D i The number of moving pixels num greater than the threshold 0.2 i ; S1-2-4: For movement distance D = {D i Each of them is represented by the optical flow feature OF i The generated movement distance D i , call step S1-2-3, get the number of moving pixels num={num i }, i = 1, 2, ..., L-1; S1-2-5: Get the frame with the maximum amplitude of micro-expression peak ∈R 3×H×W .
3. The method for comparing micro-expression based on text position attention according to claim 2, characterized in that: The calculation of the optical flow feature OF described in step S1-2-1 specifically includes the following steps: S1-2-1-1: Given a video frame i A pixel point (a, b) in the video frame is calculated i and frame i+1 The pixel optical flow features OF i (a,b), is: OF i (a,b)=(u i (a,b),v i (a,b)) Where a=1,2,…,W,b=1,2,…,H,i=1,2,…,L-1,W represents the width of the image, H represents the height of the image, L represents the total number of frames of the video, u i (a,b) represents frame i The optical flow feature OF at the pixel point (a, b) i The horizontal component of (a,b), v i (a,b) represents frame i The optical flow feature OF at the pixel point (a, b) i The vertical component of (a, b); S1-2-1-2: video frame i For each pixel point (a, b) in the image, the optical flow feature OF is calculated. i =OF i (a,b)}, a=1,2,…,W, b=1,2,…,H, i=1,2,…,L-1,OF i ∈R 2×H×W ; S1-2-1-3: i } each of which consists of a video frame frame i and frame i+1 Generated optical flow features OF i , call S1-2-1-2, calculate the optical flow feature OF = {OF i }; The calculation of the movement distance D described in step S1-2-2 specifically includes the following steps: S1-2-2-1: Given the optical flow feature OF i A pixel point (a, b) in , calculate its pixel point movement distance D i (a,b), is: S1-2-2-2: Optical flow features OF i For each pixel point (a, b), call S1-2-2-1 to calculate the movement distance D i ={D i (a,b)}, a=1,2,…,W, b=1,2,…,H, i=1,2,…,L-1; S1-2-2-3: Yes {D i Each of them is composed of optical flow features OF i The generated movement distance D i , call S1-2-2-2, calculate the movement distance D = {D i }, i = 1, 2, ..., L-1; The calculation D described in step S1-2-3 i The number of moving pixels num greater than the threshold 0.2 i , specifically including the following steps: S1-2-3-1: Set the number of moving pixels num i The initial value is set to 0; S1-2-3-2: If D i The movement distance D at the pixel point (a, b) i (a,b) is greater than the threshold 0.2, the number of moving pixels num i Add 1, otherwise the number of moving pixels num i The value remains unchanged; S1-2-3-3: Move the pixel point to a distance D i For all the pixels (a, b) in , a=1,2,…,W, b=1,2,…,H, call S1-2-3-2 to calculate num i ; Step S1-2-5 of obtaining the micro-expression maximum amplitude frame peak ∈R 3×H×W , specifically including the following steps: S1-2-5-1: Get the number of moving pixels num = {num i }, the largest num i , use the value of i+1 as the subscript index peak; S1-2-5-2: Get {frame i The maximum amplitude frame of micro-expression in} peak .
4. The method for comparing micro-expression based on text position attention according to claim 3, characterized in that: The feature of the optical flow change rate of micro-expression extracted in step S1-5 e , specifically including the following steps: S1-5-1: Using micro-expression optical flow features [u e ,v e ], calculate the optical flow change rate z e , the calculation formula is as follows: in, is the micro-expression optical flow feature [u e ,v e ]; S1-5-2: Generate optical flow change rate features, OF e =[u e ,v e ,z e ],OF e ∈R 3×H×W ; Step S1-6 extracts the optical flow position attention A o , specifically including the following steps: S1-6-1: Optical flow change rate feature OF e , input into a convolutional layer, and get the feature F e ∈R 3×h×w , where the convolution kernel size is [k,k] and the step size is k. S1-6-2: For feature F e ∈R 3×h×w , perform maximum pooling in the first dimension to obtain the optical flow position attention A o ∈R 1 ×h×w .
5. The method for comparing micro-expression based on text position attention according to claim 4, characterized in that: The extraction of facial optical flow features described in step S2 specifically includes the following steps: S2-1: Get the frame with the maximum micro-expression movement amplitude in S1-2-5 peak ∈R 3×H×W ; S2-2: The frame with the largest micro-expression movement amplitude peak ∈R 3×H×W , divided into image blocks [k′,k′] with a height of k′ and a width of k′, and the local image block frame is obtained peak = {LI s }, LI s ∈R 3×k′×k′ , s represents the subscript index of the local block, s = 1, ..., hw, S2-3: Extracting facial visual features F V ; S2-4: The facial visual features F V With optical flow position attention A o Point product to generate the face optical flow feature F VO ∈R D×h×w ,for: F VO =F V ⊙A o Here, ⊙ represents the dot product.
6. The method for comparing micro-expression based on text position attention according to claim 5, characterized in that: Step S2-3 described in extracting facial visual features F V , specifically including the following steps: S2-3-1: Given a local image block LI s ∈R 3×k′×k′ , flatten the sth block to obtain a local block vector S2-3-2: For each local image block LI s , call S2-3-1 to get the local block vector S2-3-3: The local block vector {LV s }, merge along the first dimension to generate the feature with the maximum amplitude of micro-expression movement S2-3-4: The feature F with the maximum micro-expression movement amplitude peak , input a linear layer, and obtain the features Where D is the embedding dimension; S2-3-5: In the characteristics On the top, add position feature F p ∈R hw×D , get the features S2-3-6: In the characteristics Add a token vector V along the first dimension token ∈R 1×D , get the features S2-3-7: Features Input into the Vision Transformer with CLIP pre-trained weights to obtain the Transformer feature F T ∈R (hw+1)×D ; S2-3-8: Along the Transformer feature F T The first dimension of F is to take the first hw features and obtain F T ′∈R hw×D ; S2-3-9: For F T ′∈R hw×D , perform matrix dimension flipping to obtain F T ”∈R D×hw ; S2-3-10: For F T ”∈R D×hw , perform matrix dimension transformation to obtain the face visual feature F V ∈R D×h×w .
7. The method for comparing micro-expression based on text position attention according to claim 1, characterized in that: Step S3-2 generates a text prompt t f , specifically including the following steps: S3-2-1: Input the micro-expression category text set Text, and tokenize it with CLIP pre-trained weights to get the word tag w t ∈R 7×n , 7 is the number of micro-expression categories, and n is the number of tokens; S3-2-2: Calculate text tag weight t t ; S3-2-3: Input word mark w t , into the encode_token with CLIP pre-trained weights, and get the word embedding w e ∈R 7×n×D , D is the embedding dimension; S3-2-4: Calculate text embedding weight t e , S3-2-5: Input text label weight t t and text embedding weight t e , get the text prompt t from encode_text with CLIP pre-trained weights f ∈R 7×D ; The calculation text tag weight t described in step S3-2-2 t , specifically including the following steps: S3-2-2-1: Initialize text tag weight t t ∈R 7×n , the initial value of the weight is all zero; S3-2-2-2: Mark the word w t ∈R 7×n , extract w t The subscript ind of the maximum value in the j-th row, j = 1, ..., 7, ind = 1, ..., n; S3-2-2-3: Set p r To indicate the prefix subscript, p r =10; S3-2-2-4: Set p o To indicate the suffix subscript, p o =10; S3-2-2-5: Mark the word w t ∈R 7×n , extract the value of the jth row and the indth column, and assign it to the text tag weight t t ∈R 7×n In the jth row, the (p r +p o +ind) column, is: t t [j,p r +p o +ind]=w t [j,ind] S3-2-2-6: For word tag w t ∈R 7×n For each line in , call steps S3-2-2-2, S3-2-2-3, S3-2-2-4, S3-2-2-5 to get the text tag weight t t ∈R 7×n ; The calculated text embedding weight t described in step S3-2-4 e , specifically including the following steps: S3-2-4-1: Initialize text embedding weight t e ∈R 7×n×D , the weight is initialized to a random value that follows a standard normal distribution; S3-2-4-2: Call step S3-2-2-2 to obtain word tag w t The subscript ind of the maximum value in the j-th row, j = 1, ..., 7, ind = 1, ..., n; S3-2-4-3: embedding word w e ∈R 7×n×D , extract the vector of the jth row and the first column, and assign it to the text embedding weight t e ∈R 7 ×n×D In , the vector in the jth row and the first column is: t e [j,1,1:D]=w e [j,1,1:D] S3-2-4-4: embedding word w e , extract the vector from row j, column 1 to row j, column ind, and assign it to the text embedding weight t e In the jth row, the (p r +1) column to the jth row (p r +ind) columns, which is: t e [j,p r +1:p r +ind,1:D]=w e [j,1:ind,1:D] S3-2-4-5: embedding word w e , extract the vector of the jth row and indth column and assign it to the text embedding weight t e In the jth row, the (p r +p i +ind) columns, which is: t e [j,p r +p o +ind,1:D]=w e [j,ind,1:D] S3-2-4-6: For word tag w t ∈R 7×n For each row in , call steps S3-2-4-2, S3-2-4-3, S3-2-4-4, S3-2-4-5 to get the text embedding weight t e ∈R 7×n×D ; The extracted text position feature F in step S3-5 TP , specifically including the following steps: S3-5-1: Focus on the text position A T ∈R 7×h×w , divided into text position local attention A T ={LAT j }, LAT j ∈R h×w , j represents the subscript index of the micro-expression category, j = 1,…,7; S3-5-2: Local attention LAT given text position j , the face optical flow feature F obtained by step S2-4 VO ∈R D ×h×w , with LAT j Perform dot multiplication to obtain the local feature LFT of the text position j ∈R D×h×w ; LFT j =F VO ⊙LAT j Among them, ⊙ represents the dot product; S3-5-3: Local attention LAT for each text position j , call step S3-5-2, and obtain the text position local feature LFT = {LFT j }, LFT j ∈R D×h×w , j = 1,…,7; S3-5-4: Given a text position local feature LFT j , perform dimension expansion and obtain LFT j ∈R 1×D×h×w ; S3-5-5: For each text position local feature LFT j , call step S3-5-4, get {LFT j }, LFT j ∈R 1 ×D×h×w , j = 1,…,7; S3-5-6: Local feature of text position {LFT j }, merge along the first dimension to obtain the text position feature F TP ∈R 7×D×h×w S3-5-7: Text position feature F TP , input a fully connected layer to transform the dimension to F TP ∈R D×h×w .
8. The method for comparing micro-expression based on text position attention according to claim 7, characterized in that: The training micro-expression comparison model described in step S4 specifically includes the following steps: S4-1: Calculate micro-expression contrast loss con ; S4-2: Calculate normal / abnormal position loss ab ; S4-3: Calculate micro-expression prediction score S pre ; S4-4: Calculate micro-expression feature loss fea ; S4-5: Calculate the total loss of the model; S4-6: Use the total loss of the model to train and obtain the optimal model parameters; Calculate the micro-expression contrast loss loss described in step S4-1 con , specifically including the following steps: S4-1-1: Text prompt t f ∈R 7×D , divided into text prompt vector t f ={tv j }, tv j ∈R 1×D , j represents the subscript index of the micro-expression category, j = 1,…,7; S4-1-2: Calculate micro-expression contrast loss con ,for: Calculation of normal / abnormal position loss loss in step S4-2 ab , specifically including the following steps: S4-2-1: Attention to the text position A T ∈R 7×h×w , transform the matrix dimension to obtain the text position attention matrix A TM ∈R 7×hw ; S4-2-2: Text position attention matrix A TM , input a fully connected layer to get the text position attention vector A TV ∈R 1×hw ; S4-2-3: Calculate the average value S of normal / abnormal micro-expression scores ab ; S4-2-4: Obtain the micro-expression category label y corresponding to the micro-expression video x, and calculate the binary label y ab ,for: S4-2-5: Calculate micro-expression position loss ab ,for: loss ab =-[y ab log(S ab )+(1-y ab )log(1-S ab )] The calculation of the micro-expression prediction score S in step S4-3 pre , specifically including the following steps: S4-3-1: The text position feature F obtained by S3-5 TP ∈R D×h×w , transform the matrix dimension to f TP ′∈R D ×hw ; S4-3-2: Text position feature F TP ′ Input a fully connected layer to obtain the micro-expression prediction feature F pre ∈R 7×hw , 7 is the number of micro-expression categories; S4-3-3: Micro-expression prediction feature F pre , divided into micro-expression prediction local features F pre ={LFpef j }, LFpre j ∈R 1×hw , j represents the subscript index of the micro-expression category, j = 1,…,7; S4-3-4: Given micro-expression prediction local features LFpre j ∈R 1×hw , take LFpre j The first K blocks of the numerical size in the , and merge the first K blocks to generate the micro-expression prediction K value feature LFKpre j ∈R 1×K ; S4-3-5: Predict local features LFpre for each micro-expression j , call step S4-3-4 to generate micro-expression prediction K value features {LFKpre j }, LFKpre j ∈R 1×K , j = 1,…,7; S4-3-6: Predicting micro-expression K value features {LFKpre j }, merge along the first dimension to generate the micro-expression prediction score feature F s , F s ∈R 7×K ; S4-3-7: Micro-expression prediction score feature F s , average pooling is performed along the second dimension, and the softmax function is used for normalization to obtain the micro-expression prediction score S pre ∈R 7×1 ; Calculate the micro-expression feature loss loss described in step S4-4 fea , specifically including the following steps: S4-4-1: Convert the micro-expression category label y corresponding to the micro-expression video x into a one-hot vector y h ; S4-4-2: Score the micro-expression prediction pre ∈R 7×1 , traverse each category one by one, and get S pre = {Spre j }, Spre j ∈R 1×1 , j represents the subscript index of the micro-expression category, j = 1,…,7; S4-4-3: Calculate micro-expression feature loss fea ,for: The micro-expression category label y corresponding to the micro-expression video x described in step S4-4-1 is converted into a one-hot vector y h , specifically including the following steps: S4-4-1-1: Given a yh j ∈R 1×1 , j represents the subscript index of the micro-expression category, j = 1, ..., 7, and its onehot value is calculated as S4-4-1-2: Yes j Each yh in j , j = 1, ..., 7, call step S4-4-1-1, calculate its onehot value, and obtain the onehot vector y h ={yh j },y h ∈R 7×1 , yh j ∈R 1×1 .
9. The method for comparing micro-expression based on text position attention according to claim 8, characterized in that: The test model described in step S5 specifically includes the following steps: S5-1: Input micro-expression video frame; S5-2: According to the optimal parameters of the model, call step S1 to extract the optical flow position attention A o ; S5-3: Based on the optimal model parameters, call step S2 to extract the face optical flow feature F VO ; S5-4: Based on the optimal model parameters, call step S3 to extract the text position feature F TP ; S5-5: Based on the optimal model parameters, call step S4-3 to calculate the micro-expression prediction score S pre ; S5-6: Identify micro-expression categories based on micro-expression prediction scores.
Citation Information
Patent Citations
A micro-expression detection method based on optical flow
CN111461021B
Micro-expression recognition method, device and medium based on bidirectional recurrent neural network
CN113723287B
Micro-expression recognition method based on Transform motion feature fusion
CN118366202A
Cross-library micro-expression recognition method and device based on optical flow attention neural network
CN110516571A
Micro-expression recognition method and device based on video time domain dynamic attention model
CN114550272A