A micro-expression analysis method based on facial key point recognition

Through the micro-expression analysis method based on facial key point recognition, the problem of insufficient real-time and accuracy of micro-expression recognition in the prior art is solved, and efficient and accurate recognition and analysis of micro-expression is achieved.

CN119964223BActive Publication Date: 2025-06-10LESHAN NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510438216.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-06-10
Estimated Expiration
2045-04-09

Smart Images

  • Figure CN119964223B_ABST
    Figure CN119964223B_ABST
Patent Text Reader

Abstract

The present invention discloses a micro-expression analysis method based on facial key point recognition, which relates to the technical field of image processing. The method includes obtaining facial video information, performing sampling, setting time stamps, and frame division; extracting key frames and matching the target facial detection area to obtain an enhanced facial detection area; performing key point detection and time domain interpolation, and fusing to obtain a full-face image feature vector; performing window function filtering, Mel filtering, and Fourier transform in sequence to obtain spectral data, and performing smoothing processing on the spectral data to obtain a reference image feature vector; comparing the differences between two adjacent reference image feature vectors in sequence. When the difference is greater than a preset threshold, the two adjacent reference image feature vectors are input into the micro-expression synthesis module, and then feature extraction and recognition analysis are performed to obtain micro-expression analysis data. The present invention can effectively improve the real-time performance and accuracy of facial recognition, and improve the accuracy of capturing micro-expressions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a micro-expression analysis method based on facial key point recognition. Background Art

[0002] With the rapid development of artificial intelligence technology, means of analyzing expressions through human faces have been applied in multiple fields, such as human-computer interaction, business negotiation, security protection, judicial criminal investigation, and medical treatment. Facial expressions are an important carrier of human emotions and social communication, and their analysis and understanding play a key role in revealing a person's emotional state, social intention, and psychological state. According to the duration, facial expressions can be divided into macro-expressions and micro-expressions. Among them, in social interactions, people often use some false macro-expressions to deceive others to achieve their own purposes. However, some very subtle micro-expression changes can express real inner feelings. A "micro-expression" lasts only 0.04 seconds at the shortest. Although it is only a very short period of time, because it is unconscious, it often reveals real emotions. "Micro-expressions" often flash by. Compared with the expressions made consciously by people, "micro-expressions" can better reflect people's real feelings and motives. Micro-expressions have a short duration, small fluctuations, and are difficult to be recognized by the naked eye, which brings many challenges to the automatic recognition and analysis of micro-expressions.

[0003] Currently, for the problem of micro-expression recognition, the mainstream methods of computer vision technology are divided into two categories: traditional methods and deep learning methods. These methods cannot naturally learn subtle spatio-temporal changes, resulting in the loss of some facial information. On the other hand, with the rapid development of computer vision and graphics processing technology, convolutional neural networks and long short-term memory neural networks are combined to extract temporal and spatial features, but this will lead to an increase in the number of network parameters and running time, and it is easy to cause overfitting for small-sample data sets such as micro-expressions. In addition, three-dimensional convolutional neural networks (3DCNNs) have gradually replaced two-dimensional convolutions in the field of micro-expression recognition by virtue of the advantage of jointly extracting spatio-temporal features. However, there are still deficiencies in feature extraction: First, the effective information of micro-expressions only exists in specific regions at specific times, and most of the spatio-temporal information extracted by three-dimensional convolutional neural networks is unimportant; Second, simply stacking 3DCNN blocks will not only ignore the details of shallow-layer images but may also lead to overfitting; Finally, when three-dimensional convolutional neural networks perform facial recognition, they often lack real-time performance and accuracy, and it is difficult to accurately capture facial micro-expressions, resulting in the inability to accurately and objectively analyze people's real emotions. Summary of the Invention

[0004] Aiming at the above deficiencies in the prior art, the present invention provides a micro-expression analysis method based on facial key point recognition, which can effectively improve the real-time performance and accuracy of facial recognition, and enhance the accuracy of capturing micro-expressions.

[0005] A micro-expression analysis method based on facial key point recognition provided by the present invention includes the following steps:

[0006] Obtain the facial video information of the target person, and sample, set time stamps and frame the facial video information according to a preset time period;

[0007] Extract key frames from each frame respectively, match the target facial detection area according to the key frames, and expand the target facial detection area within a preset radius range to obtain a plurality of enhanced facial detection areas;

[0008] Perform key point detection and time-domain interpolation on the enhanced facial detection areas through a convolutional neural network model to obtain eye image feature vectors, eyebrow image feature vectors, mouth image feature vectors and optical flow image feature vectors, and fuse the eye image feature vectors, eyebrow image feature vectors, mouth image feature vectors and optical flow image feature vectors through a pyramid fusion model to obtain a plurality of full-face image feature vectors;

[0009] Calculate the spectral energy of each full-face image feature vector respectively, perform window function filtering, Mel filtering and Fourier transform in sequence to obtain spectral data, and perform smoothing processing on the spectral data to obtain a plurality of reference image feature vectors;

[0010] Compare the differences between adjacent two reference image feature vectors in sequence through a residual model according to the preset time period and time stamps. When the difference is greater than a preset threshold, input the two adjacent reference image feature vectors into a micro-expression synthesis module to obtain a micro-expression feature attribute vector, and perform feature extraction and recognition and analysis on the micro-expression feature attribute vector to obtain micro-expression analysis data.

[0011] Further, the framing includes: obtaining a start frame, a peak frame, an end frame and an offset frame and removing invalid frames.

[0012] Further, the step of extracting key frames from each frame specifically includes:

[0013] Compare whether there is a difference between the image in the current frame and the image in the previous frame. If there is, extract the current frame as a key frame.

[0014] Further, the step of expanding the target facial detection area within a preset radius range specifically includes: determining an expansion ratio threshold of the polar radius based on the key frame, the polar coordinate information of the target facial detection area and a preset overlap rate.

[0015] Further, the step of fusing the eye image feature vectors, eyebrow image feature vectors, mouth image feature vectors and optical flow image feature vectors through a pyramid fusion model specifically includes:

[0016] Weightedly splice the eye image feature vector, the eyebrow image feature vector, and the mouth image feature vector, and perform optical flow feature extraction on the optical flow image feature vector;

[0017] Perform correlation fusion based on the data after the weighted splicing and optical flow feature extraction.

[0018] Further, the steps of training the convolutional neural network model specifically include:

[0019] Input the target training sample into a preset convolutional neural network model for key point detection to obtain predicted key point information;

[0020] Calculate the difference between the predicted key point information and the key point annotation information corresponding to the target training sample to obtain a loss value;

[0021] Adjust the parameters of the convolutional neural network model and continue iterative training until the loss value is within a preset accuracy range, end the iterative training, and update the convolutional neural network model.

[0022] Further, the step of inputting the two adjacent reference image feature vectors into the micro-expression synthesis module to obtain a micro-expression feature attribute vector specifically includes:

[0023] Perform expression extraction and label extraction on the two adjacent reference image feature vectors;

[0024] Generate corresponding expression parameters and label parameters according to the results of the expression extraction and label extraction;

[0025] Perform correlation fusion on the expression parameters and the label parameters to obtain a micro-expression feature attribute vector.

[0026] The beneficial effects of the present invention are as follows:

[0027] The present invention first extracts key frames from facial video information, matches the target facial detection area, then enlarges the target facial detection area to obtain multiple enhanced facial detection areas, and then performs key point detection and temporal interpolation through a convolutional neural network model, and can efficiently and real-time identify key image information from facial video information.

[0028] The present invention fuses the obtained eye image feature vector, eyebrow image feature vector, mouth image feature vector, and optical flow image feature vector through a pyramid fusion model to obtain a full-face image feature vector, enhances the accuracy of capturing image detail information, and improves the performance and efficiency of identifying the full-face image.

[0029] By calculating the spectral energy of each full-face image feature vector, and performing window function filtering, Mel filtering, and Fourier transform in sequence to obtain the corresponding reference image feature vector, the present invention can quickly extract the effective image feature information of the full-face image. When the difference between two adjacent reference image feature vectors is greater than a preset threshold, the two adjacent reference image feature vectors are input into the micro-expression synthesis module, and then feature extraction and recognition analysis are performed on the obtained micro-expression feature attribute vector to obtain micro-expression analysis data, thereby effectively improving the real-time performance and accuracy of face recognition. Description of the Drawings

[0030] Figure 1 is a schematic flowchart of a micro-expression analysis method based on facial key point recognition provided by an embodiment of the present invention. Detailed Embodiments

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0032] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations. In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other. Embodiment

[0033] As Figure 1 shown, a micro-expression analysis method based on facial key point recognition provided by an embodiment of the present invention includes the following steps S10 to S50:

[0034] Step S10: Obtain the facial video information of the target person, and sample, set time stamps, and frame the facial video information according to a preset time period.

[0035] Since the shortest duration of a micro-expression is only 0.04 seconds. In this embodiment, the facial video information of the target person is captured by a high-speed high-definition camera. At the same time, the longest period of the preset time is set to 0.04 seconds, and it can also be set to 0.02 seconds or 0.01 seconds. The sampling frequency is set to the reciprocal of the preset time period. Time stamps are set synchronously during the sampling process, and then the facial video information with time stamps is framed.

[0036] Among them, the frame segmentation processing includes: obtaining a starting frame, a peak frame, an ending frame, and an offset frame, and removing invalid frames. Specifically, by performing energy extraction and zero-crossing rate extraction on each frame of the image, the invalid frames are removed according to threshold comparison.

[0037] Step S20: Extract key frames from each frame respectively, match the target face detection area according to the key frames, and expand the target face detection area within a preset radius range to obtain a plurality of enhanced face detection areas.

[0038] In one implementation manner, the step of extracting key frames from each frame specifically includes: comparing whether there is a difference between the image in the current frame and the image in the previous frame. If there is a difference, extract the current frame as a key frame.

[0039] For each frame of the image, first obtain the previous frame corresponding to the current frame, then compare the current frame with the previous frame in terms of pixels to determine whether there is a difference between the current frame and the previous frame. If there is a difference between the current frame and the previous frame, extract the current frame as a key frame.

[0040] Further, the step of expanding the target face detection area within a preset radius range specifically includes: determining an expansion ratio threshold of the polar radius based on the polar coordinate information of the key frame and the target face detection area and a preset overlapping rate.

[0041] Among them, the target face detection area is obtained by performing face detection on the training samples based on a face detection algorithm. For example, the training samples can be input into a face detection model for face detection to obtain the face detection box information output by the model, and this face detection box information indicates the face detection area in the training samples. Taking a face detection model as an example, the training samples are input into the face detection model for face detection to obtain the face detection box output by the model. In this embodiment, the face detection box is set to be circular or elliptical.

[0042] Since the accuracy of the face detection algorithm cannot guarantee that all face detections are accurate values, and some face areas are too small to play a positive role in subsequent key point detections, therefore, in this embodiment, an expansion ratio threshold of the polar radius is determined based on the polar coordinate information of the key frame and the target face detection area and a preset overlapping rate. The expansion ratio threshold is preferably 1.1 or 1.2.

[0043] Step S30: Perform key point detection and temporal interpolation on the enhanced face detection areas through a convolutional neural network model to obtain an eye image feature vector, an eyebrow image feature vector, a mouth image feature vector, and an optical flow image feature vector, and fuse the eye image feature vector, the eyebrow image feature vector, the mouth image feature vector, and the optical flow image feature vector through a pyramid fusion model to obtain a plurality of full-face image feature vectors.

[0044] In one implementation, the step of training the convolutional neural network model specifically includes:

[0045] Input a target training sample into a preset convolutional neural network model for key point detection to obtain predicted key point information;

[0046] Calculate the difference between the predicted key point information and the key point annotation information corresponding to the target training sample to obtain a loss value;

[0047] Adjust the parameters of the convolutional neural network model and continue iterative training until the loss value is within a preset accuracy range, end the iterative training, and update the convolutional neural network model.

[0048] Among them, only the facial key points need to be annotated for the training data, which greatly reduces the training data collection cost. Moreover, in the training of the convolutional neural network model, there is no need to introduce weight hyperparameters between different tasks, and stable improvement of the model training can be achieved.

[0049] The step of time domain interpolation can increase the number of images included in key point detection, thereby prolonging the micro-expression duration. This method first uses a graph embedding algorithm to embed the graph into a low-dimensional manifold, and finally substitutes the image vector to calculate this high-dimensional continuous curve. Resampling on the curve can obtain the interpolated eye image feature vector, eyebrow image feature vector, mouth image feature vector, and optical flow image feature vector.

[0050] This embodiment first extracts key frames from the facial video information and matches the target facial detection area, then expands the target facial detection area to obtain multiple enhanced facial detection areas, and then performs key point detection and time domain interpolation through a convolutional neural network model, and can efficiently and real-time identify key image information from the facial video information.

[0051] Further, the step of fusing the eye image feature vector, eyebrow image feature vector, mouth image feature vector, and optical flow image feature vector through a pyramid fusion model specifically includes:

[0052] Perform weighted splicing on the eye image feature vector, eyebrow image feature vector, and mouth image feature vector, and perform optical flow feature extraction on the optical flow image feature vector;

[0053] Perform correlation fusion based on the data after the weighted splicing and optical flow feature extraction.

[0054] In this embodiment, the eye image feature vector, the eyebrow image feature vector, and the mouth image feature vector are fused through a pyramid fusion model to complete the joint learning of spatio-temporal features and detailed information, achieving the capture of image detailed information and improving the performance and efficiency of image recognition. Then, by using the high resolution and motion information of the optical flow image features and the weighted micro-action semantic information for the eye and mouth regions, enhanced correlation fusion is performed, enhancing the correction and update of the original prediction by local details and optical flow motion information.

[0055] In the embodiment of the present invention, the obtained eye image feature vector, eyebrow image feature vector, mouth image feature vector, and optical flow image feature vector are fused through a pyramid fusion model to obtain a full-face image feature vector, enhancing the accuracy of capturing image detailed information and improving the performance and efficiency of recognizing the full-face image.

[0056] Step S40: Calculate the spectral energy of each full-face image feature vector respectively, and obtain spectral data after window function filtering, Mel filtering, and Fourier transform in sequence. The spectral data is smoothed to obtain a plurality of reference image feature vectors.

[0057] In one implementation, first, the spectral energy of the full-face image feature vector corresponding to each frame calculated is multiplied by the window function discrete value through a look-up table; then, the spectral energy of each frame is summed through a look-up table operation using a Mel filter bank to obtain a Mel spectrum output. After that, a logarithmic operation is performed on the Mel spectrum output to obtain Mel energy features; finally, spectral data is obtained through Fourier transform, which can reduce the computational amount and time in the feature extraction stage, thus better meeting the real-time requirements.

[0058] Step S50: Compare the differences between adjacent two reference image feature vectors in sequence through a residual model according to the preset time period and time stamp. When the difference is greater than the preset threshold, the two adjacent reference image feature vectors are input into a micro-expression synthesis module to obtain a micro-expression feature attribute vector, and feature extraction and recognition analysis are performed on the micro-expression feature attribute vector to obtain micro-expression analysis data.

[0059] Among them, the micro-expression analysis data includes happiness, curiosity, fear, anger, etc.

[0060] In one implementation, the step of inputting the two adjacent reference image feature vectors into a micro-expression synthesis module to obtain a micro-expression feature attribute vector specifically includes the following sub-steps:

[0061] Perform expression extraction and label extraction on the two adjacent reference image feature vectors;

[0062] Generate corresponding expression parameters and label parameters according to the results of the expression extraction and label extraction;

[0063] Associate and fuse the expression parameters and label parameters to obtain a micro-expression feature attribute vector.

[0064] In this embodiment, by calculating the spectral energy of each full-face image feature vector and performing window function filtering, Mel filtering, and Fourier transform in sequence to obtain the corresponding reference image feature vector, the effective image feature information of the full-face image can be quickly extracted. When the difference between two adjacent reference image feature vectors is greater than a preset threshold, the two adjacent reference image feature vectors are input into the micro-expression synthesis module, and then feature extraction and recognition analysis are performed on the obtained micro-expression feature attribute vector to obtain micro-expression analysis data, thereby effectively improving the real-time performance and accuracy of face recognition.

[0065] In summary, in the embodiment of the present invention, key frames are first extracted from the facial video information and the target face detection area is matched, then the target face detection area is enlarged to obtain multiple enhanced face detection areas, and then key point detection and time-domain interpolation are performed through a convolutional neural network model, so that key image information can be efficiently and real-time recognized from the facial video information.

[0066] In the embodiment of the present invention, the obtained eye image feature vector, eyebrow image feature vector, mouth image feature vector, and optical flow image feature vector are fused through a pyramid fusion model to obtain a full-face image feature vector, enhancing the accuracy of capturing image detail information and improving the performance and efficiency of recognizing the full-face image.

[0067] In this embodiment, by calculating the spectral energy of each full-face image feature vector and performing window function filtering, Mel filtering, and Fourier transform in sequence to obtain the corresponding reference image feature vector, the effective image feature information of the full-face image can be quickly extracted. When the difference between two adjacent reference image feature vectors is greater than a preset threshold, the two adjacent reference image feature vectors are input into the micro-expression synthesis module, and then feature extraction and recognition analysis are performed on the obtained micro-expression feature attribute vector to obtain micro-expression analysis data, thereby effectively improving the real-time performance and accuracy of face recognition.

[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A micro-expression analysis method based on facial key point recognition, characterized in that: The following steps are involved: Acquire facial video information of a target person, and sample, set timestamps and divide frames for the facial video information according to a preset time period; Extracting key frames from each frame respectively, matching a target facial detection area according to the key frames, and expanding the target facial detection area within a preset radius to obtain a plurality of enhanced facial detection areas; Performing key point detection and time domain interpolation on the enhanced facial detection area through a convolutional neural network model to obtain eye image feature vectors, eyebrow image feature vectors, mouth image feature vectors, and optical flow image feature vectors, and fusing the eye image feature vectors, eyebrow image feature vectors, mouth image feature vectors, and optical flow image feature vectors through a pyramid fusion model to obtain multiple full-face image feature vectors; The step of performing key point detection on the enhanced facial detection area by using a convolutional neural network model specifically includes: Input the target training sample into the preset convolutional neural network model to detect key points and obtain the predicted key point information; Calculating the difference between the predicted key point information and the key point annotation information corresponding to the target training sample to obtain a loss value; Adjust the parameters of the convolutional neural network model and continue iterative training until the loss value is within a preset accuracy range, end the iterative training and update the convolutional neural network model; The time domain interpolation step specifically includes: Use graph embedding algorithm to embed key point detection image into a low-dimensional manifold, substitute image vector, and calculate high-dimensional continuous curve; Re-sampling is performed on the continuous curve to obtain an interpolated eye image feature vector, an eyebrow image feature vector, a mouth image feature vector and an optical flow image feature vector; Calculating the spectrum energy of each full-face image feature vector respectively, performing window function filtering, Mel filtering and Fourier transform in sequence to obtain spectrum data, and performing smoothing processing on the spectrum data to obtain multiple reference image feature vectors; The difference between the feature vectors of two adjacent reference images is compared in sequence through a residual model according to the preset time period and timestamp. When the difference is greater than a preset threshold, the feature vectors of the two adjacent reference images are input into a micro-expression synthesis module to obtain a micro-expression feature attribute vector. Feature extraction and recognition analysis are performed on the micro-expression feature attribute vector to obtain micro-expression analysis data.

2. The micro-expression analysis method based on facial key point recognition according to claim 1, characterized in that: The framing includes: obtaining a start frame, a peak frame, an end frame and an offset frame, and removing an invalid frame.

3. The micro-expression analysis method based on facial key point recognition according to claim 1, characterized in that: The step of extracting key frames from each frame specifically includes: Compare the image in the current frame with the image in the previous frame to see if there is any difference. If so, extract the current frame as a key frame.

4. The micro-expression analysis method based on facial key point recognition according to claim 1, characterized in that: The step of expanding the target facial detection area within a preset radius specifically includes: determining an expansion ratio threshold of the polar radius based on polar coordinate information of the key frame and the target facial detection area and a preset overlap ratio.

5. The micro-expression analysis method based on facial key point recognition according to claim 1, characterized in that: The step of fusing the eye image feature vector, the eyebrow image feature vector, the mouth image feature vector and the optical flow image feature vector through a pyramid fusion model specifically includes: Perform weighted concatenation on the eye image feature vector, the eyebrow image feature vector and the mouth image feature vector, and perform optical flow feature extraction on the optical flow image feature vector; The data after weighted splicing and optical flow feature extraction are associated and fused.

6. The micro-expression analysis method based on facial key point recognition according to claim 1, characterized in that: The step of training the convolutional neural network model specifically includes: Input the target training sample into the preset convolutional neural network model to detect key points and obtain the predicted key point information; Calculating the difference between the predicted key point information and the key point annotation information corresponding to the target training sample to obtain a loss value; Adjust the parameters of the convolutional neural network model and continue iterative training until the loss value is within a preset accuracy range, end the iterative training and update the convolutional neural network model.

7. The micro-expression analysis method based on facial key point recognition according to claim 1, characterized in that: The step of inputting the two adjacent reference image feature vectors into a micro-expression synthesis module to obtain a micro-expression feature attribute vector specifically includes: Performing expression extraction and label extraction on the feature vectors of the two adjacent reference images; Generate corresponding expression parameters and label parameters according to the results of expression extraction and label extraction; The expression parameters and label parameters are associated and fused to obtain a micro-expression feature attribute vector.

Citation Information

Patent Citations

  • Psychological state analysis method based on facial micro-expression

    AU2020102556A4

  • Micro-expression recognition method and device

    CN116543440A