Micro-expression analysis method based on facial key point recognition

Through the micro-expression analysis method based on facial key point recognition, the problem of insufficient real-time and accuracy of micro-expression recognition in the prior art is solved, and efficient and accurate capture and analysis of micro-expression can be achieved, and people's true emotions can be accurately analyzed.

CN119964223AActive Publication Date: 2025-05-09LESHAN NORMAL UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510438216.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-09
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The prior art has problems with insufficient real-time and accuracy in micro-expression recognition, and it is difficult to accurately capture and analyze micro-expression, resulting in the inability to accurately and objectively analyze people's true emotions.

Method used

The micro-expression analysis method based on facial key point recognition is adopted. By obtaining facial video information, keyframes are extracted and matched with the target facial detection area, the keypoint detection and time-domain interpolation of the convolution neural network model are performed after the detection area is expanded. Combined with the pyramid fusion model and spectrum analysis technology, the image feature vector is extracted and fused, and finally feature extraction and recognition is performed through the micro-expression synthesis module.

Benefits of technology

It improves the real-time and accuracy of facial recognition, enhances the accuracy of capturing micro-expressions, and can effectively analyze people's true emotions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964223A_ABST
    Figure CN119964223A_ABST
Patent Text Reader

Abstract

The invention discloses a micro-expression analysis method based on facial key point recognition, and relates to the technical field of image processing. The method comprises the following steps: acquiring face video information, sampling, setting a timestamp and framing; extracting a key frame and a matched target face detection area to obtain an enhanced face detection area; carrying out key point detection and time domain interpolation, and fusing to obtain a full face image feature vector; performing window function filtering, Mel filtering and Fourier transform in sequence to obtain spectrum data, and performing smoothing processing on the spectrum data to obtain a reference image feature vector; and sequentially comparing the difference values of the feature vectors of the two adjacent reference images, inputting the feature vectors of the two adjacent reference images into a micro-expression synthesis module when the difference values are greater than a preset threshold value, and then carrying out feature extraction and recognition analysis to obtain micro-expression analysis data. According to the invention, the real-time performance and accuracy of face recognition can be effectively improved, and the precision of capturing micro expressions is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a micro-expression analysis method based on facial key point recognition. Background Art

[0002] With the rapid development of artificial intelligence technology, facial expression analysis has been applied to many fields, such as human-computer interaction, business negotiations, security protection, judicial investigation and medical treatment. Facial expressions are an important carrier of human emotions and social communication. Their analysis and understanding play a key role in revealing people's emotional state, social willingness and psychological state. According to the duration, facial expressions can be divided into macro-expressions and micro-expressions. Among them, in social interactions, people often use some false macro-expressions to confuse the other party in order to achieve their own goals. However, some very subtle changes in micro-expressions can express the true feelings in the heart. The shortest "micro-expression" is only 0.04 seconds. Although it is only a very short time, it often exposes the real emotions because it is unconscious. "Micro-expressions" often flash by. Compared with the expressions that people consciously make, "micro-expressions" can better reflect people's true feelings and motivations. Micro-expressions have a short duration, small fluctuations, and are difficult to recognize with the naked eye, which brings many challenges to the automatic recognition and analysis of micro-expressions.

[0003] At present, the mainstream methods of computer vision technology for micro-expression recognition are divided into two categories: traditional methods and deep learning methods. These methods cannot naturally learn subtle spatiotemporal changes, resulting in the loss of some facial information. On the other hand, with the rapid development of computer vision and graphics processing technology, the combination of convolutional neural networks and long short-term memory neural networks to extract temporal and spatial features will increase the number of network parameters and running time, and it is easy to cause overfitting for small sample data sets such as micro-expressions. In addition, three-dimensional convolutional neural networks (3DCNN) have gradually replaced two-dimensional convolution in the field of micro-expression recognition by taking advantage of the joint extraction of spatiotemporal features. However, there are still deficiencies in feature extraction: first, the effective information of micro-expressions only exists in a specific area at a specific time, while the spatiotemporal information extracted by three-dimensional convolutional neural networks is mostly unimportant; second, simply superimposing 3DCNN blocks will not only ignore the details of shallow images, but may also lead to overfitting; finally, three-dimensional convolutional neural networks often lack real-time and accuracy in facial recognition, making it difficult to accurately capture facial micro-expressions, resulting in the inability to accurately and objectively analyze people's true emotions. Summary of the invention

[0004] In view of the above-mentioned deficiencies in the prior art, the present invention provides a micro-expression analysis method based on facial key point recognition, which can effectively improve the real-time and accuracy of facial recognition, and enhance the precision of capturing micro-expressions.

[0005] The present invention provides a micro-expression analysis method based on facial key point recognition, comprising the following steps: Acquire facial video information of a target person, and sample, set timestamps and divide frames for the facial video information according to a preset time period; Extracting key frames from each frame respectively, matching a target facial detection area according to the key frames, and expanding the target facial detection area within a preset radius to obtain a plurality of enhanced facial detection areas; Performing key point detection and time domain interpolation on the enhanced facial detection area through a convolutional neural network model to obtain eye image feature vectors, eyebrow image feature vectors, mouth image feature vectors, and optical flow image feature vectors, and fusing the eye image feature vectors, eyebrow image feature vectors, mouth image feature vectors, and optical flow image feature vectors through a pyramid fusion model to obtain multiple full-face image feature vectors; Calculating the spectrum energy of each full-face image feature vector respectively, performing window function filtering, Mel filtering and Fourier transform in sequence to obtain spectrum data, and performing smoothing processing on the spectrum data to obtain multiple reference image feature vectors; The difference between the feature vectors of two adjacent reference images is compared in sequence through a residual model according to the preset time period and timestamp. When the difference is greater than a preset threshold, the feature vectors of the two adjacent reference images are input into a micro-expression synthesis module to obtain a micro-expression feature attribute vector. Feature extraction and recognition analysis are performed on the micro-expression feature attribute vector to obtain micro-expression analysis data.

[0006] Furthermore, the framing includes: acquiring a start frame, a peak frame, an end frame and an offset frame, and removing invalid frames.

[0007] Furthermore, the step of extracting a key frame from each frame specifically includes: Compare the image in the current frame with the image in the previous frame to see if there is any difference. If so, extract the current frame as a key frame.

[0008] Further, the step of expanding the target facial detection area within a preset radius specifically includes: determining an expansion ratio threshold of the polar radius based on polar coordinate information of the key frame and the target facial detection area and a preset overlap ratio.

[0009] Furthermore, the step of fusing the eye image feature vector, the eyebrow image feature vector, the mouth image feature vector and the optical flow image feature vector through a pyramid fusion model specifically includes: Perform weighted concatenation on the eye image feature vector, the eyebrow image feature vector and the mouth image feature vector, and perform optical flow feature extraction on the optical flow image feature vector; The data after weighted splicing and optical flow feature extraction are associated and fused.

[0010] Furthermore, the step of training the convolutional neural network model specifically includes: Input the target training sample into the preset convolutional neural network model to detect key points and obtain the predicted key point information; Calculating the difference between the predicted key point information and the key point annotation information corresponding to the target training sample to obtain a loss value; Adjust the parameters of the convolutional neural network model and continue iterative training until the loss value is within a preset accuracy range, end the iterative training and update the convolutional neural network model.

[0011] Furthermore, the step of inputting the two adjacent reference image feature vectors into a micro-expression synthesis module to obtain a micro-expression feature attribute vector specifically includes: Performing expression extraction and label extraction on the feature vectors of the two adjacent reference images; Generate corresponding expression parameters and label parameters according to the results of expression extraction and label extraction; The expression parameters and label parameters are associated and fused to obtain a micro-expression feature attribute vector.

[0012] The beneficial effects of the present invention are: The present invention first extracts key frames from facial video information and matches the target facial detection area, then expands the target facial detection area to obtain multiple enhanced facial detection areas, and then performs key point detection and time domain interpolation through a convolutional neural network model, which can efficiently and real-timely identify key image information from facial video information.

[0013] The present invention fuses the obtained eye image feature vector, eyebrow image feature vector, mouth image feature vector and optical flow image feature vector through a pyramid fusion model to obtain a full-face image feature vector, thereby enhancing the accuracy of capturing image detail information and improving the performance and efficiency of recognizing full-face images.

[0014] The present invention calculates the spectral energy of each full-face image feature vector, and then performs window function filtering, Mel filtering and Fourier transform in sequence to obtain the corresponding reference image feature vector, which can quickly extract the effective image feature information of the full-face image. When the difference between two adjacent reference image feature vectors is greater than a preset threshold, the two adjacent reference image feature vectors are input into the micro-expression synthesis module, and then the obtained micro-expression feature attribute vector is subjected to feature extraction and recognition analysis to obtain micro-expression analysis data, thereby effectively improving the real-time and accuracy of facial recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a flow chart of a micro-expression analysis method based on facial key point recognition provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0017] It should be understood that when used in this specification and the appended claims, the terms "comprise" and "include" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their collections. In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other. Example

[0018] like Figure 1 As shown, a micro-expression analysis method based on facial key point recognition provided by an embodiment of the present invention includes the following steps S10 to S50: Step S10: Acquire facial video information of the target person, and sample, set timestamps and divide frames for the facial video information according to a preset time period.

[0019] Since the shortest duration of micro-expressions is only 0.04 seconds, in this embodiment, the facial video information of the target person is captured by a high-speed high-definition camera. At the same time, the longest period of the preset time is set to 0.04 seconds, and can also be set to 0.02 seconds or 0.01 seconds. The sampling frequency is set to the reciprocal of the preset time period. A timestamp is synchronously set during the sampling process, and then the facial video information with the timestamp is frame-processed.

[0020] The frame processing includes: obtaining a start frame, a peak frame, an end frame and an offset frame and removing invalid frames. Specifically, by performing energy extraction and zero-crossing rate extraction on each frame of the image, invalid frames are removed according to threshold comparison.

[0021] Step S20: extract key frames from each frame respectively, match the target facial detection area according to the key frames, expand the target facial detection area within a preset radius, and obtain multiple enhanced facial detection areas.

[0022] In one embodiment, the step of extracting a key frame from each frame specifically includes: comparing whether an image in a current frame is different from an image in a previous frame, and if so, extracting the current frame as a key frame.

[0023] For each frame of the image, first obtain the previous frame corresponding to the current frame, then compare the current frame with the previous frame in units of pixels to determine whether there is a difference between the current frame and the previous frame. If there is a difference between the current frame and the previous frame, extract the current frame as a key frame.

[0024] Further, the step of expanding the target facial detection area within a preset radius specifically includes: determining an expansion ratio threshold of the polar radius based on polar coordinate information of the key frame and the target facial detection area and a preset overlap ratio.

[0025] The target face detection area refers to the face detection obtained by performing face detection on the training sample based on the face detection algorithm. For example, the training sample can be input into the face detection model for face detection to obtain the face detection frame information output by the model, and the face detection frame information indicates the face detection area in the training sample. Taking the face detection model as an example, the training sample is input into the face detection model for face detection to obtain the face detection frame output by the model. In this embodiment, the face detection frame is set to be circular or elliptical.

[0026] Since the accuracy of the face detection algorithm cannot guarantee that all face detections are accurate values, and some face areas are too small to play a positive role in subsequent key point detection, this embodiment determines the enlargement ratio threshold of the polar radius based on the polar coordinate information of the key frame and the target face detection area and the preset overlap rate. The enlargement ratio threshold is preferably 1.1 or 1.2.

[0027] Step S30: performing key point detection and time domain interpolation on the enhanced facial detection area through a convolutional neural network model to obtain eye image feature vectors, eyebrow image feature vectors, mouth image feature vectors and optical flow image feature vectors, and fusing the eye image feature vectors, eyebrow image feature vectors, mouth image feature vectors and optical flow image feature vectors through a pyramid fusion model to obtain multiple full face image feature vectors.

[0028] In one embodiment, the step of training the convolutional neural network model specifically includes: Input the target training sample into the preset convolutional neural network model to detect key points and obtain the predicted key point information; Calculating the difference between the predicted key point information and the key point annotation information corresponding to the target training sample to obtain a loss value; Adjust the parameters of the convolutional neural network model and continue iterative training until the loss value is within a preset accuracy range, end the iterative training and update the convolutional neural network model.

[0029] Among them, only facial key points need to be labeled for training data, which greatly reduces the cost of collecting training data. Moreover, in the training of convolutional neural network models, there is no need to introduce weight hyperparameters between different tasks, which can achieve stable improvement of model training.

[0030] The step of time domain interpolation can increase the number of images included in key point detection, thereby extending the duration of micro-expressions. The method first uses a graph embedding algorithm to embed the graph into a low-dimensional manifold, and finally substitutes the image vector to calculate this high-dimensional continuous curve. Resampling on the curve can obtain the interpolated eye image feature vector, eyebrow image feature vector, mouth image feature vector and optical flow image feature vector.

[0031] This embodiment first extracts key frames from facial video information and matches the target facial detection area, then expands the target facial detection area to obtain multiple enhanced facial detection areas, and then performs key point detection and time domain interpolation through a convolutional neural network model, which can efficiently and real-timely identify key image information from facial video information.

[0032] Furthermore, the step of fusing the eye image feature vector, the eyebrow image feature vector, the mouth image feature vector and the optical flow image feature vector through a pyramid fusion model specifically includes: Perform weighted concatenation on the eye image feature vector, the eyebrow image feature vector and the mouth image feature vector, and perform optical flow feature extraction on the optical flow image feature vector; The data after weighted splicing and optical flow feature extraction are associated and fused.

[0033] In this embodiment, the eye image feature vector, eyebrow image feature vector and mouth image feature vector are fused by the pyramid fusion model to achieve joint learning of spatiotemporal features and detail information, and achieve the capture of image detail information, thus improving the performance and efficiency of image recognition. The high resolution and motion information of the optical flow image features are then used, and the weighted micro-motion semantic information of the eye and mouth regions is simultaneously enhanced and associated to enhance the correction and update of the original prediction by the local details and optical flow motion information.

[0034] The embodiment of the present invention fuses the obtained eye image feature vector, eyebrow image feature vector, mouth image feature vector and optical flow image feature vector through a pyramid fusion model to obtain a full face image feature vector, thereby enhancing the accuracy of capturing image detail information and improving the performance and efficiency of recognizing full face images.

[0035] Step S40: Calculate the spectrum energy of each full-face image feature vector respectively, perform window function filtering, Mel filtering and Fourier transform in sequence to obtain spectrum data, and perform smoothing processing on the spectrum data to obtain multiple reference image feature vectors.

[0036] In one implementation, the spectral energy of the calculated full-face image feature vector corresponding to each frame is first multiplied by a lookup table with a discrete value of a window function; the spectral energy of each frame is then summed by a lookup table operation through a Mel filter array to obtain a Mel spectrum output, and then a Mel energy feature is obtained by performing a logarithmic operation on the Mel spectrum output; finally, the spectrum data is obtained after a Fourier transform, which can reduce the amount of calculation and the time in the feature extraction stage, thereby better meeting real-time requirements.

[0037] Step S50: comparing the difference between the feature vectors of two adjacent reference images in sequence through the residual model according to the preset time period and timestamp; when the difference is greater than a preset threshold, inputting the feature vectors of the two adjacent reference images into the micro-expression synthesis module to obtain a micro-expression feature attribute vector; performing feature extraction and recognition analysis on the micro-expression feature attribute vector to obtain micro-expression analysis data.

[0038] The micro-expression analysis data includes happiness, curiosity, fear, anger, etc.

[0039] In one embodiment, the step of inputting the two adjacent reference image feature vectors into a micro-expression synthesis module to obtain a micro-expression feature attribute vector specifically includes the following sub-steps: Performing expression extraction and label extraction on the feature vectors of the two adjacent reference images; Generate corresponding expression parameters and label parameters according to the results of expression extraction and label extraction; The expression parameters and label parameters are associated and fused to obtain a micro-expression feature attribute vector.

[0040] This embodiment calculates the spectral energy of each full-face image feature vector, and then performs window function filtering, Mel filtering and Fourier transform in sequence to obtain the corresponding reference image feature vector, which can quickly extract the effective image feature information of the full-face image. When the difference between two adjacent reference image feature vectors is greater than a preset threshold, the two adjacent reference image feature vectors are input into the micro-expression synthesis module, and then the obtained micro-expression feature attribute vector is subjected to feature extraction and recognition analysis to obtain micro-expression analysis data, thereby effectively improving the real-time and accuracy of facial recognition.

[0041] To summarize, the embodiment of the present invention first extracts key frames from facial video information and matches the target facial detection area, then expands the target facial detection area to obtain multiple enhanced facial detection areas, and then performs key point detection and time domain interpolation through a convolutional neural network model, so as to efficiently and real-time identify key image information from facial video information.

[0042] The embodiment of the present invention fuses the obtained eye image feature vector, eyebrow image feature vector, mouth image feature vector and optical flow image feature vector through a pyramid fusion model to obtain a full face image feature vector, thereby enhancing the accuracy of capturing image detail information and improving the performance and efficiency of recognizing full face images.

[0043] The embodiment of the present invention calculates the spectral energy of each full-face image feature vector, and then performs window function filtering, Mel filtering and Fourier transform in sequence to obtain the corresponding reference image feature vector, which can quickly extract the effective image feature information of the full-face image. When the difference between two adjacent reference image feature vectors is greater than a preset threshold, the two adjacent reference image feature vectors are input into the micro-expression synthesis module, and then the obtained micro-expression feature attribute vector is subjected to feature extraction and recognition analysis to obtain micro-expression analysis data, thereby effectively improving the real-time performance and accuracy of facial recognition.

[0044] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A micro-expression analysis method based on facial key point recognition, characterized in that: The following steps are involved: Acquire facial video information of a target person, and sample, set timestamps and divide frames for the facial video information according to a preset time period; Extracting key frames from each frame respectively, matching a target facial detection area according to the key frames, and expanding the target facial detection area within a preset radius to obtain a plurality of enhanced facial detection areas; Performing key point detection and time domain interpolation on the enhanced facial detection area through a convolutional neural network model to obtain eye image feature vectors, eyebrow image feature vectors, mouth image feature vectors, and optical flow image feature vectors, and fusing the eye image feature vectors, eyebrow image feature vectors, mouth image feature vectors, and optical flow image feature vectors through a pyramid fusion model to obtain multiple full-face image feature vectors; Calculating the spectrum energy of each full-face image feature vector respectively, performing window function filtering, Mel filtering and Fourier transform in sequence to obtain spectrum data, and performing smoothing processing on the spectrum data to obtain multiple reference image feature vectors; The difference between the feature vectors of two adjacent reference images is compared in sequence through a residual model according to the preset time period and timestamp. When the difference is greater than a preset threshold, the feature vectors of the two adjacent reference images are input into a micro-expression synthesis module to obtain a micro-expression feature attribute vector. Feature extraction and recognition analysis are performed on the micro-expression feature attribute vector to obtain micro-expression analysis data.

2. The micro-expression analysis method based on facial key point recognition according to claim 1, characterized in that: The framing includes: obtaining a start frame, a peak frame, an end frame and an offset frame, and removing an invalid frame.

3. The micro-expression analysis method based on facial key point recognition according to claim 1, characterized in that: The step of extracting key frames from each frame specifically includes: Compare the image in the current frame with the image in the previous frame to see if there is any difference. If so, extract the current frame as a key frame.

4. The micro-expression analysis method based on facial key point recognition according to claim 1, characterized in that: The step of expanding the target facial detection area within a preset radius specifically includes: determining an expansion ratio threshold of the polar radius based on polar coordinate information of the key frame and the target facial detection area and a preset overlap ratio.

5. The micro-expression analysis method based on facial key point recognition according to claim 1, characterized in that: The step of fusing the eye image feature vector, the eyebrow image feature vector, the mouth image feature vector and the optical flow image feature vector through a pyramid fusion model specifically includes: Perform weighted concatenation on the eye image feature vector, the eyebrow image feature vector and the mouth image feature vector, and perform optical flow feature extraction on the optical flow image feature vector; The data after weighted splicing and optical flow feature extraction are associated and fused.

6. The micro-expression analysis method based on facial key point recognition according to claim 1, characterized in that: The step of training the convolutional neural network model specifically includes: Input the target training sample into the preset convolutional neural network model to detect key points and obtain the predicted key point information; Calculating the difference between the predicted key point information and the key point annotation information corresponding to the target training sample to obtain a loss value; Adjust the parameters of the convolutional neural network model and continue iterative training until the loss value is within a preset accuracy range, end the iterative training and update the convolutional neural network model.

7. The micro-expression analysis method based on facial key point recognition according to claim 1, characterized in that: The step of inputting the two adjacent reference image feature vectors into a micro-expression synthesis module to obtain a micro-expression feature attribute vector specifically includes: Performing expression extraction and label extraction on the feature vectors of the two adjacent reference images; Generate corresponding expression parameters and label parameters according to the results of expression extraction and label extraction; The expression parameters and label parameters are associated and fused to obtain a micro-expression feature attribute vector.

Citation Information

Patent Citations

  • Micro-expression recognition method and device

    CN116543440A

  • Micro-expression recognition method based on parameter migration and optical flow feature extraction

    CN117058735A

  • Micro-expression recognition method and system based on regional weighted optical flow features

    CN117197877A

  • AU2020102556A4