Facial micro-expression detection method and device, electronic equipment, storage medium and program product
By extracting and processing the video set data of facial micro-expressions, multi-dimensional spectrum features are generated, and optimization processing based on reconstruction loss function, the problem of the inability to accurately identify facial micro-expressions in the prior art is solved, and efficient and accurate facial micro-expressions detection is achieved.
Patent Information
- Application Number
- CN202510198022.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
The existing facial expression recognition method based on LTP features is poor in the recognition effect when processing micro-expressions, and it is impossible to accurately identify facial expressions.
By obtaining the video set data of facial micro-expressions, converting it into image frame data, and pre-processing the image frame data to extract multi-dimensional timing features. These features are then processed using the Mel frequency cepspectral coefficients to generate multidimensional spectral features. Finally, the multi-dimensional spectrum features are input into the initial facial micro-expression detection model, and the optimized facial micro-expression detection model is obtained through optimization processing based on the reconstruction loss function.
It improves the accuracy of facial micro-expression detection, enhances the generalization ability of the model, and achieves the effect of efficient and accurate detection of facial micro-expression.
Smart Images

Figure CN120126196A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular, to a method, device, electronic device, storage medium and program product for detecting facial micro-expressions. Background Art
[0002] Facial micro-expressions are an important non-verbal way to convey emotions and intentions. Therefore, facial expression detection plays a key role in social interaction and human-computer interaction.
[0003] Existing face expression recognition methods based on LTP (Local Ternary Pattern) features can capture the features of facial expressions to a certain extent, but when dealing with micro-expressions, which are subtle and rapidly changing expressions, their recognition effect is not good, and they cannot accurately recognize facial micro-expressions.
[0004] Therefore, there is an urgent need for a solution that can accurately detect facial micro-expressions. Summary of the Invention
[0005] The present application provides a method, device, electronic device, storage medium and program product for detecting facial micro-expressions, so as to solve the technical problem of being unable to accurately detect changes in facial micro-expressions.
[0006] In a first aspect, the present application provides a method for detecting facial micro-expressions, including:
[0007] Obtaining video set data of facial micro-expressions;
[0008] Inputting the video set data into a facial micro-expression detection model to obtain a facial micro-expression detection result output by the facial micro-expression detection model; wherein, the facial micro-expression detection model is obtained by training on video set data of facial micro-expressions;
[0009] Among them, the video set data is used to be converted into image frame data; the image frame data is used to perform data preprocessing on the image frame data to obtain multi-dimensional time series features; wherein, the multi-dimensional time series features represent the time series features of the image frame data in different dimensions; the multi-dimensional time series features are used to use Mel Frequency Cepstral Coefficients to process the multi-dimensional time series features to obtain multi-dimensional spectral features; the multi-dimensional spectral features are used to input the multi-dimensional spectral features into an initial facial micro-expression model, and based on a reconstruction loss function, process the facial micro-expression detection model to obtain the facial micro-expression detection model.
[0010] In a second aspect, the present application provides a device for detecting facial micro-expressions, including:
[0011] A data acquisition module, configured to acquire video set data of facial micro-expressions;
[0012] An expression detection module, configured to input the video set data into a facial micro-expression detection model to obtain a facial micro-expression detection result output by the facial micro-expression detection model; wherein, the facial micro-expression detection model is obtained by training on video set data of facial micro-expressions;
[0013] Among them, the video set data is used to be converted into image frame data; the image frame data is used to perform data preprocessing on the image frame data to obtain multi-dimensional time series features; wherein, the multi-dimensional time series features represent the time series features of the image frame data in different dimensions; the multi-dimensional time series features are used to process the multi-dimensional time series features using Mel-frequency cepstral coefficients to obtain multi-dimensional spectral features; the multi-dimensional spectral features are used to input the multi-dimensional spectral features into an initial facial micro-expression model, and based on a reconstruction loss function, process the facial micro-expression detection model to obtain the facial micro-expression detection model.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;
[0015] The memory stores computer-executable instructions.
[0016] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the methods provided in the first aspect and any implementation manner of the first aspect.
[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the above first aspect and / or various possible implementation manners of the first aspect.
[0018] In a fifth aspect, an embodiment of the present application provides a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the methods provided in the first aspect and any implementation manner of the first aspect.
[0019] A method, device, electronic device, storage medium, and program product for detecting facial microexpressions provided by an embodiment of the present application obtain video set data of facial microexpressions and input the data into a specially trained facial microexpression detection model to obtain facial microexpression detection results. The facial microexpression model is obtained by training video set data of facial microexpressions. The training process involves converting the video set data into image frame data, then performing data preprocessing on the image frame data, and extracting multi-dimensional temporal features that can characterize the temporal changes of the image frame data in different dimensions. Subsequently, the multi-dimensional temporal features are processed using Mel Frequency Cepstral Coefficients to generate multi-dimensional spectral features. Finally, the multi-dimensional spectral features are input into the initial facial microexpression model, and through optimization processing based on the reconstruction loss function, an optimized facial microexpression detection model is obtained. This method not only improves the accuracy of facial microexpression detection but also enhances the generalization ability of the model. It achieves the effect of efficiently and accurately detecting facial microexpressions. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0021] Figure 1 Schematic flowchart of a method for detecting facial microexpressions provided by the present application Figure 1 ;
[0022] Figure 2 Schematic flowchart of a method for detecting facial microexpressions provided by the present application Figure 2 ;
[0023] Figure 3 Schematic flowchart of step S202 in a method for detecting facial microexpressions provided by an embodiment of the present application;
[0024] Figure 4 Schematic structure diagram of a device for detecting facial microexpressions provided by an embodiment of the present application Figure 1 ;
[0025] Figure 5 Schematic structure diagram of a device for detecting facial microexpressions provided by an embodiment of the present application Figure 2 ;
[0026] Figure 6 Schematic structure diagram of an electronic device provided by an embodiment of the present application.
[0027] Through the above-mentioned accompanying drawings, specific embodiments of the present application have been shown, and will be described in more detail hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Description of Specific Embodiments
[0028] Here, exemplary embodiments will be described in detail, and examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0029] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, complies with the relevant laws, regulations, and standards of relevant countries and regions, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or refuse.
[0030] In addition, the present application involves big data analysis of user information (including but not limited to personal biometric characteristics, identity data, consumption data, asset data, electronic terminal operation data, etc.), and uses artificial intelligence technology for automated decision-making. For technical solutions that make decisions having a significant impact on personal rights and interests based on the results of automated decision-making, corresponding operation entrances are provided for users to choose to agree or refuse the results of automated decision-making; if the user chooses to refuse, the expert decision-making process will be entered.
[0031] It should be noted that the methods, devices, equipment, storage media, and products for facial microexpressions provided by the present application can be used in the field of artificial intelligence or any field other than artificial intelligence. The application fields of the facial microexpression detection methods, devices, equipment, storage media, and products in the present application are not limited.
[0032] The development of existing facial expression analysis technologies is of extremely important significance. In human communication, facial expressions are important non-verbal ways to convey emotions and intentions, and they can convey more delicate and real emotional information than words. Therefore, facial expression analysis technologies play a key role in social interaction and human-computer interaction.
[0033] However, although certain progress has been made in facial expression analysis technology, the existing technology still faces a significant technical problem: the inability to accurately recognize facial micro-expressions. Micro-expressions are subtle changes in facial expressions that last for a short time and flash by, and they are often more difficult to detect and recognize than regular expressions. Existing face micro-expression recognition methods based on LTP (Local Ternary Pattern) features, although able to capture the features of facial expressions to a certain extent, have poor recognition effects when dealing with such subtle and rapidly changing micro-expressions and often cannot accurately recognize face micro-expressions. This is mainly because the subtle features and rapid changes of micro-expressions pose extremely high requirements on the accuracy and speed of the recognition algorithm, and the existing recognition algorithms still have deficiencies in capturing these features.
[0034] A facial micro-expression detection method, device, electronic device, storage medium, and program product provided by this application aim to solve the above technical problems of the existing technology.
[0035] The technical solutions of this application and how the technical solutions of this application solve the above technical problems will be described in detail below with specific embodiments. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below in conjunction with the accompanying drawings.
[0036] Figure 1 Flow schematic of a facial micro-expression detection method provided by this application Figure 1 , as Figure 1 shown, this method includes:
[0037] S101. Obtain video set data of facial micro-expressions.
[0038] Exemplarily, obtain video set data of facial micro-expressions.
[0039] S102. Input the video set data into a facial micro-expression detection model to obtain a facial micro-expression detection result output by the facial micro-expression detection model; wherein, the facial micro-expression detection model is obtained by training on the video set data of facial micro-expressions.
[0040] Exemplarily, input the video set data into a facial micro-expression detection model to obtain a facial micro-expression detection result output by the facial micro-expression detection model; wherein, the facial micro-expression detection model is the facial micro-expression detection model trained in the following embodiments.
[0041] In an exemplary embodiment, the training process of the facial micro-expression detection model includes the following steps:
[0042] Use video set data to convert it into image frame data; the image frame data is used to perform data preprocessing on the image frame data to obtain multi-dimensional temporal features; wherein, the multi-dimensional temporal features represent the temporal features of the image frame data in different dimensions; the multi-dimensional temporal features are used to process the multi-dimensional temporal features using Mel Frequency Cepstral Coefficients to obtain multi-dimensional spectral features; the multi-dimensional spectral features are used to input the multi-dimensional spectral features into an initial facial micro-expression detection model, and process the initial facial micro-expression detection model based on a reconstruction loss function to obtain a facial micro-expression detection model.
[0043] Exemplarily, use a high-speed camera with a frame rate above 200fps to capture the changing video of human facial micro-expressions, and screen and convert the collected video into image frames to obtain the facial micro-expression data of the image frames. Facial micro-expressions include six basic emotions: anger, disgust, fear, happiness, sadness, and surprise. Based on these six basic emotions, label the facial micro-expression data of the image frames. Exemplarily, perform three types of processing on the image frame data respectively, convert the image data into temporal features; after obtaining the multi-dimensional temporal features, use Mel Frequency Cepstral Coefficients to process the multi-dimensional temporal features to obtain multi-dimensional spectral features; the multi-dimensional spectral features are used to input the multi-dimensional spectral features into an initial facial micro-expression detection model, and process the initial facial micro-expression detection model based on a reconstruction loss function to obtain a facial micro-expression detection model.
[0044] A facial micro-expression detection method provided in this embodiment obtains the video set data of pre-facial micro-expressions and inputs these data into a specially trained facial micro-expression detection model to obtain the facial micro-expression detection result. The facial micro-expression detection model is obtained by training the video set data of facial micro-expressions. The training process involves converting the video set data into image frame data, and then performing data preprocessing on the image frame data to extract multi-dimensional temporal features, which can represent the temporal changes of the image frame data in different dimensions. Subsequently, use Mel Frequency Cepstral Coefficients to process these multi-dimensional temporal features to generate multi-dimensional spectral features. Finally, input the multi-dimensional spectral features into the initial facial micro-expression detection model, and through the optimization process based on the reconstruction loss function, obtain an optimized facial micro-expression detection model. This method not only improves the accuracy of facial micro-expression detection but also enhances the generalization ability of the model. It achieves the effect of efficiently and accurately detecting the facial micro-expression state.
[0045] Figure 2 It is a flowchart of a facial micro-expression detection method provided by this application Figure 2 , such as Figure 2 shown, based on the Figure 1 embodiment, a facial micro-expression detection method is described in detail. The method includes:
[0046] S201. Obtain the video set data of facial micro-expressions; use the video set data to convert it into image frame data.
[0047] Exemplarily, a video is composed of a series of consecutive static images, and each frame is actually a picture. Therefore, converting a video into image frame data means decomposing the video into its individual constituent frames. Perform processing such as resizing, color space conversion, denoising, and normalization on the image frame data.
[0048] S202. Use the image frame data to perform data preprocessing on the image frame data to obtain multi-dimensional temporal features; among them, the multi-dimensional temporal features characterize the temporal features of the image frame data in different dimensions. The multi-dimensional temporal features include a head pose sequence, a gaze direction sequence, and a gaze coordinate sequence. Among them, the head pose sequence characterizes the head pose coordinate values in a three-dimensional coordinate system; the gaze direction sequence characterizes the gaze direction of the eyes in a spherical coordinate system; the gaze coordinate sequence characterizes the coordinates of the eyes looking at the display screen in a two-dimensional coordinate system.
[0049] Exemplarily, use different data processing methods to process the image frame data to obtain multi-dimensional temporal features; among them, the multi-dimensional temporal features represent the temporal features of the image frame data in different dimensions. The multi-dimensional temporal features include a head pose sequence, a gaze direction sequence, and a gaze coordinate sequence. Among them, the head pose sequence is the head pose coordinate value in a three-dimensional coordinate system; the gaze direction sequence is the gaze direction of the eyes in a spherical coordinate system; the gaze coordinate sequence is the coordinates of the eyes looking at the display screen in a two-dimensional coordinate system.
[0050] In one example, Figure 3 is a schematic flowchart of step S202 in a facial micro-expression detection method provided by an embodiment of the present application, as Figure 3 shown, S202 includes:
[0051] S2021. Input the image frame data into a head pose estimation model to obtain a head pose angle; among them, the head pose angle characterizes the head pose offset angle with the head facing forward as the benchmark; convert the head pose angle into a head pose coordinate value through a normalization exponential function to obtain a head pose sequence.
[0052] Exemplarily, the image frame data is input into an open-source head pose estimation model, such as the HopeNet model, which can accurately detect and estimate the position and orientation of a person's head in an image or video. Through this process, the head pose angles are obtained based on the pose offset angles of the head relative to the standard straight-ahead position. To obtain more discriminative feature categories, these head pose angles are further subdivided into different discrete angle intervals. Subsequently, by applying the softmax function, these feature categories are converted into probability values, thus forming a series of continuous and ordered head pose sequences. This series of data not only accurately records the pose changes of the head but also provides strong support for subsequent analysis and processing.
[0053] By continuously capturing and analyzing the head pose sequences in the image frames, the model can finely quantify the degree of facial micro-expressions. Minor changes in head pose, such as slight nodding, frequent blinking, or unconscious head shaking, can be accurately captured and recorded, providing more detailed data support for detecting the micro-expression state.
[0054] S2022. Based on a neural network, perform image segmentation on the image frame data to obtain head segmentation data; among them, the head segmentation data needs to retain complete head information; use a head detection model to verify the head segmentation data; if the verification passes, input the head segmentation data into a residual network to obtain facial feature data; input the facial feature data into a bidirectional long short-term memory network to obtain a gaze direction sequence; among them, the facial feature data includes facial geometric structures and local eye features; if the verification fails, re-perform image segmentation on the image frame data.
[0055] Exemplarily, first, the image frame data is taken as input and input into the open-source Faster-RCNN model. The model is used to segment the image, and then the head segmentation data is obtained. Next, the head detection model is downloaded and loaded from open-source resources. The main function of this model is to verify the integrity of the head segmentation data. Among them, the head detection model can be one of the following models, and this embodiment does not limit: MTCNN (Multi-task Cascaded Convolutional Networks), YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector). Subsequently, the head detection model is used to verify the head segmentation data. Specifically, it is to check whether the head information is completely retained, and the verification criteria cover multiple aspects, such as whether the head contour is complete and whether the internal features of the head are retained (such as eyes, nose, mouth, etc.). If the verification fails, that is, the head segmentation data is incomplete or key information is lost, then it is necessary to readjust the image preprocessing parameters or select other segmentation algorithms and perform the image segmentation operation again. If the verification passes, according to the input requirements of the residual network, further preprocessing of the head segmentation data is required. Operations such as resizing and cropping are necessary. After the preprocessing is completed, the head segmentation data is input into the open-source ResNet-18 model, and 256-dimensional facial features are extracted through this model. On this basis, the extracted facial feature data is input into the bidirectional long short-term memory network for training, and finally the gaze direction sequence is obtained. The gaze direction includes pitch angle, yaw angle, and offset number. Among them, the pitch angle is the change angle of the eye's line of sight along the vertical axis (up and down direction). When the eyes look up, the pitch angle is positive; when looking down, the pitch angle is negative. It reflects the inclination degree of the line of sight relative to the horizontal plane. The yaw angle reflects the change angle of the eye's line of sight along the horizontal axis (left and right direction). When the eyes look to the right, the yaw angle is positive; when looking to the left, the yaw angle is negative. This indicates the rotation degree of the line of sight on the horizontal plane. The offset number is the offset index for gaze direction prediction.
[0056] Initial segmentation is performed using Faster-RCNN, and then a head detection model (such as MTCNN, YOLO, or SSD) is used for verification, ensuring the integrity and accuracy of the head segmentation data. This double-check mechanism reduces the false alarm rate and improves the ability to capture real head information. As a powerful convolutional neural network, ResNet-18 can extract rich 256-dimensional facial features from the head segmentation data. These features not only contain facial structure information but may also imply subtle clues such as expression changes, which helps to detect microexpressions more accurately. Through a bidirectional long short-term memory network (BiLSTM), the system can not only consider the static features at a single time point but also combine historical and future information to model the time series of the gaze direction. This method is particularly suitable for capturing the attention drift phenomenon over a long period and is an important basis for judging facial microexpressions.
[0057] S2023. Extract eye data from the image frame data based on the eye tracker, where the eye data includes eye position, eye direction, eye size, and eye pupil morphology; generate an eye vector according to the eye data; establish a mapping function between the eye and the display according to the eye vector; and obtain the gaze coordinate sequence according to the mapping function and the head pose angle.
[0058] Exemplarily, first, the image frame data needs to be preprocessed. Specifically, the color image therein is converted into a grayscale image, and denoising algorithms such as Gaussian filtering are used to reduce the noise in the image, so as to improve the image quality and lay a good foundation for subsequent operations. Then, a Haar cascade classifier or a deep learning model (such as SSD, YOLO, etc.) is used to detect the face region in the image. After successfully detecting the face region, the specific positions of the eyes are further located using eye features (such as color, shape, etc.). This step is an important prerequisite for subsequent precise analysis of eye-related data. Then, more detailed processing is carried out on the eye region. Methods such as Canny edge detection or Hough transform are used to extract the contour of the eye region. After obtaining the eye contour, the pupil position is located by methods such as gray value analysis or template matching. On this basis, according to the geometric features of the pupil position and the eye contour, the motion vector of the eye (such as the offset of the pupil center relative to the eye center) is calculated. Through this series of steps, the key data of the eye vector is obtained. After that, in order to establish a connection between the eye vector and the screen fixation point, a mapping model needs to be established based on the geometric relationship between the eye vector and the screen fixation point. Common models include the corneal reflection model, the pupil-corneal center model, etc. This model is a bridge for subsequent accurate estimation of the fixation point. At the same time, the head pose also affects the fixation point estimation. Therefore, a deep learning model (such as a convolutional neural network) or a traditional method (such as feature point-based pose estimation) is used to detect the head pose, and then, according to the detected head pose, the eye vector is corrected to eliminate the influence of head pose changes on the fixation point estimation and ensure the accuracy and effectiveness of the eye vector data. Finally, the corrected eye vector is combined with the screen fixation point mapping model to achieve more accurate fixation point estimation, and the estimated fixation point coordinates are arranged in chronological order to generate a fixation coordinate sequence.
[0059] By using a Haar cascade classifier or a deep learning model (SSD, YOLO, etc.) to detect the face region and then locating the specific positions of the eyes based on eye features, this step-by-step and targeted detection and location method can accurately lock the target region, find the entry point for more detailed subsequent eye-related analysis, and improve the overall location accuracy. Using methods such as Canny edge detection or Hough transform to extract the eye contour, then locating the pupil position by gray value analysis or template matching, and further calculating the eye motion vector, the progressive and detailed processing makes the acquisition of key eye data more accurate, which helps to accurately estimate the fixation point based on the eye condition subsequently.
[0060] S203. Perform frame segmentation on the multi-dimensional time series features to obtain unit time series features; multiply the unit time series features by a Hamming window and then perform a fast Fourier transform to obtain the spectral signal of the unit time series features; perform modulus square processing on the spectral signal of the unit time series features to obtain a power spectrum; input the power spectrum into a filter bank to obtain logarithmic energy; perform a discrete cosine transform on the logarithmic energy to obtain the first spectral parameter; splice the first spectral parameters to obtain multi-dimensional spectral features.
[0061] Exemplarily, the MFCC (Mel-Frequency Cepstral Coefficients) method is used to process the multi-dimensional time series features. The following is the specific processing process:
[0062] Let the original multi-dimensional time series feature be X(t), where t is the time index. Divide X(t) into frames of length N, usually with a certain overlap (such as 50% or 20%). Each frame can be represented as X i =X(t i-N / 2 ,…t i+N / 2 ), where i is the center point of the i-th frame.
[0063] Apply each frame X i to the Hamming window W(n) to obtain the windowed frame X i ′ ; that is, X i ′ =X i ·W(n), where n = 0, 1, … N - 1; W(n) = 0.54 - 0.46 × cos[2πn / (N - 1)].
[0064] Perform a fast Fourier transform on each frame X i ′ to obtain the frequency domain representation X i ″ (K); The calculation formula is as follows:
[0065]
[0066] where K = 0, 1, …, N - 1.
[0067] Take the modulus square of the frequency domain representation X i ″ (K) to obtain the power spectrum P i (k); that is, P i (k)=|X i ″ (K)| 2 .
[0068] Input the power spectrum P i(k), the input filter bank H m (f), to obtain the logarithmic energy E im ;
[0069]
[0070] where m = 1, 2, …, M; m represents the number of filters; f represents the frequency.
[0071] For the logarithmic energy E im perform a discrete cosine transform (DCT) to obtain the first spectral parameter C im ;
[0072]
[0073] where n = 1, 2, …, L; the L - order refers to the order of MFCC coefficients, usually taking 12 - 16; M is the number of filters.
[0074] Finally, concatenate the first spectral parameters of all frames in chronological order to form a multi - dimensional spectral feature vector.
[0075] The features extracted by using MFCC are highly robust to noise and linear filtering. MFCC can help reduce fluctuations caused by external factors (such as light changes, camera jitter, etc.), thereby improving the stability of features. Through MFCC processing, stable and representative features can be extracted from complex head movements.
[0076] S204. Input the multi - dimensional spectral features into the initial facial micro - expression detection model, and process the initial facial micro - expression detection model based on the reconstruction loss function to obtain the reconstructed spectral features; calculate the reconstruction loss value according to the multi - dimensional spectral features, the reconstructed spectral features, and the reconstruction loss function; process the initial facial micro - expression detection model based on the reconstruction loss value to obtain the facial micro - expression detection model.
[0077] Exemplarily, for the reconstruction loss function, its calculation formula is as follows:
[0078]
[0079] where N is the number of points of the multi - dimensional spectral features, is the point corresponding to a certain moment in the sequence of the multi - dimensional spectral feature reconstruction; x i is the point corresponding to a certain moment in the multi - dimensional spectral feature sequence.
[0080] Input the multi-dimensional spectral features into the autoencoder Transformer model of the self-attention mechanism for training to obtain the reconstructed spectral features; the reconstructed spectral features are the features predicted by the model. Use the reconstruction loss function to calculate the difference between the original multi-dimensional spectral features and the spectral features reconstructed by the model. This loss value reflects the ability of the model to reconstruct the input data and is also an important basis for subsequent model optimization and threshold adjustment. Input the multi-dimensional spectral features into the initial model, and through forward propagation, obtain the reconstructed spectral features output by the model. Calculate the loss value according to the reconstruction loss function and use the backpropagation algorithm to update the model parameters. Repeat the above process until the model converges or reaches the preset number of training epochs to obtain the final facial micro-expression detection model. This model can accurately identify and reconstruct multi-dimensional spectral features, thereby realizing the effective detection of facial micro-expressions.
[0081] By minimizing the difference between the input data and the model output (i.e., the reconstructed data), it can be ensured that the model learns the essential features of the input data, which helps the model to more accurately identify facial micro-expressions.
[0082] Figure 4 The structural schematic of a facial micro-expression detection device provided by an embodiment of the present application Figure 1 , as Figure 4 shown, a facial micro-expression detection device 40 provided in this embodiment includes:
[0083] A data acquisition module 401, configured to acquire video set data of facial micro-expressions.
[0084] An expression detection module 402, configured to input the video set data into the facial micro-expression detection model to obtain the facial micro-expression detection result output by the facial micro-expression detection model; wherein, the facial micro-expression detection model is obtained by training on the video set data of facial micro-expressions.
[0085] Among them, the video set data is used to be converted into image frame data; the image frame data is used to perform data preprocessing on the image frame data to obtain multi-dimensional temporal features; wherein, the multi-dimensional temporal features characterize the temporal features of the image frame data in different dimensions; the multi-dimensional temporal features are used to process the multi-dimensional temporal features using mel-frequency cepstral coefficients to obtain multi-dimensional spectral features; the multi-dimensional spectral features are used to input the multi-dimensional spectral features into the initial facial micro-expression detection model to process the initial facial micro-expression detection model based on the reconstruction loss function to obtain the facial micro-expression detection model.
[0086] The device provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.
[0087] Figure 5Structural schematic of a facial micro-expression detection device provided by an embodiment of the present application Figure 2 , such as Figure 5 shown, a facial micro-expression detection device 50 provided in this embodiment includes:
[0088] A data acquisition module 501, configured to acquire video set data of facial micro-expressions;
[0089] An expression detection module 502, configured to input the video set data into a facial micro-expression detection model to obtain a facial micro-expression detection result output by the facial micro-expression detection model; wherein, the facial micro-expression detection model is obtained by training on the video set data of facial micro-expressions;
[0090] Among them, the video set data is used to be converted into image frame data; the image frame data is used to perform data preprocessing on the image frame data to obtain multi-dimensional temporal features; wherein, the multi-dimensional temporal features represent the temporal features of the image frame data in different dimensions; the multi-dimensional temporal features are used to use Mel Frequency Cepstral Coefficients to process the multi-dimensional temporal features to obtain multi-dimensional spectral features; the multi-dimensional spectral features are used to input the multi-dimensional spectral features into an initial facial micro-expression detection model, and to process the initial facial micro-expression detection model based on a reconstruction loss function to obtain a facial micro-expression detection model.
[0091] In one example, the expression module 502 is specifically configured to:
[0092] Perform frame segmentation on the multi-dimensional temporal features to obtain unit temporal features;
[0093] Multiply the unit temporal features by a Hamming window and then perform a fast Fourier transform to obtain a spectral signal of the unit temporal features;
[0094] Perform modulus square processing on the spectral signal of the unit temporal features to obtain a power spectrum; input the power spectrum into a filter bank to obtain logarithmic energy; perform a discrete cosine transform on the logarithmic energy to obtain a first spectral parameter;
[0095] Perform splicing on the first spectral parameters to obtain multi-dimensional spectral features.
[0096] In one example, the expression module 502 further includes:
[0097] The multi-dimensional temporal features include a head pose sequence, a gaze direction sequence, and a gaze coordinate sequence, wherein the head pose sequence represents the head pose coordinate values in a three-dimensional coordinate system; the gaze direction sequence represents the gaze direction of the eyes in a spherical coordinate system; the gaze coordinate sequence represents the coordinates of the eyes gazing at the display screen in a two-dimensional coordinate system.
[0098] In one example, the expression module 502 is specifically configured to:
[0099] Input the image frame data into the head pose estimation model to obtain the head pose angles; wherein, the head pose angles represent the head pose offset angles with respect to the benchmark when the human head is upright and facing forward.
[0100] Convert the head pose angles into head pose coordinate values through the sigmoid function to obtain the head pose sequence.
[0101] In one example, the expression module 502 is specifically configured to:
[0102] Based on a neural network, perform image segmentation on the image frame data to obtain head segmentation data; wherein, the head segmentation data needs to retain complete head information.
[0103] Use a head detection model to verify the head segmentation data.
[0104] If the verification passes, input the head segmentation data into a residual network to obtain facial feature data; input the facial feature data into a bidirectional long short-term memory network to obtain the gaze direction sequence; wherein, the facial feature data includes facial geometric structures and local eye features.
[0105] If the verification fails, re-perform image segmentation on the image frame data.
[0106] In one example, the expression module 502 is specifically configured to:
[0107] Based on an eye tracker, extract eye data from the image frame data, wherein the eye data includes eye position, eye direction, eye size, and eye pupil shape.
[0108] Generate an eye vector based on the eye data; establish a mapping function between the eye and the display based on the eye vector; obtain the gaze coordinate sequence based on the mapping function and the head pose angles.
[0109] In one example, the expression module 502 is specifically configured to:
[0110] Input the multi-dimensional spectral features into an initial facial micro-expression detection model to process the initial facial micro-expression detection model based on a reconstruction loss function to obtain reconstructed spectral features.
[0111] Calculate a reconstruction loss value based on the multi-dimensional spectral features, the reconstructed spectral features, and the reconstruction loss function.
[0112] Process the initial facial micro-expression detection model based on the reconstruction loss value to obtain a facial micro-expression detection model.
[0113] The device provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0114] Figure 6 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 6 shown, the electronic device 60 provided in this embodiment includes: at least one processor 601 and a memory 602. Optionally, the device 60 further includes a communication component 603. Among them, the processor 601, the memory 602, and the communication component 603 are connected through a bus 604.
[0115] In a specific implementation process, at least one processor 601 executes computer-executable instructions stored in the memory 602, so that at least one processor 601 executes the above-mentioned method.
[0116] For the specific implementation process of the processor 601, reference can be made to the above method embodiment. The implementation principle and technical effect are similar, and will not be elaborated here in this embodiment.
[0117] In the above embodiment, it should be understood that the processor may be a central processing unit (English: Central Processing Unit, abbreviated: CPU), and may also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated: DSP), application specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0118] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0119] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.
[0120] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the above-mentioned method.
[0121] The present application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-mentioned method.
[0122] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disk. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0123] An exemplary readable storage medium is coupled to the processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.
[0124] The division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0125] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0126] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0127] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0128] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When this program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0129] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0130] Furthermore, it should be noted that although the steps in the flowchart are displayed in sequence according to the indication of the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0131] It should be understood that the above device embodiments are merely illustrative, and the devices of the present application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units, modules, or components can be combined, or integrated into another system, or some features can be ignored or not executed.
[0132] In addition, without special instructions, in each embodiment of the present application, each functional unit / module can be integrated in one unit / module, or each unit / module can exist physically alone, or two or more units / modules can be integrated together. The above integrated unit / module can be implemented in the form of hardware or in the form of a software program module.
[0133] When the integrated unit / module is implemented in the form of hardware, the hardware can be a digital circuit, an analog circuit, etc. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc. Without special instructions, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Without special instructions, the storage unit can be any suitable magnetic storage medium or magneto-optical storage medium, such as a resistive random access memory (RRAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), an enhanced dynamic random access memory (EDRAM), a high-bandwidth memory (HBM), a hybrid memory cube (HMC), etc.
[0134] When the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a memory and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.
[0135] In the above embodiments, the descriptions of the various embodiments each have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.
[0136] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0137] It should be understood that the present application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A facial micro-expression detection method, characterized in that: include: Obtain video data of facial micro-expressions; Inputting the video set data into a facial micro-expression detection model to obtain a facial micro-expression detection result output by the facial micro-expression model; wherein the facial micro-expression detection model is obtained by training the video set data of facial micro-expressions; Wherein, the video set data is used to convert into image frame data; the image frame data is used to perform data preprocessing on the image frame data to obtain multi-dimensional time series features; wherein the multi-dimensional time series features characterize the time series features of the image frame data in different dimensions; the multi-dimensional time series features are used to process the multi-dimensional time series features using Mel-frequency cepstral coefficients to obtain multi-dimensional spectrum features; the multi-dimensional spectrum features are used to input the multi-dimensional spectrum features into an initial facial micro-expression detection model, so as to process the initial facial micro-expression detection model based on a reconstruction loss function to obtain the facial micro-expression detection model.
2. The method according to claim 1, characterized in that: The multi-dimensional time series features are used to process the multi-dimensional time series features using Mel-frequency cepstral coefficients to obtain multi-dimensional spectrum features, including: Performing frame processing on the multi-dimensional time series features to obtain unit time series features; Multiplying the unit time series feature by a Hamming window, and then performing a fast Fourier transform to obtain a frequency spectrum signal of the unit time series feature; Performing modulo square processing on the frequency spectrum signal of the unit time series feature to obtain a power spectrum; inputting the power spectrum into a filter bank to obtain logarithmic energy; performing discrete cosine transform on the logarithmic energy to obtain a first spectrum parameter; The first spectrum parameters are spliced to obtain multi-dimensional spectrum features.
3. The method according to claim 1, characterized in that Also includes: The multi-dimensional time series features include a head posture sequence, a gaze direction sequence and a gaze coordinate sequence, wherein the head posture sequence represents the head posture coordinate value in a three-dimensional coordinate system; The gaze direction sequence represents the gaze direction of the eyes in a spherical coordinate system; the gaze coordinate sequence represents the coordinates of the display screen where the eyes are gazed in a two-dimensional coordinate system.
4. The method according to claim 3, characterized in that The image frame data is used to perform data preprocessing on the image frame data to obtain multi-dimensional time series features, including: Inputting the image frame data into a head posture estimation model to obtain a head posture angle; wherein the head posture angle represents a head posture deviation angle based on a person's head being straight and facing forward; The head posture angle is converted into a head posture coordinate value through a normalized exponential function to obtain a head posture sequence.
5. The method according to claim 3, characterized in that: The image frame data is used to perform data preprocessing on the image frame data to obtain multi-dimensional time series features, including: Based on a neural network, the image frame data is segmented to obtain head segmentation data; wherein the head segmentation data needs to retain complete head information; Using a head detection model, verifying the head segmentation data; If the verification is passed, the head segmentation data is input into the residual network to obtain facial feature data; the facial feature data is input into the bidirectional long short-term memory network to obtain the gaze direction sequence; wherein the facial feature data includes facial geometric structure and local features of the eyes; If the verification fails, the image frame data is re-segmented.
6. The method according to claim 4, characterized in that The image frame data is used to perform data preprocessing on the image frame data to obtain multi-dimensional time series features, including: Extracting eye data from the image frame data based on a gaze tracker, wherein the eye data includes eye position, eye direction, eye size, and eye pupil morphology; An eye vector is generated according to the eye data; a mapping function between the eye and the display is established according to the eye vector; and the gaze coordinate sequence is obtained according to the mapping function and the head posture angle.
7. The method according to any one of claims 1 to 6, characterized in that The multi-dimensional spectrum feature is used to input the multi-dimensional spectrum feature into an initial facial micro-expression detection model, so as to process the initial facial micro-expression detection model based on a reconstruction loss function to obtain the facial micro-expression detection model, including: Inputting the multi-dimensional spectrum feature into an initial facial micro-expression detection model to process the initial facial micro-expression detection model based on a reconstruction loss function to obtain a reconstructed spectrum feature; Calculating a reconstruction loss value according to the multi-dimensional spectrum feature, the reconstructed spectrum feature and the reconstruction loss function; The initial facial micro-expression detection model is processed based on the reconstruction loss value to obtain a facial micro-expression detection model.
8. A facial micro-expression detection device, comprising: A data acquisition module, used to acquire video set data of facial micro-expressions; The facial expression detection module is used to input the video set data into the facial micro-expression detection model to obtain the facial micro-expression detection results output by the facial micro-expression detection model; wherein the facial micro-expression detection model is obtained by training the video set data of facial micro-expressions; Wherein, the video set data is used to convert into image frame data; the image frame data is used to perform data preprocessing on the image frame data to obtain multi-dimensional time series features; wherein the multi-dimensional time series features characterize the time series features of the image frame data in different dimensions; the multi-dimensional time series features are used to process the multi-dimensional time series features using Mel-frequency cepstral coefficients to obtain multi-dimensional spectrum features; the multi-dimensional spectrum features are used to input the multi-dimensional spectrum features into an initial facial micro-expression detection model, so as to process the initial facial micro-expression detection model based on a reconstruction loss function to obtain the facial micro-expression detection model.
9. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.
11. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 7 when being executed by a processor.