Expression detection method and device based on multilayer convolutional neural network, and storage medium
Through the expression detection method based on multi-layer convolutional neural network, the preprocessing and deep learning network model is used to solve the problem of inaccurate range of expression frames in the prior art, and high-accurate expression recognition is achieved.
Patent Information
- Application Number
- CN202510169050.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, it is not accurate to determine the occurrence range of expression frames in videos, especially the duration of micro-expressions is short and the intensity is low, making it difficult to accurately identify them.
The expression detection method based on a multi-layer convolutional neural network is adopted. By obtaining image data from the video data to be detected, preprocessing and optical flow feature extraction, the expression score prediction is performed using the trained deep learning network model to determine the occurrence range of expression frames.
It improves the accuracy of expression recognition and can accurately determine the range of expression frames in the video stream, especially the detection of micro-expressions.
Smart Images

Figure CN120260094A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of facial expression detection, and particularly to a facial expression detection method, device and storage medium based on a multi-layer convolutional neural network. Background Art
[0002] In related technologies, facial expression analysis can be applied to various scenarios, such as lie detection, psychology, and healthcare. Facial expressions are conveyed and perceived through the movement of facial muscles, and are a form of non-verbal communication used to provide visual information about an individual's emotional state. According to the intensity and duration, facial expressions can be divided into two categories: macro expressions (MaEs) and micro expressions (MEs). MaEs have a higher intensity and a longer duration, usually lasting 0.5 - 4s. In contrast, MEs have a very short duration, usually within 0.5s, and a very low intensity, and are more unconscious and difficult to detect with the naked eye. Existing technologies usually use artificially designed features to analyze the feature differences between different frames in a video, and then determine the expression frames in the video. However, due to the characteristics of the intensity and duration of facial expressions, the accuracy of the appearance range of the determined expression frames in the existing technologies is not high.
[0003] In summary, the technical problems existing in the related technologies need to be improved. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to propose a facial expression detection method, device and storage medium based on a multi-layer convolutional neural network, which can accurately determine the appearance range of expression frames in a video stream.
[0005] To achieve the above purpose, on the one hand, an embodiment of the present application proposes a facial expression detection method based on a multi-layer convolutional neural network, and the method includes the following steps:
[0006] Obtain image data to be processed from the video data to be detected, and each frame of image in the image data to be processed includes a face image;
[0007] Preprocess each frame of image in the image data to be processed to obtain a target image data set;
[0008] Input the target image data set into a trained preset deep learning network model for facial expression score prediction to obtain a target facial expression score;
[0009] Determine the appearance range of the expression frames in the video data to be detected according to the target facial expression score;
[0010] Wherein, the preset deep learning network model includes four convolutional units, a global pooling layer, a convolutional layer, and a connection layer, and the four convolutional units, the global pooling layer, the convolutional layer, and the connection layer are connected in sequence;
[0011] Each of the convolutional units includes a first convolutional layer, a second convolutional layer, a feed-forward network, a first layer of normalization, and a second layer of normalization; the first layer of normalization, the first convolutional layer, the second layer of normalization, the second convolutional layer, and the feed-forward network are connected in sequence; the first convolutional layer is used for performing dynamic convolution operations, and the second convolutional layer is used for performing spatial channel attention convolution operations.
[0012] In some embodiments, obtaining the image data to be processed from the video data to be detected includes:
[0013] Obtaining the video data to be detected;
[0014] Performing face detection on the video data to be detected frame by frame through a face detector to obtain regions to be cropped;
[0015] Cropping all the regions to be cropped to obtain a number of first face images;
[0016] Performing pixel processing on the first face images to obtain second face images in the image data to be processed.
[0017] In some embodiments, preprocessing each frame of the image data to be processed to obtain a target image dataset includes:
[0018] Performing optical flow calculation on each frame of the image data to be processed to obtain a first optical flow feature image;
[0019] Processing the face feature points in the first optical flow feature image to obtain a second optical flow feature image;
[0020] Generating target pseudo-labels for each frame of the image data to be processed;
[0021] Combining the second optical flow feature image and the target pseudo-labels to form the target image dataset.
[0022] In some embodiments, processing the face feature points in the first optical flow feature image to obtain a second optical flow feature image includes:
[0023] Extracting a number of face feature points in the first optical flow feature image;
[0024] Performing blackening processing on the eye regions in the first optical flow feature image according to the face feature points to obtain a third optical flow feature image;
[0025] Determining a first upper half region of interest, a second upper half region of interest, and a lower half region of interest in the third optical flow feature image according to the face feature points;
[0026] Resize the first upper half region of interest, the second upper half region of interest, and the lower half region of interest;
[0027] Horizontally stack the resized first upper half region of interest and the second upper half region of interest to obtain the upper half region of interest;
[0028] Stitch the upper half region of interest and the resized lower half region of interest to obtain the second optical flow feature image.
[0029] In some embodiments, generating the target pseudo-label for each frame of the to-be-processed image data includes:
[0030] Obtain the expression label of the to-be-processed image data, where the expression label includes an expression start frame and an expression end frame, and the interval between the expression start frame and the expression end frame is used as the expression interval;
[0031] Generate a first linear pseudo-label corresponding to the target pseudo-label for each frame of the image within the expression interval through a linear function;
[0032] Generate a first Gaussian pseudo-label corresponding to the target pseudo-label for each frame of the image within the expression interval through a Gaussian function.
[0033] In some embodiments, the training process of the preset deep learning network model includes the following steps:
[0034] Obtain to-be-trained image data from the to-be-trained video data, where each frame of the to-be-trained image data includes a face image;
[0035] Preprocess each frame of the to-be-trained image data to obtain a target training set, where the target training set includes to-be-trained optical flow feature images, to-be-trained linear pseudo-labels, and to-be-trained Gaussian pseudo-labels;
[0036] Divide the target training set into a to-be-used training set and a test set;
[0037] Train the preset deep learning network model through the to-be-used training set, and calculate the total loss function during the training process, where the total loss function includes a linear loss function and a Gaussian loss function;
[0038] After determining that the preset deep learning network model is completed training according to the total loss function, enhance the output result of the trained preset deep learning network model through the test set.
[0039] In some embodiments, determining the occurrence range of expression frames in the to-be-detected video data according to the target expression score includes:
[0040] Calculate a standard threshold according to the target expression score;
[0041] Determine peak frames in the video data to be detected according to the standard threshold;
[0042] Calculate the minimum distance between the peak frames;
[0043] Determine the occurrence range of expression frames in the video data to be detected according to the peak frames and the minimum distance.
[0044] To achieve the above object, another aspect of the embodiments of the present application provides an expression detection device based on a multi-layer convolutional neural network. The device includes:
[0045] A first module, configured to obtain image data to be processed from video data to be detected, where each frame of image in the image data to be processed includes a face image;
[0046] A second module, configured to preprocess each frame of image in the image data to be processed to obtain a target image data set;
[0047] A third module, configured to input the target image data set into a trained preset deep learning network model for expression score prediction to obtain a target expression score;
[0048] A fourth module, configured to determine the occurrence range of expression frames in the video data to be detected according to the target expression score;
[0049] Wherein, the preset deep learning network model includes four convolutional units, a global pooling layer, a convolutional layer, and a connection layer, and the four convolutional units, the global pooling layer, the convolutional layer, and the connection layer are connected in sequence;
[0050] Each convolutional unit includes a first convolutional layer, a second convolutional layer, a feedforward network, a first normalization layer, and a second normalization layer; the first normalization layer, the first convolutional layer, the second normalization layer, the second convolutional layer, and the feedforward network are connected in sequence; the first convolutional layer is used for dynamic convolution operations, and the second convolutional layer is used for spatial channel attention convolution operations.
[0051] To achieve the above object, another aspect of the embodiments of the present application provides a computer device, including:
[0052] At least one processor;
[0053] At least one memory, configured to store at least one program;
[0054] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.
[0055] To achieve the above object, another aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above method is implemented.
[0056] The embodiments of the present application at least include the following beneficial effects: The present application provides an expression detection method, device and storage medium based on a multi-layer convolutional neural network. After obtaining the image data to be processed including face images from the video data to be detected, the method preprocesses each frame of the image data to be processed to obtain a target image data set; inputs the target image data set into a trained preset deep learning network model to predict the expression score to obtain the target expression score, and then determines the occurrence range of the expression frames in the video data to be detected according to the target expression score, so that it is possible to accurately determine the occurrence range of the expression frames in the video stream without relying on artificially designed features. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 is a flowchart of the expression detection method based on a multi-layer convolutional neural network provided by the embodiments of the present application;
[0058] Figure 2 is a schematic structural diagram of the preset deep learning network model provided by the embodiments of the present application;
[0059] Figure 3 is a schematic structural diagram of the TVPC block provided by the embodiments of the present application;
[0060] Figure 4 is a schematic structural diagram of the TVConv provided by the embodiments of the present application;
[0061] Figure 5 is a schematic structural diagram of the PRConv provided by the embodiments of the present application;
[0062] Figure 6 is a schematic diagram of the target pseudo-label provided by the embodiments of the present application;
[0063] Figure 7 is a curve graph of the expression prediction score before enhancement corresponding to the preset deep learning network model provided by the embodiments of the present application;
[0064] Figure 8 is a curve graph of the expression prediction score after enhancement corresponding to the preset deep learning network model provided by the embodiments of the present application;
[0065] Figure 9It is a schematic diagram of the standard threshold and the enhanced expression prediction score provided by the embodiments of the present application;
[0066] Figure 10 It is a schematic structural diagram of an expression detection device based on a multi-layer convolutional neural network provided by the embodiments of the present application;
[0067] Figure 11 It is a schematic hardware structure diagram of a computer device provided by the embodiments of the present application. Detailed implementation manners
[0068] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are only examples of devices and methods that are consistent with some aspects of the embodiments of the present application.
[0069] It can be understood that the terms "first", "second", etc. used in the present application can be used in this document to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be called the second information. Similarly, the second information can also be called the first information. Depending on the context, as used herein, the words "if", "when" can be interpreted as "when...", "when...", or "in response to determining".
[0070] The terms "at least one", "a plurality", "each", "any one", etc. used in the present application, at least one includes one, two or more than two, a plurality includes two or more than two, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.
[0071] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0072] In the related art, facial expression analysis can be applied to various scenarios, such as lie detection, psychology, and healthcare. Facial expressions are conveyed and perceived through the movement of facial muscles and are a form of non-verbal communication used to provide visual information about an individual's emotional state. According to the intensity and duration, facial expressions can be classified into two categories: macro expressions (MaEs) and micro expressions (MEs). MaEs have a higher intensity and a longer duration, usually lasting from 0.5 to 4 seconds. In contrast, MEs have a very short duration, usually within 0.5 seconds, and a very low intensity, being more unconscious and difficult to detect with the naked eye. Existing technologies usually use artificially designed features to analyze the feature differences between different frames in a video, and then determine the expression frames in the video. However, due to the characteristics of the intensity and duration of facial expressions, the accuracy of the appearance range of the determined expression frames in the existing technologies is not high, thereby reducing the accuracy of expression recognition.
[0073] In view of this, in the embodiments of the present application, an expression detection method, device, and storage medium based on a multi-layer convolutional neural network are provided. The present application does not need to rely on artificially designed features and can accurately determine the appearance range of expression frames in a video stream, thereby improving the accuracy of expression recognition.
[0074] The expression detection method based on a multi-layer convolutional neural network provided by the embodiments of the present application relates to the technical field of expression detection. The expression detection method based on a multi-layer convolutional neural network provided by the embodiments of the present application can be applied to a terminal, a server, or software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing the expression detection method based on a multi-layer convolutional neural network, etc., but is not limited to the above forms.
[0075] This application can be used in numerous general - purpose or special - purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi - processor systems, microprocessor - based systems, set - top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer - executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0076] The embodiments of this application will be specifically described below with reference to the accompanying drawings:
[0077] Figure 1 It is an alternative flowchart of the facial expression detection method based on a multi - layer convolutional neural network provided by an embodiment of this application. Figure 1 The method in may include but is not limited to steps S110 to S140:
[0078] Step S110: Obtain image data to be processed from the video data to be detected, where each frame of the image data to be processed includes a face image.
[0079] Step S120: Pre - process each frame of the image data to be processed to obtain a target image data set.
[0080] Step S130: Input the target image data set into a trained preset deep - learning network model for facial expression score prediction to obtain a target facial expression score.
[0081] Step S140: Determine the occurrence range of the facial expression frames in the video data to be detected according to the target facial expression score.
[0082] It can be understood that the process of obtaining the image data to be processed from the video data to be detected can be achieved by first obtaining the video data to be detected, then performing face detection on the video data to be detected frame by frame through a face detector to obtain the regions to be cropped, then cropping all the regions to be cropped to obtain several first face images, and then performing pixel processing on the first face images to obtain the second face images in the image data to be processed. Specifically, the face detector can be the Deep Neural Network (DNN) face detector of the open-source computer vision and machine learning software library (OpenCV). The cropping process in this embodiment is to crop off the redundant background content and only retain the image containing the face as the first face image. Then, the first face image is subjected to pixel conversion, that is, converted into an image with a pixel size of (128, 128) as the second face image.
[0083] In the embodiment of the present application, the process of preprocessing each frame of image in the image data to be processed to obtain the target image dataset includes, but is not limited to, the following steps:
[0084] Perform optical flow calculation on each frame of image in the image data to be processed to obtain the first optical flow feature image;
[0085] Process the face feature points in the first optical flow feature image to obtain the second optical flow feature image;
[0086] Generate the target pseudo-label for each frame of image in the image data to be processed;
[0087] Combine the second optical flow feature image and the target pseudo-label to form the target image dataset.
[0088] It can be understood that in this embodiment, OpenCV can be used to calculate the TV-L1 optical flow between the i-th frame and the (i + k)-th frame in the image data to be processed. Among them, the TV-L1 optical flow method combines the Total Variation and the L1 norm, which can be used to process images with noise or dynamic scenes and can effectively restore the details in the images. Exemplarily, assuming that the gray value at the pixel point (x, y) at time t is I(x, y, t), after a time dt, this pixel point moves to (x + dx, y + dy), and according to the gray value conservation, we get: I(x, y, t) = I(x + dx, y + dy, t + dy). Expand the right side of the equation by Taylor's formula to get: where β represents the derivative term above the second order. After ignoring the higher-order derivatives, we can get: where and respectively represent the partial derivatives of the pixel gray value in the image along the three directions of x, y, and t. Specifically, is the velocity vector of the optical flow along the horizontal direction. is the velocity vector of the optical flow in the vertical direction. Take (u, v) as the first two dimensions of the TV-L1 optical flow feature and use it to solve the optical strain in the following formula:
[0089]
[0090] In the formula, ∈ represents the optical strain, respectively represent the derivatives of u with respect to x and y, and respectively represent the derivatives of v with respect to x and y.
[0091] In this embodiment, the amplitude of the optical strain |∈| is used as the third dimension of the optical flow feature. The amplitude of the optical strain can be obtained by the equation Finally, (u, v, |∈|) as the final optical flow feature can be organized like an RGB image to obtain the first optical flow feature image. Among them, the dimension of the final optical flow feature is (128, 128, 3).
[0092] It can be understood that after obtaining the first optical flow feature image in this embodiment, the face feature points in the first optical flow feature image are processed. Among them, the processing includes but is not limited to the following steps:
[0093] Extract several face feature points in the first optical flow feature image;
[0094] Fill the eye region in the first optical flow feature image black according to the face feature points to obtain the third optical flow feature image;
[0095] Determine the first upper half region of interest, the second upper half region of interest and the lower half region of interest in the third optical flow feature image according to the face feature points;
[0096] Adjust the sizes of the first upper half region of interest, the second upper half region of interest and the lower half region of interest;
[0097] Horizontally stack the resized first upper half region of interest and the second upper half region of interest to obtain the upper half region of interest;
[0098] Stitch the upper half region of interest and the resized lower half region of interest to obtain the second optical flow feature image.
[0099] Exemplarily, in this embodiment, a cross-platform C++ library (Dlib) can be used to extract 68 facial feature points in the first optical flow feature image. Then, after determining the eye region based on the extracted facial feature points, the eye region is filled in black to avoid the influence of eye blinking on the optical flow features. Next, three facial regions are determined in the third optical flow feature image. The left eye and the left eyebrow region can be used as the first upper half region of interest ROI1, the right eye and the right eyebrow region can be used as the second upper half region of interest ROI2, and the mouth region can be used as the lower half region of interest ROI3. After adjusting the sizes of the first upper half region of interest ROI1 and the second upper half region of interest ROI2 to a region size of 21*21, these two regions are horizontally stacked together to form an upper half region of interest with a size of 21*42. At the same time, after adjusting the size of the lower half region of interest ROI3 to a region size of 21*42, it is stitched with the upper half region of interest to obtain a region image with a size of 42*42 as the second optical flow feature image. Specifically, since the dimension of the optical flow features of the first optical flow feature image is (128, 128, 3), the dimension of the second optical flow feature image obtained after processing is (42, 42, 3).
[0100] It can be understood that this embodiment also generates target pseudo-labels for each frame of the image data to be processed, and its generation process includes but is not limited to the following steps:
[0101] Obtain the expression label of the image data to be processed, where the expression label includes the expression start frame and the expression end frame, and the interval between the expression start frame and the expression end frame is used as the expression interval;
[0102] Generate the first linear pseudo-label of the target pseudo-label corresponding to each frame of the image within the expression interval through a linear function;
[0103] Generate the first Gaussian pseudo-label of the target pseudo-label corresponding to each frame of the image within the expression interval through a Gaussian function.
[0104] Specifically, in this embodiment, for each frame Fi located within the expression interval [onset, offset] from onset (expression start frame) to offset (expression end frame), a linear function and a Gaussian function are used to re-label each frame Fi with a value between [0, 1] as shown in Figure 6 as the target pseudo-label. Among them, the formula of the linear function is as follows:
[0105]
[0106] where F i , F on , F off and F apexrespectively represent the current frame, start frame, end frame, and vertex frame (the frame with the highest expression intensity) of the expression duration. For the i-th frame Fi in the real expression interval [Fonset, Fapex] from the start frame to the vertex frame, its label is si = (Fapex - Fi) / (Faset - Fonset). The same strategy is also applied to the interval [Faset, Foffset] from the vertex frame to the end frame, such that Specifically, frames not within the expression interval are labeled as '0'. Through the above method, a linear pseudo-label set L1 = {l i |i = 1,..., end} including the first linear pseudo-labels can be generated, and these pseudo-labels are obtained by frame-by-frame calculation.
[0107] The formula of the Gaussian function is as follows:
[0108]
[0109] where μ is set to the value of the vertex frame, and σ is set to 10. The vertex frame is set to (μ, 1), and the start frame and end frame are set to (0, 0.1) and (2μ, 0.1) respectively. Therefore, the scores of the start frame and end frame are 0.1, while the score of the vertex frame is 1, ensuring that the scores of the frames within the expression interval of the Gaussian function gradually increase from 0.1 to 1 and then gradually decrease to 0.1. Then, the Gaussian function is used to assign labels to each frame within the [Fonseet, Foffset] expression interval, and the labels of other non-expression frames are all set to 0. In this way, a Gaussian pseudo-label set G2 = {gi|i = 1,..., end} including the first Gaussian pseudo-labels can be obtained.
[0110] It can be understood that before applying the preset deep learning network model in this embodiment, the preset deep learning network model will also be trained through the following steps. Among them, the training process includes but is not limited to the following steps:
[0111] Obtain the training image data from the video data to be trained, and each frame of the training image data includes a face image;
[0112] Preprocess each frame of the training image data to obtain a target training set, and the target training set includes the training optical flow feature images to be trained, the training linear pseudo-labels, and the training Gaussian pseudo-labels;
[0113] Divide the target training set into a training set to be used and a test set;
[0114] Train the preset deep learning network model through the training set to be used, and calculate the total loss function during the training process. The total loss function includes a linear loss function and a Gaussian loss function;
[0115] After determining that the preset deep learning network model is completed training according to the total loss function, the output result of the trained preset deep learning network model is enhanced through a test set.
[0116] In the embodiment of the present application, the preset deep learning network model includes four convolutional units, a global pooling layer, a convolutional layer, and a connection layer, and the four convolutional units, the global pooling layer, the convolutional layer, and the connection layer are connected in sequence. Each convolutional unit includes a first convolutional layer, a second convolutional layer, a feed-forward network, a first layer of normalization, and a second layer of normalization; the first layer of normalization, the first convolutional layer, the second layer of normalization, the second convolutional layer, and the feed-forward network are connected in sequence. Among them, the first convolutional layer is used to perform dynamic convolution operations, and the second convolutional layer is used to perform spatial channel attention convolution operations.
[0117] Specifically, as Figure 2 shown, the first convolutional unit starts with an embedding layer, uses a 4×4 convolutional kernel, and performs convolution with a stride of 4; the latter three convolutional units start with a Merging layer, use a 2×2 convolutional kernel, and perform convolution with a stride of 2 to achieve spatial downsampling and channel expansion. A plurality of TVPC blocks are stacked inside each convolutional unit, and the numbers of TVPC blocks in the four stages are {1, 2, 8, 2} respectively. As Figure 3 shown, the internal structure of each TVPC block is a first convolutional layer (TVConv), a second convolutional layer (PRConv), a Feed Forward Network (FFN), a first layer of normalization (Layer Norm), and a second layer of normalization. Finally, the last three layers of the network are a global average pooling layer, a 1×1 convolutional layer, and a fully connected layer, and these layers jointly complete feature transformation and finally generate an expression score.
[0118] Specifically, convolution has demonstrated strong performance in various tasks, and various neural network architectures improve efficiency by enhancing convolutional layers. These improvements include increasing the size of convolutional kernels, changing the shape of convolutional kernels, and introducing new convolutional operation methods, etc. However, although these improvements have enhanced efficiency in certain tasks, their basic convolutional operation mechanism remains unchanged, which makes them have limited effectiveness when dealing with tasks with specific layouts (such as face detection and medical image segmentation). In these specific tasks, the model needs to consume a large amount of computing resources to learn for feature matching. For example, facial expressions can be decomposed into specific movements of the muscles in each part of the face, called Action Units (AUs). To accurately identify whether a macro or micro expression appears, the model must be able to effectively capture local features, global features from various facial regions, and the relationships between these local features. This requires that the convolutional operation not only be efficient but also be able to handle complex relationships between different regions.
[0119] Therefore, in this embodiment, Figure 4 the shown Translation Variant Convolution (TVConv) is used as the first convolutional layer. Among them, TVConv is an efficient dynamic convolution specifically for layout-aware visual processing. Due to its translational invariance and the property of being shared among images, TVConv is applicable to expression detection. First, the learnable affinity maps are fed into a weight generation module to generate weights as shown in the formula: where A is the learnable affinity map, and its initial value is set to 1. The weight generation block consists of multiple layers, including standard Conv, layer normalization, and Relu activation functions. These layers can be stacked multiple times, and finally, one more layer of standard Conv is stacked to form the complete structure of the weight generation module . Since the operations involved are all differentiable, the affinity map A can be trained end-to-end through standard backpropagation. Then, the weights W generated by are convolved with the input I to produce the final output.
[0120] In the facial expression detection task, efficient feature extraction and processing are crucial. Although TVConv enhances the model performance by dynamically adjusting the convolutional kernel weights, in practical applications, it still faces challenges of high computational resource consumption and long training time when dealing with large-scale video data. To address this issue, the core structure PConv in the FasterNet model is improved to PRConv and used as the next convolutional layer of TVConv. PConv is a lightweight convolution that only applies conventional Conv to some input channels for spatial feature learning, while the remaining channels remain unchanged. And the Figure 5 PRConv shown is based on the original PConv and performs RFAConv convolution operations on the channels that have not undergone convolution. RFAConv is a spatial channel attention convolution. Specifically, the input is divided into two parts in the C dimension and Then, apply convolution with a kernel of k to I1 to get O1, and at the same time perform RFAConv convolution operation on I2 to get O2. Finally, concatenate O1 and O2 in the C dimension to get the output where C P and C-C P are in a ratio of 1:3.
[0121] RFAConv learns the attention map by interacting with the receptive field feature information to improve the network performance. The input first aggregates the global information of each receptive field feature through AvgPool. Next, use 1×1 grouped convolution for information interaction. Then, highlight the importance of each receptive field feature through Softmax to generate the attention map A rf . Subsequently, the input X quickly extracts the receptive field spatial features through k×k Group Conv, and through batch normalization (BatchNorm) and ReLU activation function, obtains the receptive field spatial feature F rf . Finally, the attention map A rf is multiplied by the transformed receptive field spatial feature F rf to get the output F, and this process is shown as follows:
[0122]
[0123] It can be understood that after determining the structure of the preset deep learning network model, the above method is used to process the image data to be trained, so that the target training set includes the optical flow feature image to be trained, the linear pseudo-label to be trained, and the Gaussian pseudo-label to be trained. Then, the preset deep learning network model is trained using the training set to be used in the target training set, and the mean square error is used as the loss function, and the optimizer algorithm uses the stochastic gradient descent algorithm with the learning rate set to 0.0010. Specifically, the loss function of this embodiment includes a linear loss function and a Gaussian loss function. Among them, the linear loss function is as follows:
[0124]
[0125] The Gaussian loss function is as follows:
[0126]
[0127] where l i represents the linear pseudo-label, g i represents the Gaussian pseudo-label, represents the predicted value of the model, and N represents the number of samples.
[0128] The total loss function of the model is as follows:
[0129]
[0130] In the formula, λ is 0.3.
[0131] During the testing process, the preset deep learning network model will output Figure 7 the expression prediction score s as shown. After completing the training of the preset deep learning network model, the test set is input into the preset deep learning network model to enhance the output result s of the preset deep learning network model to obtain the label prediction score as shown in Figure 8
[0132]
[0133] In the formula, s j and respectively represent the j-th value in the original prediction score sequence s and the i-th value in the smoothed score sequence . In the smoothing scheme, the smoothed score value i of the current frame F is the average value taken from the interval [F (i-k) , F (i+k-1) . Each smoothed score value now represents the probability that the current frame F i is within the expression interval.
[0134] It can be understood that after the training and testing of the preset deep learning network model are completed in this embodiment, the target image dataset corresponding to the video data to be detected is input into the trained preset deep learning network model for predicting the expression score to obtain the target expression score. After that, the occurrence range of the expression frames in the video data to be detected is determined according to the target expression score, and the process includes but is not limited to the following steps:
[0135] Calculate the standard threshold according to the target expression score;
[0136] Determine the peak frames in the video data to be detected according to the standard threshold;
[0137] Calculate the minimum distance between the peak frames;
[0138] Determine the occurrence range of the expression frames in the video data to be detected according to the peak frames and the minimum distance.
[0139] Among them, the calculation formula of the standard threshold is as follows:
[0140]
[0141] In the formula, and are respectively the average value and the maximum value of the smoothed score sequence , and t is a percentage adjustment parameter ranging from 0 to 1.
[0142] As Figure 9 shown, in this embodiment, the peak frames F with peak s in the video data to be detected are determined through the standard threshold T p . After that, the minimum distance between the peaks is calculated as k. The detected peak frames F p are used as the detected vertex frames, and by expanding k frames forward and backward, the detection start - end interval [F p , F (p-k) , F (p+k) for evaluating the occurrence range of the expression frames is determined, so that the expression detection can be performed within the occurrence range of the expression frames, effectively improving the accuracy of the expression detection.
[0143] In some embodiments, the method of this application embodiment is applied to the CAS(ME)^2 database for micro - expression detection. Among them, the variables and parameters involved in the process are shown in Table 1:
[0144] Table 1
[0145] Variable Parameter k (Half of the average frame length of the expression) 6 λ (Parameter of the loss function) 0.3 t (Percentage adjustment parameter of the threshold) 0.97 Number of training epochs 10 batch_size 128 Depth of the network architecture (1,2,8,2) Learning rate of the SGD optimizer <![CDATA[1×10 -3 >
[0146] Apply the method of the embodiment of the present application to the CAS(ME)^2 database for macro-expression detection. Among them, the variables and parameters involved in the process are shown in Table 2:
[0147] Table 2
[0148] Variable Parameter k (Half of the average frame length of the expression) 18 λ (Parameter of the loss function) 0.3 t (Percentage adjustment parameter of the threshold) 0.65 Number of training epochs 10 batch_size 128 Depth of the network architecture (1,2,8,2) Learning rate of the SGD optimizer <![CDATA[1×10 -3 >
[0149] Through testing, it can be known that the detection results of the method of the present application for micro-expressions are shown in Table 3:
[0150] Table 3
[0151]
[0152]
[0153] Through testing, it can be known that the detection results of the method of the present application for micro-expressions are shown in Table 4:
[0154] Table 4
[0155] Item Standard metric Example 1 Expression category Macro-expression Macro-expression True Positive 88 False Positive 240 False Negative 212 Precision 0.2683 Recall 0.2933 F1-Score 0.2161 0.2803
[0156] It can be seen from Table 3 and Table 4 that the detection results of the method of the present application for micro-expressions and macro-expressions in the database are both highly accurate.
[0157] In summary, the method of this embodiment solves the problems that the expression duration is too short and the change of each frame of expression is too small by building a TVPR network (preset deep learning network model). And through the TVConv convolutional layer, it solves the problem that the convolutional operation has limited effect in face detection; through the PRConv convolutional layer, it solves the problem of too long training model time and at the same time solves the problem of unsatisfactory PConv performance, so as to effectively improve the accuracy of the output result of the preset deep learning network model, and further improve the accuracy of the expression detection result.
[0158] Refer to Figure 10 , the embodiment of the present application provides an expression detection device based on a multi-layer convolutional neural network. The device includes:
[0159] The first module 1010 is used to obtain the image data to be processed from the video data to be detected, and each frame of the image data to be processed includes a face image;
[0160] The second module 1020 is used to preprocess each frame of the image data to be processed to obtain a target image data set;
[0161] The third module 1030 is used to input the target image data set into the trained preset deep learning network model for expression score prediction to obtain a target expression score;
[0162] The fourth module 1040 is configured to determine the occurrence range of expression frames in the video data to be detected according to the target expression score;
[0163] The preset deep learning network model includes four convolutional units, a global pooling layer, a convolutional layer, and a connection layer, and the four convolutional units, the global pooling layer, the convolutional layer, and the connection layer are connected in sequence;
[0164] Each convolutional unit includes a first convolutional layer, a second convolutional layer, a feedforward network, a first normalization layer, and a second normalization layer; the first normalization layer, the first convolutional layer, the second normalization layer, the second convolutional layer, and the feedforward network are connected in sequence; the first convolutional layer is used to perform dynamic convolution operations, and the second convolutional layer is used to perform spatial channel attention convolution operations.
[0165] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0166] An embodiment of the present application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above method is implemented. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.
[0167] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0168] Please refer to Figure 11 , Figure 11 which shows the hardware structure of a computer device according to another embodiment. The computer device includes:
[0169] A processor 1110, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute related programs to implement the technical solutions provided by the embodiments of the present application;
[0170] The memory 1120 can be implemented in the form of a Read Only Memory (ROM), a static storage device, a dynamic storage device, or a Random Access Memory (RAM), etc. The memory 1120 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1120 and are called by the processor 1110 to execute the above-mentioned method of the embodiments of this application;
[0171] The input / output interface 1130 is used to implement information input and output;
[0172] The communication interface 1140 is used to implement communication and interaction between this device and other devices. It can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.);
[0173] The bus 1150 transmits information between the various components of the device (such as the processor 1110, the memory 1120, the input / output interface 1130, and the communication interface 1140);
[0174] Among them, the processor 1110, the memory 1120, the input / output interface 1130, and the communication interface 1140 achieve communication connections with each other inside the device through the bus 1150.
[0175] The embodiments of this application also provide a computer-readable storage medium. This computer-readable storage medium stores a computer program, and when this computer program is executed by a processor, it implements the above-mentioned method.
[0176] It can be understood that the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0177] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory can optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0178] The embodiments described in the embodiments of this application are to more clearly illustrate the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.
[0179] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0180] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0181] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware and their appropriate combinations.
[0182] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0183] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single items (items) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0184] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0185] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0186] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0187] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0188] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, and thus do not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of the rights of the embodiments of this application.
Claims
1. An expression detection method based on a multi-layer convolutional neural network, characterized in that, The method includes the following steps: Obtain image data to be processed from the video data to be detected, where each frame of the image data to be processed includes a face image; Preprocess each frame of the image data to be processed to obtain a target image dataset; Input the target image dataset into a trained preset deep learning network model for predicting an expression score to obtain a target expression score; Determine the occurrence range of expression frames in the video data to be detected according to the target expression score; Among them, the preset deep learning network model includes four convolutional units, a global pooling layer, a convolutional layer, and a connection layer, and the four convolutional units, the global pooling layer, the convolutional layer, and the connection layer are connected in sequence; Each convolutional unit includes a first convolutional layer, a second convolutional layer, a feedforward network, a first normalization layer, and a second normalization layer; the first normalization layer, the first convolutional layer, the second normalization layer, the second convolutional layer, and the feedforward network are connected in sequence; the first convolutional layer is used for performing dynamic convolution operations, and the second convolutional layer is used for performing spatial channel attention convolution operations.
2. The method according to claim 1, characterized in that, The obtaining of the image data to be processed from the video data to be detected includes: Obtain the video data to be detected; Perform face detection on the video data to be detected frame by frame through a face detector to obtain regions to be cropped; Crop all the regions to be cropped to obtain a number of first face images; Perform pixel processing on the first face images to obtain second face images in the image data to be processed.
3. The method according to claim 1, wherein The preprocessing of each frame of the image data to be processed to obtain a target image dataset includes: Perform optical flow calculation on each frame of the image data to be processed to obtain a first optical flow feature image; Process the face feature points in the first optical flow feature image to obtain a second optical flow feature image; Generate target pseudo-labels for each frame of the image data to be processed; Combine the second optical flow feature image and the target pseudo-labels to form the target image dataset.
4. The method according to claim 3, characterized in that, The processing of the face feature points in the first optical flow feature image to obtain a second optical flow feature image includes: Extract a number of face feature points in the first optical flow feature image; Perform blackening processing on the eye regions in the first optical flow feature image according to the face feature points to obtain a third optical flow feature image; Determine a first upper half region of interest, a second upper half region of interest, and a lower half region of interest in the third optical flow feature image according to the face feature points; Adjust the sizes of the first upper half region of interest, the second upper half region of interest, and the lower half region of interest; Horizontally stack the size-adjusted first upper half region of interest and the second upper half region of interest to obtain an upper half region of interest; Stitch the upper half region of interest and the size-adjusted lower half region of interest to obtain the second optical flow feature image.
5. The method according to claim 3, characterized in that, The generating of the target pseudo-labels for each frame of the image data to be processed includes: Obtain the expression label of the to-be-processed image data, where the expression label includes an expression start frame and an expression end frame, and the interval between the expression start frame and the expression end frame is used as the expression interval; Generate a first linear pseudo-label corresponding to the target pseudo-label for each frame of the image within the expression interval through a linear function; Generate a first Gaussian pseudo-label corresponding to the target pseudo-label for each frame of the image within the expression interval through a Gaussian function.
6. The method according to claim 1, characterized in that, The training process of the preset deep learning network model includes the following steps: Obtain the to-be-trained image data from the to-be-trained video data, where each frame of the to-be-trained image data includes a face image; Preprocess each frame of the to-be-trained image data to obtain a target training set, where the target training set includes to-be-trained optical flow feature images, to-be-trained linear pseudo-labels, and to-be-trained Gaussian pseudo-labels; Divide the target training set into a to-be-used training set and a test set; Train the preset deep learning network model through the to-be-used training set, and calculate the total loss function during the training process, where the total loss function includes a linear loss function and a Gaussian loss function; After determining that the preset deep learning network model is trained according to the total loss function, enhance the output result of the trained preset deep learning network model through the test set.
7. The method according to claim 1, characterized in that, The determining the occurrence range of the expression frames in the to-be-detected video data according to the target expression score includes: Calculate a standard threshold according to the target expression score; Determine the peak frames in the to-be-detected video data according to the standard threshold; Calculate the minimum distance between the peak frames; Determine the occurrence range of the expression frames in the to-be-detected video data according to the peak frames and the minimum distance.
8. An expression detection device based on a multi-layer convolutional neural network, characterized in that, The device includes: A first module, configured to obtain to-be-processed image data from the to-be-detected video data, where each frame of the to-be-processed image data includes a face image; A second module, configured to preprocess each frame of the to-be-processed image data to obtain a target image data set; A third module, configured to input the target image data set into a trained preset deep learning network model for expression score prediction to obtain a target expression score; A fourth module, configured to determine the occurrence range of the expression frames in the to-be-detected video data according to the target expression score; Wherein, the preset deep learning network model includes four convolutional units, a global pooling layer, a convolutional layer, and a connection layer, and the four convolutional units, the global pooling layer, the convolutional layer, and the connection layer are connected in sequence; Each convolutional unit includes a first convolutional layer, a second convolutional layer, a feed-forward network, a first normalization layer, and a second normalization layer; the first normalization layer, the first convolutional layer, the second normalization layer, the second convolutional layer, and the feed-forward network are connected in sequence; the first convolutional layer is used to perform dynamic convolution operations, and the second convolutional layer is used to perform spatial channel attention convolution operations.
9. A computer device, characterized in that, Includes: At least one processor; At least one memory, configured to store at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.