Face detection method and device, computer equipment and storage medium
By constructing triple samples and combining global facial expression feature extraction and similarity modules, the existing AU detection methods have solved the problem of high hardware requirements and low speed and accuracy, and achieved more efficient AU detection.
Patent Information
- Application Number
- CN202410066199.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-18
AI Technical Summary
The existing AU detection methods have high hardware requirements and are difficult to improve detection speed and accuracy at the same time.
By constructing triple samples, the global facial expression feature extraction module and similarity module are used to combine AU detection and AU combination similarity learning tasks to improve the effect of AU detection in face images.
Improve the accuracy and efficiency of AU detection and reduce hardware requirements.
Smart Images

Figure CN120340078A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a face detection method, apparatus, computer device, and storage medium. Background Art
[0002] Facial expressions can be divided into dozens of action units according to the movements of facial muscle groups, forming a complete Facial Action Coding System (FACS). Any human expression can be represented as a combination of a set of action units and their different intensities. The AU (Action Unit, (facial) action unit) detection task is to detect the AU categories that appear in the input face image based on the input face image. AU detection can be regarded as a multi-label classification problem, so some conventional classification loss functions can be used to supervise the training of the detection algorithm.
[0003] In related technologies, most deep learning solutions for AU detection are classified and trained on some AU datasets using existing AU labels. Existing AU detection using traditional machine algorithms and deep learning algorithms not only requires high hardware conditions but also makes it difficult to improve the detection speed and accuracy simultaneously. Summary of the Invention
[0004] Embodiments of this application provide a face detection method, apparatus, computer device, and storage medium to improve the effect of AU detection in face images.
[0005] Embodiments of this application provide a face detection method, including:
[0006] Obtain the first face feature of the first face image and the second face feature of the second face image;
[0007] Perform a similarity detection on the first face feature and the second face feature to obtain the similarity information between the first face feature and the second face feature;
[0008] Perform facial action recognition processing on the first face feature and the second face feature respectively to obtain the first facial action information corresponding to the first face feature and the second facial action information corresponding to the second face feature;
[0009] Based on the similarity information, the first facial action information, and the second facial action information, determine the facial action similarity between the first face image and the second face image.
[0010] Correspondingly, embodiments of this application also provide a face detection apparatus, including:
[0011] A first acquisition unit, configured to acquire a first facial feature of a first facial image and a second facial feature of a second facial image;
[0012] A first processing unit, configured to perform a similarity detection on the first facial feature and the second facial feature to obtain similarity information between the first facial feature and the second facial feature;
[0013] A second processing unit, configured to perform facial action recognition processing on the first facial feature and the second facial feature respectively to obtain first facial action information corresponding to the first facial feature and second facial action information corresponding to the second facial feature;
[0014] A determination unit, configured to determine a facial action similarity between the first facial image and the second facial image based on the similarity information, the first facial action information, and the second facial action information.
[0015] Correspondingly, an embodiment of the present application further provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the face detection method provided in any embodiment of the present application.
[0016] Correspondingly, an embodiment of the present application further provides a storage medium, which stores multiple instructions adapted to be loaded by a processor to execute the above face detection method.
[0017] In an embodiment of the present application, multiple anchor samples are determined according to the AU label combination categories of facial images in a dataset. Samples with the same AU information as the anchor samples are used as positive samples, and samples with different AU information from the anchor samples are used as negative samples. A triplet sample is constructed based on the anchor sample, the positive sample, and the negative sample. The triplet sample is input into a global facial expression feature extraction module for feature extraction to obtain feature vectors corresponding to each sample image. Further, the feature vectors are respectively sent to a similarity module and a detection module to learn the AU combination similarity information and AU category information in the input triplet sample respectively. By organically combining AU detection and AU combination similarity learning, and using the AU combination similarity learning task to assist the AU detection task, the effect of AU detection in facial images can be improved. Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained without creative efforts based on these drawings.
[0019] Figure 1 Schematic flowchart of a face detection method provided by an embodiment of the present application.
[0020] Figure 2 Schematic diagram of an application scenario of a face detection method provided by an embodiment of the present application.
[0021] Figure 3 Schematic diagram of an application scenario of another face detection method provided by an embodiment of the present application.
[0022] Figure 4 Schematic diagram of an application scenario of another face detection method provided by an embodiment of the present application.
[0023] Figure 5 Schematic diagram of an application scenario of another face detection method provided by an embodiment of the present application.
[0024] Figure 6 Schematic diagram of an application scenario of another face detection method provided by an embodiment of the present application.
[0025] Figure 7 Schematic diagram of an application scenario of another face detection method provided by an embodiment of the present application.
[0026] Figure 8 Schematic diagram of an application scenario of another face detection method provided by an embodiment of the present application.
[0027] Figure 9 Block diagram of the structure of a face detection device provided by an embodiment of the present application.
[0028] Figure 10 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0029] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0030] The embodiments of the present application provide an information recommendation method, apparatus, storage medium, and computer device. Specifically, the information recommendation method in the embodiments of the present application can be executed by a computer device, where the computer device can be a terminal or a server, etc. The terminal can be a terminal device such as a smart phone, a tablet computer, a laptop computer, a touch screen, a personal computer (PC), or a personal digital assistant (PDA). The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0031] For example, the computer device can be a server, and the server can obtain the first face feature of the first face image and the second face feature of the second face image; perform a similarity detection on the first face feature and the second face feature to obtain the similarity information between the first face feature and the second face feature; perform facial action recognition processing on the first face feature and the second face feature respectively to obtain the first facial action information corresponding to the first face feature and the second facial action information corresponding to the second face feature; and determine the facial action similarity between the first face image and the second face image based on the similarity information, the first facial action information, and the second facial action information.
[0032] Based on the above problems, the embodiments of the present application provide a first face detection method, apparatus, computer device, and storage medium to improve the effect of AU detection in face images.
[0033] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.
[0034] The embodiments of the present application provide a face detection method, which can be executed by a terminal or a server. The embodiments of the present application will be described by taking the face detection method executed by a server as an example.
[0035] Please refer to Figure 1 , Figure 1 , which is a schematic flowchart of a face detection method provided by the embodiments of the present application. The specific process of the face detection method can be as follows:
[0036] 101. Obtain the first face feature of the first face image and the second face feature of the second face image.
[0037] Among them, the first face image and the second face image can be two face images for which facial expression comparison is required.
[0038] In the embodiments of the present application, after obtaining the first face image and the second face image, image preprocessing can be performed on the first face image and the second face image first.
[0039] Among them, image preprocessing can include various processing methods, such as face alignment, cropping, scaling, etc.
[0040] First, face key point detection can be performed on the first face image and the second face image to determine the first face key points in the first face image and the second face key points in the second image.
[0041] Among them, face key points represent the positions of the facial features of the face in the image, such as eyes, nose, mouth, etc.
[0042] Furthermore, face alignment processing can be performed on the first face image according to the first face key points, and face alignment processing can be performed on the second face according to the second face key points.
[0043] Among them, face alignment refers to the similarity transformation based on facial features such as the nose, both eyes, and lips, and aligning the face to a reference face without changing the shapes of the key facial features.
[0044] For example, please refer to Figure 2 , Figure 2 which is a schematic diagram of the application scenario of a face detection method provided by the embodiments of the present application. In the Figure 2 shown initial face image, the face has an offset. By performing face alignment processing on the initial face image, the face offset is corrected to obtain the aligned face image. In the aligned face image, the facial features are in the positions of the reference face facial features.
[0045] Then, image cropping can be performed on the aligned first face image and the second face image to remove the parts related to the face in the image, which can improve the subsequent face comparison detection efficiency.
[0046] Among them, feature extraction from the first face image and the second face image can include respectively performing feature extraction on the preprocessed first face image and the preprocessed second face image to obtain the first face features of the first face image and the second face features of the second face image.
[0047] Among them, face features can at least include face expression features. Face expression features refer to the geometric relationships between facial features such as eyes, nose, and mouth in the face, such as distance, area, and angle, etc.
[0048] In some embodiments, the step of "obtaining the first face feature of the first face image and the second face feature of the second face image" may include the following operations:
[0049] Input the first face image and the second face image into the target detection model;
[0050] Respectively extract features from the first face image and the second face image through the face feature extraction module of the target detection model to obtain the first face feature and the second face feature.
[0051] In the embodiments of the present application, the target detection model can be used to detect the facial action similarity between face images. The target detection model may include a face feature extraction module, and the face feature extraction module can be used to extract the face feature in the face image.
[0052] Specifically, the preprocessed first face image and the preprocessed second face image can be respectively input into the target detection model, and the face feature extraction module of the target detection model is used to extract features from the first face image to obtain the first face feature, and to extract features from the second face image to obtain the second face feature.
[0053] In some embodiments, in order to improve the detection effect of the target detection model, before the step of "inputting the first face image and the second face image into the target detection model", the method may further include the following steps:
[0054] Obtain a triple sample of face images;
[0055] Construct a target detection model based on each sample face image in the triple sample and the facial action information of each sample face image.
[0056] Among them, the triple sample includes an anchor sample face image, a positive sample face image, and a negative sample face image. Among them, the actual facial action information of the positive sample face image is the same as the actual facial action information of the anchor sample face image, and the actual facial action information of the negative sample face image is different from the actual facial action information of the anchor sample face image.
[0057] In the embodiments of the present application, the triple sample refers to a sample used to train the target detection model.
[0058] Among them, obtaining a triple sample of face images may include: first collecting a plurality of sample face images and determining the actual facial action information of each sample face image.
[0059] Among them, the face image can be collected by taking a picture of the face or obtaining it from an existing face image database, etc.
[0060] Among them, the actual facial action information refers to the true existing combination of facial action units of the face in the sample face image. The combination of facial action units includes at least one facial action unit, that is, AU (Action Unit).
[0061] Among them, human facial expressions can be divided into dozens of facial action units according to the movement of facial muscle groups, forming a complete facial action unit coding system. Any human expression can be represented as a combination of a set of facial action units and their different intensities.
[0062] For example, please refer to Figure 3 , Figure 3 , which is a schematic diagram of the application scenario of another face detection method provided by the embodiment of the present application. In the Figure 3 shown face image, the facial expression of the face can be determined as happy. The facial action units included in the face may include: cheek raising and mouth corner raising. Then, the combination of facial action units of cheek raising and mouth corner raising can be defined as a happy expression.
[0063] In the embodiment of the present application, corresponding labels are set for each facial action unit. For example, the label for cheek raising can be: AU6, the label for mouth corner raising can be AU12, the label for frowning can be AU4, etc. Different facial action units are distinguished by different labels.
[0064] Among them, the actual facial action information of each sample face image can be marked manually.
[0065] For example, please continue to refer to Figure 3 , Figure 3 The shown face image can be a collected sample face image. By manually judging that the facial action units included in the sample face image are: cheek raising and mouth corner raising, the actual facial action information of the sample face image can be obtained as: cheek raising and mouth corner raising.
[0066] Among them, the anchor sample face image refers to the sample used as the target of the anchor facial action information.
[0067] In the embodiment of the present application, AU combinations corresponding to multiple facial expressions are preset in advance, and these AU combinations are used as the anchor facial action information.
[0068] For example, the preset facial expressions can include: happy, smiling, crying, sad, etc. According to the multiple AUs corresponding to each facial expression, the AU combinations corresponding to each facial expression are obtained.
[0069] Among them, by presetting the AU combinations corresponding to multiple facial expressions respectively, the AU combination categories and the total number of AU combinations can be counted from the existing AU dataset, and then a triple sample can be constructed based on each AU combination. The triple sample includes an anchor sample face image, a positive sample face image, and a negative sample face image. The constructed multiple triple samples are used as the dataset for subsequent model training.
[0070] For example, please refer to Figure 4 , Figure 4 which is a schematic diagram of the application scenario of another face detection method provided by the embodiment of the present application. In Figure 4 the shown table, AU represents the facial action unit. In the first column of the table, different AUs are represented respectively, such as AU1, AU2, AU3, AU4, AU5, AU6, AU7, AU8, AU9. Among them, A, B, C, D, and E can represent different sample face images. Among them, the 0s and 1s in other cells of the table indicate the occurrence of AUs in the sample face images. Among them, 1 indicates occurrence, and 0 indicates non-occurrence. For example, in the sample face image A, AU2 and AU5 appear; in the sample face image B, AU2 and AU5 appear; in the sample face image C, AU3 and AU7 appear; in the sample face image D, AU2 and AU9 appear; in the sample face image E, AU1 and AU3 appear. Among them, Figure 4 only some AUs and the AU combinations that appear in some sample face images are shown.
[0071] Furthermore, select the sample face image containing the anchor facial action information from multiple sample face images as the anchor sample face image, then select the sample face image with the same anchor facial action information as the positive sample face image from multiple sample face images, and select the sample face image with different anchor facial action information as the negative sample face image from multiple sample face images.
[0072] According to the above steps, multiple combinations of anchor sample face images, positive sample face images, and negative sample face images can be selected from the sample face images to obtain multiple triple samples.
[0073] Furthermore, a target detection model can be constructed according to the obtained multiple triple samples, which can be used to detect the facial action similarity of face images.
[0074] In some embodiments, in order to improve the accuracy of image feature extraction, the step of "constructing a target detection model based on each sample face image in the triple sample and the facial action information of each sample face image" may include the following operations:
[0075] Preprocess each sample face image to obtain the processed sample face image corresponding to each sample face image;
[0076] Train a preset detection model based on the processed sample face image and the actual facial motion information to obtain a target detection model.
[0077] In the embodiments of the present application, in order to eliminate irrelevant information in the image, restore useful real information, enhance the detectability of relevant information, and simplify the data to the greatest extent, thereby improving the reliability of feature extraction, image segmentation, matching, and recognition, the sample face images in the triple samples can be preprocessed respectively.
[0078] Among them, the preprocessing performed on the sample face image may include face alignment, cropping, scaling, etc.
[0079] In some embodiments, the step of "preprocessing each sample face image to obtain the processed sample face image corresponding to each sample face image" may include the following operations:
[0080] Perform key point detection on each sample face image to determine the face key points in each sample face image;
[0081] Perform alignment processing on the sample face image based on the face key points to obtain an aligned face image;
[0082] Perform cropping processing on the aligned face image based on the face region in the aligned face image to obtain the processed sample face image.
[0083] Among them, key point detection on each sample image can be performed through face detection technology to detect the face key points in the sample image, including each key point indicating the positions of the facial features.
[0084] For example, please refer to Figure 5 , Figure 5 which is a schematic diagram of an application scenario of another face detection method provided by the embodiments of the present application. Among them, the original image may be the initial sample face image. The preprocessing process may include the following:
[0085] First, use a face detection method to detect the original image, detect the face region from the original image, and at the same time perform key point detection to detect the face key points. Then, based on the detected face key points, perform face alignment processing on the original image. Further, obtain the facial frame in the face image, adjust the size of the facial frame so that the facial frame includes the entire face region, and then perform image cropping according to the adjusted facial frame to retain the face region in the image. Finally, the cropped image can be scaled and adjusted so that the size of the face image meets the image size required for model processing. The finally obtained face image is used as the processed sample face image.
[0086] Among them, for face alignment processing based on face key points, the positions of the left eye corner key point and the right eye corner key point can be obtained from the face key points. According to the positions of the left eye corner key point and the right eye corner key point, calculate the center point position between the two eyes, as well as calculate the coordinate difference Δx between the two eye corners on the X-axis and the coordinate difference Δy between the two eye corners on the Y-axis. Further, according to the center point position, the coordinate difference Δx and the coordinate difference Δy, calculate the rotation angle required for face alignment.
[0087] Among them, the formula for calculating the center point position between the two eyes according to the positions of the left eye corner key point and the right eye corner key point can be as follows:
[0088]
[0089] where x c is the X coordinate of the center point position, and y c is the Y coordinate of the center point position; x r is the X coordinate of the right eye corner key point, and y r is the Y coordinate of the right eye corner key point; x l is the X coordinate of the left eye corner key point, and x l is the Y coordinate of the left eye corner key point.
[0090] Among them, the formula for calculating the rotation angle according to the center point position, the coordinate difference Δx and the coordinate difference Δy can be as follows:
[0091]
[0092] where θ is the rotation angle, Δy represents the coordinate difference between the left eye corner key point and the right eye corner key point on the Y-axis; Δx represents the coordinate difference between the left eye corner key point and the right eye corner key point on the X-axis.
[0093] For example, please refer to Figure 6 , Figure 6 which is a schematic diagram of the application scenario of another face detection method provided by the embodiment of the present application. Figure 6In the shown face image, the left eye corner key point l and the right eye corner key point r are determined from the detected face key points, the position coordinates of the left eye corner key point l and the position coordinates of the right eye corner key point r are obtained, the center point position c between the two eye corner key points is determined according to the two position coordinates, and the coordinate difference Δx and Δy of the two eye corner key points are determined. Then, the rotation angle θ can be calculated according to the above formula.
[0094] Further, the face in the face image is rotated according to the rotation angle θ, so that an image after face alignment can be obtained.
[0095] Among them, training the preset detection model based on the processed sample face image and the actual facial action information to obtain the target detection model may include inputting the processed sample face image and the actual facial action information corresponding to each sample face image into the preset detection model, and performing iterative training on the preset detection model to obtain the target detection model.
[0096] In the embodiments of the present application, the preset detection model may at least include three modules, namely a face feature extraction module, a similarity module, and a detection module.
[0097] For example, please refer to Figure 7 , Figure 7 which is a schematic diagram of an application scenario of another face detection method provided by the embodiments of the present application. Figure 7 The preset detection model designed by the present solution is shown, including a face feature extraction module, a detection module, and a similarity detection module.
[0098] In some embodiments, in order to improve the detection efficiency of the target detection model, the step of "training the preset detection model based on the processed sample face image and the actual facial action information to obtain the target detection model" may include the following operations:
[0099] Input the processed sample face image into the preset detection model, and respectively extract the sample face features corresponding to each sample face image through the face feature extraction module of the preset detection model;
[0100] Train the similarity module in the preset detection model based on the sample face features to obtain a trained similarity module;
[0101] Train the detection module in the preset detection model based on the sample face features to obtain a trained detection module;
[0102] Based on the trained similarity module and the trained detection module, obtain the target detection model.
[0103] Among them, inputting the processed original face image of the sample into the preset detection model means inputting the processed anchor sample face image, positive sample face image, and negative sample face image in a triple into the preset detection model.
[0104] In the embodiment of the present application, the face feature extraction module can be a global face expression feature extraction module (GEE), and the global face expression feature extraction module can be used to extract global face features.
[0105] Among them, global face expression feature extraction is a technology for identifying face expressions and extracting their features. In practical applications, the global face expression feature extraction module uses some local feature extraction methods, such as HOG, LBP, SIFT, SURF, etc., to concatenate the obtained local feature vectors as the global face representation. In addition, multiple features including human physiological features (such as gender, age, skin color, hair color, beard) and expression features (smile, normal, angry) as well as facial wearing accessory features can also be used for extraction.
[0106] For example, please continue to refer to Figure 7 , Figure 7 In, A refers to the anchor sample face image in the triple sample, P refers to the positive sample face image in the triple sample, and N refers to the negative sample face image in the triple sample. Input the anchor sample face image A, positive sample face image P, and negative sample face image N into the face feature extraction module, and extract the global face expression feature of the anchor sample face image A through the face feature extraction module, that is, hA, extract the global face expression feature of the positive sample face image P, that is, hP, and extract the global face expression feature of the negative sample face image N, that is, hN. Then the sample face features hA, hP, and hN can be obtained.
[0107] Specifically, the face feature extraction module is composed of a deep neural network (Inception-Net). Input a set of triple samples, and this module can map the triple samples from the image space to a d-dimensional latent variable feature space. This latent variable feature, denoted as (hA, hP, hN), will be used as the input of the subsequent similarity module and detection module.
[0108] Among them, the deep neural network is a multi-layer unsupervised neural network, and the output features of the previous layer are used as the input of the next layer for feature learning. After layer-by-layer feature mapping, the features of the existing space samples are mapped to another feature space to learn better feature expressions for the existing inputs. The deep neural network has multiple non-linear mapping feature transformations and can fit highly complex functions.
[0109] In some embodiments, in order to improve the processing effect of the similarity module in the target detection model, the step of "training the similarity module in the preset detection model based on the sample face features to obtain the trained similarity module" may include the following operations:
[0110] The similarity module is used to perform transformation processing on each sample face feature respectively to obtain the transformed features corresponding to each sample face feature;
[0111] Calculate the first difference between the transformed feature corresponding to the anchor sample face image and the transformed feature corresponding to the positive sample face image;
[0112] Calculate the second difference between the transformed feature corresponding to the anchor sample face image and the transformed feature corresponding to the negative sample face image;
[0113] Adjust the similarity module based on the first difference, the second difference, and the first loss function to obtain the trained similarity module.
[0114] Among them, the similarity module can be used to detect the similarity between different AU combinations, that is, to detect the similarity between the AU combinations corresponding to different face images, so as to detect the expression similarity between different face images.
[0115] In the embodiments of the present application, each sample face feature is input into the similarity module, and the similarity module is used to perform transformation on the input sample face features respectively to obtain the transformed features with the same dimension as the sample face features.
[0116] Among them, the similarity module can be composed of a multi-layer neural network. Specifically, in the embodiments of the present application, a three-layer fully connected network is used as a structure of the module.
[0117] For example, please continue to refer to Figure 7 , the triple features (hA, hP, hN) output by the face feature extraction module are input into the similarity module, and after passing through the similarity module, the output is also a feature vector (qA, qP, qN) of dimension d.
[0118] Among them, during the processing of the similarity module, the input triple features (hA, hP, hN) can be respectively subjected to linear + non-linear transformation through each layer of the multi-layer neural network, and the final output result is the feature vector (qA, qP, qN).
[0119] Among them, calculating the first difference between the transformed feature corresponding to the anchor sample face image and the transformed feature corresponding to the positive sample face image can be to calculate the difference between qA and qP, which can be used to characterize the difference in the AU combination between the anchor sample face image A and the positive sample face image P.
[0120] Among them, calculating the second difference between the transformed features corresponding to the anchor sample face image and the transformed features corresponding to the negative sample face image can be calculating the difference between qA and qN, which can be used to characterize the difference in AU combinations between the anchor sample face image A and the negative sample face image N.
[0121] Furthermore, for the output feature vectors (qA, qP, qN), a first loss function can be used for supervised learning. Among them, the first loss function L tri is as follows:
[0122]
[0123] Among them, q A refers to the transformed features corresponding to the anchor sample face image, q P refers to the transformed features corresponding to the positive sample face image, q N refers to the transformed features corresponding to the negative sample face image; m is a preset threshold, whose function is to judge whether the difference between the distance between A and P and the distance between A and N is less than this threshold. If it is less than this threshold, it is considered that the distance between A and P has been optimized to be close enough, that is, it means that the processing effect of the similarity module is optimized enough.
[0124] In the embodiments of the present application, the first loss function L uri enables the similarity module to learn an expressive latent variable space, which can map similar latent variables with similar AU combinations closer, and map latent variables with different AU combinations farther.
[0125] For example, please continue to refer to Figure 7 , for the output feature vectors (qA, qP, qN), calculate the difference between qA and qP to obtain the first difference between A and P, and calculate the difference between qA and qN to obtain the second difference between A and N. At this time, the first difference is greater than the second difference. Then, use the first loss function L tri for supervised learning, and calculate that the difference between qA and qP after supervised learning is less than the difference between qA and qN, that is, it accurately expresses that the AU combination similarity between A and P is higher.
[0126] In some embodiments, in order to improve the processing effect of the detection module in the target detection model, the step of "training the detection module in the preset detection model based on the sample face features to obtain the trained detection module" may include the following operations:
[0127] Extract facial action features from each sample face feature through the detection module;
[0128] Determine the predicted facial action information corresponding to each facial action feature;
[0129] Adjust the detection module based on the predicted facial action information, the actual facial action information, and the second loss function to obtain the trained detection module.
[0130] Among them, the detection module can be used to detect the AU categories included in the human face.
[0131] In the embodiments of the present application, the detection module may include multiple sub-modules, and each sub-module performs different processes respectively.
[0132] In some embodiments, the step of "extracting facial action features from each sample facial feature through the detection module" may include the following operations:
[0133] Extract local facial action features from each sample facial feature through the first sub-module;
[0134] Extract global facial action features from each sample facial feature through the second sub-module;
[0135] Determine the facial action features based on the local facial action features and the global facial action features.
[0136] Among them, the detection module may include a first sub-module and a second sub-module. The first sub-module may be an AU attention map extraction module, and the second sub-module may be an AU feature extraction module.
[0137] For example, please continue to refer to Figure 7 , where the AU attention map extraction module may be composed of a multi-layer deconvolution structure. The triple features (hA, hP, hN) output by the facial feature extraction module are input into the AU attention map extraction module. Through the AU attention map extraction module, the features (hA, hP, hN) are deconvolved to obtain a single-channel matrix (mA, mP, mN) with the same size as the input sample image, and the value range of each element in this single-channel matrix is (0, 1). This single-channel matrix is used as the AU local attention map. (mA, mP, mN) can be used as the local facial action features of each sample facial feature.
[0138] Among them, (mA, mP, mN) are three attention maps, where 0 in the value range represents the lowest attention and 1 represents the highest attention; the higher the attention, the greater the proportion will be input into the subsequent neural network for processing, and vice versa, the smaller the proportion.
[0139] Among them, the AU feature extraction module can also be composed of a multi-layer transposed convolution structure, whose structure is similar to that of the AU attention map extraction module, but deeper than that of the AU attention map extraction module. The triplet features (hA, hP, hN) output by the face feature extraction module are input into the AU feature extraction module. Through the AU feature extraction module, the features (hA, hP, hN) are subjected to transposed convolution operations to obtain a three-channel feature matrix with the same size as the input sample image. This three-channel matrix is the AU global feature map (fA, fP, fN) of the input triplet samples, and (fA, fP, fN) can be used as the global facial action features of each sample face image.
[0140] Among them, based on the local facial action features and the global facial action features, the facial action features can be determined by calculating the product of the local facial action features and the global facial action features to obtain the facial action features.
[0141] Specifically, the (mA, mP, mN) output by the AU attention map extraction module is multiplied by the (fA, fP, fN) output by the AU feature extraction module to obtain the AU local feature map of each sample face image, denoted as (lA, lP, lN), that is, mA * fA to obtain lA, mP * fP to obtain lP, and mN * fN to obtain lN.
[0142] In some embodiments, the step of "determining the predicted facial action information corresponding to each facial action feature" may include the following operations:
[0143] The third sub-module calculates the probability values of each facial action feature on each preset facial action category to obtain the predicted facial action information.
[0144] Among them, the detection module may include a third sub-module, and this third sub-module may be an AU classifier module. The AU classifier module can be used to predict the probabilities of facial action features belonging to each AU category.
[0145] Among them, the AU classifier module may be composed of N convolutional networks, where N represents the number of AU categories. For example, the AU categories may include: AU1, AU2,..., AUN.
[0146] For example, please continue to refer to Figure 7 ., the facial action features (lA, lP, lN) are input into the AU classifier module, and each convolutional network of the AU classifier module calculates the probability values of each facial action feature on each AU category as the predicted facial action information.
[0147] Among them, adjusting the detection module based on the predicted facial action information, the actual facial action information, and the second loss function means performing supervised learning on the predicted facial action information and the actual facial action information corresponding to each individual face image using the second loss function.
[0148] Among them, the second loss function L au is as follows:
[0149]
[0150] Among them, p i represents the probability that the predicted facial action feature belongs to each AU category, that is, the predicted facial action information. represents the true probability (0 or 1, 0 means not belonging, 1 means belonging) that the facial action feature belongs to each AU category, that is, the true facial action information.
[0151] In the embodiments of the present application, using the second loss function can improve the accuracy of the detection module in detecting AU categories in face images.
[0152] In the embodiments of the present application, different loss weights are configured for the detection module and the similarity module respectively as the total loss function of the preset detection model, as follows:
[0153] L Total = αL au + βL tri
[0154] Among them, L Total is the total loss function of the preset detection model, α is the weight value of the second loss function L au and β is the weight value of the first loss function L tri . The total loss function is the weighted sum of the first loss function and the second loss function.
[0155] Furthermore, according to the trained detection module and the trained similarity module, a target detection module is obtained.
[0156] In the embodiments of the present application, the AU detection is organically combined with the AU combination similarity. Through a multi-task framework, the AU combination similarity learning task is used to assist the AU detection task, which is more comprehensive and accurate than the existing AU detection tasks.
[0157] 102. Perform a similarity detection on the first face feature and the second face feature to obtain the similarity information between the first face feature and the second face feature.
[0158] Among them, the similarity information refers to the similarity between the AU combination in the first face image and the AU combination in the second face image.
[0159] In some embodiments, the step of "detecting the similarity between the first facial feature and the second facial feature to obtain the similarity information between the first facial feature and the second facial feature" may include the following operations:
[0160] Detect the similarity between the first facial feature and the second facial feature through the similarity module of the target detection model to obtain the similarity information between the first facial feature and the second facial feature.
[0161] Specifically, input the first facial feature and the second facial feature into the similarity module of the target detection network. Through the similarity module, perform linear + non-linear transformations on the first facial feature and the second facial feature to obtain the AU combination feature corresponding to the first facial feature and the AU combination feature corresponding to the second facial feature. Then, calculate the similarity between the two AU combination features as the similarity information of the AU combination.
[0162] 103. Perform facial action recognition processing on the first facial feature and the second facial feature respectively to obtain the first facial action information corresponding to the first facial feature and the second facial action information corresponding to the second facial feature.
[0163] Among them, the first facial action information refers to the probability of the first face image predicted in each AU category, and the second facial action information refers to the probability of the second face image predicted in each AU category.
[0164] In some embodiments, the step of "performing facial action recognition processing on the first facial feature and the second facial feature respectively to obtain the first facial action information corresponding to the first facial feature and the second facial action information corresponding to the second facial feature" may include the following operations:
[0165] Identify the facial action corresponding to the first facial feature through the detection module of the target detection model to obtain the first facial action information;
[0166] Identify the facial action corresponding to the second facial feature through the detection module of the target detection model to obtain the second facial action information.
[0167] Specifically, input the first facial feature into the AU attention map extraction module in the detection module. Through the AU attention map extraction module, extract the local facial action features in the first facial feature, and input the first facial feature into the AU feature extraction module in the detection module. Through the AU feature extraction module, extract the global facial action features in the first facial feature.
[0168] Furthermore, multiply the local facial action features by the global facial action features to obtain the facial action features corresponding to the first facial feature.
[0169] Then, the facial action features can be input into the AU classifier module in the detection module, and the AU classifier module calculates the probability values of the facial action features for each AU category, thereby obtaining the first facial action information.
[0170] Among them, the processing of the second face feature is the same as the above processing steps, including: inputting the second face feature into the AU attention map extraction module in the detection module, extracting the local facial action features in the second face feature through the AU attention map extraction module, and inputting the second face feature into the AU feature extraction module in the detection module, and extracting the global facial action features in the second face feature through the AU feature extraction module.
[0171] Further, multiplying the local facial action features by the global facial action features can obtain the facial action features corresponding to the second face feature.
[0172] Then, the facial action features can be input into the AU classifier module in the detection module, and the AU classifier module calculates the probability values of the facial action features for each AU category, thereby obtaining the second facial action information.
[0173] 104. Based on the similarity information, the first facial action information, and the second facial action information, determine the facial action similarity between the first face image and the second face image.
[0174] In some embodiments, the step of "based on the similarity information, the first facial action information, and the second facial action information, determine the facial action similarity between the first face image and the second face image" includes:
[0175] Based on the similarity information and a preset first weight, determine the first similarity;
[0176] According to the similarity between the first facial action information and the second facial action information, and a preset second weight, determine the second similarity;
[0177] Based on the first similarity and the second similarity, determine the facial action similarity between the first face image and the second face image.
[0178] In the embodiments of the present application, a preset first weight can be assigned to the similarity information output by the similarity module in the target detection model, and a preset second weight can be assigned to the similarity between the first facial action information and the second facial action information output by the detection module in the target detection model.
[0179] Among them, the similarity information includes the AU combination similarity between the first face image and the second face image, and the AU combination similarity can be multiplied by the preset first weight to obtain the first similarity of the AU information between the first face image and the second face image.
[0180] Among them, the similarity between the first facial action information and the second facial action information is calculated, and the similarity is multiplied by a preset second weight to obtain a second similarity of the AU information of the first face image and the second face image.
[0181] Furthermore, by adding the first similarity and the second similarity, the similarity of the AU information of the first face image and the second face image, that is, the facial action similarity, can be obtained finally.
[0182] In the embodiments of the present application, by using the AU combination similarity detected by the similarity module to assist the AU category information detected by the detection module, the AU detection accuracy of the target detection model can be improved.
[0183] In some embodiments, for the target detection network trained by this solution, the face feature extraction module and the similarity module in it can be used alone to perform expression similarity perception supervision on downstream tasks such as expression transfer, so as to further improve the effect of expression transfer.
[0184] For example, please refer to Figure 8 , Figure 8 which is a schematic diagram of an application scenario of another face detection method provided by the embodiments of the present application. Among them, the expression transfer module can be used to transfer the expression on one face image to another face image.
[0185] Among them, Iin refers to the face image used to extract the expression, and Iout refers to the image of the face expression applied with Iin. That is, Iin and another face image are input into the expression transfer module, and the face expression of Iin is transferred to another face image through the expression transfer module, and Iout is output.
[0186] Among them, the face feature extraction module is the face feature extraction module in the target detection model. Iin and Iout are input into the face feature extraction module, and the global face expression features of Iin are extracted through the face feature extraction module, and the global face expression features of Iout are extracted.
[0187] Among them, the similarity module refers to the similarity module in the target detection model. The global face expression features of Iin and the global face expression features of Iout are input into the similarity module. The AU combination features of Iin are extracted through the similarity module to obtain Qin, and the AU features of Iout are extracted to obtain Qout.
[0188] Furthermore, supervised learning is performed on Qin and Qout through the distance loss, so as to supervise the training of the expression transfer module to improve the expression transfer accuracy of the expression transfer module. Among them, the distance loss can be a preset distance loss function.
[0189] An embodiment of the present application discloses a face detection method, which includes: obtaining a first face image and a second face image; respectively extracting features from the first face image and the second face image to obtain a first face feature of the first face image and a second face feature of the second face image; performing a similarity detection on the first face feature and the second face feature to obtain similarity information between the first face feature and the second face feature; respectively performing facial action recognition processing on the first face feature and the second face feature to obtain first facial action information corresponding to the first face feature and second facial action information corresponding to the second face feature; and determining the facial action similarity between the first face image and the second face image based on the similarity information, the first facial action information, and the second facial action information. In this way, the effect of AU detection in face images can be improved.
[0190] To facilitate better implementation of the pop-up window management method provided by the embodiments of the present application, the embodiments of the present application further provide a face detection device based on the above face detection method. The meanings of the nouns are the same as those in the above face detection method, and the specific implementation details can be referred to the description in the method embodiments.
[0191] Please refer to Figure 9 , Figure 9 which is a structural block diagram of a face detection device provided by an embodiment of the present application. The device includes:
[0192] A first acquisition unit 301, which acquires a first face feature of a first face image and a second face feature of a second face image;
[0193] A first processing unit 302, which is used to perform a similarity detection on the first face feature and the second face feature to obtain similarity information between the first face feature and the second face feature;
[0194] A second processing unit 303, which is used to respectively perform facial action recognition processing on the first face feature and the second face feature to obtain first facial action information corresponding to the first face feature and second facial action information corresponding to the second face feature;
[0195] A determination unit 304, which is used to determine the facial action similarity between the first face image and the second face image based on the similarity information, the first facial action information, and the second facial action information.
[0196] In some embodiments, the determination unit 304 may include:
[0197] A first determination subunit, which is used to determine a first similarity based on the similarity information and a preset first weight;
[0198] A second determination subunit, configured to determine a second similarity based on a similarity between the first facial action information and the second facial action information, and a preset second weight;
[0199] A third determination subunit, configured to determine a facial action similarity between the first face image and the second face image based on the first similarity and the second similarity.
[0200] In some embodiments, the first acquisition unit 301 may include:
[0201] An input subunit, configured to input the first face image and the second face image into a target detection model;
[0202] An extraction subunit, configured to respectively perform feature extraction on the first face image and the second face image through a face feature extraction module of the target detection model to obtain the first face feature and the second face feature.
[0203] In some embodiments, the apparatus may further include:
[0204] A second acquisition unit, configured to acquire a triple sample of face images, where the triple sample includes an anchor face image, a positive face image, and a negative face image, and where actual facial action information of the positive face image is the same as actual facial action information of the anchor face image, and actual facial action information of the negative face image is different from actual facial action information of the anchor face image;
[0205] A construction unit, configured to construct the target detection model based on each sample face image in the triple sample and facial action information of each sample face image.
[0206] In some embodiments, the construction unit may include:
[0207] A processing subunit, configured to respectively perform preprocessing on each sample face image to obtain a processed sample face image corresponding to each sample face image;
[0208] A training subunit, configured to train a preset detection model based on the processed sample face images and the actual facial action information to obtain the target detection model.
[0209] In some embodiments, the processing subunit may specifically be configured to:
[0210] Perform key point detection on each sample face image to determine face key points in each sample face image;
[0211] Align the sample face image based on the face key points to obtain an aligned face image;
[0212] Crop the aligned face image based on the face region in the aligned face image to obtain the processed sample face image.
[0213] In some embodiments, the training subunit may specifically be used for:
[0214] Input the processed sample face image into the preset detection model, and respectively extract the sample face features corresponding to each sample face image through the face feature extraction module of the preset detection model;
[0215] Train the similarity module in the preset detection model based on the sample face features to obtain a trained similarity module;
[0216] Train the detection module in the preset detection model based on the sample face features to obtain a trained detection module;
[0217] Obtain the target detection model based on the trained similarity module and the trained detection module.
[0218] In some embodiments, the training subunit may specifically be used for:
[0219] Input the processed sample face image into the preset detection model, and respectively extract the sample face features corresponding to each sample face image through the face feature extraction module of the preset detection model;
[0220] Perform transformation processing on each sample face feature through the similarity module to obtain the transformed features corresponding to each sample face feature;
[0221] Calculate the first difference between the transformed features corresponding to the anchor sample face image and the transformed features corresponding to the positive sample face image;
[0222] Calculate the second difference between the transformed features corresponding to the anchor sample face image and the transformed features corresponding to the negative sample face image;
[0223] Adjust the similarity module based on the first difference, the second difference, and the first loss function to obtain the trained similarity module;
[0224] Train the detection module in the preset detection model based on the sample face features to obtain a trained detection module;
[0225] Obtain the target detection model based on the trained similarity module and the trained detection module.
[0226] In some embodiments, the training subunit may specifically be configured to:
[0227] Input the processed sample face images into the preset detection model, and respectively extract the sample face features corresponding to each sample face image through the face feature extraction module of the preset detection model;
[0228] Train the similarity module in the preset detection model based on the sample face features to obtain a trained similarity module;
[0229] Extract facial action features from each of the sample face features through the detection module;
[0230] Determine the predicted facial action information corresponding to each facial action feature;
[0231] Adjust the detection module based on the predicted facial action information, the actual facial action information, and the second loss function to obtain the trained detection module;
[0232] Obtain the target detection model based on the trained similarity module and the trained detection module.
[0233] In some embodiments, the training subunit may specifically be configured to:
[0234] Input the processed sample face images into the preset detection model, and respectively extract the sample face features corresponding to each sample face image through the face feature extraction module of the preset detection model;
[0235] Train the similarity module in the preset detection model based on the sample face features to obtain a trained similarity module;
[0236] Extract local facial action features from each of the sample face features through the first sub-module;
[0237] Extract global facial action features from each of the sample face features through the second sub-module;
[0238] Determine the facial action features based on the local facial action features and the global facial action features;
[0239] Determine the predicted facial action information corresponding to each facial action feature;
[0240] Adjust the detection module based on the predicted facial action information, the actual facial action information, and the second loss function to obtain the trained detection module;
[0241] Based on the trained similarity module and the trained detection module, the target detection model is obtained.
[0242] In some embodiments, the training subunit may specifically be configured to:
[0243] Input the processed sample face images into the preset detection model, and respectively extract the sample face features corresponding to each sample face image through the face feature extraction module of the preset detection model;
[0244] Train the similarity module in the preset detection model based on the sample face features to obtain the trained similarity module;
[0245] Extract facial action features from each of the sample face features through the detection module;
[0246] Calculate the probability values of each facial action feature on each preset facial action category through the third sub-module to obtain the predicted facial action information;
[0247] Adjust the detection module based on the predicted facial action information, the actual facial action information, and the second loss function to obtain the trained detection module;
[0248] Based on the trained similarity module and the trained detection module, the target detection model is obtained.
[0249] An embodiment of the present application discloses a face detection device, which obtains the first face feature of the first face image and the second face feature of the second face image through the first acquisition unit 301; the first processing unit 302 is a detection unit, which is used to perform a similarity detection on the first face feature and the second face feature to obtain the similarity information between the first face feature and the second face feature; the second processing unit 303 is used to perform facial action recognition processing on the first face feature and the second face feature respectively to obtain the first facial action information corresponding to the first face feature and the second facial action information corresponding to the second face feature; the determination unit 304 determines the facial action similarity between the first face image and the second face image based on the similarity information, the first facial action information, and the second facial action information. In this way, the effect of AU detection in face images can be improved.
[0250] Correspondingly, an embodiment of the present application further provides a computer device, and this computer device may be a server. As Figure 10 shown, Figure 10Schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 500 includes a processor 501 having one or more processing cores, a memory 502 having one or more computer-readable storage media, and a computer program stored on the memory 502 and executable on the processor. Among them, the processor 501 is electrically connected to the memory 502. Those skilled in the art can understand that the computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0251] The processor 501 is the control center of the computer device 500, connecting various parts of the entire computer device 500 through various interfaces and lines. By running or loading software programs and / or modules stored in the memory 502, and calling data stored in the memory 502, it executes various functions of the computer device 500 and processes data, thereby monitoring the computer device 500 as a whole.
[0252] In the embodiment of the present application, the processor 501 in the computer device 500 will load the instructions corresponding to the processes of one or more application programs into the memory 502 according to the following steps, and the processor 501 will run the application programs stored in the memory 502 to implement various functions:
[0253] Obtain the first facial feature of the first face image and the second facial feature of the second face image; perform a similarity detection on the first facial feature and the second facial feature to obtain the similarity information between the first facial feature and the second facial feature; perform facial action recognition processing on the first facial feature and the second facial feature respectively to obtain the first facial action information corresponding to the first facial feature and the second facial action information corresponding to the second facial feature; based on the similarity information, the first facial action information, and the second facial action information, determine the facial action similarity between the first face image and the second face image.
[0254] In this embodiment, multiple anchor samples are determined according to the AU label combination categories of face images in the dataset. Samples with the same AU information as the anchor samples are used as positive samples, and samples with different AU information from the anchor samples are used as negative samples. A triplet sample is constructed based on the anchor sample, the positive sample, and the negative sample, and the triplet sample is input into the global facial expression feature extraction module for feature extraction to obtain the feature vectors corresponding to each sample image. Further, the feature vectors are respectively sent to a similarity module and a detection module to learn the AU combination similarity information and AU category information in the input triplet sample, respectively. By organically combining AU detection and AU combination similarity learning, and using the AU combination similarity learning task to assist the AU detection task, the effect of AU detection in face images can be improved.
[0255] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated herein.
[0256] Optionally, as Figure 10 shown, the computer device 500 further includes: a touch display screen 503, a radio frequency circuit 504, an audio circuit 505, an input unit 506, and a power supply 507. Among them, the processor 501 is electrically connected to the touch display screen 503, the radio frequency circuit 504, the audio circuit 505, the input unit 506, and the power supply 507 respectively. Those skilled in the art can understand that Figure 10 the structure of the computer device shown in
[0257] does not limit the computer device, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0258] The radio frequency circuit 504 can be used to receive and transmit radio frequency signals to establish wireless communication with a network device or other computer devices through wireless communication, and to receive and transmit signals between the network device or other computer devices.
[0259] The audio circuit 505 can be used to provide an audio interface between the user and the computer device through a speaker and a microphone. The audio circuit 505 can transmit the electrical signal converted from the received audio data to the speaker, and the speaker converts it into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 505 and then converted into audio data. After the audio data is output to the processor 501 for processing, it is sent through the radio frequency circuit 504 to, for example, another computer device, or the audio data is output to the memory 502 for further processing. The audio circuit 505 may also include an earphone jack to provide communication between the peripheral earphone and the computer device.
[0260] The input unit 506 can be used to receive input digital, character information or user characteristic information (such as fingerprint, iris, facial information, etc.), and to generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0261] The power supply 507 is used to supply power to each component of the computer device 500. Optionally, the power supply 507 can be logically connected to the processor 501 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 507 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0262] Although Figure 10 not shown in the figure, the computer device 500 may also include a camera, a sensor, a Wi-Fi module, a Bluetooth module, etc., which will not be elaborated here.
[0263] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0264] As can be seen from the above, the computer device provided in this embodiment can obtain the first face feature of the first face image and the second face feature of the second face image; perform a similarity detection on the first face feature and the second face feature to obtain the similarity information between the first face feature and the second face feature; perform facial action recognition processing on the first face feature and the second face feature respectively to obtain the first facial action information corresponding to the first face feature and the second facial action information corresponding to the second face feature; and determine the facial action similarity between the first face image and the second face image based on the similarity information, the first facial action information, and the second facial action information.
[0265] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0266] Therefore, an embodiment of the present application provides a computer-readable storage medium in which multiple computer programs are stored. The computer programs can be loaded by a processor to execute the steps in any of the face detection methods provided in the embodiments of the present application. For example, the computer program can execute the following steps:
[0267] Obtain the first face feature of the first face image and the second face feature of the second face image;
[0268] Perform a similarity detection on the first face feature and the second face feature to obtain the similarity information between the first face feature and the second face feature;
[0269] Perform facial action recognition processing on the first face feature and the second face feature respectively to obtain the first facial action information corresponding to the first face feature and the second facial action information corresponding to the second face feature;
[0270] Determine the facial action similarity between the first face image and the second face image based on the similarity information, the first facial action information, and the second facial action information.
[0271] In the present embodiment, multiple anchor samples are determined by combining the AU label categories of face images in the dataset. Samples with the same AU information as the anchor samples are used as positive samples, and samples with different AU information from the anchor samples are used as negative samples. A triplet sample is constructed based on the anchor samples, positive samples, and negative samples, and the triplet sample is input into the global face expression feature extraction module for feature extraction to obtain the feature vectors corresponding to each sample image. Further, the feature vectors are respectively fed into a similarity module and a detection module to learn the AU combination similarity information and AU category information in the input triplet samples, respectively. By organically combining AU detection and AU combination similarity learning and using the AU combination similarity learning task to assist the AU detection task, the effect of AU detection in face images can be improved.
[0272] For the specific implementation of each of the above operations, reference can be made to the previous embodiments and will not be elaborated here.
[0273] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0274] Since the computer program stored in the storage medium can execute the steps in any of the face detection methods provided by the embodiments of the present application, the beneficial effects achievable by any of the face detection methods provided by the embodiments of the present application can be realized. For details, refer to the previous embodiments and will not be elaborated here.
[0275] The above has introduced in detail a face detection method, device, storage medium, and computer device provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A face detection method, characterized in that, The method includes: Obtaining the first facial feature of the first face image and the second facial feature of the second face image; Performing a similarity detection on the first facial feature and the second facial feature to obtain similarity information between the first facial feature and the second facial feature; Performing facial action recognition processing on the first facial feature and the second facial feature respectively to obtain first facial action information corresponding to the first facial feature and second facial action information corresponding to the second facial feature; Based on the similarity information, the first facial action information, and the second facial action information, determining the facial action similarity between the first face image and the second face image.
2. The method according to claim 1, wherein The determining the facial action similarity between the first face image and the second face image based on the similarity information, the first facial action information, and the second facial action information includes: Determining a first similarity based on the similarity information and a preset first weight; Determining a second similarity according to the similarity between the first facial action information and the second facial action information and a preset second weight; Based on the first similarity and the second similarity, determining the facial action similarity between the first face image and the second face image.
3. The method according to claim 1, characterized in that, The obtaining the first facial feature of the first face image and the second facial feature of the second face image includes: Inputting the first face image and the second face image into a target detection model; Performing feature extraction on the first face image and the second face image respectively through a facial feature extraction module of the target detection model to obtain the first facial feature and the second facial feature.
4. The method according to claim 3, wherein Before inputting the first face image and the second face image into the target detection model, the method further includes: Obtaining a triplet sample of face images, where the triplet sample includes an anchor face image, a positive face image, and a negative face image, and wherein the actual facial action information of the positive face image is the same as the actual facial action information of the anchor face image, and the actual facial action information of the negative face image is different from the actual facial action information of the anchor face image; Constructing the target detection model based on each face image in the triplet sample and the facial action information of each face image.
5. The method according to claim 4, wherein The constructing the target detection model based on each face image in the triplet sample and the facial action information of each face image includes: Performing preprocessing on each face image respectively to obtain a processed sample face image corresponding to each face image; Training a preset detection model based on the processed sample face image and the actual facial action information to obtain the target detection model.
6. The method according to claim 5, wherein The performing preprocessing on each face image respectively to obtain a processed sample face image corresponding to each face image includes: Performing key point detection on each face image to determine the face key points in each face image; Align the sample face image based on the face key points to obtain an aligned face image; Crop the aligned face image based on the face region in the aligned face image to obtain the processed sample face image.
7. The method according to claim 5, characterized in that, The training of the preset detection model based on the processed sample face image and the actual facial motion information to obtain the target detection model includes: Input the processed sample face image into the preset detection model, and extract the sample face features corresponding to each sample face image through the face feature extraction module of the preset detection model; Train the similarity module in the preset detection model based on the sample face features to obtain a trained similarity module; Train the detection module in the preset detection model based on the sample face features to obtain a trained detection module; Obtain the target detection model based on the trained similarity module and the trained detection module.
8. The method according to claim 7, characterized in that, The training of the similarity module in the preset detection model based on the sample face features to obtain a trained similarity module includes: Perform transformation processing on each sample face feature through the similarity module to obtain the transformed features corresponding to each sample face feature; Calculate the first difference between the transformed features corresponding to the anchor sample face image and the transformed features corresponding to the positive sample face image; Calculate the second difference between the transformed features corresponding to the anchor sample face image and the transformed features corresponding to the negative sample face image; Adjust the similarity module based on the first difference, the second difference, and the first loss function to obtain the trained similarity module.
9. The method according to claim 7, wherein The training of the detection module in the preset detection model based on the sample face features to obtain a trained detection module includes: Extract facial motion features from each sample face feature through the detection module; Determine the predicted facial motion information corresponding to each facial motion feature; Adjust the detection module based on the predicted facial motion information, the actual facial motion information, and the second loss function to obtain the trained detection module.
10. The method according to claim 9, wherein The detection module includes a first sub-module and a second sub-module; The extraction of facial motion features from each sample face feature through the detection module includes: Extract local facial motion features from each sample face feature through the first sub-module; Extract global facial motion features from each sample face feature through the second sub-module; Determine the facial motion feature based on the local facial motion feature and the global facial motion feature.
11. The method according to claim 9, wherein The detection module includes a third sub-module; The determination of the predicted facial motion information corresponding to each facial motion feature includes: Calculate the probability values of each facial motion feature in each preset facial motion category through the third sub-module to obtain the predicted facial motion information.
12. A face detection device, characterized in that, The device includes: A first acquisition unit for acquiring the first face feature of the first face image and the second face feature of the second face image; A first processing unit, configured to perform a similarity detection on the first face feature and the second face feature to obtain similarity information between the first face feature and the second face feature; A second processing unit, configured to perform facial action recognition processing on the first face feature and the second face feature respectively to obtain first facial action information corresponding to the first face feature and second facial action information corresponding to the second face feature; A determination unit, configured to determine a facial action similarity between the first face image and the second face image based on the similarity information, the first facial action information, and the second facial action information.
13. A computer device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, wherein, When the processor executes the program, it implements the face detection method according to any one of claims 1 to 11.
14. A storage medium, characterized in that, The storage medium stores a plurality of instructions, and the instructions are adapted to be loaded by the processor to execute the face detection method according to any one of claims 1 to 11.