Facial action recognition method, apparatus, device, and medium
The automatic recognition of facial movements through a self-supervised learning framework solves the low efficiency problem of relying on manual recognition in existing technologies and achieves efficient facial movement recognition.
Patent Information
- Application Number
- CN202310183993.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-21
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2043-02-21
AI Technical Summary
In the existing technology, facial action recognition mainly relies on manual methods, which is inefficient especially when the number of images is large, affecting the execution of subsequent operations.
Through a self-supervised learning framework, the backbone network, contrastive learning components and predictive learning components are used to automatically recognize facial actions. Facial actions are recognized by dividing the facial area and calculating the loss value, reducing manual intervention.
It realizes automatic recognition of facial movements, improves recognition efficiency, reduces the need for manual labeling, and improves recognition accuracy and efficiency.
Smart Images

Figure CN116110107B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a facial action recognition method, apparatus, device and medium. Background Art
[0002] When performing certain operations based on facial action images, such as training models, it is necessary to identify and label facial actions in facial images. Currently, this is mainly achieved manually, and the recognition efficiency is low, especially when the number of images is large.
[0003] Application Contents
[0004] The embodiments of the present application provide a facial action recognition method, apparatus, device, and medium, which can improve the recognition efficiency of facial actions.
[0005] In a first aspect, an embodiment of the present application provides a facial action recognition method, comprising:
[0006] Acquire 2N first facial images, where the first facial image includes at least one first facial feature point;
[0007] For each first facial image, dividing the first facial image according to positions of the first facial feature points to obtain P facial regions;
[0008] determining a first loss value between the first facial region and the second facial region based on a first eigenvector corresponding to a first facial feature point in the first facial region and a second eigenvector corresponding to the first facial feature point in the second facial region, where the first facial region and the second facial region are at least one facial region in the 2N first facial images;
[0009] determining a second loss value between the second facial image and the third facial image based on a third eigenvector of the second facial image and a fourth eigenvector of the third facial image, where the second facial image and the third facial image are any two images from the two first facial images;
[0010] determining a third loss value between the third facial region and the fourth facial region based on a similarity between the third facial region and the fourth facial region, where the third facial region and the fourth facial region are correlated facial regions in the same first facial image;
[0011] Facial actions of the first facial image are identified based on the first loss value, the second loss value, and the third loss value.
[0012] In a second aspect, an embodiment of the present application provides a facial action recognition device, comprising an acquisition module, a division module, a determination module, and a recognition module;
[0013] An acquisition module, configured to acquire 2N first facial images, where the first facial image includes at least one first facial feature point;
[0014] a segmentation module, configured to segment each first facial image according to positions of the first facial feature points to obtain P facial regions;
[0015] a determining module configured to determine a first feature vector corresponding to a first facial feature point in a first facial region and a second feature vector corresponding to the first facial feature point in a second facial region, and to determine a first loss value between the first facial region and the second facial region, where the first facial region and the second facial region are at least one facial region in the 2N first facial images;
[0016] The determining module is further configured to determine a second loss value between the second facial image and the third facial image based on a third eigenvector of the second facial image and a fourth eigenvector of the third facial image, where the second facial image and the third facial image are any two images among the 2N first facial images;
[0017] The determining module is further configured to determine a third loss value between the third facial region and the fourth facial region based on a similarity between the third facial region and the fourth facial region, where the third facial region and the fourth facial region are correlated facial regions in the same first facial image;
[0018] The recognition module is used to recognize the facial action of the first facial image according to the first loss value, the second loss value and the third loss value.
[0019] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0020] processor;
[0021] a memory for storing computer program instructions;
[0022] When the computer program instructions are executed by a processor, the method according to the first aspect is implemented.
[0023] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon, which implement the method described in the first aspect when the computer program instructions are executed by a processor.
[0024] In a fifth aspect, an embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the method described in the first aspect.
[0025] In the embodiment of the present application, for each acquired first facial image, the first facial image is divided according to the position of the first facial feature point to obtain P facial regions; a first loss value between the first facial region and the second facial region is determined based on the feature vector corresponding to the first facial feature point; a second loss value between the second facial image and the third facial image is determined based on the feature vector of the second facial image and the third facial image; a third loss value between the third facial region and the fourth facial region is determined based on the similarity between the third facial region and the fourth facial region; and facial actions of the first facial image are identified based on the first loss value, the second loss value, and the third loss value. That is, by determining the loss values between facial regions and the loss values between facial images, the embodiment of the present application realizes automatic recognition of facial actions, eliminating the need for manual recognition, thereby improving the efficiency of facial action recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0027] Figure 1 A schematic diagram of a facial action recognition method provided in an embodiment of the present application;
[0028] Figure 2 A flowchart of a facial action recognition method provided in an embodiment of the present application;
[0029] Figure 3 A schematic diagram of displaying a facial area provided in an embodiment of the present application;
[0030] Figure 4 A schematic diagram of the relationship between facial regions provided in an embodiment of the present application;
[0031] Figure 5 A structural diagram of a facial action recognition device provided in an embodiment of the present application;
[0032] Figure 6 A structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and Examples. It should be understood that the specific embodiments described herein are only configured to explain the present application and are not configured to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.
[0034] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.
[0035] In scenarios where facial movements in facial images need to be recognized, they are currently mainly identified and labeled manually. Especially when the number of facial images is large, this greatly reduces the recognition efficiency and affects the execution of subsequent operations.
[0036] To this end, embodiments of the present application provide a facial action recognition method, apparatus, device, and medium that can improve the efficiency of facial action recognition.
[0037] The facial action recognition method provided in the embodiment of the present application can be applied to Figure 1 The self-supervised learning framework shown may include a backbone network 110 , a contrastive learning component 120 , a predictive learning component 130 , and a facial action recognition model 140 .
[0038] Among them, the backbone network 110 is used to enhance the acquired original facial images respectively, then extract the local representation of each facial area, map it to a low-dimensional latent space, and use the contrast learning component 120, the predictive learning component 130 and the facial action recognition model 140 to recognize facial actions in the low-dimensional latent space.
[0039] Among them, the contrast learning component 120 is used to determine the differences within the region, and the predictive learning component 130 is used to further enhance the differences obtained by the contrast learning component 120 based on the differences between the correlation regions. By combining the contrast component 120 and the predictive learning component 130, not only can different facial regions be distinguished, but also the co-occurrence relationship between facial regions can be obtained.
[0040] The differences between the facial regions obtained by the comparison component 120 and the relationships between the facial regions obtained by the prediction learning component 130 are input as supervisory signals into the facial action recognition model 140 to obtain facial action recognition results.
[0041] pass Figure 1 The framework shown can automatically recognize facial actions in facial images without the need for manual work, thereby improving recognition efficiency.
[0042] The facial action recognition method provided by the embodiment of the present application is described below in combination with the above framework and specific embodiments. Figure 2 This is a flowchart of a facial action recognition method provided in an embodiment of the present application. The method can be applied to electronic devices, including but not limited to mobile phones, tablet computers, laptop computers, PDAs, etc.
[0043] like Figure 2 As shown, the facial action recognition method may include the following steps:
[0044] S210: Acquire 2N first facial images.
[0045] The first facial image includes at least one first facial feature point.
[0046] S220 . For each first facial image, divide the first facial image according to positions of the first facial feature points to obtain P facial regions.
[0047] S230: Determine a first loss value between the first facial region and the second facial region based on a first feature vector corresponding to a first facial feature point in the first facial region and a second feature vector corresponding to the first facial feature point in the second facial region.
[0048] The first facial region and the second facial region are at least one facial region in the 2N first facial images.
[0049] S240: Determine a second loss value between the second facial image and the third facial image based on the third eigenvector of the second facial image and the fourth eigenvector of the third facial image.
[0050] The second facial image and the third facial image are any two images among the 2N first facial images.
[0051] S250: Determine a third loss value between the third facial region and the fourth facial region based on the similarity between the third facial region and the fourth facial region.
[0052] The third facial region and the fourth facial region are correlated facial regions in the same first facial image.
[0053] S260: Identify facial actions of the first facial image according to the first loss value, the second loss value, and the third loss value.
[0054] The embodiment of the present application divides each acquired first facial image into P facial regions based on the positions of the first facial feature points; determines a first loss value between the first facial region and the second facial region based on the feature vectors corresponding to the first facial feature points; determines a second loss value between the second facial image and the third facial image based on the feature vectors of the second facial image and the third facial image; determines a third loss value between the third facial region and the fourth facial region based on the similarity between the third facial region and the fourth facial region; and identifies facial actions in the first facial image based on the first loss value, the second loss value, and the third loss value. That is, the embodiment of the present application achieves automatic recognition of facial actions by determining the loss values between facial regions and the loss values between facial images, eliminating the need for manual recognition, thereby improving the efficiency of facial action recognition.
[0055] The above steps are explained in detail below:
[0056] In S210, the first facial image may be an image containing the user's face, that is, an image containing a human face. For example, the first facial image may be an enhanced image, that is, an image obtained by performing enhancement processing on the original image.
[0057] The first facial image may include at least one first facial feature point. The first facial feature point may be a feature point with iconic significance for each facial area, that is, a feature point that can represent the facial area. For example, for the facial area corresponding to the mouth, the first facial feature point may include the left lip corner, the right lip corner, the midpoint of the upper lip, and the midpoint of the lower lip.
[0058] Exemplarily, the first facial image may be obtained in the following manner:
[0059] Obtain W original facial images, W = N;
[0060] For each original facial image, performing enhancement processing according to a first enhancement method and a second enhancement method, respectively, to obtain a first enhanced image and a second enhanced image;
[0061] The first enhanced image and the second enhanced image are determined as a first facial image.
[0062] The embodiment of the present application does not limit the method for obtaining the original facial image. For example, the original facial image can be obtained from the Internet or from a pre-built image dataset.
[0063] Taking into account the possible presence of noise and distortion in the original facial image, in order to improve the accuracy of the facial action recognition results, the original facial image can be enhanced, for example. The embodiment of the present application takes the enhancement of the original facial image according to two enhancement methods as an example. In actual application, one or more enhancement methods can also be used.
[0064] Specifically, the original facial image may be enhanced according to a first enhancement method and a second enhancement method, respectively, to obtain a first enhanced image and a second enhanced image, which provide a basis for subsequent recognition of facial actions.
[0065] Exemplarily, the first enhancement method may be a low-pass filtering method, that is, only low-frequency signals of the original facial image are allowed to pass, thereby removing noise in the original facial image.
[0066] The second enhancement method can be a high-pass filtering method, that is, enhancing high-frequency signals such as edges to improve the clarity of the original facial image. Of course, the first enhancement method and the second enhancement method can also be other methods, which are not limited in this embodiment of the application.
[0067] In S220 , considering that different facial regions in the facial image correspond to different facial actions, the facial image may be divided in order to more accurately identify the user's facial actions.
[0068] For example, in an embodiment of the present application, the first facial image may be divided based on the positions of the first facial feature points to obtain P facial regions.
[0069] For example, refer to Figure 3 Based on the position of the first facial feature point, the first facial image can be divided into 8 facial areas, namely area 1 (forehead), area 2 (left eye), area 3 (right eye), area 4 (nose), area 5 (left cheek), area 6 (right cheek), area 7 (mouth), and area 8 (chin), each of which can be a rectangular area.
[0070] The appearance changes and facial actions corresponding to each facial region can be found in Table 1, which exemplifies the relationship between facial appearance changes associated with certain facial actions and the corresponding facial regions. Each facial action can be determined by jointly observing the appearance changes of multiple different facial regions. For example, the facial regions associated with AU1 include Region 1, Region 2, and Region 3.
[0071] Table 1
[0072]
[0073]
[0074] When different facial actions are activated, the corresponding appearance changes of different facial regions are also different, and there are also correlations between different facial regions. For example, as shown in Table 1, AU7 and AU8 both cause appearance changes in regions 5, 6, and 7. When AU7 is activated, the lip at the corner of region 7 elongates and tilts, and the lower nasolabial groove deepens in regions 5 and 6 simultaneously. When AU8 is activated, the inner side of the upper lip in region 7 is raised, and the upper nasolabial groove deepens in regions 5 and 6 simultaneously. This shows that the appearance changes of regions 5, 6, and 7 activated by different facial actions are highly consistent.
[0075] In some embodiments, each facial region can be mapped to a low-dimensional space, and facial action recognition can be completed in the low-dimensional space. For example, for a first facial image of 224*224RGB, a 4096-dimensional local representation of each facial region is obtained by processing, and then the local representation is passed through a multi-layer perceptron (the size of the hidden layer is 2048), and a 128-dimensional vector is output for subsequent calculations, which reduces the amount of calculation and improves efficiency. The above process can be performed in Figure 1 This is done in the backbone network shown.
[0076] In S230, the first facial area and the second facial area can be at least one facial area in the 2N first facial images. Exemplarily, the first facial area and the second facial area can be the same facial area in the same first facial image, for example, both are area 1 in the i-th first facial image.
[0077] Exemplarily, the first facial region and the second facial region may also be different facial regions in the same first facial image. For example, the first facial region may be region 1 in the i-th first facial image, and the second facial region may be region 3 in the i-th first facial image.
[0078] Exemplarily, the first facial region and the second facial region may also be the same facial region in different first facial images. For example, the first facial region may be region 1 in the i-th first facial image, and the second facial region may be region 1 in the j-th first facial image, where i is not equal to j.
[0079] Exemplarily, the first facial region and the second facial region may also be different facial regions in different first facial images. For example, the first facial region may be region 1 in the i-th first facial image, and the second facial region may be region 2 in the j-th first facial image.
[0080] The first feature vector is used to represent the position information of the first facial feature point, and the loss value between the two facial regions can be determined according to the position information of the first facial feature point.
[0081] Exemplarily, the above S230 may include the following steps:
[0082] determining a first similarity between the first facial region and the second facial region based on a first feature vector corresponding to a first facial feature point in the first facial region and a second feature vector corresponding to the first facial feature point in the second facial region;
[0083] A first loss value between the first facial region and the second facial region is determined based on the first similarity and the association relationship between the first facial region and the second facial region.
[0084] The first similarity is used to indicate the degree of similarity between the first facial region and the second facial region. The greater the first similarity, the more similar the first facial region and the second facial region are. For example, the similarity between the first facial region and the second facial region can be measured using a cosine model, which is simple and convenient. The similarity between u and v can be expressed as follows using the cosine model:
[0085]
[0086] Among them, u and v are two parameters whose similarity needs to be calculated. For example, in the embodiment of the present application, they can represent the feature vector of the feature point in the first facial region and the feature vector of the feature point in the second facial region respectively.
[0087] For example, the first similarity between the first facial region and the second facial region can be expressed as in, represents the feature vector of the pth first feature point in the i-th first facial image, that is, the first facial area in the i-th first facial image, The feature vector representing the qth first feature point in the jth first facial image, that is, the second facial region in the jth first facial image, is taken as an example here to represent one facial region by one first feature point.
[0088] The association relationship of the first face region and the second face region can include whether the first face region and the second face region are the same region or a symmetric region (if the symmetric region exists), and according to the association relationship of the first face region and the second face region and the first similarity of the first face region and the second face region, a loss value of the first face region and the second face region can be determined.
[0089] Exemplarily,
[0090]
[0091]
[0092] wherein, L a is a loss value between the first face region and the second face region, that is, a first loss value, p represents a pth face region, that is, a pth first face feature point, is a positive pair number related to For example, if the first face region and the second face region are the same or a symmetric region, they are considered as a positive data pair, otherwise, they are considered as a negative data pair. []· is a function belonging to 0 or 1, for example, when the condition is true, II []· = 1, otherwise II []· = 0, for example, in the embodiment of the present application, i≠j and p≠q are true, then II [i≠j V p≠q] = 1. represents the similarity of the first face region and the second face region . Φ(p) represents a set of regions, including p and its symmetric region (if the symmetric region exists), and τ is a temperature parameter.
[0093] The above loss function pulls the positive data pair closer and pushes the negative data pair away, so that the diversity of features in the region can be maintained.
[0094] In S240, in addition to considering the difference of each face region, the difference of the face images can also be considered, for example, in the embodiment of the present application, the difference between the two face images can be determined based on the feature vectors of the two face images, and a second loss value is obtained.
[0095] Exemplarily, the above S240 can include the following steps:
[0096] According to the third feature vector of the fifth face region in the second face image and the sixth feature vector of the sixth face region in the third face image, a second similarity of the second face image and the third face image is determined, and the fifth face region and the sixth face region are the same face region.
[0097] A second loss value between the second facial image and the first facial image is determined based on the second similarity.
[0098] The fifth facial region and the sixth facial region are the same facial region in different facial images. For example, the fifth facial region and the sixth facial region may both be Figure 3 Area 1 shown.
[0099] For example,
[0100]
[0101] Among them, L b is the loss value of the second facial image and the third facial image, that is, the second loss value, i and k represent the i-th facial image and the k-th facial image, that is, the second facial image and the third facial image, respectively. represents another enhanced image of the same original facial image, i.e. It represents the similarity between two enhanced images corresponding to the same original facial image. Represents the similarity between different enhanced images. k≠i is true, that is, II [k≠i]· =1.
[0102] In addition to considering the differences between two facial regions, the embodiments of the present application also consider the differences between two facial images, maintaining the diversity of features, thereby improving the accuracy of facial action recognition results.
[0103] The above process can be Figure 1 The contrastive learning component shown in FIG5 is used to obtain the difference within the region as a supervision signal while maintaining the diversity of features within the region.
[0104] In S250, as shown in Table 1, when a facial action is activated, it may cause the appearance of multiple facial regions to change. Due to the co-occurrence relationship between regions of appearance change, the appearance of a facial region should be predicted from its related facial regions. In other words, each facial action requires the joint observation of multiple facial regions. Based on this, the embodiment of the present application considers the similarity between related facial regions in the same facial image, and determines the loss value of two facial regions based on this similarity, providing a basis for subsequent facial action recognition.
[0105] The third facial region and the fourth facial region are facial regions with correlation in the same first facial image. The correlation between the facial regions can be seen in Figure 4For example, region 1 is correlated with region 2 and region 3, that is, region 1 and region 2 are correlated facial regions. Region 1 and region 7 are uncorrelated facial regions, that is, region 1 and region 7 are not correlated.
[0106] Based on the third facial region, illustratively, before S250 , a fourth facial region correlated with the third facial region may be determined in the following manner:
[0107] determining, from the candidate facial regions, a target facial region having correlation with the third facial region based on the fifth eigenvector of the third facial region, the correlation between the third facial region and the candidate facial region, and a predictor between the third facial region and the candidate facial region, the candidate facial region being a facial region other than the third facial region in the first facial image;
[0108] The target facial region is determined as the fourth facial region.
[0109] For example,
[0110] in, is the prediction factor between two facial regions. In this embodiment of the application, 19 prediction factors are preset for the relationship between two facial regions to represent the correlation between the two facial regions. K is the number of facial regions related to region p. Here, region p can be the third facial region, and region q can be the fourth facial region. The feature vector of region p can be expressed as The eigenvector of region q can be expressed as or II [q~p] is a function that evaluates to 1 if region p is correlated with region q and 0 otherwise.
[0111] For example, the relationship between two facial regions can be predicted by a predictor, e.g. Figure 4 Each arrow in can correspond to a predictor for predicting the relevant area corresponding to the facial area.
[0112] For region p, the embodiment of the present application uses the average value of multiple prediction results from other related facial regions from the same facial image as the feature vector of the fourth facial region. For example, if there are two facial regions related to region p, the average value of the feature vectors of these two facial regions can be used as the feature vector of the fourth facial region.
[0113] After the fourth facial region is determined, a third loss value between the third facial region and the fourth facial region may be determined based on the similarity between the third facial region and the fourth facial region.
[0114] Exemplarily, “determining a third loss value between the third facial region and the fourth facial region based on the similarity between the third facial region and the fourth facial region” may include the following steps:
[0115] The third loss value between the third facial region and the fourth face is determined by the following relationship:
[0116]
[0117] Among them, L pre represents the loss value between the third facial region and the fourth facial region, that is, the third loss value, represents the similarity between the third facial region and the fourth facial region, (can also be represented by q) represents the fourth facial area, and the explanation of other parameters can refer to the above embodiment.
[0118] The embodiment of the present application utilizes the correlation between different facial areas of the same facial image to determine the loss value between two related facial areas, so that the diversity of features can be further enhanced on the basis of the above embodiment, thereby improving the accuracy of facial action recognition results.
[0119] The above process can be Figure 1 The prediction learning component shown is performed.
[0120] In S260 , facial actions of the first facial image may be recognized based on the first loss value, the second loss value, and the third loss value.
[0121] Exemplarily, the above S260 may include the following steps:
[0122] weighting the first loss value, the second loss value, and the third loss value according to a first weight coefficient of the first loss value, a second weight coefficient of the second loss value, and a third weight coefficient of the third loss value to obtain a first weighted result;
[0123] The first weighted result is input into a pre-trained facial action recognition model to obtain a facial action label of the first facial image, where the facial action label is used to characterize the facial action of the first facial image.
[0124] For example, L con =L a +λL b , λ is the weight of the diversity of features in the balance area, L = αL con +βL pre =α(L a +λL b )+βL pre =αL a +αλL b +βLpre , where the first weight coefficient can be α, the second weight coefficient can be αλ, and the third weight coefficient can be β, α and β are used to balance L con and L pre .
[0125] The first weighted result L is input into the pre-trained facial action recognition model to obtain the facial action label of the first facial image to characterize the facial action of the first facial image. Figure 1 The facial action recognition model shown, Figure 1 The framework shown needs to be trained with training samples before practical application.
[0126] The embodiments of this application are combined Figure 1 The framework shown in the figure can more completely and comprehensively identify the typical facial actions shown in Table 1 by dividing the face into 8 regions, thereby improving the performance of facial action recognition. The entire process does not require manual labeling and annotation, thus saving manpower and improving recognition efficiency.
[0127] Based on the same inventive concept, the embodiment of the present application also provides a facial action recognition device. Figure 5 The facial action recognition device provided in the embodiment of the present application is described in detail.
[0128] Figure 5 This is a structural diagram of a facial action recognition device provided in an embodiment of the present application.
[0129] like Figure 5 As shown, the facial action recognition device may include an acquisition module 510, a division module 520, a determination module 530 and a recognition module 540;
[0130] An acquisition module 510 is configured to acquire 2N first facial images, where the first facial image includes at least one first facial feature point;
[0131] a segmentation module 520 for segmenting each first facial image according to positions of the first facial feature points to obtain P facial regions;
[0132] a determination module 530 for determining a first loss value between a first facial region and a second facial region based on a first feature vector corresponding to a first facial feature point in the first facial region and a second feature vector corresponding to the first facial feature point in the second facial region, where the first facial region and the second facial region are at least one facial region in the 2N first facial images;
[0133] Determining module 530 is further configured to determine a second loss value between the second facial image and the third facial image based on a third eigenvector of the second facial image and a fourth eigenvector of the third facial image, where the second facial image and the third facial image are any two images from the 2N first facial images;
[0134] The determining module 530 is further configured to determine a third loss value between the third facial region and the fourth facial region based on a similarity between the third facial region and the fourth facial region, where the third facial region and the fourth facial region are related facial regions in the same first facial image;
[0135] The recognition module 540 is configured to recognize the facial action of the first facial image according to the first loss value, the second loss value, and the third loss value.
[0136] The embodiment of the present application divides each acquired first facial image into P facial regions according to the positions of the first facial feature points; determines a first loss value between the first facial region and the second facial region based on the feature vectors corresponding to the first facial feature points; determines a second loss value between the second facial image and the third facial image based on the feature vectors of the second facial image and the third facial image; determines a third loss value between the third facial region and the fourth facial region based on the similarity between the third facial region and the fourth facial region; and identifies facial actions in the first facial image based on the first loss value, the second loss value, and the third loss value. That is, the embodiment of the present application achieves automatic recognition of facial actions by determining the loss values between facial regions and the loss values between facial images, eliminating the need for manual recognition, thereby improving the efficiency of facial action recognition.
[0137] In some embodiments, the determination module 530 is specifically configured to:
[0138] determining a first similarity between the first facial region and the second facial region based on a first feature vector corresponding to a first facial feature point in the first facial region and a second feature vector corresponding to the first facial feature point in the second facial region;
[0139] A first loss value between the first facial region and the second facial region is determined based on the first similarity and an association relationship between the first facial region and the second facial region and the first facial image.
[0140] In some embodiments, the determination module 530 is specifically configured to:
[0141] determining, based on a third eigenvector of a fifth facial region in the second facial image and a fourth eigenvector of a sixth facial region in the third facial image, a second similarity between the second facial image and the third facial image, the fifth facial region and the sixth facial region being the same facial region;
[0142] A second loss value between the second facial image and the first facial image is determined based on the second similarity.
[0143] In some embodiments, the determination module 530 is specifically configured to:
[0144] The third loss value between the third facial region and the fourth facial region is determined by the following relationship:
[0145]
[0146] Among them, L pre represents the third loss value between the third facial region and the fourth facial region, represents the similarity between the third facial region and the fourth facial region, p represents the third facial region, represents the fourth facial region, The feature vector representing the third facial region, A feature vector representing the fourth facial region.
[0147] In some embodiments, the determination module 530 is further configured to determine, before determining the third loss value between the third facial region and the fourth facial region based on the similarity between the third facial region and the fourth facial region, a target facial region having correlation with the third facial region from among the candidate facial regions based on the fifth eigenvector of the third facial region, the correlation between the third facial region and the candidate facial region, and the prediction factor between the third facial region and the candidate facial region, where the candidate facial region is a facial region other than the third facial region in the first facial image;
[0148] The target facial region is determined as the fourth facial region.
[0149] In some embodiments, the identification module 540 is specifically configured to:
[0150] weighting the first loss value, the second loss value, and the third loss value according to a first weight coefficient of the first loss value, a second weight coefficient of the second loss value, and a third weight coefficient of the third loss value to obtain a first weighted result;
[0151] The first weighted result is input into a pre-trained facial action recognition model to obtain a facial action label of the first facial image, where the facial action label is used to characterize the facial action of the first facial image.
[0152] In some embodiments, the acquisition module 510 is specifically configured to:
[0153] Obtain W original facial images, W = N;
[0154] For each original facial image, performing enhancement processing according to a first enhancement method and a second enhancement method, respectively, to obtain a first enhanced image and a second enhanced image;
[0155] The first enhanced image and the second enhanced image are determined as a first facial image.
[0156] Figure 5 Each module in the device shown has the function of realizing Figure 1-Figure 4 The functions of each step in the process can achieve the corresponding technical effects, so for the sake of brevity, they will not be described here in detail.
[0157] Based on the same inventive concept, the embodiment of the present application also provides an electronic device, which may be, for example, a mobile phone, a tablet computer, a notebook computer, a PDA, etc. Figure 6 The electronic device provided in the embodiments of the present application is described in detail.
[0158] like Figure 6 As shown, the electronic device may include a processor 610 and a memory 620 for storing computer program instructions.
[0159] The processor 610 may include a central processing unit (CPU) or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0160] The memory 620 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 620 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In one example, the memory 620 may include a removable or non-removable (or fixed) medium, or the memory 620 may be a non-volatile solid-state memory. In one example, the memory 620 may be a read-only memory (ROM). In one example, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.
[0161] The processor 610 reads and executes the computer program instructions stored in the memory 620 to implement Figure 1-Figure 4The method in the embodiment shown in FIG. Figure 1-Figure 4 The corresponding technical effects achieved by executing the method in the illustrated embodiment are not described in detail here for the sake of brevity.
[0162] In one example, the electronic device may further include a communication interface 630 and a bus 640. Figure 6 As shown, the processor 610 , the memory 620 , and the communication interface 630 are connected via a bus 640 and communicate with each other.
[0163] The communication interface 630 is mainly used to implement communication between various modules, devices and / or equipment in the embodiments of the present application.
[0164] Bus 640 includes hardware, software or both, and each component of electronic equipment is coupled to each other.For example, but not limitation, bus 440 may include accelerated graphics end (Accelerated Graphics Port, AGP) or other graphics buses, enhanced industry standard architecture (Extended Industry Standard Architecture, EISA) bus, front side bus (Front Side Bus, FSB), hyper transport (Hyper Transport, HT) interconnection, industry standard architecture (Industry Standard Architecture, ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 640 may include one or more buses. Although the present application embodiment describes and shows specific bus, the application considers any suitable bus or interconnection.
[0165] After acquiring 2N first facial images, the electronic device can execute the facial action recognition method in the embodiment of the present application, thereby realizing the combination of Figure 1-Figure 4 The facial action recognition method described and Figure 5 A facial action recognition device is described.
[0166] In addition, in conjunction with the facial action recognition method in the above embodiments, embodiments of the present application may provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any of the facial action recognition methods in the above embodiments is implemented.
[0167] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.
[0168] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0169] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0170] Aspects of the present invention are described above with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present invention.It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions.These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed via the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram.Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit.It is also understood that each box in the block diagram and / or flowchart and the combination of the boxes in the block diagram and / or flowchart can also be implemented by the dedicated hardware that performs the specified function or action, or can be implemented by the combination of dedicated hardware and computer instructions.
[0171] The above description is only a specific embodiment of the present application. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the scope of protection of the present application.
Claims
1. A facial action recognition method, characterized in that: include: Acquire 2N first facial images, where the first facial image includes at least one first facial feature point; For each of the first facial images, dividing the first facial image according to positions of the first facial feature points to obtain P facial regions; determining a first loss value between the first facial region and a second facial region based on a first eigenvector corresponding to a first facial feature point in the first facial region and a second eigenvector corresponding to the first facial feature point in the second facial region, where the first facial region and the second facial region are at least one facial region in the 2N first facial images; determining a second loss value between the second facial image and the third facial image based on a third eigenvector of the second facial image and a fourth eigenvector of the third facial image, where the second facial image and the third facial image are any two images among the 2N first facial images; determining a third loss value between the third facial region and the fourth facial region based on a similarity between the third facial region and the fourth facial region, the third facial region and the fourth facial region being correlated facial regions in the same first facial image; Facial actions of the first facial image are identified based on the first loss value, the second loss value, and the third loss value.
2. The method according to claim 1, characterized in that The determining, based on a first feature vector corresponding to a first facial feature point in the first facial region and a second feature vector corresponding to the first facial feature point in the second facial region, a first loss value between the first facial region and the second facial region includes: determining a first similarity between the first facial region and the second facial region based on a first feature vector corresponding to a first facial feature point in the first facial region and a second feature vector corresponding to the first facial feature point in the second facial region; A first loss value between the first facial region and the second facial region is determined based on the first similarity and the association relationship between the first facial region and the second facial region.
3. The method according to claim 1, characterized in that The determining, based on the third eigenvector of the second facial image and the fourth eigenvector of the third facial image, a second loss value between the second facial image and the third facial image includes: determining a second similarity between the second facial image and the third facial image based on a third eigenvector of a fifth facial region in the second facial image and a fourth eigenvector of a sixth facial region in the third facial image, wherein the fifth facial region and the sixth facial region are the same facial region; A second loss value between the second facial image and the first facial image is determined based on the second similarity.
4. The method according to claim 1, wherein The determining, based on the similarity between the third facial region and the fourth facial region, a third loss value between the third facial region and the fourth facial region includes: The third loss value between the third facial region and the fourth facial region is determined by the following relationship: Among them, L pre represents the third loss value between the third facial region and the fourth facial region, represents the similarity between the third facial region and the fourth facial region, p represents the third facial region, represents the fourth facial region, The feature vector representing the third facial region, A feature vector representing the fourth facial region.
5. The method according to claim 1, wherein Before determining a third loss value between the third facial region and the fourth facial region based on the similarity between the third facial region and the fourth facial region, the method further includes: determining, from the candidate facial regions, a target facial region having a correlation with the third facial region based on a fifth eigenvector of the third facial region, a correlation between the third facial region and a candidate facial region, and a prediction factor between the third facial region and the candidate facial region, the candidate facial region being a facial region other than the third facial region in the first facial image; The target facial region is determined as a fourth facial region.
6. The method according to claim 1, characterized in that The identifying the facial action of the first facial image according to the first loss value, the second loss value, and the third loss value includes: weighting the first loss value, the second loss value, and the third loss value according to a first weight coefficient of the first loss value, a second weight coefficient of the second loss value, and a third weight coefficient of the third loss value to obtain a first weighted result; The first weighted result is input into a pre-trained facial action recognition model to obtain a facial action label of the first facial image, where the facial action label is used to characterize the facial action of the first facial image.
7. The method according to claim 1, characterized in that The obtaining of N first facial images includes: Obtain W original facial images, W = N; For each of the original facial images, performing enhancement processing according to a first enhancement method and a second enhancement method, respectively, to obtain a first enhanced image and a second enhanced image; The first enhanced image and the second enhanced image are determined as the first facial image.
8. A facial action recognition device, characterized in that: It includes an acquisition module, a division module, a determination module and an identification module; The acquisition module is configured to acquire 2N first facial images, where the first facial image includes at least one first facial feature point; The segmentation module is configured to segment each of the first facial images according to the positions of the first facial feature points to obtain P facial regions; The determining module is configured to determine a first loss value between the first facial region and a second facial region based on a first feature vector corresponding to a first facial feature point in the first facial region and a second feature vector corresponding to the first facial feature point in the second facial region, where the first facial region and the second facial region are at least one facial region in the 2N first facial images; The determining module is further configured to determine a second loss value between the second facial image and the third facial image based on a third eigenvector of the second facial image and a fourth eigenvector of the third facial image, where the second facial image and the third facial image are any two images among the 2N first facial images; The determining module is further configured to determine a third loss value between the third facial region and the fourth facial region based on a similarity between the third facial region and the fourth facial region, wherein the third facial region and the fourth facial region are correlated facial regions in the same first facial image; The recognition module is used to recognize the facial action of the first facial image according to the first loss value, the second loss value and the third loss value.
9. An electronic device, characterized in that: include: processor; a memory for storing computer program instructions; When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Systems and methods for shaking action recognition based on facial feature points
CN110770742A
Facial recognition method and device, computer equipment and storage medium
CN114333021A