A method, device, electronic device and storage medium for recognizing human facial expressions
By extracting and stitching the features of face images, the corresponding full-face expression vectors are determined, which solves the problem of inaccurate facial expression recognition in the prior art, and achieves higher recognition accuracy.
Patent Information
- Application Number
- CN202210225170.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-09
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-03-09
AI Technical Summary
The existing facial expression recognition method cannot accurately describe facial expression information because the facial image features cannot accurately describe facial expression information.
By extracting the target face sub-picture corresponding to at least one face organ from the target face image to be identified, the target global face features and local face features are obtained, and the spliced face features are obtained, so as to determine the corresponding full-face expression vector to achieve accurate recognition of face expressions.
It improves the recognition accuracy of facial expressions and can more accurately recognize facial expressions, thereby improving the recognition performance of the system.
Smart Images

Figure CN114581993B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of artificial intelligence, and in particular, to a method, apparatus, electronic device, and storage medium for recognizing human facial expressions. Background Art
[0002] Human facial expressions are a form of non-verbal communication and are the main means of expressing social information among humans. With the development of technology, human facial expression recognition technology has been widely used in fields such as human-computer interaction, intelligent control, security, medical treatment, and communication.
[0003] Existing human facial expression recognition methods usually input a human face image into a machine learning algorithm network, extract the features of the human face image through the machine learning algorithm network, and recognize the human facial expression information through the features of the human face image, thereby outputting an expression coefficient to achieve human facial expression recognition. However, in existing human facial expression recognition methods, since the features of the human face image cannot accurately describe the human facial expression information, the human facial expression cannot be accurately recognized. Summary of the Invention
[0004] Embodiments of the present invention provide a method, apparatus, electronic device, and storage medium for recognizing human facial expressions, which can accurately recognize human facial expressions and thus improve the accuracy of human facial expression recognition.
[0005] According to one aspect of the present invention, there is provided a method for recognizing human facial expressions, including:
[0006] extracting a target human face sub-image corresponding to at least one human face organ from a target human face image to be recognized;
[0007] acquiring a target global human face feature and each target local human face feature respectively corresponding to the target human face image and each target human face sub-image;
[0008] performing feature splicing on the target global human face feature and each target local human face feature to obtain a target spliced human face feature;
[0009] determining a target full-face expression vector corresponding to the target spliced human face feature, and recognizing a human facial expression matching the target human face image according to the target full-face expression vector.
[0010] According to another aspect of the present invention, there is provided a device for recognizing human facial expressions, including:
[0011] a target human face sub-image acquisition module, configured to extract a target human face sub-image corresponding to at least one human face organ from a target human face image to be recognized;
[0012] A face feature acquisition module, configured to acquire a target global face feature and each target local face feature respectively corresponding to a target face image and each target face sub-image;
[0013] A spliced face feature acquisition module, configured to splice the target global face feature and each target local face feature to obtain a target spliced face feature;
[0014] A face expression recognition module, configured to determine a target full-face expression vector corresponding to the target spliced face feature, and recognize a face expression matching the target face image according to the target full-face expression vector.
[0015] According to another aspect of the present invention, there is provided an electronic device, which includes:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the face expression recognition method according to any embodiment of the present invention.
[0019] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the face expression recognition method according to any embodiment of the present invention when executed.
[0020] The technical solution of the embodiment of the present invention extracts target face sub-images corresponding to at least one face organ from a target face image to be recognized, and acquires a target global face feature and each target local face feature respectively corresponding to the target face image and each target face sub-image, so as to splice the target global face feature and each target local face feature to obtain a target spliced face feature, thereby determining a target full-face expression vector corresponding to the target spliced face feature, and recognizing a face expression matching the target face image according to the target full-face expression vector, solving the problem that the existing face expression recognition method cannot accurately recognize the face expression due to the fact that the face image features cannot accurately describe the face expression information, and being able to accurately recognize the face expression, thereby improving the recognition accuracy of the face expression.
[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0023] Figure 1 It is a flowchart of a method for recognizing human face expressions provided in the first embodiment of the present invention;
[0024] Figure 2 It is a flowchart of a method for recognizing human face expressions provided in the second embodiment of the present invention;
[0025] Figure 3 It is a schematic diagram of the dimensions of a human face image provided in the third embodiment of the present invention;
[0026] Figure 4 It is an example flowchart of a method for recognizing human face expressions provided in the third embodiment of the present invention;
[0027] Figure 5 It is a schematic diagram of a device for recognizing human face expressions provided in the fourth embodiment of the present invention;
[0028] Figure 6 It is a schematic diagram of the structure of an electronic device for implementing the method for recognizing human face expressions in the embodiments of the present invention. Detailed implementation manners
[0029] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some, rather than all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0030] It should be noted that the terms "target", "standard", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0031] Embodiment 1
[0032] Figure 1 is a flowchart of a method for recognizing human face expressions provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of accurately recognizing human face expressions. This method can be executed by a human face expression recognition device, which can be implemented in software and / or hardware, and is generally directly integrated in the electronic device that executes this method. The electronic device can be a terminal device or a server device. The embodiments of the present invention do not limit the type of the electronic device that executes the human face expression recognition method. Specifically, as Figure 1 shown, the method for recognizing human face expressions can specifically include the following steps:
[0033] S110. Extract a target human face sub-graph corresponding to at least one human face organ from the target human face image to be recognized.
[0034] Among them, the target human face image can be any human face image that needs to perform human face expression recognition. The target human face sub-graph can be an image corresponding to the human face organ in the target human face image. For example, it can be an image corresponding to the left eye in the target human face image, or an image corresponding to the nose in the target human face image, etc. The embodiments of the present invention do not limit this.
[0035] In the embodiments of the present invention, a target human face sub-graph corresponding to at least one human face organ is extracted from the target human face image to be recognized to obtain a target global human face feature and each target local human face feature corresponding to the target human face image and each target human face sub-graph respectively. It can be understood that there can be multiple target human face sub-graphs. Exemplarily, extracting a target human face sub-graph corresponding to at least one human face organ from the target human face image to be recognized can be extracting a target human face sub-graph corresponding to the left eye from the target human face image to be recognized, or extracting a target human face sub-graph corresponding to the left eye, a target human face sub-graph corresponding to the right eye, and a target human face sub-graph corresponding to the nose from the target human face image to be recognized, etc. The embodiments of the present invention do not limit this.
[0036] S120. Obtain a target global face feature and each target local face feature corresponding to the target face image and each target face sub - image respectively.
[0037] Among them, the target global face feature may be a global face feature extracted from the target face image. The target local face feature may be a local face feature extracted from the target face sub - image.
[0038] In an embodiment of the present invention, after extracting a target face sub - image corresponding to at least one facial organ from the target face image to be recognized, a target global face feature corresponding to the target face image may be further obtained, and each target local face feature corresponding to each target face sub - image may be obtained. It can be understood that different target face sub - images correspond to different target local face features.
[0039] S130. Perform feature splicing on the target global face feature and each target local face feature to obtain a target spliced face feature.
[0040] Among them, the target spliced face feature may be a spliced face feature obtained by performing feature splicing on the target global face feature and each target local face feature.
[0041] In an embodiment of the present invention, after obtaining the target global face feature and each target local face feature corresponding to the target face image and each target face sub - image respectively, the target global face feature and each target local face feature may be further subjected to feature splicing to obtain a target spliced face feature.
[0042] S140. Determine a target full - face expression vector corresponding to the target spliced face feature, and identify a face expression matching the target face image according to the target full - face expression vector.
[0043] Among them, the target full - face expression vector may be a feature vector of the full - face expression corresponding to the target spliced face feature. Optionally, the target full - face expression vector may include: a plurality of vector elements corresponding to conventional expressions, and a plurality of vector elements corresponding to fine expressions.
[0044] In an embodiment of the present invention, after performing feature splicing on the target global face feature and each target local face feature to obtain a target spliced face feature, a target full - face expression vector corresponding to the target spliced face feature may be further determined to identify a face expression matching the target face image according to the target full - face expression vector. It should be noted that the specific implementation manner of determining the target full - face expression vector corresponding to the target spliced face feature in the embodiment of the present invention is not limited, as long as the determination of the target full - face expression vector can be achieved.
[0045] In the technical solution of this embodiment, by extracting target face sub-images corresponding to at least one facial organ from the target face image to be recognized, and obtaining the target global face features and each target local face feature corresponding to the target face image and each target face sub-image respectively, the target global face features and each target local face feature are feature-stitched to obtain the target stitched face feature, so as to determine the target full-face expression vector corresponding to the target stitched face feature, and according to the target full-face expression vector, the facial expression matching the target face image is recognized, which solves the problem that the existing facial expression recognition method cannot accurately recognize the facial expression due to the inability of the face image features to accurately describe the facial expression information, and can accurately recognize the facial expression, thereby improving the recognition accuracy of the facial expression.
[0046] Embodiment 2
[0047] Figure 2 It is a flowchart of a method for recognizing a facial expression provided in Embodiment 2 of the present invention. This embodiment further refines the above technical solutions, and gives various specific optional implementation manners for extracting target face sub-images corresponding to at least one facial organ from the target face image to be recognized, obtaining the target global face features and each target local face feature corresponding to the target face image and each target face sub-image respectively, and recognizing the facial expression matching the target face image according to the target full-face expression vector. The technical solutions in this embodiment can be combined with each optional solution in one or more of the above embodiments. As Figure 2 shown, the method may include the following steps:
[0048] S210. Recognize facial feature points in the target face image, and obtain a plurality of target facial feature points included in the target face image.
[0049] Among them, the facial feature points may be any feature points in the face image. The target facial feature points may be target feature points among the recognized facial feature points.
[0050] In the embodiment of the present invention, facial feature points are recognized in the target face image to obtain a plurality of target facial feature points included in the target face image. It can be understood that the target face image may include multiple facial feature points, and by recognizing each facial feature point in the target face image, a plurality of target facial feature points in the target face image can be obtained.
[0051] S220. Extract target face sub-images corresponding to at least one facial organ from the target face image according to the semantic features of each of the target facial feature points.
[0052] In an embodiment of the present invention, after obtaining a plurality of target face feature points included in a target face image, target face sub-images corresponding to at least one face organ can be extracted from the target face image according to the semantic features of the respective target face feature points. It can be understood that a target face sub-image corresponding to a face organ can be extracted according to the semantic features of the face feature points corresponding to the face organ.
[0053] S230. Input the target face image into a global feature extraction network to obtain the target global face feature.
[0054] Among them, the global feature extraction network can be a network for extracting the global face feature of the target face image. For example, it can be a CNN (Convolutional Neural Networks) feature extraction network, or an RNN (Recurrent Neural Network) feature extraction network, etc. The embodiments of the present invention do not limit this.
[0055] In an embodiment of the present invention, by inputting the target face image into the global feature extraction network, the target global face feature can be obtained. Optionally, the global feature extraction network can be built into the expression vector generation model. Among them, the expression vector generation model can be a model that can generate a face expression vector through the recognition of a face image.
[0056] It should be noted that there is no sequential relationship between steps S210 - S220 and step S230. Step S210 - S220 can be implemented first, followed by step S230; step S230 can be implemented first, followed by step S210 - S220; or the two can be implemented in parallel.
[0057] S240. Input each of the target face sub-images into a corresponding local feature extraction network to obtain each target local face feature; different local feature extraction networks correspond to different face organs.
[0058] Among them, the local feature extraction network can be a network for extracting the local face feature of the target face sub-image. For example, it can be a CNN feature extraction network, or an RNN feature extraction network, etc. The embodiments of the present invention do not limit this. Optionally, the local feature extraction network can be built into the expression vector generation model.
[0059] In an embodiment of the present invention, after extracting target face sub - images corresponding to at least one facial organ in a target face image according to the semantic features of each target facial feature point, each target face sub - image can be further input into a matching local feature extraction network to obtain each target local facial feature. It can be understood that different local feature extraction networks can correspond to different facial organs, that is, the target face sub - images corresponding to different facial organs are input into different local feature extraction networks to obtain the target local facial features corresponding to different facial organs. Exemplarily, if the facial organ includes the left - eye area, the target face sub - image corresponding to the left - eye area is input into the local feature extraction network corresponding to the left - eye area to obtain the target local facial feature corresponding to the left - eye area.
[0060] It should be noted that there is no sequential relationship between step S230 and step S240. Step S230 can be implemented first, followed by step S240; step S240 can be implemented first, followed by step S230; or both can be implemented in parallel.
[0061] S250. Determine target correction residual sub - vectors respectively corresponding to each target local facial feature according to the difference values between the local expression sub - vectors in the full - face expression vector determined from the standard face image and the labeled local expression sub - vectors in the labeled full - face expression vector of the standard face image.
[0062] Among them, the standard face image can be any image capable of performing facial expression recognition. The full - face expression vector can be a feature vector of the full - face expression corresponding to the standard face image. The local expression sub - vector can be a feature vector of the local facial expression in the full - face expression vector. The labeled full - face expression vector can be a feature vector of the full - face expression obtained through labeling. The labeled local expression sub - vector can be a feature vector of the local facial expression obtained through labeling. It should be noted that the specific implementation manner of obtaining the labeled local expression sub - vector in the embodiment of the present invention is not limited, as long as the labeled local expression sub - vector can be realized. The target correction residual sub - vector can be a feature vector capable of correcting the target full - face expression vector. It can be understood that due to the target local facial features corresponding to different facial organs, different facial organs correspond to different target correction residual sub - vectors. Exemplarily, the target correction residual sub - vector corresponding to the target local facial feature corresponding to the left - eye area is different from the target correction residual sub - vector corresponding to the target local facial feature corresponding to the right - eye area.
[0063] In an embodiment of the present invention, after obtaining the target local face features corresponding to each target face sub - image, the target correction residual sub - vectors corresponding to each target local face can be further determined according to the difference values between the local expression sub - vectors in the full - face expression vector determined from the standard face image and the labeled local expression sub - vectors in the labeled full - face expression vector of the standard face image.
[0064] Specifically, the determination of the local expression sub - vectors in the full - face expression vector determined from the standard face image may include: extracting standard face sub - images corresponding to at least one face organ in the standard face image, obtaining the global face features and the local face features corresponding to the standard face image and each standard face sub - image respectively, performing feature splicing on the global face features and the local face features to obtain the spliced face features, determining the full - face expression vector according to the spliced face features, and thus determining the local expression sub - vectors according to the full - face expression vector.
[0065] Optionally, according to the difference values between the local expression sub - vectors in the full - face expression vector determined from the standard face image and the labeled local expression sub - vectors in the labeled full - face expression vector of the standard face image, determining the target correction residual sub - vectors corresponding to each target local face feature may include: inputting each target local face feature into a matching residual extraction network to determine the target correction residual sub - vectors corresponding to each target local face feature; different face organs correspond to different residual extraction networks; wherein, each residual extraction network is trained using the difference value between the local expression sub - vector of the matching face organ in the full - face expression vector determined from the standard face image and the labeled local expression sub - vector of the matching face organ in the labeled full - face expression vector of the standard face image as the loss function.
[0066] Among them, the residual extraction network may be a network for extracting the correction residual sub - vector. It can be understood that different face organs correspond to different residual extraction networks, that is, different face organs correspond to different target local face features, and different local face features correspond to different residual extraction networks. Exemplarily, the residual extraction network corresponding to the target local face feature corresponding to the left - eye region is different from the residual extraction network corresponding to the target local face feature corresponding to the nose region.
[0067] Specifically, after obtaining the target local face features corresponding to each target face sub - image, the target local face features can be further input into the corresponding residual extraction network respectively to determine the target corrected residual sub - vectors corresponding to the target local face features respectively. Among them, the loss function during the training of the residual extraction network can be: the difference value between the local expression sub - vector of the face organs matched in the full - face expression vector determined by the standard face image and the labeled local expression sub - vector of the face organs matched in the labeled full - face expression vector in the standard face image. It can be understood that during the training of the residual extraction network, when the loss function reaches the preset threshold, the training of the residual extraction network stops, so as to use the trained residual extraction network to determine the target corrected residual sub - vector.
[0068] Optionally, each residual extraction network can be built into the expression vector generation model; in the expression vector generation model, each local feature extraction network can be connected in sequence with the corresponding residual extraction network; the output ends of the global feature extraction network and each local feature extraction network can be respectively connected to the feature splicing module; the output end of the feature splicing module and the output ends of each residual extraction network can be respectively connected to the adder; among them, each network in the expression vector generation module can be trained uniformly using the training sample set, and the training samples in the training sample set can include: standard face images, the labeled full - face expression vectors corresponding to the standard face images, and the labeled local expression sub - vectors corresponding to each face organ respectively.
[0069] Specifically, in the expression vector generation model, the output ends of each local feature extraction network can be connected in sequence with the input ends of the corresponding residual extraction network to obtain the target corrected residual sub - vectors corresponding to each target local face feature respectively. The output end of the global feature extraction network and the output ends of each local feature extraction network can be respectively connected to the feature splicing module to splice the target global face feature and each target local face feature through the feature splicing model to obtain the target spliced face feature. The output end of the feature splicing module and the output ends of each residual extraction network can be respectively connected to the adder to add the target full - face expression vector and the target corrected residual sub - vector to obtain the corrected full - face expression vector.
[0070] Specifically, each network in the expression vector generation module can be trained uniformly using the training sample set. Among them, the training samples in the training sample set can include: standard face images, the labeled full - face expression vectors corresponding to the standard face images, and the labeled local expression sub - vectors corresponding to each face organ respectively.
[0071] S260. Splice the target global face feature and each target local face feature to obtain the target spliced face feature.
[0072] It should be noted that there is no sequence relationship between step S250 and step S260. Step S250 can be implemented first, followed by step S260; step S260 can be implemented first, followed by step S250; or both steps can be implemented in parallel.
[0073] S270. Determine a target full-face expression vector corresponding to the target spliced face feature, and use each target correction residual sub-vector to correct the target full-face expression vector to obtain a corrected full-face expression vector, and identify a face expression matching the target face image according to the corrected full-face expression vector.
[0074] Among them, the corrected full-face expression vector can be a feature vector of the full-face expression obtained after being corrected by the target correction residual sub-vector.
[0075] In an embodiment of the present invention, after performing feature splicing on the target global face feature and each target local face feature to obtain the target spliced face feature, a target full-face expression vector corresponding to the target spliced face feature can be further determined, and each target correction residual sub-vector is used to correct the target full-face expression vector to obtain a corrected full-face expression vector, so as to identify a face expression matching the target face image according to the corrected full-face expression vector. Optionally, the target full-face expression vector may include target local expression sub-vectors respectively corresponding to each face organ in the target face image. Among them, the target local expression sub-vector can be a feature vector of the local face expression in the target full-face expression vector.
[0076] In the technical solution of this embodiment, by identifying facial feature points in the target face image, a plurality of target facial feature points included in the target face image are obtained. Then, according to the semantic features of each target facial feature point, target face sub-images corresponding to at least one facial organ are extracted from the target face image. Next, the target face image is input into a global feature extraction network to obtain a target global facial feature. Each target face sub-image is respectively input into a matching local feature extraction network to obtain each target local facial feature. Based on the difference values between the local expression sub-vectors in the full-face expression vector determined from the standard face image and the annotated local expression sub-vectors of the standard face image, target correction residual sub-vectors corresponding to each target local facial feature are determined. The target global facial feature and each target local facial feature are concatenated to obtain a target concatenated facial feature, and a target full-face expression vector corresponding to the target concatenated facial feature is determined. Thus, each target correction residual sub-vector is used to correct the target full-face expression vector to obtain a corrected full-face expression vector. Furthermore, based on the corrected full-face expression vector, the facial expression matching the target face image is recognized, which solves the problem that the existing facial expression recognition method cannot accurately recognize facial expressions because the facial image features cannot accurately describe the facial expression information, and can accurately recognize facial expressions, thereby improving the recognition accuracy of facial expressions.
[0077] Embodiment III
[0078] The application scenario of subtle facial expression recognition is taken as an example in the embodiments of the present invention for specific illustration. Figure 3 It is a schematic diagram of the dimensions of the face image provided in Embodiment III of the present invention. As Figure 3 shown, each face has different expressions, which are different expression bases. A person's expression can be considered as a linear combination of all expression bases.
[0079] In an embodiment of the present invention, the expression vector generation model receives a face image as input, which can be a blendshape. Among them, the dimension of the face image can be 240, that is, the input of the expression vector generation model can be 240 blendshapes. Then the output of the expression vector generation model can include 240 blendshape coefficients. For example, the output of the expression vector generation model can be [left eye closing degree, right eye closing degree, mouth opening degree...], and the range of each vector value is 0-1. When the coefficient corresponding to the blendshape is 1, it means that this blendshape is fully manifested. Among the 240 blendshapes of the input, many blendshapes are extremely small changes, such as on the lower eyelid, nose twitching, etc. Since the changes in these tiny areas are too small and insignificant, it is difficult to identify the blendshapes, and finally the machine learning algorithm network will fit the blendshapes corresponding to the micro-expressions to the blendshapes with significant changes, so that the micro-expressions will be ignored during the micro-face expression recognition process, resulting in poor accuracy of micro-face expression recognition. Among them, expressions with significant changes, that is, expressions that are easy to learn, can include mouth opening, eye closing, etc.
[0080] In the traditional face expression recognition method, the face image is input into the machine learning algorithm network, and the machine learning algorithm network outputs the expression coefficients. However, there are two problems in the traditional face expression recognition method. On the one hand, the feature area of the micro-expression is very small; on the other hand, since the micro-expression is difficult to predict, the loss is very large when the machine learning algorithm network is learning, resulting in the machine learning algorithm network leaning towards difficult-to-learn labels, and finally it is difficult to predict even the expressions that are easy to learn.
[0081] Figure 4 It is an example flowchart of a face expression recognition method provided in the third embodiment of the present invention, as Figure 4 shown, and the method may specifically include:
[0082] (1) Identify the facial feature points in the target face image to obtain multiple target facial feature points included in the target face image; according to the semantic features of each target facial feature point, extract target face sub-images corresponding to at least one facial organ in the target face image.
[0083] (2) Input the target face image into the global feature extraction network (i.e., the backbone network) to obtain the target global face feature (i.e., the whole face feature); input each target face sub-image into the matching local feature extraction network (i.e., the right eye region feature extractor, CNN feature extractor, etc.) to obtain each target local face feature (i.e., the right eye feature, left eye feature, nose region feature, mouth region feature).
[0084] (3) Input each target local facial feature into the matching residual extraction network to determine the target corrected residual sub-vectors (i.e., right eye expression vector residual, left eye expression vector residual, nose expression vector residual, mouth expression vector residual) corresponding to each target local facial feature. Specifically, when training each residual extraction network, the sum of the absolute values of the differences between the local expression sub-vectors of the matching facial organs in the full face expression vector determined by the standard face image and the labeled local expression sub-vectors of the matching facial organs in the standard face image is used as the loss function for training.
[0085] Specifically, by inputting the target face image into the global feature extraction network (i.e., the backbone network), and inputting each target local face feature into the matching residual extraction network (i.e., the residual branch), the residual branch and the backbone network do not share inputs, making the ROI (region of interest) resolution of each part of the face larger, solving the problem of too small ROI feature resolution when sharing inputs, thereby providing more information for the residual extraction network.
[0086] (4) The target global facial features and each target local facial features are spliced (i.e., feature merged) through the feature splicing module to obtain the target spliced facial features, and the target full-face expression vector (i.e., expression vector) corresponding to the target spliced facial features is determined.
[0087] Specifically, by determining the target corrected residual subvector, the problem of manually selecting difficult and easy expression bases can be avoided, thereby improving the accuracy of facial expression recognition.
[0088] (5) Using each target corrected residual sub-vector to correct the target full-face expression vector, a corrected full-face expression vector (i.e., output expression vector) is obtained, and based on the corrected full-face expression vector, the facial expression matching the target face image is identified.
[0089] Specifically, by using each target corrected residual sub-vector to correct the target full-face expression vector, the expression vector generation model can be used to mine the symbiotic relationship of the whole-face expression base. This is because the expression base has a regularity of appearing at the same time (for example, the expression of surprise will appear with an open mouth and raised eyebrows at the same time). At the same time, the residual branch network outputs the expression residual from the regional features, which allows the expression vector generation model to focus on the mutually exclusive relationship of the expression bases in the local region (for example, if the eyes look to the left, they will not be closed at the same time). If the backbone network and the residual network are learned separately, this connection cannot be utilized.
[0090] Furthermore, by incorporating the global feature extraction network, each local feature extraction network, and each residual extraction network into the expression vector generation model for joint training, the symbiotic relationship between different expression bases can be deeply mined, and the mutually exclusive relationship of expression bases can be avoided as much as possible, thereby improving the training efficiency and recognition accuracy of the entire expression vector generation model.
[0091] The above technical solution proposes a multi-branch expression recognition network (i.e., the expression vector generation model) to solve the problem of unrecognizable micro-expression bases. It can include a backbone network and a fine-grained prediction branch (i.e., the residual branch). Specifically, the backbone network is responsible for predicting all expressions (including micro-expressions), and the fine-grained branch makes residual predictions for micro-expression bases, and then adds the residuals to the prediction results of the backbone network to obtain the finally output expression vector, thus realizing the accurate recognition of subtle human facial expressions.
[0092] Embodiment 4
[0093] Figure 5 is a schematic diagram of a human face expression recognition device provided in Embodiment 4 of the present invention. As Figure 5 shown, the device includes: a target human face sub-graph acquisition module 510, a human face feature acquisition module 520, a spliced human face feature acquisition module 530, and a human face expression recognition module 540, where:
[0094] The target human face sub-graph acquisition module 510 is configured to extract a target human face sub-graph corresponding to at least one human face organ from a target human face image to be recognized;
[0095] The human face feature acquisition module 520 is configured to acquire a target global human face feature and each target local human face feature corresponding to the target human face image and each target human face sub-graph respectively;
[0096] The spliced human face feature acquisition module 530 is configured to perform feature splicing on the target global human face feature and each target local human face feature to obtain a target spliced human face feature;
[0097] The human face expression recognition module 540 is configured to determine a target full-face expression vector corresponding to the target spliced human face feature, and recognize a human face expression matching the target human face image according to the target full-face expression vector.
[0098] The technical solution of this embodiment extracts target face sub - images corresponding to at least one facial organ from the target face image to be recognized, and obtains the target global face feature and each target local face feature corresponding to the target face image and each target face sub - image respectively. Then, the target global face feature and each target local face feature are feature - stitched to obtain the target stitched face feature, thereby determining the target full - face expression vector corresponding to the target stitched face feature. According to the target full - face expression vector, the facial expression matching the target face image is recognized, which solves the problem that the existing facial expression recognition method cannot accurately recognize facial expressions because the facial image features cannot accurately describe the facial expression information. It can accurately recognize facial expressions, thereby improving the recognition accuracy of facial expressions.
[0099] Optionally, the target face sub - image acquisition module 510 can be specifically used for: recognizing facial feature points in the target face image to obtain multiple target facial feature points included in the target face image; extracting target face sub - images corresponding to at least one facial organ from the target face image according to the semantic features of each target facial feature point.
[0100] Optionally, the face feature acquisition module 520 can be specifically used for: inputting the target face image into the global feature extraction network to obtain the target global face feature; inputting each target face sub - image into the corresponding local feature extraction network respectively to obtain each target local face feature; different local feature extraction networks correspond to different facial organs.
[0101] Optionally, the target full - face expression vector may include target local expression sub - vectors corresponding to each facial organ in the target face image; correspondingly, the face feature acquisition module 520 can also be specifically used for: determining the target correction residual sub - vectors corresponding to each target local face feature according to the difference values between the local expression sub - vectors in the full - face expression vector determined by the standard face image and the labeled local expression sub - vectors in the labeled full - face expression vector of the standard face image. Correspondingly, the facial expression recognition module 540 can be specifically used for: using each target correction residual sub - vector to correct the target full - face expression vector to obtain the corrected full - face expression vector, and recognizing the facial expression matching the target face image according to the corrected full - face expression vector.
[0102] Optionally, the facial feature acquisition module 520 may specifically be configured to: input each target local facial feature into a matching residual extraction network respectively to determine a target corrected residual sub-vector corresponding to each target local facial feature; different facial organs correspond to different residual extraction networks; wherein, each residual extraction network is trained by using a difference value between a local expression sub-vector of a full-face expression vector determined by a standard facial image and matching the facial organ and a labeled local expression sub-vector of the labeled full-face expression vector in the standard facial image as a loss function.
[0103] Optionally, the global feature extraction network, each local feature extraction network, and each residual extraction network may be built into the expression vector generation model; in the expression vector generation model, each local feature extraction network may be connected in sequence with a matching residual extraction network; the output ends of the global feature extraction network and each local feature extraction network may be respectively connected with the feature splicing module; the output end of the feature splicing module and the output ends of each residual extraction network may be respectively connected with an adder; wherein, each network in the expression vector generation module is uniformly trained by using a training sample set, and the training samples in the training sample set may include: a standard facial image, a labeled full-face expression vector corresponding to the standard facial image, and labeled local expression sub-vectors corresponding to each facial organ respectively.
[0104] Optionally, the target full-face expression vector may include: a plurality of vector elements corresponding to conventional expressions, and a plurality of vector elements corresponding to fine expressions.
[0105] The facial expression recognition device provided by the embodiments of the present invention may execute the facial expression recognition method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0106] Embodiment Five
[0107] Figure 6 FIG. shows a schematic structural diagram of an electronic device 10 that may be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0108] As Figure 6As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0109] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0110] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for recognizing facial expressions: extracting a target facial sub-graph corresponding to at least one facial organ from the target facial image to be recognized; obtaining a target global facial feature and each target local facial feature corresponding to the target facial image and each target facial sub-graph respectively; performing feature stitching on the target global facial feature and each target local facial feature to obtain a target stitched facial feature; determining a target full-face expression vector corresponding to the target stitched facial feature, and recognizing the facial expression matching the target facial image according to the target full-face expression vector.
[0111] In some embodiments, the method for recognizing human facial expressions can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method for recognizing human facial expressions described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the method for recognizing human facial expressions by any other suitable means (e.g., by means of firmware).
[0112] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0113] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when the computer programs are executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0114] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0115] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0116] The systems and techniques described herein can be implemented in a computing system that includes backend components (such as, for example, a data server), or a computing system that includes middleware components (such as, for example, an application server), or a computing system that includes frontend components (such as, for example, a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (such as, for example, a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0117] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0118] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0119] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for recognizing human face expressions, characterized in that, it includes: extracting target face sub - images corresponding to at least one human face organ from the target face image to be recognized; obtaining a target global face feature and each target local face feature corresponding to the target face image and each target face sub - image respectively; determining target correction residual sub - vectors corresponding to each target local face feature according to the difference values between the local expression sub - vectors in the full - face expression vector determined by the standard face image and the labeled local expression sub - vectors in the labeled full - face expression vector of the standard face image; performing feature splicing on the target global face feature and each target local face feature to obtain a target spliced face feature; determining a target full - face expression vector corresponding to the target spliced face feature, using each target correction residual sub - vector to correct the target full - face expression vector to obtain a corrected full - face expression vector, and recognizing the human face expression matching the target face image according to the corrected full - face expression vector; wherein, the corrected full - face expression vector includes expression basis coefficients, and the target full - face expression vector includes target local expression sub - vectors corresponding to each human face organ in the target face image.
2. The method according to claim 1, characterized in that, extracting target face sub - images corresponding to at least one human face organ from the target face image to be recognized includes: recognizing human face feature points in the target face image to obtain a plurality of target human face feature points included in the target face image; extracting target face sub - images corresponding to at least one human face organ from the target face image according to the semantic features of each of the target human face feature points.
3. The method according to claim 1, characterized in that, obtaining a target global face feature and each target local face feature corresponding to the target face image and each target face sub - image respectively includes: inputting the target face image into a global feature extraction network to obtain the target global face feature; inputting each of the target face sub - images into a matching local feature extraction network respectively to obtain each of the target local face features; different local feature extraction networks correspond to different human face organs.
4. The method according to claim 3, characterized in that, determining target correction residual sub - vectors corresponding to each target local face feature according to the difference values between the local expression sub - vectors in the full - face expression vector determined by the standard face image and the labeled local expression sub - vectors in the labeled full - face expression vector of the standard face image includes: inputting each of the target local face features into a matching residual extraction network respectively to determine target correction residual sub - vectors corresponding to each target local face feature; different human face organs correspond to different residual extraction networks; wherein, each of the residual extraction networks is trained using the difference value between the local expression sub - vector of the matching human face organ in the full - face expression vector determined by the standard face image and the labeled local expression sub - vector of the matching human face organ in the labeled full - face expression vector of the standard face image as a loss function.
5. The method according to claim 4, characterized in that, The global feature extraction network, each local feature extraction network, and each residual extraction network are built into the expression vector generation model; In the expression vector generation model, each of the local feature extraction networks is connected in sequence with the matching residual extraction networks; the output ends of the global feature extraction network and each of the local feature extraction networks are respectively connected to the feature splicing module; the output end of the feature splicing module and the output ends of each of the residual extraction networks are respectively connected to an adder; Among them, each network in the expression vector generation module is obtained by unified training using a training sample set, and the training samples in the training sample set include: a standard face image, a labeled full-face expression vector corresponding to the standard face image, and labeled local expression sub-vectors corresponding to each face organ respectively.
6. The method according to claim 1, wherein, The target full-face expression vector includes: a plurality of vector elements corresponding to conventional expressions, and a plurality of vector elements corresponding to fine expressions.
7. A face expression recognition device, wherein, including: A target face subgraph acquisition module, configured to extract a target face subgraph corresponding to at least one face organ from a target face image to be recognized; A face feature acquisition module, configured to acquire a target global face feature and each target local face feature corresponding to the target face image and each target face subgraph respectively; Determine a target corrected residual sub-vector corresponding to each target local face feature according to the difference value between each local expression sub-vector in the full-face expression vector determined by the standard face image and each labeled local expression sub-vector in the labeled full-face expression vector of the standard face image; A spliced face feature acquisition module, configured to perform feature splicing on the target global face feature and each target local face feature to obtain a target spliced face feature; A face expression recognition module, configured to determine a target full-face expression vector corresponding to the target spliced face feature, and use each target corrected residual sub-vector to correct the target full-face expression vector to obtain a corrected full-face expression vector, and recognize the face expression matching the target face image according to the corrected full-face expression vector; Among them, the corrected full-face expression vector includes expression basis coefficients, and the target full-face expression vector includes target local expression sub-vectors corresponding to each face organ in the target face image respectively.
8. An electronic device, wherein, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the face expression recognition method according to any one of claims 1-6.
9. A computer-readable storage medium, wherein, The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute the face expression recognition method according to any one of claims 1-6 when executed.
Citation Information
Patent Citations
A face multi-area fusion expression recognition method based on depth learning
CN109344693A
Facial expression determination method, expression parameter determination model, medium and equipment
CN112614213A