Expression information processing method, expression recognition method, device, equipment and medium
By using a facial expression recognition network, the region to be recognized is automatically identified and sub-expression information is extracted from it, which solves the problem of noise interference in the whole face image and improves the accuracy of expression recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-17
- Publication Date
- 2026-03-17
AI Technical Summary
In existing technologies, extracting motion unit information from a whole face image can easily introduce noise, resulting in low accuracy of expression recognition.
By using a facial expression recognition network, the region to be recognized corresponding to each action unit is automatically determined, and sub-expression information is extracted only within these regions to avoid noise interference.
It improves the accuracy of facial expression information extraction, reduces the impact of noise, and enhances the accuracy of AU detection.
Smart Images

Figure CN115331293B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and more specifically, to a method for processing facial expression information, a method for recognizing facial expressions, an apparatus, a device, and a medium. Background Technology
[0002] Facial expressions, as a primary indicator of emotion, play a crucial role in communication. For example, facial expressions can be used to infer the psychological state of an individual, enabling intelligent human-computer interaction. Therefore, accurately recognizing facial expressions is a problem that needs to be solved.
[0003] In existing technologies, facial expressions can be divided into dozens of facial action units (AUs) based on the movement of muscle groups. Any expression of an object can be decomposed into a combination of AUs of different intensities. Generally, AU information is obtained directly from the entire facial region of a face image through manual design or deep convolutional networks, and facial expressions of the object are recognized based on the AU information.
[0004] However, extracting AU information from the entire face image can easily introduce more noise, which is not conducive to AU discrimination. Summary of the Invention
[0005] The purpose of this application includes, for example, providing a method for processing facial expression information, a method for recognizing facial expressions, an apparatus, a device, and a medium, which can automatically acquire the regions to be recognized in a facial image corresponding to each action unit using a facial expression recognition network, and further acquire the sub-expressions of the target object based on the regions to be recognized, thereby avoiding the introduction of too much irrelevant noise.
[0006] The embodiments of this application can be implemented as follows:
[0007] In a first aspect, embodiments of this application provide a method for processing facial expression information, the method comprising:
[0008] Acquire a face image to be identified, wherein the face image to be identified includes: face image information of the target object;
[0009] Based on a pre-trained facial expression recognition network and multiple preset action units, the regions to be recognized in the face image to be recognized corresponding to each action unit are determined, the matching information between each region to be recognized and the corresponding action unit is determined, and at least one sub-expression of the face image to be recognized is determined according to the matching information. The combination of the at least one sub-expression is used to represent the expression of the target object, and each action unit is used to represent one sub-expression of the target object.
[0010] Secondly, embodiments of this application provide an expression recognition method, including:
[0011] Obtain at least one sub-expression from the facial image to be identified;
[0012] The facial expression of the face image to be identified is determined based on the at least one sub-expression.
[0013] Thirdly, embodiments of this application provide a facial expression information extraction device, comprising:
[0014] The image acquisition module is used to acquire a face image to be identified, which includes: face image information of the target object.
[0015] The expression extraction module is used to determine the regions to be recognized in the face image to be recognized corresponding to each of the action units based on a pre-trained facial expression recognition network and a set of preset action units, determine the matching information between each region to be recognized and the corresponding action unit, and determine at least one sub-expression of the face image to be recognized based on the matching information. The combination of the at least one sub-expression is used to represent the expression of the target object, and each action unit is used to represent a sub-expression of the target object.
[0016] Fourthly, embodiments of this application also provide a facial expression recognition device, comprising:
[0017] The action unit acquisition module is used to acquire at least one sub-expression of the facial image to be recognized.
[0018] An expression determination module is used to determine the facial expression of the face image to be identified based on the at least one sub-expression.
[0019] Fifthly, embodiments of this application provide a processing device, the processing device comprising: a processor, a storage medium and a bus, the storage medium storing machine-readable instructions executable by the processor, and when the processing device is running, the processor communicates with the storage medium via the bus, the processor executing the machine-readable instructions to perform the steps of the facial expression information processing method as described in any one of the first aspects or the facial expression recognition method as described in the second aspect.
[0020] In a sixth aspect, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the facial expression information processing method as described in any one of the first aspects or the facial expression recognition method as described in the second aspect.
[0021] The beneficial effects of the embodiments of this application include:
[0022] Using the facial expression information processing method, facial expression recognition method, apparatus, device, and medium provided in this application, firstly, a facial expression recognition network determines multiple regions to be recognized corresponding to action units. Then, within each region to be recognized, sub-expressions in the facial image to be recognized are determined. This method of facial expression information extraction refines the granularity to the regions to be recognized corresponding to action units. By matching the regions to be recognized with preset action units, the sub-expressions of the object are determined, avoiding the noise introduced by directly extracting information from the entire facial image of the object. Secondly, the regions to be recognized are automatically determined by the facial expression recognition network, allowing for more flexible AU region division and closer alignment with AU regions of interest, thus improving the accuracy of facial expression information extraction. Attached Figure Description
[0023] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart illustrating the steps of the facial expression information processing method provided in the embodiments of this application;
[0025] Figure 2 A flowchart of the facial expression information processing method provided in the embodiments of this application;
[0026] Figure 3 This is a flowchart illustrating another step of the facial expression information processing method provided in an embodiment of this application.
[0027] Figure 4 This is a flowchart illustrating another step of the facial expression information processing method provided in an embodiment of this application.
[0028] Figure 5 This is a flowchart illustrating another step of the facial expression information processing method provided in an embodiment of this application.
[0029] Figure 6 This is a flowchart illustrating another step of the facial expression information processing method provided in an embodiment of this application.
[0030] Figure 7 This is a flowchart illustrating another step of the facial expression information processing method provided in an embodiment of this application.
[0031] Figure 8 This is a flowchart illustrating another step of the facial expression information processing method provided in an embodiment of this application.
[0032] Figure 9This is a flowchart illustrating another step of the facial expression information processing method provided in an embodiment of this application.
[0033] Figure 10 This is a flowchart illustrating the steps of the facial expression recognition method provided in the embodiments of this application;
[0034] Figure 11 This is a schematic diagram of the structure of the facial expression information processing device provided in the embodiments of this application;
[0035] Figure 12 This is a schematic diagram of the structure of the facial expression recognition device provided in the embodiments of this application;
[0036] Figure 13 This is a schematic diagram of the processing device provided in an embodiment of this application.
[0037] Icons: 201-Face image to be identified; 202-Encoder; 203-Mask decoder; 204-Facial feature decoder; 205-Face mask matrix; 2051-First face mask matrix; 2052-Second face mask matrix; 2053-Nth face mask matrix; 206-Foreground feature matrix; 2061-First foreground feature matrix; 2062-Second foreground feature matrix; 2063-Nth foreground feature matrix; 207-Region to be identified; 2071-First region to be identified; 2072-Second region to be identified; 2073-Nth region to be identified; 208- Recognition network; 2081-First recognition module; 2082-Second recognition module; 2083-Nth recognition module; 2084-First matching information; 2085-Second matching information; 2086-Nth matching information; 110-Expression information processing device; 1101-Image acquisition module; 1102-Expression extraction module; 1103-Network training module; 1104-Encoder pre-training module; 120-Expression recognition device; 1201-Action unit acquisition module; 1202-Expression determination module; 2001-Processor; 2002-Storage medium; 2003-Bus. Detailed Implementation
[0038] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0039] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0040] It should be noted that, where there is no conflict, the features in the embodiments of this application can be combined with each other.
[0041] AU (Active Expression) is the cornerstone of facial expressions, and much research has been conducted in recent years to extract AU information from faces. Early methods mostly constructed facial features through manually defined rules, such as local binary patterns and wavelet features, and then used classification algorithms to extract and detect AU information based on these features. This approach not only has low extraction efficiency and quality, but also poor accuracy.
[0042] To improve the accuracy of AU (Active Aspect) information extraction, deep learning has also been applied. The typical process involves using convolutional neural networks to extract global facial features from a face image, and then performing classification or regression based on these features to determine the AU detection result. However, directly extracting global facial features can easily result in excessive interference noise in the extracted features, in addition to the target AU, putting more pressure on the convolutional neural network and reducing the accuracy of AU detection.
[0043] Based on this, the applicant has proposed a facial expression information processing method, facial expression recognition method, device, equipment and medium, which can automatically determine the region to be recognized corresponding to each action unit by using a facial expression recognition network, and obtain AU information in the region to be recognized, avoiding interference from other AUs and improving the accuracy of AU detection.
[0044] It should be noted that the objects can be human or humanoid game characters, all of which have corresponding facial features and can adapt their expressions.
[0045] The Facial Expression Coding System (FACS) is a system for analyzing micro-expressions on the face. Based on the characteristics of facial anatomy, FACS divides the face into several independent yet interconnected Units (AUs) and analyzes the motion characteristics of these AUs, the main areas they control, and the related expressions.
[0046] To accurately extract AU information from a face image to be recognized, this application provides a method for processing facial expression information. Based on the obtained AU information, the facial expression recognition method provided in this application can be used to further determine the facial expression of the target object in the face image to be recognized.
[0047] The following explanation, in conjunction with several specific application examples, illustrates the facial expression information processing method, facial expression recognition method, apparatus, device, and medium provided in the embodiments of this application.
[0048] Figure 1 The diagram shown is a flowchart illustrating the steps of an expression information processing method provided in an embodiment of this application. The executing entity of this method can be a computer device with computing and processing capabilities. Figure 1 As shown, the method includes the following steps:
[0049] S101, Obtain the image of the face to be recognized.
[0050] The face image to be identified includes: facial image information of the target object.
[0051] The face image to be identified can be a processed image containing only a single target face image and a small amount of background information, and its size can be 256×256.
[0052] The facial images to be identified can be obtained from publicly available facial datasets, such as the YALE database, the FERET database, etc., but are not limited to these.
[0053] Optionally, the target objects in different facial images to be identified can be different, and the expressions they contain can also be different.
[0054] S102, based on the pre-trained facial expression recognition network and multiple preset action units, determine the regions to be recognized in the facial image to be recognized that correspond to each action unit, determine the matching information between each region to be recognized and the corresponding action unit, and determine at least one sub-expression of the facial image to be recognized based on the matching information.
[0055] In this context, at least one combination of sub-expressions is used to represent the expression of the target object, and each action unit is used to represent a sub-expression of the target object.
[0056] The action unit can be at least one of the 45 action units included in the aforementioned facial expression coding system, and each action unit can serve as a sub-expression that the facial expression recognition network needs to recognize. Based on the combination of at least one action unit recognized on the target object, the expression of the target object can be determined. Optionally, the specific number can be set as needed, and this application does not limit it.
[0057] Facial expression recognition networks can be deep convolutional neural networks. After training, they can identify the action units contained in the target object in the input facial image to be recognized.
[0058] For example, in FACS, AU4 represents frowning, AU5 represents upper eyelid elevation, and AU7 represents orbicularis oculi muscle contraction. If the facial expression recognition network only recognizes AU5 on the target object, the target object's expression may be disgust or surprise. If the facial expression recognition network recognizes both AU5 and AU7 on the target object, the target object's expression may be fear. And if the facial expression recognition network recognizes both AU4 and AU7 on the target object, the target object's expression may be anger or concealed anger.
[0059] Optionally, when recognizing action units of a target object, the facial expression recognition network first determines the region to be recognized corresponding to each action unit. The region to be recognized can be an image containing only information about the target object within the corresponding action unit region. For example, if the action unit is AU17, which represents pushing the lower lip upwards, then the region to be recognized could be the region including the target object's lips.
[0060] It is worth noting that the relationship between the action unit and the region to be identified can also be many-to-one, with multiple action units potentially corresponding to the same region to be identified. For example, AU12 pulls the corners of the mouth upwards, AU13 sharply pulls the lips, and AU14 tightens the corners of the mouth; all these correspond to the lip region of the target object, and their positional information in the labeled face image is similar.
[0061] Then, the facial expression recognition network identifies each region to be recognized, detecting whether the region matches the sub-expression represented by the action unit, and outputs matching information based on the matching results. This matching information can be used to describe the number and names of action units contained in the target object.
[0062] In this embodiment, the facial expression recognition network detects the regions to be recognized corresponding to each action unit, determines the action units contained in the target object, avoids introducing noise from other action unit regions, and has a finer granularity in AU definition, thus improving the accuracy of expression information extraction.
[0063] The aforementioned facial expression recognition network may include: an encoder, a mask decoder and a facial feature decoder connected to the encoder respectively, and a recognition network, wherein the recognition network includes: a recognition module corresponding to each action unit.
[0064] The encoder can be a deep convolutional neural network, specifically the first three convolutional modules of a ResNet50. ResNet50 is a deep residual network for feature extraction containing four convolutional modules. In this embodiment, only the first three modules are retained, allowing for the preservation of more texture information of the target object.
[0065] The mask decoder can consist of a multi-layer convolutional module, each consisting of a convolutional layer and an InstanceNorm, used to extract the attention region corresponding to the action unit from the input facial feature vector.
[0066] The facial feature decoder can be implemented by a convolutional module similar to the ResNet residual structure. Each convolutional module consists of two convolutions and InstanceNorms that alternate to obtain facial information that focuses on the corresponding action unit.
[0067] It should be noted that the mask decoder, facial feature decoder, and recognition network described above are all equipped with AU labels. Mask decoders, facial feature decoders, and recognition networks with the same AU label have a corresponding relationship.
[0068] Figure 2 The diagram shown is a flowchart of the facial expression information processing method provided in an embodiment of this application. Figure 2 As shown, firstly, the face image 201 to be recognized is input into the encoder 202 of the facial expression recognition network to obtain a facial feature vector. This facial feature vector can be used to describe the expression-related feature information of the target object in the face image 201 to be recognized.
[0069] Next, the facial feature vector can be input into the mask decoder 203 and the facial feature decoder 204 to obtain multiple face mask matrices 205 and foreground feature matrices 206 corresponding to each preset action unit. In this embodiment, taking N preset action units as an example, we can obtain the first face mask matrix 2051 and the first foreground feature matrix 2061, the second face mask matrix 2052 and the second foreground feature matrix 2062, and so on, the Nth face mask matrix 2053 and the Nth foreground feature matrix 2063.
[0070] Furthermore, the face mask matrix 205 is multiplied by the corresponding foreground feature matrix 206 to obtain the first region to be identified 2071, the second region to be identified 2072, ... the Nth region to be identified 2073. The region to be identified 207 only includes the face image of the region where the corresponding action unit is located.
[0071] Finally, the aforementioned regions to be identified 207 are input into the recognition network 208. The first recognition module 2081, the second recognition module 2082…the Nth recognition module 2083, corresponding to each action unit, extract features from the regions to be identified 207, obtaining feature information for each region. This feature information is then compared with the feature information of the sub-expressions represented by the action units corresponding to each region 207, resulting in first matching information 2084, second matching information 2085…the Nth matching information 2086. Each of these matching information entries can be a binary classification result of the comparison process, used to describe whether the sub-expressions corresponding to each action unit exist in the face image to be identified.
[0072] Optionally, before determining the matching result of the face image 201 to be identified, the initial network for facial expression recognition can be trained to improve the accuracy of the matching information. The training process is detailed in the following embodiments.
[0073] Optionally, such as Figure 3 As shown, in step S202 above, based on the pre-trained facial expression recognition network and multiple preset action units, the region to be recognized in the face image to be recognized corresponding to each action unit is determined, the matching information between each region to be recognized and the corresponding action unit is determined, and at least one sub-expression of the face image to be recognized is determined according to the matching information. This can be achieved by the following steps S301 to S303.
[0074] S301, Input the face image to be recognized into the encoder to obtain the face feature vector.
[0075] Facial feature vectors are used to describe the facial expression-related feature information of a target object.
[0076] The resolution of the input face image to be recognized is reduced from 256×256 to 16×16, and the three channels contained in the face image to be recognized are converted into 512 feature channels to obtain the face feature vector.
[0077] This facial feature vector retains various feature information related to the target object, such as expression, action unit, texture, shape, etc.
[0078] S302, the facial feature vector is input into the mask decoder and the facial feature decoder respectively, and the mask decoder and the facial feature decoder obtain the region to be recognized in the face image to be recognized corresponding to each action unit based on multiple preset action units.
[0079] Furthermore, the aforementioned facial feature images are input into the mask decoder and facial feature decoder for decoding, respectively. As described above, the number of mask decoders and facial feature decoders is the same as the number of action units, and they correspond to the action units according to the set AU labels.
[0080] Further processing of the decoding results of the mask decoder and facial feature decoder can yield regions to be identified corresponding to multiple action units. Each region to be identified only includes the region contained in the corresponding action unit.
[0081] S303, each region to be identified is input into the corresponding recognition module in the recognition network. The recognition module determines the matching information between the region to be identified and the corresponding action unit, and determines whether the face image to be identified has the sub-expression represented by the corresponding action unit based on the matching information.
[0082] The recognition network can be a multi-layer convolutional network composed of multiple recognition modules. Each recognition module corresponds to an AU label and is used to recognize the sub-expression represented by the action unit corresponding to the AU label.
[0083] The recognition network can establish a correspondence between AU tags and regions to be recognized that share the same characteristics. It then identifies these regions to determine if any sub-expressions corresponding to the AU tag appear, and outputs matching information based on the recognition results. For example, if the AU tag is AU12 (pulling the corners of the mouth upwards), the region to be recognized is the lip area. The recognition network identifies this lip area to determine if any sub-expressions of pulling the corners of the mouth upwards appear, and outputs corresponding matching information based on the recognition results.
[0084] In this embodiment, the encoder first obtains the facial feature vector from the face image to be recognized, then the mask decoder and facial feature decoder determine the region to be recognized, and finally the matching information is determined. This process ensures that the facial texture information is not lost while acquiring the region to be recognized, thus improving the accuracy of AU recognition.
[0085] Optionally, such as Figure 4 As shown, in step S302 above, the facial feature vector is input into the mask decoder and the facial feature decoder respectively. The mask decoder and the facial feature decoder obtain the region to be identified in the face image to be identified corresponding to each action unit based on multiple preset action units. This can be achieved by the following steps S401 to S403.
[0086] S401, input the facial feature vector into the mask decoder to obtain multiple facial mask matrices corresponding to each action unit. The facial mask matrix is used to describe the facial position of the target object corresponding to the action unit.
[0087] As can be seen from the above embodiments, the mask decoder can be composed of multiple convolutional modules, wherein the last convolutional module can use the Sigmoid activation function to activate the output face mask matrix to the range of [0, 1].
[0088] Optionally, the face mask matrix can be 256×256 in size. The value of each pixel represents the importance of that location; the closer to 0, the lower the importance, and the closer to 1, the higher the importance. It can be understood that the pixel value is highest in the facial region where the corresponding action unit (AU) is located. In other words, the face mask matrix acts like a mask, retaining only the facial region of interest to its corresponding AU, while weakening or eliminating other facial regions.
[0089] S402, input the facial feature vector into the facial feature decoder to obtain multiple foreground feature matrices corresponding to each action unit. The foreground feature matrices are used to describe the facial expression unit features corresponding to the action unit, as well as the relationship between the facial expression unit features and the face of the target object.
[0090] The facial feature decoder can be a convolutional module with a ResNet residual structure. The last convolutional module can use LeakyReLU as the activation function and add residual connections on the basis of the last convolutional module to map the 512 feature channels of facial feature vectors into a 3 feature channels, RGB mode foreground image.
[0091] Based on this, the foreground image is further activated using Sigmoid activation, which activates the output within the range of [0, 1], resulting in a foreground feature matrix identical to the action unit corresponding to each facial feature decoder. This foreground feature matrix can be 256×256 in size, where the value of each pixel represents the importance of that location; values closer to 0 indicate lower importance, while values closer to 1 indicate higher importance.
[0092] Since each foreground feature matrix is generated by a facial feature decoder corresponding to a different action unit, each foreground feature matrix also has a corresponding relationship with an action unit. It contains facial information extracted from the entire face of the target object, but mainly focuses on the feature information of the action unit corresponding to the foreground feature matrix.
[0093] S403, based on each face mask matrix and the corresponding foreground feature matrix, determine the region to be identified in the face image corresponding to each action unit.
[0094] Therefore, based on the face mask matrix and foreground feature matrix corresponding to the same AU label, the region to be identified corresponding to that AU label can be determined. Thus, each region to be identified only includes the area containing the corresponding AU, and does not include other areas of the target object's face.
[0095] In this embodiment, the face mask matrix and foreground feature matrix corresponding to each AU are obtained by the mask decoder and the face feature decoder, respectively, and the region to be identified is further determined. In this way, the importance of the regions of AUs that are not related to each region to be identified is reduced, so that the subsequent recognition network can pay more attention to the regions related to AUs and improve the accuracy of AU information extraction.
[0096] Optionally, in step S403 above, determining the region to be recognized in the face image corresponding to each action unit based on each face mask matrix and the corresponding foreground feature matrix may include:
[0097] The first face mask matrix corresponding to the first action unit is multiplied by the first foreground feature matrix to obtain the region to be identified corresponding to the first action unit.
[0098] Wherein, the first action unit is any one of the multiple action units, and the first foreground feature matrix is the foreground feature matrix corresponding to the first action unit among the multiple foreground feature matrices.
[0099] The first face mask matrix corresponding to the first action unit is multiplied by the first foreground feature matrix. Since the first face mask matrix is equivalent to a mask, the multiplication ensures that the first foreground feature matrix retains only the enhanced facial features of the region where the first action unit is located, which serves as the region to be identified for the first action unit.
[0100] For the face mask matrix and foreground feature matrix corresponding to the same action unit, the two can be multiplied to obtain the region to be identified, which includes only the region where the action unit corresponding to the face mask matrix and foreground feature matrix is located.
[0101] For example, if the first action unit is AU24 (lips pressing together), the facial mask matrix output by the mask decoder with AU24 will have a value closer to 1 in the region of the target object's lips. The foreground feature matrix output by the facial feature decoder with AU24 contains the foreground features of the target object's face, where the lip features corresponding to AU24 are highlighted and have higher values than other regions. Then, the facial mask matrix corresponding to AU24 and the foreground feature matrix are multiplied to obtain an image of the region to be identified that contains only the lip region of the target object.
[0102] In this embodiment, the corresponding region to be identified is obtained by multiplying the face mask matrix and the foreground feature matrix corresponding to the same action unit. This process completely eliminates the influence of features of AU that are unrelated to the region to be identified, avoids the introduction of noise, and improves the accuracy of expression information extraction.
[0103] Optionally, such as Figure 5 As shown, in step S303 above, each region to be identified is input into the corresponding recognition module in the recognition network. The recognition module determines the matching information between the region to be identified and the corresponding action unit, and determines whether the face image to be identified has the sub-expression represented by the corresponding action unit based on the matching information. This can be achieved by the following steps S501 to S502:
[0104] S501, the region to be identified corresponding to the second action unit is input into the second recognition module corresponding to the second action unit in the recognition network. The second recognition module extracts the feature information of the region to be identified and the feature information of the sub-expression represented by the second action unit, and compares the feature information of the region to be identified and the feature information of the sub-expression represented by the second action unit to obtain matching information.
[0105] The second action unit is any one of the multiple action units.
[0106] As can be seen from the above embodiments, the identification network contains multiple identification modules, and the number of identification modules is equal to the number of preset action units. Each module is used to identify the action unit corresponding to its AU tag.
[0107] Each recognition network can pre-store the feature information of the sub-expression represented by its corresponding action unit. After inputting the corresponding region to be recognized, each recognition network extracts features from the input region to be recognized, compares the extracted feature information with the pre-stored feature information of the sub-expression, and outputs matching information.
[0108] For example, if the second action unit is an AU9 nose wrinkle, then the second recognition area includes the facial area on both sides of the nose, around the lower eyelids, and between the eyebrows. The second recognition module pre-stores the feature information of the AU9 representing a wrinkle, compares it with the feature information extracted from the second recognition area, and determines the similarity between the two.
[0109] S502, if the matching information is a preset matching value, then it is determined that the face image to be recognized has the sub-expression represented by the second action unit.
[0110] Optionally, the matching information can be the binary classification result of the comparison process, with "1" or "0" representing the sub-expression represented by the action unit corresponding to the "existence" or "non-existence" in the region to be identified.
[0111] For example, after the comparison process, if the similarity between the feature information of the sub-expression represented by the second action unit and the feature information extracted from the second recognition region is greater than a preset threshold, the matching result can output a preset matching value "1"; otherwise, it outputs "0".
[0112] In this embodiment, the recognition module corresponding to each action unit performs feature extraction and comparison processing on the face image to be recognized, determines the matching information, and thus focuses on accurately detecting whether the sub-representation corresponding to each action unit appears.
[0113] The following steps S601 to S603 can be performed before step S201 or before step S202, and are used to train the facial expression recognition network. The specific execution process is not limited here.
[0114] Optionally, such as Figure 6 As shown, the facial expression information processing method provided in this application embodiment may further include the following steps:
[0115] S601, construct the initial network for facial expression recognition.
[0116] The initial network for facial expression recognition includes: an initial encoder, an initial mask decoder, an initial facial feature decoder, and an initial recognition network.
[0117] The initial encoder, initial mask decoder, initial facial feature decoder, and initial recognition network in the initial facial expression recognition network have not yet had their parameters optimized, resulting in an inaccurate mapping relationship between the input face image to be recognized and the matching information. Therefore, it is necessary to initially construct the initial facial expression recognition network and then train it.
[0118] S602, Obtain the facial expression label dataset, which includes: multiple labeled face images and label information for each labeled face image.
[0119] The label information includes: the identifiers of multiple action units and the position information of each action unit in the labeled face image.
[0120] The initial training process of the facial expression recognition network is supervised, and it can be trained using labeled facial images and corresponding label information.
[0121] As described in the previous embodiments, different action units may correspond to the same facial area.
[0122] S603. Based on the facial expression label dataset, train the initial facial expression recognition network to obtain the facial expression recognition network.
[0123] The labeled face images in the facial expression labeling dataset are used as input face images to be recognized to obtain matching information. Based on this, the initial facial expression recognition network is trained multiple times. When the similarity between the action units corresponding to the matching information and the label information corresponding to the input labeled face image is greater than a preset threshold, the network can be considered to be trained. The trained initial facial expression recognition network is then used as the facial expression recognition network.
[0124] In this embodiment, an initial facial expression recognition network is constructed and trained using a facial expression labeled dataset, which improves the accuracy of the facial expression recognition network in extracting facial expression information.
[0125] Optionally, such as Figure 7 As shown, in step S603 above, the initial facial expression recognition network is trained based on the facial expression label dataset to obtain the facial expression recognition network, which can be achieved by the following steps S701 to S702.
[0126] S701, each labeled face image is input into the initial network for facial expression recognition. Based on the initial encoder, initial mask decoder, initial facial feature decoder and initial recognition network in the initial network for facial expression recognition, multiple training matching information are obtained.
[0127] The labeled face images in the facial expression labeling dataset are used as input to the face images to be identified. After passing through the initial encoder, the labeled face feature vectors are obtained.
[0128] The labeled facial feature vectors are then input into the initial mask decoder and the initial facial feature decoder to obtain the labeled facial mask matrix and foreground feature matrix corresponding to each action unit.
[0129] Furthermore, the labeled face mask matrix and the foreground feature matrix corresponding to the same action unit are multiplied together to obtain the labeled region to be identified corresponding to each action unit.
[0130] Finally, each labeled region to be identified is input into the initial recognition network for feature extraction and recognition, and the training matching information corresponding to the input labeled face image is output, which is used to characterize the action unit name contained in the labeled face image.
[0131] S702, based on the training matching information and the label information of the labeled face images, the initial encoder, initial mask decoder, initial facial feature decoder and initial recognition network in the initial facial expression recognition network are corrected to obtain the facial expression recognition network.
[0132] Based on the training matching information output from the labeled face image and the label information of the labeled face image, the initial network for facial expression recognition is modified. For example, if the action unit corresponding to a certain recognition module is AU16 pulling the lower lip down, and the matching result is "existing", and the labeled face image also corresponds to the identifier and position information of the action unit, then the parameters of the initial mask decoder and the initial facial feature decoder corresponding to the same action unit are modified according to the position information so that they focus more on the facial region where the action unit is located.
[0133] Alternatively, if the action unit corresponding to a certain recognition module is AU17 pushing the lower lip upward, and the matching result is "not found", but the action unit corresponding to the marked face image contains AU17, then the parameters of the initial encoder and the initial recognition network are modified to make their feature extraction and recognition more accurate.
[0134] In this embodiment, based on the training matching information output from the labeled face image and the label information of the labeled face image, the initial network for facial expression recognition is modified, which strengthens the attention of each initial mask decoder and initial facial feature decoder to their respective action units and improves the accuracy of facial expression information extraction.
[0135] The following steps S801 to S804 can be performed before step S601 above to pre-train a more robust facial expression feature extractor. They can also be performed after step S702 above to further improve the accuracy of feature extraction.
[0136] Optionally, such as Figure 8 As shown, the facial expression information processing method provided in this application embodiment may further include the following steps:
[0137] S801, construct a pre-trained encoder.
[0138] As described in the previous embodiments, the pre-trained encoder can be the first three convolutional modules of a ResNet50, used to extract features from the input face image to be recognized. In this embodiment, the facial features extracted by the pre-trained encoder are not yet accurate, and their parameters need to be corrected through pre-training.
[0139] S802, Obtain the facial expression dataset.
[0140] The facial expression dataset includes multiple frames of facial images with different expressions.
[0141] The facial expression dataset may include multiple frames of facial images labeled with facial expression information. For example, the facial expressions of multiple facial images in the facial expression database may include: happy, sad, angry, disgusted, neutral, surprised, afraid, etc., but are not limited to this.
[0142] S803 uses the facial expression dataset to generate multiple triples, each triple consisting of three frames of facial images with different expressions.
[0143] Choose any three facial images from the above expression dataset that have different or the same expressions to form a triplet. For example: <sad, sad, happy>, <angry, disgusted, neutral>, etc., but not limited to these.
[0144] Optionally, the face corresponding to each triple can be the same object to reduce the impact of differences between faces on the training results.
[0145] S804, based on the triples, trains the pre-trained encoder on facial expression similarity to obtain the initial encoder.
[0146] Optionally, a similarity comparison method can be used to input the triples into the pre-trained encoder, so that the pre-trained encoder can be trained by comparing the images within the triples, learning fine-grained features between facial expression changes, and thus obtaining the initial encoder.
[0147] In this embodiment, the pre-trained encoder is trained using the constructed triples to further improve the initial encoder's ability to capture facial feature information.
[0148] Optionally, such as Figure 9 As shown, in step S201 above, obtaining the face image to be identified can be achieved by the following steps S901 to S903.
[0149] S901, Obtain the input face image.
[0150] The input facial image can be obtained from the public dataset in the above embodiments, or it can be captured by a camera or extracted from a video. The specific acquisition method is not limited here.
[0151] S902, perform face detection on the input face image to determine the target face region of the input face image.
[0152] Face detection methods, such as the hear model in OpenCV and dlib detection, are used to detect the input face image to determine whether the input face image contains the face of the target object and the region where the target object's face is located.
[0153] Optionally, the area where the target object's face is located can be represented by the coordinates of the box.
[0154] If a given input face image does not contain a recognizable face image, then that input face image is deleted.
[0155] S903: Based on the target face region, perform face alignment and cropping on the input face image to obtain the face image to be recognized.
[0156] Furthermore, the facial region of the target object is detected to obtain facial key point information, which is used to mark the positions of various irrelevant facial contours, including eyebrows, eyes, nose, mouth, etc. For example, there can be 68 facial key points, numbered sequentially from 0 to 67.
[0157] After obtaining facial key point information, the face can be aligned to achieve spatial normalization of the face of each target object, so that the features extracted by the subsequent facial expression recognition network are independent of the position of the facial features.
[0158] For example, if the face in the target object's facial region is tilted, the rotation angle θ required for face alignment can be calculated by using the positions of the left eye corner l (key point number 36) and the right eye corner r (key point number 45) to calculate the position cc of the center point of the two eyes and the coordinate difference (Δx, Δy) between the two eye corners.
[0159] Alternatively, the rotation angle θ and the positions of the center points of the two eyes cc can be calculated by the following formulas:
[0160]
[0161]
[0162] Among them, (x r y r (x) represents the coordinates of the right corner of the target object's face. l y l (x) represents the coordinates of the left corner of the target object's face. c y c ) represents the coordinates of the center point of the two eyes.
[0163] In this way, the face of the target object is rotated by an angle θ around the coordinates of the center point of the two eyes to obtain the aligned face image.
[0164] Based on this, facial key point detection can be performed again on the aligned facial image to determine the bounding box of the target object's face.
[0165] Finally, the facial image outside the target object's face bounding box is cropped to obtain the face image to be recognized.
[0166] In this embodiment, a series of preprocessing steps were performed on the input facial image to make the facial information obtained by the subsequent model clearer and reduce the impact of the target object's position and background information on the extraction of facial expression information.
[0167] like Figure 10 As shown in the embodiments of this application, an expression recognition method is also provided, which may include the following steps:
[0168] S1001, Obtain at least one sub-expression from the facial image to be recognized.
[0169] Using the facial expression processing method in the above embodiments, the facial image to be recognized can be extracted to obtain matching information. Based on the matching information, the number and names of the action units contained in the target object's face, i.e., the number of sub-expressions, can be determined.
[0170] S1002, determine the facial expression of the face image to be identified based on at least one sub-expression.
[0171] Furthermore, based on the combination of the aforementioned sub-expressions, the current expression of the target object is determined. For example, if the identified action units include: AU20 (pulling of the corners of the mouth) and AU5 (raising of the upper eyelids), the target object's expression may be fear.
[0172] Facial action units are related. To avoid detecting contradictory action units that could affect the accurate judgment of the target's expression, an expression classification network can be constructed, using the dependencies between action units as constraints during training. During recognition, the expression corresponding to multiple action units with close interdependencies can be used as the final expression type.
[0173] Optionally, each action unit and each expression type can be assigned numerical values, and a mapping relationship between action units and expression types can be established through a convolutional neural network. Finally, based on the weight of each action unit, the specific expression type can be determined after passing through the convolutional neural network.
[0174] In this embodiment, based on the above-described facial expression information processing method, the facial expression information of the target object is further determined, thereby achieving accurate recognition of the facial expression of the target object.
[0175] See Figure 11 This application embodiment also provides an expression information processing device 110, including:
[0176] The image acquisition module 1101 is used to acquire a face image to be recognized, which includes: face image information of the target object.
[0177] The expression extraction module 1102 is used to determine the regions to be recognized in the face image to be recognized corresponding to each action unit based on a pre-trained facial expression recognition network and multiple preset action units, determine the matching information between each region to be recognized and the corresponding action unit, and determine at least one sub-expression of the face image to be recognized based on the matching information. The combination of at least one sub-expression is used to represent the expression of the target object, and each action unit is used to represent a sub-expression of the target object.
[0178] The expression extraction module 1102 is further configured to: input the face image to be recognized into the encoder to obtain a facial feature vector, which is used to describe the expression-related feature information of the target object; input the facial feature vector into the mask decoder and the facial feature decoder respectively, and obtain the regions to be recognized in the face image to be recognized corresponding to each action unit based on multiple preset action units; input each region to be recognized into the corresponding recognition module in the recognition network respectively, and determine the matching information between the region to be recognized and the corresponding action unit, and determine whether the face image to be recognized has the sub-expression represented by the corresponding action unit based on the matching information.
[0179] The expression extraction module 1102 is further configured to: input facial feature vectors into a mask decoder to obtain multiple facial mask matrices corresponding to each action unit, wherein the facial mask matrices are used to describe the facial position of the target object corresponding to the action unit; input facial feature vectors into a facial feature decoder to obtain multiple foreground feature matrices corresponding to each action unit, wherein the foreground feature matrices are used to describe the expression unit features corresponding to the action unit, and the association between the expression unit features and the face of the target object; and determine the region to be identified in the face image corresponding to each action unit based on each facial mask matrix and the corresponding foreground feature matrix.
[0180] The expression extraction module 1102 is further used to multiply the first face mask matrix corresponding to the first action unit with the first foreground feature matrix to obtain the region to be identified corresponding to the first action unit, wherein the first action unit is any one of the multiple action units, and the first foreground feature matrix is the foreground feature matrix corresponding to the first action unit among the multiple foreground feature matrices.
[0181] The expression extraction module 1102 is further configured to input the region to be recognized corresponding to the second action unit into the recognition network and the second recognition module corresponding to the second action unit. The second recognition module extracts the feature information of the region to be recognized and the feature information of the sub-expression represented by the second action unit, respectively. The feature information of the region to be recognized and the feature information of the sub-expression represented by the second action unit are compared to obtain matching information. The second action unit is any one of multiple action units. If the matching information is a preset matching value, it is determined that the face image to be recognized has the sub-expression represented by the second action unit.
[0182] The network training module 1103 is used to construct an initial facial expression recognition network, which includes an initial encoder, an initial mask decoder, an initial facial feature decoder, and an initial recognition network. It also acquires a facial expression label dataset, which includes multiple labeled face images and label information for each image. The label information includes the identifiers of multiple action units and their positions within the labeled face images. Based on the facial expression label dataset, the initial facial expression recognition network is trained to obtain the facial expression recognition network.
[0183] The network training module 1103 is further configured to input each labeled face image into the initial facial expression recognition network, obtain multiple training matching information based on the initial encoder, initial mask decoder, initial facial feature decoder and initial recognition network in the initial facial expression recognition network, and correct the initial encoder, initial mask decoder, initial facial feature decoder and initial recognition network according to the training matching information and the label information of the labeled face images to obtain the facial expression recognition network.
[0184] The encoder pre-training module 1104 is used to construct a pre-trained encoder; obtain an expression dataset, which includes multiple frames of facial images with different expressions; use the expression dataset to generate multiple triples, each triple including three frames of facial images with different expressions; and train the pre-trained encoder on expression similarity based on the triples to obtain the initial encoder.
[0185] The image acquisition module 1101 is further configured to: acquire an input face image; perform face detection on the input face image to determine the target face region of the input face image; and perform face alignment and cropping processing on the input face image based on the target face region to obtain a face image to be recognized.
[0186] See Figure 12 This application embodiment also provides an expression recognition device 120, including:
[0187] The action unit acquisition module 1201 is used to acquire at least one sub-expression of the face image to be recognized.
[0188] The expression determination module 1202 is used to determine the facial expression of a face image to be recognized based on at least one sub-expression.
[0189] Figure 13 This illustration shows a schematic diagram of an electronic device according to an embodiment of the present application, including: a processor 2001, a storage medium 2002, and a bus 2003. The storage medium 2002 stores machine-readable instructions executable by the processor 2001. When the electronic device runs an expression information processing method or expression recognition method as described in the embodiment, the processor 2001 communicates with the storage medium 2002 via the bus 2003. The processor 2001 executes the machine-readable instructions, specifically the preamble of the expression information processing method item, to perform the following steps:
[0190] The process involves acquiring a facial image to be recognized, which includes facial image information of the target object; determining the regions to be recognized in the facial image corresponding to each action unit based on a pre-trained facial expression recognition network and multiple preset action units; determining the matching information between each region to be recognized and the corresponding action unit; and determining at least one sub-expression of the facial image to be recognized based on the matching information. The combination of at least one sub-expression is used to represent the expression of the target object, and each action unit is used to represent one sub-expression of the target object.
[0191] In one feasible implementation, the facial expression recognition network includes: an encoder, a mask decoder and a facial feature decoder respectively connected to the encoder, and a recognition network, wherein the recognition network includes: a recognition module corresponding to each action unit.
[0192] When the processor 2001 executes a facial expression recognition network based on a pre-trained dataset and multiple preset action units, determines the regions to be recognized in the facial image corresponding to each action unit, determines the matching information between each region to be recognized and its corresponding action unit, and determines at least one sub-expression of the facial image based on the matching information, it specifically performs the following tasks:
[0193] The face image to be recognized is input into the encoder to obtain a facial feature vector, which is used to describe the expression-related feature information of the target object. The facial feature vector is then input into the mask decoder and the facial feature decoder, respectively. The mask decoder and the facial feature decoder obtain the regions to be recognized in the face image corresponding to each action unit based on multiple preset action units. Each region to be recognized is then input into the corresponding recognition module in the recognition network. The recognition module determines the matching information between the region to be recognized and the corresponding action unit, and determines whether the face image to be recognized has the sub-expression represented by the corresponding action unit based on the matching information.
[0194] When processor 2001 executes the process of inputting facial feature vectors into the mask decoder and facial feature decoder respectively, and the mask decoder and facial feature decoder obtain the region to be recognized in the face image corresponding to each action unit based on multiple preset action units, it specifically performs the following:
[0195] The facial feature vectors are input into the mask decoder to obtain multiple facial mask matrices corresponding to each action unit. The facial mask matrices are used to describe the facial position of the target object corresponding to the action unit. The facial feature vectors are input into the facial feature decoder to obtain multiple foreground feature matrices corresponding to each action unit. The foreground feature matrices are used to describe the expression unit features corresponding to the action unit, as well as the relationship between the expression unit features and the face of the target object. Based on each facial mask matrix and the corresponding foreground feature matrices, the regions to be identified in the facial image corresponding to each action unit are determined.
[0196] When processor 2001 executes the process of determining the region to be recognized in the face image corresponding to each action unit based on each face mask matrix and the corresponding foreground feature matrix, it specifically performs the following tasks:
[0197] The first face mask matrix corresponding to the first action unit is multiplied by the first foreground feature matrix to obtain the region to be identified corresponding to the first action unit. The first action unit is any one of the multiple action units, and the first foreground feature matrix is the foreground feature matrix corresponding to the first action unit among the multiple foreground feature matrices.
[0198] When processor 2001 executes the process of inputting each region to be recognized into the corresponding recognition module in the recognition network, and the recognition module determines the matching information between the region to be recognized and the corresponding action unit, and determines whether the facial image to be recognized has the sub-expression represented by the corresponding action unit based on the matching information, it is specifically used for:
[0199] The region to be identified corresponding to the second action unit is input into the second recognition module corresponding to the second action unit in the recognition network. The second recognition module extracts the feature information of the region to be identified and the feature information of the sub-expression represented by the second action unit, respectively. The feature information of the region to be identified and the feature information of the sub-expression represented by the second action unit are compared to obtain matching information. The second action unit is any one of multiple action units. If the matching information is a preset matching value, it is determined that the face image to be identified has the sub-expression represented by the second action unit.
[0200] In one feasible implementation, the method further includes: constructing an initial facial expression recognition network, which includes an initial encoder, an initial mask decoder, an initial facial feature decoder, and an initial recognition network; obtaining a facial expression label dataset, which includes multiple labeled face images and label information for each labeled face image, the label information including the identifiers of multiple action units and the position information of each action unit in the labeled face images; and training the initial facial expression recognition network based on the facial expression label dataset to obtain the facial expression recognition network.
[0201] When processor 2001 executes the process of training the initial facial expression recognition network based on the facial expression label dataset to obtain the facial expression recognition network, it specifically performs the following tasks:
[0202] Each labeled face image is input into the initial network for facial expression recognition. Based on the initial encoder, initial mask decoder, initial facial feature decoder, and initial recognition network in the initial network for facial expression recognition, multiple training matching information is obtained. According to the training matching information and the label information of the labeled face images, the initial encoder, initial mask decoder, initial facial feature decoder, and initial recognition network in the initial network for facial expression recognition are corrected to obtain the facial expression recognition network.
[0203] In one feasible implementation, the method further includes: constructing a pre-trained encoder; obtaining an expression dataset, which includes multiple frames of facial images with different expressions; using the expression dataset, generating multiple triples, each triple including three frames of facial images with different expressions; and training the pre-trained encoder on expression similarity based on the triples to obtain an initial encoder.
[0204] When processor 2001 acquires the face image to be recognized, it specifically performs the following tasks:
[0205] Acquire the input face image; perform face detection on the input face image to determine the target face region of the input face image; based on the target face region, perform face alignment and cropping processing on the input face image to obtain the face image to be recognized.
[0206] The preamble of the processor 2001 facial expression recognition method item performs the following steps:
[0207] Obtain at least one sub-expression from the face image to be identified; determine the facial expression of the face image to be identified based on the at least one sub-expression.
[0208] When processor 2001 performs the task of determining the facial expression of a face image to be recognized based on at least one sub-expression, it specifically performs the following:
[0209] At least one sub-expression is input into a pre-trained expression classification network, and the facial expression of the face image to be identified is determined based on the dependencies between the sub-expressions.
[0210] By using the above method, the granularity of facial expression information extraction is refined to the region to be identified corresponding to the action unit. The sub-expressions of the face are determined by matching the region to be identified with the preset action unit, which avoids the noise introduced by directly extracting information from the whole face and improves the accuracy of facial expression information extraction.
[0211] This application embodiment also provides a computer-readable storage medium storing a computer program, which is executed by a processor, and the processor performs the following steps:
[0212] The process involves acquiring a facial image to be recognized, which includes facial image information of the target object; determining the regions to be recognized in the facial image corresponding to each action unit based on a pre-trained facial expression recognition network and multiple preset action units; determining the matching information between each region to be recognized and the corresponding action unit; and determining at least one sub-expression of the facial image to be recognized based on the matching information. The combination of at least one sub-expression is used to represent the expression of the target object, and each action unit is used to represent one sub-expression of the target object.
[0213] In one feasible implementation, the facial expression recognition network includes: an encoder, a mask decoder and a facial feature decoder respectively connected to the encoder, and a recognition network, wherein the recognition network includes: a recognition module corresponding to each action unit.
[0214] When the processor executes a pre-trained facial expression recognition network and multiple preset action units, determines the regions to be recognized in the facial image corresponding to each action unit, determines the matching information between each region to be recognized and the corresponding action unit, and determines at least one sub-expression of the facial image to be recognized based on the matching information, it specifically performs the following:
[0215] The face image to be recognized is input into the encoder to obtain a facial feature vector, which is used to describe the expression-related feature information of the target object. The facial feature vector is then input into the mask decoder and the facial feature decoder, respectively. The mask decoder and the facial feature decoder obtain the regions to be recognized in the face image corresponding to each action unit based on multiple preset action units. Each region to be recognized is then input into the corresponding recognition module in the recognition network. The recognition module determines the matching information between the region to be recognized and the corresponding action unit, and determines whether the face image to be recognized has the sub-expression represented by the corresponding action unit based on the matching information.
[0216] When the processor inputs facial feature vectors into the mask decoder and facial feature decoder respectively, and the mask decoder and facial feature decoder obtain the region to be recognized in the face image corresponding to each action unit based on multiple preset action units, it is specifically used for:
[0217] The facial feature vectors are input into the mask decoder to obtain multiple facial mask matrices corresponding to each action unit. The facial mask matrices are used to describe the facial position of the target object corresponding to the action unit. The facial feature vectors are input into the facial feature decoder to obtain multiple foreground feature matrices corresponding to each action unit. The foreground feature matrices are used to describe the expression unit features corresponding to the action unit, as well as the relationship between the expression unit features and the face of the target object. Based on each facial mask matrix and the corresponding foreground feature matrices, the regions to be identified in the facial image corresponding to each action unit are determined.
[0218] When the processor determines the region to be recognized in the face image corresponding to each action unit based on each face mask matrix and the corresponding foreground feature matrix, it specifically performs the following operations:
[0219] The first face mask matrix corresponding to the first action unit is multiplied by the first foreground feature matrix to obtain the region to be identified corresponding to the first action unit. The first action unit is any one of the multiple action units, and the first foreground feature matrix is the foreground feature matrix corresponding to the first action unit among the multiple foreground feature matrices.
[0220] When the processor executes the process of inputting each region to be recognized into the corresponding recognition module in the recognition network, and the recognition module determines the matching information between the region to be recognized and the corresponding action unit, and determines whether the facial image to be recognized has the sub-expression represented by the corresponding action unit based on the matching information, the specific functions are as follows:
[0221] The region to be identified corresponding to the second action unit is input into the second recognition module corresponding to the second action unit in the recognition network. The second recognition module extracts the feature information of the region to be identified and the feature information of the sub-expression represented by the second action unit, respectively. The feature information of the region to be identified and the feature information of the sub-expression represented by the second action unit are compared to obtain matching information. The second action unit is any one of multiple action units. If the matching information is a preset matching value, it is determined that the face image to be identified has the sub-expression represented by the second action unit.
[0222] In one feasible implementation, the method further includes: constructing an initial facial expression recognition network, which includes an initial encoder, an initial mask decoder, an initial facial feature decoder, and an initial recognition network; obtaining a facial expression label dataset, which includes multiple labeled face images and label information for each labeled face image, the label information including the identifiers of multiple action units and the position information of each action unit in the labeled face images; and training the initial facial expression recognition network based on the facial expression label dataset to obtain the facial expression recognition network.
[0223] When the processor trains the initial facial expression recognition network based on the facial expression label dataset to obtain the facial expression recognition network, it specifically performs the following tasks:
[0224] Each labeled face image is input into the initial network for facial expression recognition. Based on the initial encoder, initial mask decoder, initial facial feature decoder, and initial recognition network in the initial network for facial expression recognition, multiple training matching information is obtained. According to the training matching information and the label information of the labeled face images, the initial encoder, initial mask decoder, initial facial feature decoder, and initial recognition network in the initial network for facial expression recognition are corrected to obtain the facial expression recognition network.
[0225] In one feasible implementation, the method further includes: constructing a pre-trained encoder; obtaining an expression dataset, which includes multiple frames of facial images with different expressions; using the expression dataset, generating multiple triples, each triple including three frames of facial images with different expressions; and training the pre-trained encoder on expression similarity based on the triples to obtain an initial encoder.
[0226] When the processor acquires the image of the face to be recognized, it is specifically used for:
[0227] Acquire the input face image; perform face detection on the input face image to determine the target face region of the input face image; based on the target face region, perform face alignment and cropping processing on the input face image to obtain the face image to be recognized.
[0228] The preamble of the processor facial expression recognition method item performs the following steps:
[0229] Obtain at least one sub-expression from the face image to be identified; determine the facial expression of the face image to be identified based on the at least one sub-expression.
[0230] When the processor performs the task of determining the facial expression of a face image to be recognized based on at least one sub-expression, it specifically performs the following tasks:
[0231] At least one sub-expression is input into a pre-trained expression classification network, and the facial expression of the face image to be identified is determined based on the dependencies between the sub-expressions.
[0232] By using the above method, the granularity of facial expression information extraction is refined to the region to be identified corresponding to the action unit. The sub-expressions of the face are determined by matching the region to be identified with the preset action unit, which avoids the noise introduced by directly extracting information from the whole face and improves the accuracy of facial expression information extraction.
[0233] In this embodiment, the computer program, when run by the processor, can also execute other machine-readable instructions to perform other methods as described in the embodiments. For details on the specific execution steps and principles, please refer to the description of the embodiments, which will not be repeated here.
[0234] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0235] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0236] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0237] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0238] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0239] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.
Claims
1. An expression information processing method characterized by comprising: The method comprises: obtaining a face image to be identified, the face image to be identified comprising face image information of a target object; determining, based on a pre-trained face expression recognition network and a plurality of preset action units, a to-be-identified region in the face image to be identified corresponding to each of the action units, determining matching information of each of the to-be-identified regions and a corresponding action unit, and determining at least one sub-expression of the face image to be identified according to the matching information, wherein a combination of the at least one sub-expression is used to represent an expression of the target object, and each of the action units is used to represent a sub-expression of the target object; the face expression recognition network comprises an encoder, a mask decoder and a face feature decoder connected to the encoder respectively, the determining, based on the pre-trained face expression recognition network and the plurality of preset action units, of the to-be-identified region in the face image to be identified corresponding to each of the action units comprises: inputting the face image to be identified into the encoder to obtain a face feature vector, the face feature vector being used to describe expression-related feature information of the target object; inputting the face feature vector into the mask decoder and the face feature decoder respectively, and obtaining, by the mask decoder and the face feature decoder, the to-be-identified region in the face image to be identified corresponding to each of the action units based on the plurality of preset action units.
2. The expression information processing method according to claim 1, characterized by, The face expression recognition network further comprises an identification network, the identification network comprising an identification module corresponding to each of the action units; the determining, based on the matching information, of the at least one sub-expression of the face image to be identified comprises: inputting each of the to-be-identified regions into a corresponding identification module in the identification network, determining, by the identification module, the matching information of the to-be-identified region and the corresponding action unit, and determining, according to the matching information, whether the face image to be identified has a sub-expression represented by the corresponding action unit.
3. The expression information processing method according to claim 2, characterized by, the inputting the face feature vector into the mask decoder and the face feature decoder respectively, and obtaining, by the mask decoder and the face feature decoder, the to-be-identified region in the face image to be identified corresponding to each of the action units based on the plurality of preset action units comprises: inputting the face feature vector into the mask decoder to obtain a plurality of face mask matrices corresponding to each of the action units, the face mask matrices being used to describe a face position of the target object corresponding to the action unit; inputting the face feature vector into the face feature decoder to obtain a plurality of foreground feature matrices corresponding to each of the action units, the foreground feature matrices being used to describe an expression unit feature corresponding to the action unit and an association relationship between the expression unit feature and the face of the target object; determining, according to each of the face mask matrices and each of the corresponding foreground feature matrices, the to-be-identified region in the face image to be identified corresponding to each of the action units.
4. The expression information processing method according to claim 3, characterized by, The method further comprises: constructing a face expression recognition initial network, wherein the face expression recognition initial network comprises an initial encoder, an initial mask decoder, an initial facial feature decoder, and an initial recognition network; 5. The expression information processing method according to claim 2, characterized by, obtaining a face expression labeled data set, wherein the face expression labeled data set comprises a plurality of labeled face images and label information of each labeled face image, and the label information comprises identification of a plurality of action units and position information of each action unit in the labeled face image; training the face expression recognition initial network according to the face expression labeled data set to obtain the face expression recognition network. The method further comprises:
6. The expression information processing method according to any one of claims 2 to 5, characterized by, constructing a pre-training encoder; obtaining an expression data set, wherein the expression data set comprises a plurality of face images with different expressions; 7. The expression information processing method according to claim 6, characterized by, 8. The expression information processing method according to claim 6, characterized by, The expression dataset is used to generate a plurality of triplets, each of the triplets comprising three frames of facial images with different expressions; According to the triplets, the pre-trained encoder is trained for expression similarity to obtain the initial encoder.
9. The expression information processing method according to claim 1, characterized by, The obtained face image to be recognized comprises: An input face image is obtained; A target face region of the input face image is determined by performing face detection on the input face image; The input face image is aligned and cropped according to the target face region to obtain the face image to be recognized.
10. A facial expression recognition method, characterized by, It comprises: An expression information processing method according to any one of claims 1-9 is applied to obtain at least one sub-expression of the face image to be recognized; According to the at least one sub-expression, the facial expression of the face image to be recognized is determined.
11. An expression information extraction apparatus characterized by comprising: It comprises: An image acquisition module is configured to acquire a face image to be recognized, wherein the face image to be recognized comprises face image information of a target object; An expression extraction module is configured to determine, based on a pre-trained face expression recognition network and a plurality of action units, a to-be-recognized region corresponding to each of the action units in the face image to be recognized, determine matching information between each of the to-be-recognized regions and the corresponding action units, and determine at least one sub-expression of the face image to be recognized according to the matching information, wherein a combination of the at least one sub-expression is used to represent an expression of the target object, and each of the action units is used to represent a sub-expression of the target object; The face expression recognition network comprises an encoder, a mask decoder, and a face feature decoder connected to the encoder, and the expression extraction module is specifically configured to input the face image to be recognized into the encoder to obtain a face feature vector, wherein the face feature vector is used to describe expression-related feature information of the target object; The face feature vector is input into the mask decoder and the face feature decoder, respectively, and the mask decoder and the face feature decoder obtain the to-be-recognized region corresponding to each of the action units in the face image to be recognized based on the plurality of action units.
12. A facial expression recognition apparatus, characterized by comprising: It comprises: An action unit acquisition module is configured to apply the expression information extraction device of claim 11 to obtain at least one sub-expression of the face image to be recognized; An expression determination module is configured to determine a facial expression of the face image to be recognized according to the at least one sub-expression.
13. An electronic device, comprising: It comprises: A processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, the processor and the storage medium communicate through the bus when the electronic device is running, and the processor executes the machine-readable instructions to perform the steps of the expression information processing method of any one of claims 1-9 or the expression recognition method of claim 10.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is executed by the processor to perform the steps of the expression information processing method of any one of claims 1-9 or the expression recognition method of claim 10.
Citation Information
Patent Citations
Facial expression recognition method and device, electronic equipment and storage medium
CN114743241A