Facial motion unit recognition method based on convolutional fusion network
Through the method based on the convolutional fusion network, cube features are generated and multi-directional feature fusion is carried out, which solves the problem of poor recognition effect of newborn facial motor units in the prior art, and achieves higher recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510166673.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively identify neonatal facial motor units, and the convolutional neural network model has limited perspective, ignoring the correlation and mutual exclusion between faces, and failing to capture rich information.
Using a method based on a convolutional fusion network, a cube feature is generated through a representation generator, and feature fusion is performed in three directions: parallel representation of cube features, image channel and data features, and a graph convolutional network is used to capture more comprehensive context and detailed information.
It improves the accuracy and robustness of facial motor unit recognition in neonatal babies, can extract and fuse facial features more effectively, and improves the recognition effect.
Smart Images

Figure CN120108017A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of artificial intelligence, and in particular relates to a facial motion unit recognition method based on a convolutional fusion network. Background Art
[0002] Appearing before language ability is developed, facial expressions, vocalization and body movements are the means for newborns to communicate emotions and intentions and coordinate social interactions, and facial expressions are one of the most important means for newborns to communicate emotions. In order to better obtain reliable and valuable information from the expressions of newborns, such as whether they are in pain, it is necessary to judge the facial movement units (action units, AU, the following specification will refer to facial movement units as AU) of newborns. However, the cost of manually judging AU is relatively high, and due to the rapid development of deep learning in recent years, automatic recognition using deep networks has gradually become well known to those skilled in the art.
[0003] In recent years, there are significantly more AU detection methods for adults than for newborns. Compared with adults, newborn faces have different proportions, prominent fat cheek pads, smoother skin with less texture, lighter eyebrows, less pronounced chin contours, and unique facial movements. For these and related reasons, action unit (AU) recognizers trained on adult faces may be difficult to generalize to newborn faces. In addition, most AU detection methods in recent years rely on convolutional neural networks (CNNs), which limit the model's perspective to local areas, ignore the correlation and mutual exclusivity between different faces, ignore the correlation and mutual exclusivity between AUs, and fail to capture more comprehensive and rich information.
[0004] Based on this, it is urgent to improve the existing facial motion unit recognition method to solve the technical defects of the existing technology. Summary of the invention
[0005] The purpose of the present invention is to provide a facial motion unit recognition method based on a convolutional fusion network to address the deficiencies of the prior art, extract global multi-directional stereo fusion features, and train on a newborn facial motion unit database to achieve better newborn facial motion unit recognition results.
[0006] In order to achieve the above technical effects, this application implements the following technical solutions:
[0007] A facial motion unit recognition method based on a convolutional fusion network comprises the following steps:
[0008] S101, detecting facial feature points of facial images collected in the laboratory and storing them in a database, aligning and trimming the facial feature points in the database according to landmarks, and marking multiple facial motion units of different categories;
[0009] S201, using a convolutional neural network pre-trained on the large-scale image processing library ImageNet to perform initial feature extraction on the aligned face image;
[0010] S301, constructing n representations of the extracted initial features through a representation generator and splicing cube features according to the representations, wherein the representation generator is composed of n parallel fully connected networks;
[0011] S401, the cube features are fused in three directions, namely, parallel representation, image channel and data features, through different feature fusion units to obtain fused features of the three dimensions;
[0012] S501, expanding the fusion features and performing facial motion unit test through a multi-layer perceptron;
[0013] Among them, the number of facial motion units is 13.
[0014] The above technical solution produces the following technical effects:
[0015] The present application generates cube features by a representation generator composed of multiple different fully connected layers, increasing the diversity and complexity of the features, and uses a graph convolutional network to perform feature fusion in three directions: parallel representation of cube features, data features, and image channels, to capture more comprehensive context and detail information, thereby improving the accuracy and robustness of facial motion unit recognition.
[0016] As a further improvement of the facial motion unit recognition method based on a convolutional fusion network of the present application, in step S101, the laboratory preprocesses the data of the face image and recognizes the data information of the face image as 68 facial feature points through OpenFace;
[0017] The geometric transformation of the face image from the original posture to the standard posture is calculated based on the facial feature points, and the face image is aligned and cropped through affine transformation to obtain a face image in a unified format.
[0018] As a further improvement of the facial motion unit recognition method based on the convolutional fusion network of the present application, the facial motion units in the facial motion unit detection are AU4 lowered eyebrows, AU6 raised cheeks, AU7 closed eyelids, AU9 raised nose wings, AU10 raised upper lip, AU12 raised corners of the mouth, AU15 lowered corners of the mouth, AU16 protruding lower lip, AU20 stretched lips, AU25 opened mouth, AU26 relaxed chin, AU27 stretched chin, AU43 closed eyes;
[0019] Among them, the label of each facial motion unit is 0-5. When the label of the facial motion unit is 0, the facial motion unit does not appear in the face image. The label of the facial motion unit is 1-5, which indicates the strength of the facial motion unit and increases with the increase of the label value.
[0020] As a further improvement of the facial motion unit recognition method based on a convolutional fusion network in the present application, the initial feature extracted in step S201 is the initial feature E, and the pre-trained convolutional neural network is the RestNet50 network.
[0021] As a further improvement of the facial motion unit recognition method based on a convolutional fusion network of the present application, the representation generator is composed of 13 parallel fully connected networks;
[0022] The input of each fully connected network is The output is
[0023] Concatenate the outputs of the parallel fully connected networks to obtain cubic features: X∈R L×D×C , where C is the number of channels, D is the data feature obtained after the width and height of the two-dimensional image are expanded and passed through the fully connected network, and L is the number of fully connected networks in the generator.
[0024] As a further improvement of the facial motion unit recognition method based on the convolutional fusion network of the present application, the parallel representation direction of the cube feature is the L-axis direction, the data feature direction of the cube feature is the D-axis direction, and the image channel direction of the cube feature is the C-axis direction;
[0025] The feature fusion unit includes a first fusion unit, a second fusion unit and a third fusion unit;
[0026] The first fusion unit is used to perform a feature fusion operation on the cube feature in the L-axis direction, the second fusion unit is used to perform a feature fusion operation on the cube feature in the D-axis direction, and the third fusion unit is used to perform a feature fusion operation on the cube feature in the C-axis direction;
[0027] Among them, the first fusion unit, the second fusion unit and the third fusion unit are all composed of graph convolutional networks.
[0028] As a further improvement of the facial motion unit recognition method based on a convolutional fusion network in the present application, the graph convolutional network is a GCN graph convolutional network, and the formula of the graph convolutional network is:
[0029]
[0030] Among them, H (l) is the node matrix of the lth layer, the initial H 0 =Xl,d,c ; A is an N×N adjacency matrix, which is used to represent the connection relationship between nodes in the graph, and N is the number of nodes in the graph; To add self-loops to the adjacency matrix, each node is connected to itself; the matrix D is a diagonal matrix, and its diagonal elements D ii represents the degree of node i, which is expressed as:
[0031] D ii =∑ j A ij
[0032] The degree matrix after the matrix D plus the self-loop is denoted as Its diagonal elements are:
[0033]
[0034] Among them, the normalized adjacency matrix is
[0035] As a further improvement of the facial motion unit recognition method based on the convolutional fusion network of the present application, the input of the first fusion unit is a set of data X *,d,c ∈R L×H×1 The data of face images is T∈R C×H×W , where C is the number of channels, H and W are the height and width of the image;
[0036] Cube features X∈R to be fused L×D×C The second dimension data is a channel T of the image c ∈R H×W After the width and height are expanded, the data extracted by the fully connected network is divided into W parts, and the length of each part is H;
[0037] Among them, X *,d,c is the d-th piece of data after the image data in the c-th channel is expanded, (d,c)∈{(1,1),(1,2),...(2,1),(2,2),...(D / H,C)}.
[0038] As a further improvement of the facial motion unit recognition method based on the convolutional fusion network of the present application, the input of the second fusion unit is a set of vectors U l,*,c ∈R 1×D×1 , (l,c)∈{(1,1),(1,2),…,(2,1),(2,2),…,(L',C);
[0039] Each input vector of the third fusion unit is V l,d,* ∈R 1×1×C, (l,d)∈{(1,1),(1,2),…,(2,1),(2,2),…,(L',D').
[0040] As a further improvement of the facial motion unit recognition method based on the convolutional fusion network of the present application, the fusion feature is expanded to obtain the feature Z∈R S , where S is L'×D'×C';
[0041] The unfolded features are used to detect facial motion units through the recognition layer to obtain the final result;
[0042] In the process of facial motion unit detection, a weighted asymmetric loss is used to calculate the loss between the generated predicted value and the true value. The formula of weighted asymmetric loss is:
[0043]
[0044] Among them, y i ,p i ,W i are label value, predicted value and weight respectively, r i is the frequency of the ith facial motion unit in the training set, w i The calculation formula is BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0046] Figure 1 This is a schematic diagram of the process of Example 1 of the present invention;
[0047] Figure 2 This is a network structure diagram of Embodiment 1 of the present invention; DETAILED DESCRIPTION
[0048] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the implementation regulations. It should be noted that the embodiments and features in the embodiments of the present invention can be combined with each other without conflict.
[0049] The following detailed description is an exemplary description, which is intended to provide further detailed description of the present invention. Unless otherwise specified, all technical terms used in the present invention have the same meaning as those generally understood by those skilled in the art to which the present invention belongs. The terms used in the present invention are only for describing specific embodiments, and are not intended to limit the exemplary embodiments according to the present invention.
[0050] Example 1
[0051] As is known, in 1978, American psychologists Ekman et al. (Ekman et al., 1978) first proposed the facial action coding system (FACS) from the perspective of facial anatomy. FACS defines 44 facial action units (AUs, hereinafter referred to as AUs), and specifically defines the action area, movement appearance characteristics, and AU composition of various expressions of each AU.
[0052] Each AU represents a facial muscle movement with specific semantics, such as AU1 represents the movement of raising the inner corner of the eyebrow, AU2 represents the movement of raising the outer corner of the eyebrow, etc. AUs can be activated individually or in combination. Any facial event can be expressed as a combination of several AUs. For example, a smile can be expressed as a combination of the upward movement of the corners of the mouth (AU12) and the lifting of the cheeks (AU6). The introduction of FACS plays a vital role in the development of contemporary automatic facial expression analysis technology. Compared with the six types of prototype expressions (anger, disgust, fear, happiness, sadness, surprise) defined by Ekman et al. (Ekman et al., 1971) in 1971, FACS provides a more objective and fine-grained method for describing facial expressions. At the same time, in 2004, psychologist Harriet Oster extended adult FACS to newborns. Through newborn FACS, newborn facial expressions such as negative and positive expressions can be better defined and judged.
[0053] In recent years, there are significantly more AU detection methods for adults than for newborns. Compared with adults, newborn faces have different proportions, prominent fat cheek pads, smoother skin with less texture, lighter eyebrows, less pronounced chin contours, and unique facial movements. For these and related reasons, action unit (AU) recognizers trained on adult faces may be difficult to generalize to newborn faces. In addition, most AU detection methods in recent years rely on convolutional neural networks (CNNs), which limit the model's perspective to local areas, ignore the correlation and mutual exclusivity between different faces, ignore the correlation and mutual exclusivity between AUs, and fail to capture more comprehensive and rich information.
[0054] like Figure 1 As shown, the present invention proposes a facial motion unit recognition method based on a convolutional fusion network, comprising the following steps:
[0055] S101, detecting facial feature points of facial images collected in the laboratory and storing them in a database, aligning and trimming the facial feature points in the database according to landmarks, and marking multiple facial motion units of different categories;
[0056] S201, using a convolutional neural network pre-trained on the large-scale image processing library ImageNet to perform initial feature extraction on the aligned face image;
[0057] S301, constructing n representations of the extracted initial features through a representation generator and splicing cube features according to the representations, wherein the representation generator is composed of n parallel fully connected networks;
[0058] S401, the cube features are fused in three directions, namely, parallel representation, image channel and data features, through different feature fusion units to obtain fused features of the three dimensions;
[0059] S501, expanding the fusion features and performing facial motion unit test through a multi-layer perceptron;
[0060] Among them, the number of facial motion units is 13.
[0061] Furthermore, the working principle of the above technical solution is as follows: first, the facial motion unit database of newborns (in the specific implementation process, the facial images are newborn facial images) is normalized by the laboratory, and each facial image is aligned and cropped; then, the facial data is subjected to a convolutional neural network RestNet50 pre-trained on the large-scale image database ImageNet to perform preliminary feature extraction on the face; secondly, the initial features are passed through a representation generator composed of n parallel fully connected networks (FC) to obtain different representations, and these representations are connected to obtain cube features; then, through a feature fusion unit (in the specific implementation process, it is a feature fusion unit of three boss types), the spliced cube features are fused in three directions: parallel representation of the cube, data features, and image channels; finally, the fused features are expanded and recognized through a multi-layer perceptron, and a weighted asymmetric loss is used to calculate the loss between the true value and the predicted value.
[0062] In practical application, the method specifically includes the following steps:
[0063] (1) Normalize the newborn facial movement unit database.
[0064] Specifically, in this application, the facial images of newborns collected by this laboratory are preprocessed, and 68 facial feature points are identified through the OpenFace tool. In order to standardize the facial images, improve the accuracy and consistency of subsequent processing and analysis, and reduce the interference caused by changes in facial posture (such as side face, head tilt, etc.), the present invention calculates the geometric transformation required to transform the face from the original posture to the standard posture through facial feature points, including translation (moving the face to the center of the image), rotation (correcting the posture of the face), scaling (adjusting the size of the face), etc., and then aligns the face images through affine transformation, that is, unifies the direction and position, and then crops the images into a unified format.
[0065] Furthermore, each face in the experimental data includes 13 AUs, and the label range of each AU is 0-5, where 0 means that the AU does not appear on the face, and 1-5 represents the strength of the AU, 1 means that the AU appears but is not obvious, and 5 means that the AU is most obvious.
[0066] (2) The ResNet50 network is used for initial feature extraction, and the cube features are obtained through a representation generator consisting of n parallel fully connected networks.
[0067] like Figure 2 As shown in the figure, the aligned newborn faces are passed through a convolutional neural network to extract the initial features E. Here, the ResNet50 network is selected and pre-trained on the large-scale image database ImageNet. The Res Net50 network can learn more complex and abstract features and performs well in extracting image features.
[0068] Further, such as Figure 2 As shown, the extracted initial features are passed through the representation generator. This module consists of n parallel fully connected networks, where n is 13, corresponding to the number of detected AUs. These fully connected networks generate different feature representations. The outputs of these parallel fully connected networks are spliced to obtain the cube feature X∈R L×D×C , where C is the number of channels, D is the feature obtained by expanding the width and height of the two-dimensional image through the fully connected layer, and L is the number of fully connected networks representing the generator, which is also the number of AUs detected n.
[0069] (3) The cube features are passed through three different graph convolutional networks in sequence to perform feature fusion in three directions: parallel representation of features (L axis), data features (D axis), and image channels (C axis).
[0070] like Figure 2 As shown, the feature X∈R L×D×C The three dimensions of features are fused through three different fusion units. Specifically, the first fusion unit f l The feature RL×*×* Convert to R L'×*×* , perform the feature fusion operation on the L axis, L' is the size of the L axis after reduction, and so on for the second fusion unit f D :R *×D×* →R *×D'×* , the third fusion unit f C :R *×*×C →R *×*×C' Each fusion unit consists of a graph convolutional network.
[0071] Among them, the formula of graph convolutional network is:
[0072]
[0073] Specifically, H (l) is the node matrix of the lth layer, the initial H 0 The data initially fed into the network. A is the adjacency matrix, an N×N square matrix used to represent the connection relationship between nodes in the graph, 1 represents a direct connection between nodes, and 0 represents the opposite, where N is the number of nodes in the graph. To add a self-loop to the adjacency matrix, each node is connected to itself. To avoid the influence of different node degrees when aggregating features, the adjacency matrix is usually normalized. The degree matrix is defined as a diagonal matrix with diagonal elements D ii represents the degree of node i (i.e. the number of edges connected to node i).
[0074] Specifically expressed as:
[0075] D ii =∑ j A ij
[0076] Furthermore, the degree matrix after the matrix plus the self-loop is denoted as Its diagonal elements are:
[0077]
[0078] Furthermore, the normalized adjacency matrix is:
[0079]
[0080] This normalization process allows the node features to be evenly propagated and aggregated in the convolution operation, enhancing the stability and effect of training.
[0081] Example 2
[0082] like Figure 2 As shown, in the fusion unit acting on the L axis, the input quantity X is a set of data X *,d,c ∈RL×H×1 The set of (d,c)∈{(1,1),(1,2),...(2,1),)(2,2),..(D / H,C)}. The image data is T∈R C×H×W , C is the number of channels, H and W are the width and height of the image, and the feature to be fused X∈R L×D×C The second dimension of data is a channel T of the image c ∈R H×W After the width and height are expanded, the data extracted by the fully connected network is divided into W parts, each with a length of H. *,d,c It is the dth piece of data after the image data in the cth channel is expanded.
[0083] Furthermore, each element X in the input *,d,c Treated as an independent input, the adjacency matrix is fully connected and feature fusion is performed through a graph convolutional network G. The data processed by the complete first fusion unit can be expressed by mathematical formula:
[0084] U *,d,c =LN(G L (X *,d,c ))∈R L'×H×1 ,
[0085] for(d,c)∈{(1,1),(1,2),…,(2,1),(2,2),…,(D / H,C)},
[0086] Among them, LN is layer normalization, G L is the graph convolution fusion unit of the L axis, input data X *,d,c ∈R L×H×1 After the first fusion unit, it becomes U *,d,c ∈R L'×H×1 , cube feature X∈R L×D×C The output after the first fusion unit is U∈R L '×D×C .
[0087] like Figure 2 As shown, the second fusion unit acts on the D axis, and the input of the second fusion unit is a set of vectors U l,*,c ∈R 1 ×D×1 The mathematical formula of the data output by the second fusion unit can be expressed as:
[0088] V l,*,c =LN(G D (U l,*,c ))∈R 1×D'×1 ,
[0089] for(l,c)∈{(1,1),(1,2),…,(2,1),(2,2),…,(L',C)},
[0090] Among them, it is known that G D is the second graph convolution fusion unit, U l,*,c ∈R 1×D×1 After the second fusion unit, it is V l,*,c ∈R 1×D'×1 , the output of the first fusion unit U∈R L'×D×C After the second fusion unit, V∈R L'×D'×C ;(l,c)∈{(1,1),(1,2),…,(2,1),(2,2),…,(L',C)}.
[0091] like Figure 2 As shown, the third fusion unit acts on the C axis. The difference between the first fusion unit and the second fusion unit is that each input vector of the third fusion unit is V l,d,* ∈R 1×1×C Among them, (l,d)∈{(1,1),(1,2),…,(2,1),(2,2),…,(L',D')}, the mathematical formula of the output data of the third fusion unit can be expressed as:
[0092] X' l,d,* =LN(G C (V l,d,* ))∈R 1×1×C' ,
[0093] for(l,d)∈{(1,1),(1,2),…,(2,1),(2,2),…,(L',D')},
[0094] Among them, G C As the third graph convolution fusion unit, the network converts each V l,d,* ∈R 1×1×C Convert to X' l,d,* ∈R 1 ×1×C' , the feature obtained by the final fusion unit is X'∈R L'×D'×C' .
[0095] Finally, the fused features are unfolded and passed through a multi-layer perceptron for facial motion unit recognition.
[0096] Further, such as Figure 2 As shown, the obtained fusion feature X' is expanded (Flatten) to obtain the feature Z∈R S, where S is L'×D'×C'. The expanded features are passed through a multi-layer perceptron (MLP) for AU detection to obtain the final prediction result. Because the AU dataset has unbalanced labels, some AUs appear less frequently than other AUs, and most AUs are not activated for most face images. To alleviate these problems, a weighted asymmetric loss is used to calculate the loss between the generated prediction value and the true value. The weighted asymmetric loss formula is:
[0097]
[0098] y i ,p i ,W i are the true value, predicted value and weight respectively. i The calculation formula is:
[0099]
[0100] r i is the frequency of the i-th AU in the training set.
[0101] In summary, the present invention extracts a newborn facial motion unit recognition method based on a multi-directional cubic graph convolutional fusion network, generates cubic features through a representation generator composed of multiple different fully connected layers, increases the diversity and complexity of the features, and uses a graph convolutional network to perform feature fusion in three directions of the cubic features to capture more comprehensive context and detail information, thereby improving the accuracy and robustness of facial motion unit recognition.
[0102] The above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
[0103] It will be appreciated by those skilled in the art that embodiments of the present invention may provide methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0104] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0105] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
Claims
1. A facial motion unit recognition method based on a convolutional fusion network, characterized in that: The following steps are included: S101, detecting facial feature points of facial images collected in the laboratory and storing them in a database, performing facial alignment and trimming on the facial feature points in the database according to landmarks, and marking multiple facial motion units of different categories; S201, performing initial feature extraction on the aligned face image using a convolutional neural network pre-trained on the large-scale image processing library ImageNet; S301, constructing n representations of the extracted initial features through a representation generator and splicing cube features according to the representations, wherein the representation generator is composed of n parallel fully connected networks; S401, the cube features are fused in three directions, namely, parallel representation, image channel and data feature, through different feature fusion units to obtain fused features of the three dimensions; S501, unfolding the fusion features and performing facial motion unit test through a multi-layer perceptron; Among them, the number of facial motion units is 13.
2. A facial motion unit recognition method based on a convolutional fusion network according to claim 1, characterized in that: In step S101, the laboratory pre-processes the data of the face image and identifies the data information of the face image into 68 facial feature points through OpenFace; The geometric transformation of the face image from the original posture to the standard posture is calculated according to the face feature points, and the face image is aligned and cropped through affine transformation to obtain the face image in a unified format.
3. The facial motion unit recognition method based on convolutional fusion network according to claim 1, characterized in that: The facial motion units in the facial motion unit detection are AU4 lowered eyebrows, AU6 raised cheeks, AU7 closed eyelids, AU9 raised nose wings, AU10 raised upper lip, AU12 raised corners of mouth, AU15 lowered corners of mouth, AU16 protruded lower lip, AU20 stretched lips, AU25 opened mouth, AU26 relaxed chin, AU27 stretched chin, AU43 closed eyes; Among them, the label of each facial motion unit is 0-5. When the label of the facial motion unit is 0, the facial motion unit does not appear on the face image. The label 1-5 of the facial motion unit represents the intensity of the facial motion unit and increases with the increase of the label value.
4. The facial motion unit recognition method based on convolutional fusion network according to claim 1, characterized in that: The initial feature extracted in step S201 is the initial feature E, and the pre-trained convolutional neural network is the RestNet50 network.
5. The facial motion unit recognition method based on convolutional fusion network according to claim 1, characterized in that: The representation generator consists of 13 parallel fully connected networks; The input of each of the fully connected networks is The output is The outputs of the parallel fully connected networks are concatenated to obtain cubic features: X∈R L×D×C , where C is the number of channels, D is the data feature obtained after the width and height of the two-dimensional image are expanded and passed through the fully connected network, and L is the number of the fully connected networks in the representation generator.
6. The facial motion unit recognition method based on convolutional fusion network according to claim 1, characterized in that: The parallel representation direction of the cube feature is the L-axis direction, the data feature direction of the cube feature is the D-axis direction, and the image channel direction of the cube feature is the C-axis direction; The feature fusion unit includes a first fusion unit, a second fusion unit and a third fusion unit; The first fusion unit is used to perform a feature fusion operation on the cube feature in the L-axis direction, the second fusion unit is used to perform a feature fusion operation on the cube feature in the D-axis direction, and the third fusion unit is used to perform a feature fusion operation on the cube feature in the C-axis direction; Among them, the first fusion unit, the second fusion unit and the third fusion unit are all composed of graph convolutional networks.
7. A facial motion unit recognition method based on a convolutional fusion network according to claim 6, characterized in that: The graph convolution network is a GCN graph convolution network, and the formula of the graph convolution network is: Among them, H (l) is the node matrix of the lth layer, the initial H 0 =X l,d,c ; A is an N×N adjacency matrix, which is used to represent the connection relationship between nodes in the graph, and N is the number of nodes in the graph; To add self-loops to the adjacency matrix, each node is connected to itself; the matrix D is a diagonal matrix, and its diagonal elements D ii represents the degree of node i, which is expressed as: D ii =∑ j A ij The matrix D plus the degree matrix after the self-loop is denoted as Its diagonal elements are: Among them, the normalized adjacency matrix is 8. The facial motion unit recognition method based on convolutional fusion network according to claim 6, characterized in that: The input of the first fusion unit is a set of data X *,d,c ∈R L×H×1 The data of the face image is T∈R C×H×W , where C is the number of channels, H and W are the height and width of the image; The cube feature X∈R to be fused L×D×C The second dimension data is a channel T of the image c ∈R H×W The data extracted by the fully connected network after width and height expansion is divided into W parts, and the length of each part is H; Among them, X *,d,c is the d-th piece of data after the image data in the c-th channel is expanded, (d,c)∈{(1,1),(1,2),...(2,1),(2,2),...(D / H,C)}.
9. The facial motion unit recognition method based on convolutional fusion network according to claim 6, characterized in that: The input of the second fusion unit is a set of vectors U l,*,c ∈R 1×D×1 , (l,c)∈{(1,1),(1,2),…,(2,1),(2,2),…,(L',C); Each input vector of the third fusion unit is V l,d,* ∈R 1×1×C , (l,d)∈{(1,1),(1,2),…,(2,1),(2,2),…,(L',D').
10. The facial motion unit recognition method based on convolutional fusion network according to claim 1, characterized in that: The fused features are expanded to obtain the feature Z∈R S , where S is L'×D'×C'; The features after expansion are subjected to facial motion unit detection through the recognition layer to obtain the final result; In the facial motion unit detection process, a weighted asymmetric loss is used to calculate the loss between the generated predicted value and the true value. The formula of the weighted asymmetric loss is: Among them, y i ,p i ,W i are label value, predicted value and weight respectively, r i is the frequency of the ith facial motion unit in the training set, w i The calculation formula is