A method and device for recognizing facial expression
By performing multiple convolution operations on the face image sequence in the three-dimensional convolution neural network model and screening face images by genetic algorithms, the problem of high computational volume in the existing technology is solved, and efficient and accurate facial expression recognition is achieved.
Patent Information
- Application Number
- CN202111509782.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-12-10
AI Technical Summary
The prior art has a large amount of calculation in facial expression recognition, making it difficult to achieve efficient recognition.
By performing multiple convolution operations on the face image sequence in the three-dimensional convolution neural network model, the size of the convolution kernel is reduced, the calculation amount is reduced, and face images are screened through gene genetic algorithms to reduce redundant data.
It improves the accuracy and efficiency of facial expression recognition, reduces the calculation amount and reduces processing time.
Smart Images

Figure CN114360003B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method and device for recognizing facial expressions. Background Art
[0002] With the continuous development of robot technology, service robots have gradually emerged. Service robots are mainly used to provide services for people's daily lives, and they mainly rely on human-computer interaction technology. Emotion recognition is the basis of human-computer interaction technology and plays an important role in the performance of service robots. Expression recognition is an important part of human-computer interaction technology, so facial expression recognition (FER) has become an important research topic in human-computer interaction applications. In addition, facial expression recognition can also be widely used in many fields such as clinical psychology, automobile safety and multimedia. At present, expressions can be recognized through images or through continuous video frames. However, the expression of expression is a dynamic process, and the accuracy of facial expression recognition through a single image is low, so expression recognition through continuous video frames is commonly used.
[0003] At present, there are roughly two types of expression recognition methods based on videos. One is to extract the features of each frame in the video to be recognized through a two-dimensional convolutional neural network (2DCNN), and then fuse the features of each frame through feature fusion, such as attention mechanism or long short-term memory network, to extract the dynamic features of the expression, and recognize the facial expression based on the extracted dynamic features of the expression; the other is to recognize the expression through a three-dimensional convolutional neural network, such as a mobile three-dimensional convolutional neural network (mobile convolutional neural network, mobile 3DCNN). The three-dimensional convolutional neural network can directly use multiple video frames of the video stream as input for expression classification. Since 3DCNN can fuse the time dimension features of each video frame, it is conducive to more accurate recognition of expressions, but the 3DCNN input is the features of all video frames in the video stream. On the one hand, the amount of data input to 3DCNN is large. On the other hand, 3DCNN involves 3D convolution, and the size of the convolution kernel for 3D convolution is large, which undoubtedly increases the amount of calculation. Summary of the invention
[0004] The embodiments of the present application provide a facial expression recognition method and device for reducing the amount of calculation while ensuring the facial expression recognition effect.
[0005] In a first aspect, an embodiment of the present application provides a method for facial expression recognition, comprising: obtaining a facial image sequence corresponding to a target face, the facial image sequence comprising multiple first facial images, one of which contains the target face; performing a first convolution operation on the facial image sequence through a feature extraction module in a trained three-dimensional convolutional neural network model to obtain multiple first feature maps, wherein one first feature map corresponds to one of the multiple first facial images, and the one first feature map includes feature sub-maps corresponding to each of the multiple channels of the one first facial image; performing a second convolution operation on the feature sub-maps on each channel of the multiple first feature maps through the feature extraction module to obtain multiple second feature maps; performing a third convolution operation on the multiple second feature maps through the feature extraction module to obtain a third feature map, the third convolution operation is a fusion of the multiple second feature maps, and the third feature map is used to describe the expression features of the target face in the facial image sequence; performing expression recognition on the third feature map through the recognition module in the trained three-dimensional convolutional neural network model to obtain a target expression recognition result of the target face.
[0006] In the embodiment of the present application, when the first feature map of each first face image is convoluted, the first feature map is not directly 3D convolved according to the method in the prior art, but the feature sub-maps contained in the first feature map are respectively subjected to the second convolution operation, so that the size of the convolution kernel corresponding to the second convolution operation can be reduced, thereby reducing the amount of calculation. In addition, after the second convolution operation is performed on the feature sub-map, the multiple second feature maps after the convolution operation of the feature sub-map can be convoluted and fused again, which is equivalent to fusing all the features of multiple first face images. In other words, the method in the embodiment of the present application can also extract the three-dimensional features of the face image, which can extract more dimensional features of the face image compared to the two-dimensional convolutional neural network model, thereby ensuring the effect of facial expression recognition and correspondingly improving the accuracy of the facial expression recognition results.
[0007] In a possible implementation, the size of the convolution kernel corresponding to the second convolution operation is M*K*K*K, where K is a positive integer and M represents the number of channels corresponding to the first feature map; the size of the convolution kernel corresponding to the third convolution operation is M*1*1*1*N, where N represents the number of channels corresponding to the third feature map.
[0008] In an embodiment of the present application, the size of the convolution kernel corresponding to the second convolution operation can be M*K*K*K, and the size of the convolution kernel corresponding to the third convolution operation can be M*1*1*1*N. In this way, the amount of calculation can be relatively reduced while extracting more comprehensive features of the facial image.
[0009] In a possible implementation, obtaining a facial image sequence corresponding to a target face includes: obtaining a primary population based on a plurality of second facial images, the primary population including a plurality of first chromosomes, wherein one first chromosome is used to indicate a third facial image selected from the plurality of second facial images, the total number of the selected third facial images indicated by any two first chromosomes is the same, and the plurality of second facial images all include the target face; performing at least one genetic iteration operation on the primary population, wherein one genetic iteration operation includes: determining the fitness of a plurality of first chromosomes in the primary population, respectively, wherein the fitness of one first chromosome is determined according to the similarity between each two third facial images selected indicated by the one first chromosome; sampling a plurality of second chromosomes from the primary population based on the fitness of the plurality of first chromosomes; performing crossover mutation operations on the plurality of sampled second chromosomes, respectively, to obtain a next generation population; until the number of genetic iteration operations meets a preset number, determining the population obtained by the last genetic iteration operation as the target population; and determining the selected third facial image indicated by the target chromosome with the highest fitness in the target population as the facial image sequence.
[0010] In the implementation manner of the present application, a genetic algorithm can be used to further screen the multiple second face images, so as to screen out relatively redundant face images in the multiple second face images to the greatest extent, thereby reducing the processing amount in the facial expression recognition process.
[0011] In a possible implementation, the method further includes: parsing the video stream to be processed to obtain multiple video frames; segmenting from the multiple video frames multiple fourth facial images containing the target face; determining a first similarity between every two fourth facial images in the multiple fourth facial images; and based on the determined first similarity, screening out the multiple second facial images from the multiple fourth facial images, wherein the first similarity difference between any two second facial images in the multiple second facial images is greater than a preset similarity difference.
[0012] In the implementation mode of the present application, after obtaining the video stream, multiple fourth face images containing the target face can be segmented from the video stream, instead of processing each video frame in the video stream, thereby reducing the processing amount. Furthermore, redundant fourth face images are screened out from the multiple fourth face images to obtain the second face image, so that more face images with differences can be retained, and the processing amount in the facial expression recognition process can be reduced while ensuring the effect of facial expression recognition.
[0013] In one possible implementation, expression recognition is performed on the third feature map through the recognition module in the trained three-dimensional convolutional neural network model, including: multiplying the third feature map by the weights in the recognition module respectively to obtain probability values corresponding to multiple expressions of the target face; and determining the expression corresponding to the maximum probability value as the target expression recognition result corresponding to the target face.
[0014] In the implementation of the present application, the weights in the recognition module can be directly multiplied by the third feature map to obtain multiple probability values, thereby determining the expression recognition result corresponding to the target face according to the probability value, without the need to use other models for expression recognition. This method of determining the expression recognition result is simpler. In addition, the third feature map integrates the features of each face image in the face image sequence, which can help improve the accuracy of facial expression recognition.
[0015] In one possible implementation, the trained three-dimensional convolutional neural network model is trained through the following process: transferring model parameters of the trained two-dimensional convolutional neural network model as model parameters of the three-dimensional convolutional neural network model to be trained to obtain a pre-trained three-dimensional convolutional neural network model; based on a first data set, training the pre-trained three-dimensional convolutional neural network model until the pre-trained three-dimensional convolutional neural network model meets a third convergence condition to obtain the trained three-dimensional convolutional neural network model, wherein the first data set includes a first sample face image sequence corresponding to at least one first sample face, and a real facial expression recognition label corresponding to each first sample face image sequence.
[0016] In an implementation manner of the present application, the model parameters of the trained two-dimensional convolutional neural network model are migrated to the three-dimensional convolutional neural network model to be trained, and subsequently only fine-tuning and training of the three-dimensional convolutional neural network model to be trained is required. Since the amount of computation involved in the process of training the two-dimensional convolutional neural network model is less than that required for training 3D CNN, it is more conducive to improving the efficiency of the training process.
[0017] In one possible implementation, the trained two-dimensional convolutional neural network model is trained through the following process: based on a second data set, the two-dimensional convolutional neural network model to be trained is trained until the two-dimensional convolutional neural network model to be trained meets a first convergence condition, and an initially trained two-dimensional convolutional neural network model is obtained, wherein the second data set includes a first sample face image sequence corresponding to at least one second sample face, and a real face identity label corresponding to each second sample face image sequence; based on a third data set, the initially trained two-dimensional convolutional neural network module is trained until the initially trained two-dimensional convolutional neural network module meets a second convergence condition, and a trained two-dimensional convolutional neural network model is obtained, wherein the third data set includes a third sample face image sequence corresponding to at least one third sample face, and a real face expression recognition label corresponding to each third sample face image sequence.
[0018] In an implementation manner of the present application, a two-dimensional convolutional neural network model can be initially trained using a first data set corresponding to a human face to obtain an initially trained two-dimensional convolutional neural network model, and then the initially trained two-dimensional convolutional neural network model can be retrained using a second data set corresponding to an expression to obtain a trained two-dimensional convolutional neural network model. Since training is performed with the help of human face data, the demand for expression data is reduced.
[0019] In a second aspect, a facial expression recognition device is provided, including: an acquisition module, used to obtain a facial image sequence corresponding to a target face, wherein the facial image sequence includes multiple first facial images, and one of the first facial images includes the target face; a feature module, used to perform a first convolution operation on the facial image sequence through a feature extraction module in a trained three-dimensional convolutional neural network model to obtain multiple first feature maps, wherein one of the first feature maps corresponds to one of the multiple first facial images, and the one first feature map includes the one first facial image including feature sub-maps corresponding to each of the multiple channels; Through the feature extraction module, a second convolution operation is performed on the feature subgraphs in the multiple first feature graphs on each channel to obtain multiple second feature graphs; through the feature extraction module, a third convolution operation is performed on the multiple second feature graphs to obtain a third feature graph, the third convolution operation is to fuse the multiple second feature graphs, and the third feature graph is used to describe the expression features of the target face in the face image sequence; an expression recognition module is used to perform expression recognition on the third feature graph through the recognition module in the trained three-dimensional convolutional neural network model to obtain a target expression recognition result of the target face.
[0020] In a possible implementation, the size of the convolution kernel corresponding to the second convolution operation is M*K*K*K, where K is a positive integer and M represents the number of channels corresponding to the first feature map; the size of the convolution kernel corresponding to the third convolution operation is M*1*1*1*N, where N represents the number of channels corresponding to the third feature map.
[0021] In a possible implementation, the acquisition module is specifically used to: obtain a primary population based on multiple second facial images, the primary population includes multiple first chromosomes, one of the first chromosomes is used to indicate a third facial image selected from the multiple second facial images, the total number of the selected third facial images indicated by any two first chromosomes is the same, and the multiple second facial images all contain the target face; perform at least one genetic iteration operation on the primary population, wherein one genetic iteration operation includes: respectively determining the fitness of multiple first chromosomes in the primary population, the fitness of one first chromosome is determined according to the similarity between each two third facial images selected indicated by the one first chromosome; based on the fitness of the multiple first chromosomes, sampling multiple second chromosomes from the primary population; performing crossover mutation operations on the sampled multiple second chromosomes to obtain a next generation population; until the number of genetic iteration operations meets the preset number, the population obtained by the last genetic iteration operation is determined as the target population; and the selected third facial image indicated by the target chromosome with the highest fitness in the target population is determined as the facial image sequence.
[0022] In a possible implementation, the acquisition module is also used to: parse the video stream to be processed to obtain multiple video frames; segment out multiple fourth facial images containing the target face from the multiple video frames; determine a first similarity between every two fourth facial images in the multiple fourth facial images; and based on the determined first similarity, screen out the multiple second facial images from the multiple fourth facial images, wherein the first similarity difference between any two second facial images in the multiple second facial images is greater than a preset similarity difference.
[0023] In a possible implementation, the expression recognition module is specifically used to: multiply the third feature map by the weights in the recognition module respectively to obtain probability values corresponding to multiple expressions of the target face; and determine the expression corresponding to the maximum probability value as the target expression recognition result corresponding to the target face.
[0024] In one possible implementation, the trained three-dimensional convolutional neural network model is trained through the following process: transferring model parameters of the trained two-dimensional convolutional neural network model as model parameters of the three-dimensional convolutional neural network model to be trained to obtain a pre-trained three-dimensional convolutional neural network model; based on a third data set, training the pre-trained three-dimensional convolutional neural network model until the pre-trained three-dimensional convolutional neural network model meets a third convergence condition to obtain the trained three-dimensional convolutional neural network model, wherein the third data set includes a first sample face image sequence corresponding to at least one first sample face, and a real facial expression recognition label corresponding to each first sample face image sequence.
[0025] In a possible implementation, the trained two-dimensional convolutional neural network model is trained through the following process: based on a first data set, the two-dimensional convolutional neural network model to be trained is trained until the two-dimensional convolutional neural network model to be trained meets a first convergence condition, and an initially trained two-dimensional convolutional neural network model is obtained, wherein the first data set includes a first sample face image sequence corresponding to at least one second sample face, and a real face identity label corresponding to each second sample face image sequence; based on a second data set, the initially trained two-dimensional convolutional neural network module is trained until the initially trained two-dimensional convolutional neural network module meets a second convergence condition, and a trained two-dimensional convolutional neural network model is obtained, wherein the second data set includes a third sample face image sequence corresponding to at least one third sample face, and a real face expression recognition label corresponding to each third sample face image sequence.
[0026] In a third aspect, a facial expression recognition device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the at least one processor implements the facial expression recognition method as described in any one of the first aspects by executing the instructions stored in the memory.
[0027] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and when the computer instructions are executed on a computer, the computer executes the facial expression recognition method as described in any one of the first aspects.
[0028] Regarding the beneficial effects achieved in the second to fourth aspects, reference can be made to the beneficial effects discussed in the first aspect and will not be listed here again. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 A processing example diagram of a three-dimensional convolutional neural network model provided by the prior art;
[0030] Figure 2 A schematic diagram of a scenario in which the facial expression recognition method provided in an embodiment of the present application is applicable;
[0031] Figure 3 A schematic diagram of a facial expression recognition method according to an embodiment of the present invention;
[0032] Figure 4 A schematic diagram of the process flow corresponding to the genetic algorithm provided in the embodiment of the present application;
[0033] Figure 5 A schematic diagram of a process of processing a first feature map by a three-dimensional convolutional neural network model provided in an embodiment of the present application;
[0034] Figure 6 A schematic diagram of the structure of a first bottleneck module provided in an embodiment of the present application;
[0035] Figure 7 A schematic diagram of a processing process of a depthwise separable convolution in a two-dimensional convolutional neural network model provided in an embodiment of the present application;
[0036] Figure 8 A schematic diagram of the structure of a facial expression recognition device provided in an embodiment of the present application;
[0037] Fig. 9 A schematic diagram of the structure of a facial expression recognition device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0038] In order to better understand the technical solution provided by the embodiments of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0039] In order to more clearly understand the technical solutions involved in the embodiments of the present application, the technical terms involved in the embodiments of the present application are introduced below.
[0040] 1. Convolution (ConV): The convolution operation on an image can be understood as sliding a convolution kernel (or convolution template) on the image, multiplying the feature value (such as grayscale value) of the pixel on the image with the corresponding value on the convolution kernel, and then adding all the multiplied values as the feature value of the pixel corresponding to the convolution result, and so on, finally sliding to complete the image processing process.
[0041] 2. Convolutional neural network (CNN) is a feedforward neural network whose artificial neurons can respond to surrounding units within a certain coverage area, and has excellent performance in large image processing. Convolutional neural network consists of one or more convolutional layers and a fully connected layer at the top (corresponding to the classic neural network), as well as associated weights and pooling layers.
[0042] It should be noted that "A and / or B" in the embodiments of the present application may specifically include three situations: A, including B, and including A and B. "Multiple" in the embodiments of the present application refers to two or more than two.
[0043] As discussed above, when performing facial expression recognition, a two-dimensional convolutional neural network model is used to process facial images. Since the features extracted by the two-dimensional convolutional neural network model only contain two-dimensional features, the included features are not comprehensive enough, which may lead to poor recognition of facial expressions.
[0044] However, if a 3D convolutional neural network model is used for processing, the amount of calculation is relatively large. Figure 1 , which is a processing example diagram of a three-dimensional convolutional neural network model provided by the prior art. Figure 1 , the processing process of the three-dimensional convolutional neural network model in the prior art is introduced.
[0045] Specifically, the input feature map 101 of the face image sequence of the 3D convolutional neural network model includes feature sub-maps corresponding to three channels, and the size of each feature sub-map is H1×W1×T1. The size of the input feature map 101 is H1×W1×T1×Cin, where H1 can be understood as the length of each feature sub-map, W1 represents the width of each feature sub-map, T1 represents the height of each feature sub-map, and Cin represents the number of channels corresponding to the input feature map 101. Alternatively, T1 can also be understood as a time series dimension.
[0046] When the 3D convolutional neural network model performs convolution processing on the input feature map 101, it uses a convolution kernel 102 for processing, wherein the size of the convolution kernel 102 is Cin*Cout*K*K*K, so that an output feature map 103 of H2×W2×T2×Cout can be output. Among them, H2 can be understood as the length of the output feature map 103, W2 represents the width of the output feature map 103, T2 represents the height of the output feature map 103, and Cout represents the number of channels corresponding to the output feature map 103.
[0047] Based on the above, it can be seen that the size of the facial image or feature input to the three-dimensional convolutional neural network model itself is large, and the size corresponding to the convolution kernel in the convolution operation process of the three-dimensional convolutional neural network model is also large, which leads to a large amount of calculation in the process of expression recognition by the three-dimensional convolutional neural network model.
[0048] In view of this, the present application implements a facial expression recognition scheme, in which a three-dimensional convolutional neural network model is used to process the first feature maps of multiple facial images. In the process of processing the first feature map, each facial image is processed by a depth-separable convolution. For example, a second convolution operation is performed on the feature sub-maps corresponding to the multiple channels included in the first feature map to obtain multiple second feature maps, and then point convolution is performed on the multiple second feature maps. In this way, it is equivalent to splitting a 3D convolution operation in the prior art into two convolution operations, thereby reducing the data processing dimension during the convolution operation and reducing the amount of calculation in the facial expression recognition process. In addition, the three-dimensional convolutional neural network model can extract the three-dimensional features corresponding to the facial image, which is conducive to improving the accuracy of the facial expression recognition results and ensuring the effect of facial expression recognition.
[0049] The facial expression recognition scheme in the embodiment of the present application can be executed by a facial expression recognition device, which can be implemented by a device with computing capabilities, such as a terminal device or a server, such as a virtual server or a physical server. In order to simplify the description, the facial expression recognition device is referred to as a recognition device hereinafter.
[0050] Combine the following Figure 2 The application scenario diagram shown in the figure is used to introduce the application scenario of the facial expression recognition method in the embodiment of the present application. Or, Figure 2 It can also be understood as a schematic diagram of the deployment of identification equipment. Figure 2 As shown, the application scenario includes an identification device 210 , a terminal device 220 and a training device 230 . Figure 2 In the example, two terminal devices 220 are used, and the number of terminal devices 220 is not actually limited. The implementation of the identification device 210 can refer to the above. The implementation of the training device 230 can refer to the identification device 210, which will not be repeated here. The terminal device 220 is, for example, a mobile phone or a personal computer.
[0051] The training device 230 can train a three-dimensional convolutional neural network model based on the face image, thereby obtaining a trained three-dimensional convolutional neural network model. The training device 230 can also send the trained three-dimensional convolutional neural network model to the recognition device 210. The terminal device 220 can obtain multiple face images according to the user's operation, and send the multiple face images to the recognition device 210. The recognition device 210 can obtain multiple face images from the terminal device 220, and obtain expression recognition results corresponding to the multiple face images through the trained three-dimensional convolutional neural network model. Among them, the specific process of recognizing facial expressions will be introduced below.
[0052] In a possible application scenario, the recognition device 210 and the training device 230 may be the same device. In this case, it is equivalent to the recognition device 210 training the three-dimensional convolutional neural network model by itself.
[0053] It should be noted that the above Figure 2 The application scenarios shown are only examples of application scenarios applicable to the embodiments of the present application, but the facial expression recognition method in the embodiments of the present application can be applied to a variety of application scenarios, including but not limited to Figure 2 The application scenario shown.
[0054] The following is an introduction to the facial expression recognition scheme involved in the embodiment of the present application in conjunction with the accompanying drawings. Figure 3 , is a flow chart of a facial expression recognition method provided in an embodiment of the present application. Figure 2 The facial expression recognition method is introduced by taking the recognition device in FIG. 1 as an example to execute the facial expression recognition method.
[0055] Step 31, obtaining a facial image sequence corresponding to the target face.
[0056] The facial image sequence includes multiple first facial images, each of which includes a target face. The target face generally refers to a face of a person whose expression is to be recognized. There are many ways for the recognition device to obtain the facial image sequence, which are described in the following examples.
[0057] Method 1: The recognition device can obtain a facial image sequence from the terminal device.
[0058] The terminal device can obtain a facial image sequence of the target face according to the first input operation of the user. The terminal device sends the facial image sequence to the recognition device. Correspondingly, the recognition device also obtains the facial image sequence.
[0059] Method 2: The recognition device processes the plurality of second face images through a genetic algorithm to obtain a face image sequence.
[0060] Combine the following Figure 4The flowchart of the genetic algorithm provided in the embodiment of the present application is shown, which introduces a method of obtaining a facial image sequence by screening multiple second facial images based on the genetic algorithm.
[0061] Step 41, obtaining an initial generation population based on a plurality of second face images.
[0062] The primary population includes a plurality of first chromosomes, each of the plurality of first chromosomes is used to indicate a third face image selected from a plurality of second face images, and the total number of the selected third face images indicated by any two first chromosomes is the same. The plurality of second face images all contain the target face.
[0063] Optionally, each first chromosome may be represented by a character string. For example, the recognition device may preset the total number of selected third facial images to be n, where n is an integer greater than 1. The recognition device may select n third facial images from a plurality of second facial images, wherein the selected third facial images are represented by "1" and the unselected second facial images are represented by "0".
[0064] For example, assume that the number of the second facial images is N, where N is an integer greater than 1. The length of a first chromosome consists of N characters, and one second facial image corresponds to one character. Each character may be "0" or "1". Below, assuming that the value of N is 5 and the value of n is 3, the chromosomes included in the initial population may be any L of the following: "00111", "01011", "01101", "01110", "10011", "10101", "10110", "11001", "11010" and "11100", where L is an integer greater than 1. The recognition device can select L from these chromosomes, thereby obtaining L first chromosomes. Furthermore, the recognition device can select multiple first chromosomes from these L first chromosomes as the initial population.
[0065] Step 42, performing at least one genetic iteration operation on the initial population until the number of genetic iteration operations meets a preset number, and determining the population obtained by the last genetic iteration operation as the target population.
[0066] The preset number of times may be preconfigured in the recognition device, and the value of the preset number of times is an integer greater than or equal to 1. The process of each genetic iteration operation is the same, and an example is taken below to introduce one genetic iteration operation.
[0067] Step 1.1, respectively determining the fitness of multiple first chromosomes in the primary population, wherein the fitness of one first chromosome is determined according to the similarity between each two third face images selected and indicated by the first chromosome.
[0068] Optionally, the recognition device may be pre-configured with a fitness function f(m). The fitness function may be determined based on the similarity between the selected third images. For example, the fitness of a first chromosome is determined based on the sum of the similarities between every two third face images in all third face images in the first chromosome.
[0069] Step 1.2, based on the fitness of the multiple first chromosomes, multiple second chromosomes are sampled from the initial population.
[0070] The identification device determines the fitness f(m) of each first chromosome in the primary population M(t), samples the first chromosome in the primary population according to the selection probability p(f(m)), and obtains the sampled second chromosome. The probability of a first chromosome being selected is:
[0071] Step 1.3, perform crossover mutation operations on the sampled multiple second chromosomes respectively to obtain the next generation population.
[0072] The recognition device can perform a crossover and mutation operation on the sampled second chromosomes, wherein performing a crossover and mutation operation on multiple second chromosomes can be understood as first performing a crossover operation on the multiple second chromosomes and then performing a mutation operation. The crossover and mutation operations are introduced below.
[0073] 1. Cross operation.
[0074] The recognition device can divide multiple second chromosomes into two parts evenly. These two parts can be called two parent populations. The two parent populations are randomly shuffled, and then a second chromosome is taken from each of the two parent populations for crossover, thus completing a random matching process. In order to keep the chromosome feature dimension unchanged after the crossover, the recognition device can randomly select a segment (length is k) from a second chromosome, count the number of genes with 1 in the segment, and then search for a segment with the same length and the same number of genes containing "1" from the second chromosome of another matching parent population. If found, a crossover operation is performed; if not found from another parent population, it is searched from other populations until all parent populations are traversed.
[0075] 2. Mutation operation.
[0076] Gene mutation is the change of a gene at a certain site of a chromosome. Similarly, in order to keep the number of features in the chromosome unchanged after the mutation, first choose whether to perform the mutation operation according to a certain probability (called mutation probability). If yes, randomly select a second chromosome from multiple second chromosomes, and then randomly select a gene to reverse. If the gene changes from 1 to 0, then randomly select another 0 to change to 1, and vice versa. Perform the same operation until all individuals in multiple second chromosomes are traversed. In this way, it can be guaranteed that the number of features of each second chromosome will not be changed.
[0077] Similarly, after performing a crossover mutation operation on the sampled second chromosome, the next generation population can be obtained.
[0078] The recognition device repeats the above steps 1.1 to 1.3 for the next generation population until the number of genetic iteration operations reaches a preset number. The recognition device can use the population obtained by the last genetic iteration operation as the target population.
[0079] Step 43: determine the third face image selected by the target chromosome with the highest fitness in the target population as a face image sequence.
[0080] The recognition device can determine the fitness of each chromosome of the target population. The method of determining the fitness can be referred to above and will not be repeated here. The recognition device can determine the fitness of each second chromosome in the target population, and determine the target chromosome with the highest fitness in the target population, and determine the third face image indicated in the target chromosome as the face image sequence.
[0081] In a possible embodiment, the recognition device processes the video stream to be parsed to obtain a plurality of second facial images.
[0082] Step 2.1, parse the video stream to be processed to obtain multiple video frames.
[0083] The recognition device may obtain the video stream according to the second input operation of the user. Alternatively, the video stream may be obtained from the terminal device. The recognition device may parse the video stream to obtain multiple video frames.
[0084] Step 2.2, segmenting multiple fourth face images containing the target face from multiple video frames.
[0085] The recognition device obtains multiple video frames, but these multiple video frames do not necessarily contain the target face, or some video frames may contain faces that are not the target face. Therefore, in order to reduce the processing load of the subsequent recognition device, the recognition device in the embodiment of the present application can segment multiple fourth face images containing the target face from these multiple video frames, so that multiple fourth face images that all contain the target face can be obtained. In other words, each of the multiple fourth face images is segmented from one of the multiple video frames, and a face image can be part or all of the image area in the corresponding video frame. Segmentation can use, for example, an instance segmentation algorithm or a semantic segmentation algorithm, which is not limited in the embodiment of the present application.
[0086] Step 2.3: Determine a first similarity between every two fourth facial images in the plurality of fourth facial images.
[0087] The recognition device can extract the feature vectors of each of the plurality of fourth facial images, and determine the first similarity between the feature vectors corresponding to each two fourth facial images, thereby obtaining the first similarity between each two fourth facial images. The recognition device can combine the grayscale values of each pixel in each second facial image into a feature vector corresponding to the fourth facial image, or can extract the features of the fourth facial image through a feature extraction network, and use the features corresponding to the fourth facial image as the feature vector of the fourth facial image, such as a recurrent neural network (RNN).
[0088] For example, multiple fourth facial images are, for example, fourth facial image A, fourth facial image B, and fourth facial image C. The recognition device determines that the similarity between fourth facial image A and fourth facial image B is 0.5, determines that the similarity between fourth facial image A and fourth facial image C is 0.6, and determines that the similarity between fourth facial image B and fourth facial image C is 0.7, thereby obtaining the first similarity between any two of the three fourth facial images.
[0089] Step 2.4: based on the determined first similarity, select a plurality of second facial images from a plurality of fourth facial images.
[0090] After the recognition device determines the first similarity between any two fourth facial images among the plurality of fourth facial images, it can filter out, based on the preset similarity difference, a plurality of second facial images whose first similarity difference is greater than the preset similarity difference from the plurality of fourth facial images. In other words, the similarity between any two second facial images among the plurality of filtered second facial images is low. In this way, facial images with relatively high similarity can be filtered out, thereby preventing the facial image sequence from containing too many redundant images.
[0091] For example, continuing to refer to the above-mentioned fourth facial image A, fourth facial image B and fourth facial image C, the preset similarity difference is 0.1, and the terminal device determines that the similarity difference between the fourth facial image A and the fourth facial image C is 0.2, thereby determining that the multiple second facial images include the fourth facial image A and the fourth facial image C.
[0092] As an embodiment, the terminal device may directly use the screened multiple second facial images as a facial image sequence.
[0093] Step 32, performing a first convolution operation on the face image sequence through the feature extraction module in the trained three-dimensional convolutional neural network model to obtain multiple first feature maps, and performing a second convolution operation on the feature sub-maps on each channel of the multiple first feature maps to obtain multiple second feature maps, and performing a third convolution operation on the multiple second feature maps to obtain a third feature map.
[0094] The trained 3D convolutional neural network model includes a feature extraction module and a recognition module. The feature extraction module is used to extract the expression features of the face image sequence, and the recognition module is used to recognize the expression of the target face based on the expression features. The following is an example of the process of extracting expression features by the feature extraction module.
[0095] Specifically, the recognition device can perform a first convolution operation on a facial image sequence through a feature extraction module to obtain multiple first feature maps, wherein the first feature map includes a feature submap corresponding to each of the multiple channels of a first facial image. Furthermore, the feature extraction module can perform a second convolution operation on the feature submaps on each channel of the multiple first feature maps to obtain multiple second feature maps. For example, a first feature map includes feature submaps on three channels indicated by RGB, respectively. Then, the recognition device can perform a convolution operation on the feature submaps on the R channel of the multiple first feature maps through the feature extraction module, the recognition device can perform a convolution operation on the feature submaps on the G channel of the multiple first feature maps through the feature extraction module, and the recognition device can perform a convolution operation on the feature submaps on the B channel of the multiple first feature maps through the feature extraction module, thereby obtaining multiple second feature maps.
[0096] Optionally, the size of the convolution kernel corresponding to the second convolution operation is M*K*K*K, where K is a positive integer and M represents the number of channels corresponding to the first feature map. The feature extraction module can also perform a third convolution operation on multiple second feature maps to obtain a third feature map. The third feature map is used to describe the expression features of the target face in the face image sequence. The third convolution operation is used to fuse multiple second feature maps. Optionally, the size of the convolution kernel corresponding to the third convolution operation is M*1*1*1*N, where N represents the number of channels corresponding to the third feature map.
[0097] For example, see Figure 5 , which is a schematic diagram of the process of processing the first feature map by the three-dimensional convolutional neural network model provided in an embodiment of the present application. The recognition device can perform a second convolution operation on the feature sub-graph 510 on each channel of multiple first feature maps through the trained three-dimensional convolutional neural network model to obtain multiple second feature maps. The size of the convolution kernel 520 corresponding to the second convolution operation is M*K*K*K, where M represents the number of channels of the first feature map and K is a positive integer. Furthermore, the recognition device can perform a third convolution operation on the multiple second feature maps 530 to obtain a third feature map 550. The size of the convolution kernel 540 corresponding to the third convolution operation is M*1*1*1*N, where N corresponds to the number of channels of the third feature map. The feature sub-graph 560 in the third feature map 550 is obtained by processing multiple second feature maps 530.
[0098] like Figure 5 As shown, in the embodiment of the present application, Figure 5 The number of parameters of the 3D convolution operation shown is compared to Figure 1 The 3D convolution in the prior art shown in FIG. 1 is calculated as follows: Figure 1 The 3D convolution in the prior art shown is 1 / (K×K×K) times.
[0099] In a possible embodiment, the convolution feature extraction module includes at least one convolution module and at least one bottleneck module. The convolution module is used to perform a first convolution operation on the first face image. The bottleneck module can perform a second convolution operation and a third convolution operation on the feature subgraph corresponding to the first face image. Optionally, the recognition module can be implemented by a fully connected layer.
[0100] Exemplarily, the model structure of the trained three-dimensional convolutional neural network model can be specifically referred to as shown in Table 1 below.
[0101] Table 1
[0102] Module Type t c g s <![CDATA[s t ]]> Output Dimensions The first convolution module - 32 1 2 2 8*112*112*32 The first bottleneck module 1 16 1 1 1 8*112*112*16 The second bottleneck module 6 24 2 2 2 4*56*56*24 The third bottleneck module 6 32 3 2 1 2*28*28*32 The fourth bottleneck module 6 64 4 1 2 2*28*28*64 Fifth bottleneck module 6 96 3 2 1 2*14*14*96 The sixth bottleneck module 6 160 3 2 2 1*7*7*160 The seventh bottleneck module 6 320 1 1 1 1*7*7*320 The second convolution module - 1280 1 1 1 1*7*7*1280 Pooling module - - 1 - - 1*1*1280 Fully connected modules - 7 - - - 7
[0103] Where t represents the expansion factor. c represents the number of output channels. g represents the number of repetitions. s represents the spatial step size of the first layer in each module, and the step size of the remaining layers in each module is 1. t represents the time step of the first layer in each module, and the stride of the remaining layers in each module is 1.
[0104] As an embodiment, the first convolution module can implement the first convolution operation mentioned above. The first bottleneck module, the second bottleneck module, the third bottleneck module, the fourth bottleneck module, the fifth bottleneck module, the sixth bottleneck module and the seventh bottleneck module can each implement the second convolution operation and the third convolution operation mentioned above.
[0105] Optionally, the structures of the first bottleneck module, the second bottleneck module, the third bottleneck module, the fourth bottleneck module, the fifth bottleneck module, the sixth bottleneck module and the seventh bottleneck module may be the same or different. The structure of the first bottleneck module is taken as an example for description.
[0106] Please refer to Figure 6 , which is a structural schematic diagram of a first bottleneck module provided in an embodiment of the present application. The first bottleneck module includes a first convolutional layer, a depthwise separable convolutional layer, a second convolutional layer and an addition operation processing layer. Among them, the depthwise separable convolutional layer can be used to implement the second convolution operation and the third convolution operation in the foregoing text. The depthwise separable convolutional layer may specifically include two convolutional sublayers, wherein the size of the convolution kernel corresponding to one convolutional sublayer is M*K*K*K, and the meanings of M and K can be referred to in the foregoing text, and the size of the convolution kernel corresponding to the other convolutional sublayer is M*1*1*1*N, and the meanings of M and N can be referred to in the foregoing text.
[0107] Exemplarily, the size of the convolution kernel corresponding to the first convolution layer is 1*1*1, and a relu operation may be further included between the first convolution layer and the depthwise separable convolution layer. The size of the convolution kernel corresponding to the second convolution layer is 1*1*1, and a relu operation may be further included between the second convolution layer and the depthwise separable convolution layer. A linear operation may be included between the second convolution layer and the addition operation.
[0108] It should be noted that the above Table 1 and Figure 6 All of them are examples to introduce the structure of the three-dimensional convolutional neural network model in the embodiments of the present application, and do not limit the structure of the three-dimensional convolutional neural network model in the embodiments of the present application.
[0109] The trained three-dimensional convolutional neural network model in the embodiment of the present application can be trained by the recognition device itself, or the three-dimensional convolutional neural network model can be obtained by training the training device, and the recognition device obtains the trained three-dimensional convolutional neural network model from the training device. The following is an example of the recognition device training the three-dimensional convolutional neural network model by itself to obtain the trained three-dimensional convolutional neural network model.
[0110] Step 3.1, based on the second data set, train the two-dimensional convolutional neural network model to be trained until the two-dimensional convolutional neural network model to be trained meets the first convergence condition, and obtains a trained two-dimensional convolutional neural network model.
[0111] The recognition device can obtain a second data set from a network resource or other device. The second data set includes a second sample face image sequence corresponding to at least one second sample face, and a real face identity label corresponding to each second sample face image sequence, and the face identity label is used to indicate the person corresponding to the face. The recognition device uses the second data set to train the two-dimensional convolutional neural network model to be trained for face classification, so that the model parameters of the two-dimensional convolutional neural network model to be trained are adjusted under the supervision of the classification loss function, and the first convergence condition is met. The training is completed and the initially trained two-dimensional convolutional neural network model is obtained. The classification loss function is determined based on the difference between the real face expression recognition result and the real face identity label output by the two-dimensional convolutional neural network model.
[0112] Optionally, the first convergence addition is, for example, that the value of the classification loss function is less than a preset loss value and / or the number of training times for the two-dimensional convolutional neural network model to be trained is greater than or equal to a preset number of training times.
[0113] In a possible embodiment, the convolution layer in the two-dimensional convolutional neural network model may also use depth-separable convolution. For example, the convolution layer in a general two-dimensional convolutional neural network model has only one convolution operation, and the corresponding convolution parameter size is Ci*Co*K*K*. The two-dimensional convolutional neural network model in the embodiment of the present application may also use depth-separable convolution, specifically splitting one convolution operation into two convolution operations, and the sizes of the convolution kernels corresponding to these two convolution operations are Ci*K*K and Ci*1*1*Co, respectively. Among them, Ci represents the number of channels of the input features of the two-dimensional convolutional neural network model. Co represents the number of channels of the output features of the two-dimensional convolutional neural network model.
[0114] Please refer to Figure 7 , is a schematic diagram of processing a depth-separable convolution in a two-dimensional convolutional neural network model provided in an embodiment of the present application. Figure 7 As shown, the two-dimensional convolutional neural network model can perform a convolution operation on the input feature map 701. Figure 7 The convolution kernel 702 in is the convolution kernel corresponding to the convolution operation, and the intermediate feature map 703 is output. Then the convolution operation is performed on the intermediate feature map 703. Figure 7 The convolution kernel 704 in corresponds to the convolution operation, and then the feature map 706 is output. Among them, a feature sub-map 705 in the output feature map 706 is obtained by performing a convolution operation on multiple intermediate feature maps 703.
[0115] Step 3.2, based on the third data set, the initially trained two-dimensional convolutional neural network module is trained until the initially trained two-dimensional convolutional neural network module meets the second convergence condition, and a trained two-dimensional convolutional neural network model is obtained. Wherein, the third data set includes a third sample face image sequence corresponding to at least one third sample face, and a real facial expression recognition label corresponding to each third sample face image sequence. The meaning of the second convergence condition can be referred to the first convergence condition above, and will not be repeated here.
[0116] It should be noted that steps 3.1 and 3.2 are examples of the process of training a two-dimensional convolutional neural network model. The actual process of training a two-dimensional convolutional neural network model includes but is not limited to this.
[0117] Step 3.3, transfer the model parameters of the trained two-dimensional convolutional neural network model as the model parameters of the three-dimensional convolutional neural network model to be trained, and obtain a pre-trained three-dimensional convolutional neural network model.
[0118] Specifically, since the convolution layer in the two-dimensional convolutional neural network model can be regarded as two-dimensional, and the convolution layer in the three-dimensional convolutional neural network model can be regarded as three-dimensional, the model parameters in the two-dimensional convolutional neural network model can be copied and expanded, and then migrated to the three-dimensional convolutional neural network model to be trained.
[0119] For example, 2DCNN mainly includes 2D convolutional layers, ReLu and pool layers. 3DCNN mainly includes 3D convolutional layers, Relu and pool layers. The difference between the model parameters of 2DCNN and 3DCNN lies in the different dimensions of the convolutional layer. Other layers do not contain model parameters. Therefore, the copy expansion is only for the convolutional layer. The following is an example of the copy expansion process. Assume that the size of the convolution kernel of a 2D convolutional layer is 3*3, and the parameter of 3*3 is the matrix The size of the convolution kernel corresponding to a 3D convolution layer is 2*3*3, so the model parameter matrix corresponding to the copied 3D convolution kernel is:
[0120] Step 3.4, based on the first data set, train the pre-trained three-dimensional convolutional neural network model until the pre-trained three-dimensional convolutional neural network model meets the third convergence condition, and obtain a trained three-dimensional convolutional neural network model.
[0121] The first data set includes a first sample face image sequence corresponding to at least one first sample face, and a real facial expression recognition result corresponding to each first sample face image sequence. The first data set and the third data set may be the same data or different data sets. The method for the recognition device to obtain the first data set may refer to the method for obtaining the third data set in the previous text, which will not be repeated here. After the recognition device obtains the first data set, the pre-trained three-dimensional convolutional neural network model can be trained based on the first data set until the pre-trained three-dimensional convolutional neural network model meets the third convergence condition. The third convergence condition may refer to the content of the first convergence condition in the previous text, which will not be repeated here.
[0122] Step 33, performing expression recognition on the third feature map through the recognition module in the trained three-dimensional convolutional neural network model to obtain a target expression recognition result of the target face.
[0123] The recognition device can multiply the third feature map with the weights in the recognition module respectively to obtain the probability values corresponding to the target face in multiple expressions; and determine the expression corresponding to the maximum probability value as the target expression recognition result corresponding to the target face.
[0124] Based on the same inventive concept, the embodiment of the present application provides a facial expression recognition device, which can realize the functions of the facial expression recognition device discussed above. Figure 8 The facial expression recognition device includes: an acquisition module 801, which is used to obtain a facial image sequence corresponding to a target face, the facial image sequence includes multiple first facial images, one of which contains the target face; a feature module 802, which is used to perform a first convolution operation on the facial image sequence through a feature extraction module in a trained three-dimensional convolutional neural network model to obtain multiple first feature maps, wherein one first feature map corresponds to one of the multiple first facial images, and one first feature map includes a first facial image including feature sub-maps corresponding to each of the multiple channels; through the feature extraction module, a second convolution operation is performed on the feature sub-maps in the multiple first feature maps on each channel to obtain multiple second feature maps; through the feature extraction module, a third convolution operation is performed on the multiple second feature maps to obtain a third feature map, the third convolution operation is a fusion of the multiple second feature maps, and the third feature map is used to describe the expression features of the target face in the facial image sequence; an expression recognition module 803, which is used to perform expression recognition on the third feature map through a recognition module in a trained three-dimensional convolutional neural network model to obtain a target expression recognition result of the target face.
[0125] In one possible implementation, the size of the convolution kernel corresponding to the second convolution operation is M*K*K*K, where K is a positive integer and M represents the number of channels corresponding to the first feature map; the size of the convolution kernel corresponding to the third convolution operation is M*1*1*1*N, where N represents the number of channels corresponding to the third feature map.
[0126] In a possible implementation, the obtaining module 801 is specifically used to: obtain a primary population based on multiple second facial images, the primary population includes multiple first chromosomes, one of the first chromosomes is used to indicate a third facial image selected from the multiple second facial images, the total number of the selected third facial images indicated by any two first chromosomes is the same, and the multiple second facial images all contain a target face; perform at least one genetic iteration operation on the primary population, wherein one genetic iteration operation includes: respectively determining the fitness of multiple first chromosomes in the primary population, the fitness of one first chromosome is determined according to the similarity between each two third facial images selected indicated by a first chromosome; based on the fitness of the multiple first chromosomes, sampling multiple second chromosomes from the primary population; performing crossover mutation operations on the sampled multiple second chromosomes to obtain a next generation population; until the number of genetic iteration operations meets the preset number, the population obtained by the last genetic iteration operation is determined as the target population; and the selected third facial image indicated by the target chromosome with the highest fitness in the target population is determined as the facial image sequence.
[0127] In a possible implementation, the acquisition module 801 is also used to: parse the video stream to be processed to obtain multiple video frames; segment from the multiple video frames multiple fourth facial images containing the target face; determine a first similarity between every two fourth facial images in the multiple fourth facial images; based on the determined first similarity, screen out multiple second facial images from the multiple fourth facial images, wherein the first similarity difference between any two second facial images in the multiple second facial images is greater than a preset similarity difference.
[0128] In one possible implementation, the expression recognition module 803 is specifically used to: multiply the third feature map by the weights in the recognition module respectively to obtain the probability values corresponding to multiple expressions of the target face; and determine the expression corresponding to the maximum probability value as the target expression recognition result corresponding to the target face.
[0129] In one possible implementation, the trained three-dimensional convolutional neural network model is trained through the following process: transferring model parameters of the trained two-dimensional convolutional neural network model as model parameters of the three-dimensional convolutional neural network model to be trained to obtain a pre-trained three-dimensional convolutional neural network model; based on a third data set, training the pre-trained three-dimensional convolutional neural network model until the pre-trained three-dimensional convolutional neural network model meets a third convergence condition to obtain a trained three-dimensional convolutional neural network model, wherein the third data set includes a first sample face image sequence corresponding to at least one first sample face, and a real facial expression recognition label corresponding to each first sample face image sequence.
[0130] In one possible implementation, the trained two-dimensional convolutional neural network model is obtained by training through the following process: based on the first data set, the two-dimensional convolutional neural network model to be trained is trained until the two-dimensional convolutional neural network model to be trained meets the first convergence condition, and an initially trained two-dimensional convolutional neural network model is obtained, wherein the first data set includes a first sample face image sequence corresponding to at least one second sample face, and a real face identity label corresponding to each second sample face image sequence; based on the second data set, the initially trained two-dimensional convolutional neural network module is trained until the initially trained two-dimensional convolutional neural network module meets the second convergence condition, and a trained two-dimensional convolutional neural network model is obtained, wherein the second data set includes a third sample face image sequence corresponding to at least one third sample face, and a real face expression recognition label corresponding to each third sample face image sequence.
[0131] It should be noted that Figure 8 The facial expression recognition device in the present invention can also be used to implement any of the facial expression recognition methods discussed above.
[0132] Based on the same inventive concept, the present application embodiment provides a facial expression recognition device, please refer to Fig. 9 , is a schematic diagram of the structure of a facial expression recognition device provided in an embodiment of the present application. The facial expression recognition device includes: at least one processor 901, and a memory 902 that is communicatively connected to the at least one processor 901; wherein the memory 902 stores instructions that can be executed by the at least one processor 901, and the at least one processor 901 implements any of the facial expression recognition methods discussed above by executing the instructions stored in the memory 902.
[0133] Optional, Fig. 9 The facial expression recognition device shown can also realize the Figure 8 Functions of the facial expression recognition device shown.
[0134] Optionally, the processor 901 may be a central processing unit (CPU) or a digital processing unit, etc. The specific connection medium between the memory 902 and the processor 901 is not limited in the embodiment of the present application. Fig. 9 In the embodiment, the memory 902 and the processor 901 are connected via a bus 903. The bus 903 is connected to the processor 901 via a bus 903. Fig. 9 The connections between other components are shown in bold lines, which are only for illustration and are not intended to be limiting. Bus 903 can be divided into address bus, data bus, control bus, etc. For ease of illustration, Fig. 9 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0135] The memory 902 may be a volatile memory, such as a random-access memory (RAM); the memory 902 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), or the memory 902 may be any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 902 may be a combination of the above memories.
[0136] Based on the same inventive concept, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions. When the computer instructions are executed on a computer, the computer executes the facial expression recognition method as described in any one of the first aspects.
[0137] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0138] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0139] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0140] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0141] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0142] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A method for recognizing facial expressions, characterized in that: include: Obtaining a facial image sequence corresponding to a target face, wherein the facial image sequence includes a plurality of first facial images, one of which includes the target face; Performing a first convolution operation on the facial image sequence through a feature extraction module in a trained three-dimensional convolutional neural network model to obtain a plurality of first feature maps, wherein one of the first feature maps corresponds to a first facial image among the plurality of first facial images, and the first feature map includes feature sub-maps corresponding to the first facial image on a plurality of channels; By means of the feature extraction module, a second convolution operation is performed on the feature sub-graphs on each channel of the plurality of first feature graphs to obtain a plurality of second feature graphs; By means of the feature extraction module, a third convolution operation is performed on the plurality of second feature maps to obtain a third feature map, wherein the third convolution operation is a fusion of the plurality of second feature maps, and the third feature map is used to describe the expression features of the target face in the face image sequence; Performing expression recognition on the third feature map through the recognition module in the trained three-dimensional convolutional neural network model to obtain a target expression recognition result of the target face; Wherein, obtaining a facial image sequence corresponding to the target face includes: Based on the plurality of second facial images, an initial population is obtained, the initial population comprising a plurality of first chromosomes, wherein one first chromosome is used to indicate a third facial image selected from the plurality of second facial images, the total number of the selected third facial images indicated by any two first chromosomes is the same, and the plurality of second facial images all contain the target face; Performing at least one genetic iteration operation on the primary population, wherein one genetic iteration operation includes: respectively determining the fitness of a plurality of first chromosomes in the primary population, wherein the fitness of one of the first chromosomes is determined according to the similarity between each two third face images selected and indicated by the one of the first chromosomes, sampling a plurality of second chromosomes from the primary population based on the fitness of the plurality of first chromosomes, and respectively performing crossover mutation operations on the plurality of second chromosomes sampled to obtain a next generation population; Until the number of genetic iteration operations meets the preset number, the population obtained by the last genetic iteration operation is determined as the target population; The third face image selected and indicated by the target chromosome with the highest fitness in the target population is determined as the face image sequence.
2. The method according to claim 1, characterized in that The size of the convolution kernel corresponding to the second convolution operation is M*K*K*K, where K is a positive integer and M represents the number of channels corresponding to the first feature map; The size of the convolution kernel corresponding to the third convolution operation is M*1*1*1*N, where N represents the number of channels corresponding to the third feature map.
3. The method according to claim 1, characterized in that The method further comprises: Parse the video stream to be processed and obtain multiple video frames; Segmenting a plurality of fourth face images including the target face from the plurality of video frames; determining a first similarity between every two fourth facial images in the plurality of fourth facial images; Based on the determined first similarity, the plurality of second facial images are screened out from the plurality of fourth facial images, wherein a first similarity difference between any two second facial images among the plurality of second facial images is greater than a preset similarity difference.
4. The method according to claim 1, characterized in that The recognition module is implemented by a fully connected layer; and expression recognition is performed on the third feature map by using the recognition module in the trained three-dimensional convolutional neural network model, including: Multiplying the third feature map by the weights in the recognition module respectively to obtain probability values corresponding to the target face in multiple expressions respectively; The expression corresponding to the determined maximum probability value is determined as the target expression recognition result corresponding to the target face.
5. The method according to any one of claims 1 to 4, characterized in that: The trained three-dimensional convolutional neural network model is obtained by training through the following process: Transferring the model parameters of the trained two-dimensional convolutional neural network model as the model parameters of the three-dimensional convolutional neural network model to be trained, to obtain a pre-trained three-dimensional convolutional neural network model; Based on the first data set, the pre-trained three-dimensional convolutional neural network model is trained until the pre-trained three-dimensional convolutional neural network model meets the third convergence condition to obtain the trained three-dimensional convolutional neural network model, wherein the first data set includes a first sample face image sequence corresponding to at least one first sample face, and a real facial expression recognition label corresponding to each first sample face image sequence.
6. The method according to claim 5, characterized in that The trained two-dimensional convolutional neural network model is obtained by training through the following process: Based on the second data set, the two-dimensional convolutional neural network model to be trained is trained until the two-dimensional convolutional neural network model to be trained meets the first convergence condition, thereby obtaining an initially trained two-dimensional convolutional neural network model, wherein the second data set includes a first sample face image sequence corresponding to at least one second sample face, and a real face identity label corresponding to each second sample face image sequence; Based on a third data set, the initially trained two-dimensional convolutional neural network module is trained until the initially trained two-dimensional convolutional neural network module meets a second convergence condition, thereby obtaining a trained two-dimensional convolutional neural network model, wherein the third data set includes a third sample face image sequence corresponding to at least one third sample face, and a real facial expression recognition label corresponding to each third sample face image sequence.
7. A facial expression recognition device, characterized in that: include: An acquisition module, used for acquiring a facial image sequence corresponding to a target face, wherein the facial image sequence includes a plurality of first facial images, and one of the first facial images includes the target face; A feature module, used to perform a first convolution operation on the facial image sequence through a feature extraction module in a trained three-dimensional convolutional neural network model to obtain a plurality of first feature maps, wherein one of the first feature maps corresponds to a first facial image among the plurality of first facial images, and the first feature map includes feature sub-maps corresponding to the first facial image on a plurality of channels; perform a second convolution operation on the feature sub-maps in the plurality of first feature maps on each channel through the feature extraction module to obtain a plurality of second feature maps; perform a third convolution operation on the plurality of second feature maps through the feature extraction module to obtain a third feature map, wherein the third convolution operation is a fusion of the plurality of second feature maps, and the third feature map is used to describe the expression features of the target face in the facial image sequence; An expression recognition module, used to perform expression recognition on the third feature map through the recognition module in the trained three-dimensional convolutional neural network model to obtain a target expression recognition result of the target face; Wherein, obtaining a facial image sequence corresponding to the target face includes: Based on the plurality of second facial images, an initial population is obtained, the initial population comprising a plurality of first chromosomes, wherein one first chromosome is used to indicate a third facial image selected from the plurality of second facial images, the total number of the selected third facial images indicated by any two first chromosomes is the same, and the plurality of second facial images all contain the target face; Performing at least one genetic iteration operation on the primary population, wherein one genetic iteration operation includes: respectively determining the fitness of a plurality of first chromosomes in the primary population, wherein the fitness of one of the first chromosomes is determined according to the similarity between each two third face images selected and indicated by the one of the first chromosomes, sampling a plurality of second chromosomes from the primary population based on the fitness of the plurality of first chromosomes, and respectively performing crossover mutation operations on the plurality of second chromosomes sampled to obtain a next generation population; Until the number of genetic iteration operations meets the preset number, the population obtained by the last genetic iteration operation is determined as the target population; The third face image selected and indicated by the target chromosome with the highest fitness in the target population is determined as the face image sequence.
8. A facial expression recognition device, characterized in that: include: at least one processor, and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the at least one processor implements the method according to any one of claims 1 to 6 by executing the instructions stored in the memory.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Expression recognition method, device and equipment for video
CN107977634A