Method, device and equipment for generating animal expression animation based on human face expression and medium
By performing keypoint detection and optical flow field transformation on human facial expression videos and animal face images, high-quality animal expression animations were generated, solving the problem that it is difficult to convert human facial expressions into animal expressions in existing technologies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU YUNXIANG TECH CO LTD
- Filing Date
- 2022-12-14
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technology struggles to effectively convert human facial expressions into animal expressions, resulting in a limited number of videos featuring animal facial expressions.
By acquiring videos of human facial expressions and images of animal faces, keypoint detection is performed to construct a sparse optical flow field. This field is then converted into an animal sparse optical flow field using a transformation network model, ultimately generating the target animal facial animation.
It has been developed to convert human facial expressions into animal expressions, solving the problem of the limited number of animal facial expression videos and generating high-quality animal expression animations.
Smart Images

Figure CN116109739B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video image processing technology, and in particular to a method, apparatus, device and medium for generating animal facial animation based on human facial expressions. Background Technology
[0002] Using video to drive an image is a relatively new technique that has emerged in recent years. For example, you can input a video of facial movements and another human face, then drive the input face to perform similar movements. The same principle can be applied to animal faces. However, while videos of human facial expressions are relatively easy to obtain and collect, videos of animal facial expressions are not so readily available.
[0003] Existing technologies can convert human facial images into animated expressions, but struggle to convert human facial expressions into animal expressions, resulting in a limited number of animal facial expression videos. There is an urgent need for a method that can convert human facial expressions into animal expressions to address the problem of the limited number of animal facial expression videos. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, device, and medium for generating animal facial animation based on human facial expressions, so as to realize the conversion of human facial expressions into animal expressions and solve the problem of the limited number of animal facial expression videos.
[0005] To address the aforementioned technical problems, embodiments of this application provide a method for generating animal facial expression animations based on human facial expressions, including:
[0006] Acquire videos of human facial expressions and images of animal faces;
[0007] Key point detection is performed on the human facial expression video and the animal face image respectively to obtain human facial key points and animal facial key points;
[0008] Based on the aforementioned facial key points, a sparse optical flow field for the face is constructed.
[0009] The sparse optical flow field of the human face is converted into a sparse optical flow field of the animal using a trained conversion network model.
[0010] Based on the animal face image, the animal face key points, and the animal sparse optical flow field, the animal sparse optical flow field is converted into a dense optical flow field;
[0011] By performing feature transformation and feature summarization on the animal face image and the dense optical flow field, an animated expression of the target animal is generated.
[0012] To address the aforementioned technical problems, embodiments of this application provide a device for generating animal facial expression animations based on human facial expressions, comprising:
[0013] The facial expression video acquisition unit is used to acquire facial expression videos and animal face images;
[0014] A key point detection unit is used to perform key point detection on the human facial expression video and the animal face image respectively, to obtain human facial key points and animal facial key points;
[0015] A sparse optical flow field construction unit is used to construct a sparse optical flow field of the face based on the facial key points;
[0016] A sparse optical flow field conversion unit is used to convert the sparse optical flow field of the human face into the sparse optical flow field of the animal through a trained conversion network model.
[0017] A dense optical flow field generation unit is used to convert the animal sparse optical flow field into a dense optical flow field based on the animal face image, the animal face key points, and the animal sparse optical flow field.
[0018] The target animal facial expression animation generation unit is used to generate target animal facial expression animation by performing feature transformation and feature summarization on the animal face image and the dense optical flow field.
[0019] To solve the above-mentioned technical problems, one technical solution adopted by the present invention is to provide a computer device, including one or more processors; and a memory for storing one or more programs, so that the one or more processors implement the method for generating animal expression animation based on human facial expressions as described above.
[0020] To solve the above-mentioned technical problems, one technical solution adopted by the present invention is: a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a processor, implements the method for generating animal expression animation based on human facial expressions as described above.
[0021] This invention provides a method, apparatus, device, and medium for generating animal facial expression animations based on human facial expressions. The method includes: acquiring a human facial expression video and an animal face image; performing keypoint detection on the human facial expression video and the animal face image respectively to obtain human keypoints and animal face keypoints; constructing a sparse optical flow field for the face based on the facial keypoints; converting the sparse optical flow field for the face into a sparse optical flow field for the animal using a trained transformation network model; converting the sparse optical flow field for the animal into a dense optical flow field based on the animal face image, the animal face keypoints, and the animal sparse optical flow field; and generating a target animal facial expression animation by performing feature transformation and feature summarization on the animal face image and the dense optical flow field. This invention can generate target animal facial expression animations based on human facial expression videos and animal face images, realizing the conversion of human facial expressions into animal expressions, thus solving the problem of the limited number of animal facial expression videos available. Attached Figure Description
[0022] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of an implementation of the method for generating animal facial animation based on human facial expressions provided in this application embodiment;
[0024] Figure 2 This is a flowchart illustrating an implementation of a sub-process in the method for generating animal facial animation based on human facial expressions provided in this application embodiment;
[0025] Figure 3 This is another implementation flowchart of a sub-process in the method for generating animal expression animation based on human facial expressions provided in the embodiments of this application;
[0026] Figure 4 This is another implementation flowchart of a sub-process in the method for generating animal expression animation based on human facial expressions provided in the embodiments of this application;
[0027] Figure 5 This is another implementation flowchart of a sub-process in the method for generating animal expression animation based on human facial expressions provided in the embodiments of this application;
[0028] Figure 6 This is another implementation flowchart of a sub-process in the method for generating animal expression animation based on human facial expressions provided in the embodiments of this application;
[0029] Figure 7 This is another implementation flowchart of a sub-process in the method for generating animal expression animation based on human facial expressions provided in the embodiments of this application;
[0030] Figure 8 This is a schematic diagram of a device for generating animal facial animation based on human facial expressions, provided in an embodiment of this application.
[0031] Figure 9 This is a schematic diagram of the computer device provided in the embodiments of this application. Detailed Implementation
[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0033] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0034] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0035] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0036] It should be noted that the method for generating animal facial animation based on human facial expressions provided in this application embodiment is generally executed by a server, and correspondingly, the device for generating animal facial animation based on human facial expressions is generally configured in the server.
[0037] Please see Figure 1 , Figure 1 This paper illustrates a specific implementation of a method for generating animal facial animations based on human facial expressions.
[0038] It should be noted that if substantially the same result is obtained, the method of this invention is not based on... Figure 1 Limited to the order of the processes shown, this method includes the following steps:
[0039] S1: Acquire videos of human facial expressions and images of animal faces.
[0040] Specifically, facial expression video is a video featuring human facial expressions. This application's embodiments generate animal expression animations based on facial expression videos and animal face images.
[0041] S2: Perform keypoint detection on human facial expression videos and animal face images respectively to obtain human facial keypoints and animal facial keypoints.
[0042] Specifically, a facial landmark detector and a target animal facial landmark detector are trained using several human face images and animal face images, respectively. The facial landmark detector and the target animal facial landmark detector are then used to detect key points in human facial expression videos and animal face images, respectively, to obtain facial landmarks and animal facial landmarks.
[0043] This application's embodiments are primarily applicable to animal faces that roughly correspond to human faces, such as primates, felines, canines, and bears. Regarding key points, using a human face as the standard, 16 key points are obtained: the center of the left and right eyes, the left and right corners of the mouth, the bridge of the nose, the tip of the nose, the center of the left and right ears, the top of the forehead, the area above the left and right temples, the left and right cheeks, the tip of the chin, and the left and right angles of the jaw. The key points on the animal face roughly correspond to those on the human face. Because images of animal faces are easier to collect than videos of animal expressions, the target animal face key point detector and the human face key point detector are trained separately and do not participate in the training of subsequent networks.
[0044] S3: Construct a sparse optical flow field for the face based on facial key points.
[0045] Please see Figure 2 , Figure 2 A specific implementation of step S3 is shown below:
[0046] S31: Based on each key point in the face, obtain the initial frame key points in the face key points.
[0047] S32: Based on the alignment method of the same key points, calculate the displacement of each face key point relative to the key points of the initial frame, and obtain the optical flow of the key points.
[0048] S33: Constructing a sparse optical flow field for a face based on key points.
[0049] Specifically, a sparse optical flow field is composed of the optical flow of several key points; and the optical flow of a certain key point is the displacement of a certain key point in a certain frame relative to the corresponding key point in the first frame. Therefore, in this embodiment, the optical flow of each key point in the face is calculated separately. When the optical flow of all key points is calculated, all the optical flows are summarized to obtain the sparse optical flow field of the face.
[0050] S4: The sparse optical flow field of a human face is converted into a sparse optical flow field of an animal using a trained conversion network model.
[0051] Specifically, the conversion network model is an autoencoder module based on fully connected layers, comprising three interconnected fully connected layers, with an activation layer below each of the first and second fully connected layers. In this embodiment, if 16 key points in the face are used, the Linear layer in the conversion network model represents the fully connected layer, ReLU is the activation function, and the first fully connected layer Linear1 has 32 neurons, the second fully connected layer Linear2 has 16 neurons, and the third fully connected layer Linear3 has 32 neurons.
[0052] Please see Figure 3 , Figure 3 A specific implementation method prior to step S4 is shown below in detail:
[0053] S4A: Acquire multiple animal expression videos and crop the videos to obtain the animal face regions.
[0054] Specifically, acquire multiple animal facial expression videos, each video consisting of at least one frame. Crop the animal facial expression videos to ensure the animal faces are as oriented as possible, and scale the animal face regions uniformly to a fixed scale, such as 256x256.
[0055] S4B: Extract key points of the animal face region to obtain the initial animal face key points.
[0056] Specifically, use G i Represents the initial animal face keypoints in frame i+1.
[0057] S4C: Acquire an initial human face video with the same actions and expressions as the animal face expression video, and crop the face from the initial human face video to obtain the face region.
[0058] Specifically, using the actions and expressions in animal facial expression videos as a baseline, corresponding human facial videos with the same actions and expressions are collected to generate initial human facial videos. Furthermore, for the same animal expression, multiple videos can be collected from multiple people in different scenarios, thus obtaining multiple initial human facial videos, making subsequent model training more accurate. Additionally, the human facial region maintains the same size as the animal facial region.
[0059] S4D: Align the human face region with the animal face region and extract the key points in the human face region to obtain the initial human face key points.
[0060] Specifically, the human face region is aligned with the animal face region so that the number of frames corresponding to the human face region and the animal face region is the same, let's say N. Using K... i This represents the initial facial key points in the (i+1)th frame.
[0061] S4E: Training data is constructed based on initial animal facial landmarks and initial human facial landmarks, and a transformation network model is trained using the training data to obtain a trained transformation network model.
[0062] Specifically, two integers i and j are randomly generated, ranging from [0, N), thus obtaining a training data pair (K). j -K i G j -G i The former represents the input data of the current network, and the latter represents the target data. The loss function of this network is defined as... here This represents the network output. In this embodiment, training data is constructed based on initial animal facial key points and initial human facial key points, and the transformation network model is trained using the training data to obtain the trained transformation network model.
[0063] S5: Based on animal face images, animal face key points, and animal sparse optical flow fields, the animal sparse optical flow field is converted into a dense optical flow field.
[0064] Specifically, in this embodiment of the application, based on the animal face image, animal face key points and animal sparse optical flow field, the animal sparse optical flow field is converted into a dense optical flow field, so as to obtain another image from one image, i.e., imagewarping.
[0065] Please see Figure 4 , Figure 4 A specific implementation of step S5 is shown below:
[0066] S51: Downsample the animal face image to obtain a downsampled image.
[0067] Specifically, the animal face image is downsampled to reduce its size, generating a downsampled image. It should be noted that the degree of downsampling depends on the specific application scenario and is not limited here. The resulting downsampled image is typically set to 64x64. The downsampled image is denoted as Y.
[0068] S52: Generate key point heatmaps and affine transformation maps based on downsampled images, animal face key points, and animal sparse optical flow fields.
[0069] Please see Figure 5 , Figure 5 A specific implementation of step S52 is shown below:
[0070] S521: Construct target key points based on animal face key points and animal sparse optical flow field.
[0071] S522: Convert the target key points into a Gaussian distribution heatmap to obtain the key point heatmap.
[0072] S523: Generate a tensor with the same length and width as the downsampled image according to the preset depth.
[0073] S524: Add the tensor to the optical flow corresponding to the target key point to obtain the initial dense optical flow field of the current key point.
[0074] S525: Perform affine transformation processing on the downsampled image based on the initial dense optical flow field to obtain the affine transformation map.
[0075] Specifically, the coordinates of key points on the animal's face and the coordinates of key points corresponding to the animal's sparse optical flow field are obtained to generate target key points, which include the coordinates of each key point. Then, the target key points expressed in coordinates are converted into a heatmap expressed in Gaussian distribution, resulting in the key point heatmap H. i Then, according to a preset depth, a tensor with the same dimensions as the downsampled image Y is generated, representing the x and y coordinates of the corresponding pixels, respectively. The preset depth is set according to actual conditions and is not limited here; in this case, the depth can be set to 2. Then, the tensor is added to the optical flow corresponding to the target keypoint to obtain the initial dense optical flow field of the current keypoint. Finally, an affine transformation is performed on the downsampled image Y based on the initial dense optical flow field to obtain the affine transformed image Y′. i Here, 'i' represents a key point.
[0076] S53: Connect the key point heatmap and affine transformation map to obtain animal face features.
[0077] Specifically, (H0,Y′0,…,H i ,Y′ i ,,…,H M ,Y′ M The concatenation of these points in the depth dimension is the output of the feature generation module, which is the output of animal face features; here, H0 is all 0, Y′0 is equal to Y, and Y is the number of key points.
[0078] S54: A dense optical flow field is obtained by segmenting and integrating animal face features.
[0079] Please see Figure 6 , Figure 6 A specific implementation of step S54 is shown below:
[0080] S541: The UNet network model is used to segment animal face features to obtain segmentation features.
[0081] S542: Perform convolution processing on the segmentation features to obtain convolutional features.
[0082] S543: Multiply the segmentation features and convolution features element by element to obtain the multiplication result, and accumulate the multiplication result according to the depth accumulation method to obtain the dense optical flow field.
[0083] Specifically, the UNet network model is used to segment animal face features. The output of the UNet module is a tensor of depth M+1 with the same dimensions as Y. After segmentation, the segmented features are input into a convolutional layer Conv1 for convolution. The kernel of Conv1 is 7x7, and the depth and dimensions of the segmented features are not changed during convolution. Finally, the segmented features and the convolutional features are multiplied element-wise to obtain the multiplication result. This multiplication result is then accumulated along the depthwise path to obtain a dense optical flow field.
[0084] S6: Generate target animal facial animation by performing feature transformation and feature summarization on animal face images and dense optical flow fields.
[0085] Please see Figure 7 , Figure 7 A specific implementation of step S6 is shown below:
[0086] S61: The feature encoding result is obtained by performing feature decomposition and feature encoding on the animal face image.
[0087] S62: The feature transformation result is obtained by performing feature transformation processing on the dense optical flow field.
[0088] S63: Generate the target animal's facial expression animation by performing feature decoding and feature summarization on the feature transformation results.
[0089] Specifically, the animal face image is decomposed using a feature decomposition module. This module is a 7x7 convolutional layer that does not change the width and height of the input animal face image, but only increases the dimension, typically set to 64 dimensions. The decomposed animal face image is then encoded by an encoder to obtain the feature encoding result. This encoder can be designed custom-, and the number of layers is generally related to the resolution. As mentioned earlier, with a density optical flow field of 64x64, this typically involves two convolutional layers, reducing the width and height to 64x64 and increasing the depth to 256. Then, a feature transformation module uses the dense optical flow field to warp the feature encoding result (feature transformation processing) to obtain the feature transformation result. This feature transformation result is then processed by a residual network, ResNet. This residual network is set to a depth of 4 to 6 layers, without changing the width, height, and depth of the tensors. Finally, a decoder decodes the feature transformation result to obtain the decoded result. Finally, a feature summarization module summarizes the decoded result to generate the target animal facial animation. The decoder has the opposite structure to the encoder, with the same number of layers, gradually increasing the width and height while decreasing the depth. The feature summarization module, unlike the feature decomposition module, summarizes the depth to three dimensions of the color image and limits the output values to the range [0,1] through a softmax layer.
[0090] Furthermore, the loss function in the neural network learning process of this application embodiment consists of two parts: a discriminator, whose input is the output of the generator, and the discriminator has a structure similar to an encoder, consisting of 3 to 4 layers. Let the output of the discriminator be D, and the loss function be L1 = (1-D)|| 2 +D|| 2 The other part involves using the intermediate outputs of a pre-trained VGG19 algorithm. It's recommended to use the ReLU outputs of the 1st, 3rd, 5th, 9th, and 13th convolutional layers of VGG19 as features. Specifically, the real input image and the generator's output image are fed into VGG19, and the outputs of the previously recommended layers are used as features to calculate their losses. here f represents the output of the generator. i This represents the ReLU output after the i-th convolutional layer of VGG19 mentioned earlier.
[0091] The neural network learning in this embodiment is performed in three steps: first, it is trained independently; then, based on the previous step, it is trained using an artificial optical flow field; and finally, based on the previous step, it is trained normally together with a dense optical flow field module using complete animal video data. During independent training, the feature transformation module is bypassed, and its input data consists of randomly setting a portion of a normal animal face image to 0, then using the neural network of this application to generate an image that approximates the original image. The method of training using an artificial optical flow field involves artificially constructing a dense optical flow field, using this optical flow field to warp the original image, using the warped image as the input image, and inverting the artificially generated dense optical flow field as the input to the network's dense optical flow field. The output is then compared with the original image before warping to calculate the loss function.
[0092] In this embodiment, a video of a human facial expression and an image of an animal face are acquired. Keypoint detection is performed on both the video and the image to obtain facial keypoints and animal keypoints. Based on these keypoints, a sparse optical flow field for the human face is constructed. A trained transformation network model is used to convert the sparse optical flow field for the human face into a sparse optical flow field for the animal face. Based on the animal face image, the keypoints, and the sparse optical flow field, the sparse optical flow field for the animal face is converted into a dense optical flow field. Feature transformation and feature summarization are performed on the animal face image and the dense optical flow field to generate a target animal facial expression animation. This embodiment of the invention can generate a target animal facial expression animation based on a video of a human facial expression and an image of an animal face, thus solving the problem of the limited number of animal facial expression videos available.
[0093] Please refer to Figure 8 As a response to the above Figure 1 The implementation of the method shown in this application provides an embodiment of a device for generating animal expression animation based on human facial expressions. This device embodiment is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0094] like Figure 8 As shown, the animal expression animation generation device based on human facial expressions in this embodiment includes: a human facial expression video acquisition unit 71, a key point detection unit 72, a sparse optical flow field construction unit 73, a sparse optical flow field conversion unit 74, a dense optical flow field generation unit 75, and a target animal expression animation generation unit 76, wherein:
[0095] The facial expression video acquisition unit 71 is used to acquire facial expression videos and animal face images;
[0096] The key point detection unit 72 is used to perform key point detection on human face expression videos and animal face images respectively, and obtain human face key points and animal face key points;
[0097] The sparse optical flow field construction unit 73 is used to construct a sparse optical flow field of a face based on facial key points.
[0098] The sparse optical flow field conversion unit 74 is used to convert the sparse optical flow field of a human face into the sparse optical flow field of an animal through a trained conversion network model.
[0099] The dense optical flow field generation unit 75 is used to convert the animal sparse optical flow field into a dense optical flow field based on the animal face image, animal face key points and animal sparse optical flow field.
[0100] The target animal expression animation generation unit 76 is used to generate target animal expression animation by performing feature transformation and feature summarization on animal face images and dense optical flow fields.
[0101] Furthermore, the sparse optical flow field conversion unit 74 also includes:
[0102] The animal expression video acquisition unit is used to acquire multiple animal expression videos and crop the animal expression videos to obtain the animal face region;
[0103] The initial animal face keypoint generation unit is used to extract keypoints in the animal face region to obtain the initial animal face keypoints.
[0104] The face region generation unit is used to acquire an initial face video with the same actions and expressions as the animal face expression video, and to crop the face from the initial face video to obtain the face region.
[0105] The initial facial landmark generation unit is used to align the face region with the animal face region and extract the key points in the face region to obtain the initial facial landmarks.
[0106] The transformation network model training unit is used to construct training data based on the initial animal facial key points and the initial human facial key points, and to train the transformation network model using the training data to obtain the trained transformation network model.
[0107] Furthermore, the sparse optical flow field building unit 73 includes:
[0108] The initial frame key point acquisition unit is used to acquire the initial frame key points in the face key points based on each key point in the face.
[0109] The optical flow calculation unit for key points is used to calculate the displacement of each face key point relative to the key points in the initial frame based on the alignment method of the same key points, and to obtain the optical flow of the key points;
[0110] A face sparse optical flow field construction unit is used to construct a face sparse optical flow field based on key points.
[0111] Furthermore, the dense optical flow field generation unit 75 includes:
[0112] The downsampling processing unit is used to downsample the animal face image to obtain a downsampled image;
[0113] The key point heatmap generation unit is used to generate key point heatmaps and affine transformation maps based on downsampled images, animal face key points, and animal sparse optical flow fields.
[0114] The animal face feature generation unit is used to concatenate the key point heatmap and the affine transformation map to obtain animal face features;
[0115] The integration processing unit is used to obtain a dense optical flow field by segmenting and integrating animal face features.
[0116] Furthermore, the key point heatmap generation unit includes:
[0117] The target key point construction unit is used to construct target key points based on animal face key points and animal sparse optical flow field;
[0118] The heatmap conversion unit is used to convert target key points into Gaussian distributed heatmaps to obtain key point heatmaps;
[0119] Tensor generation unit is used to generate tensors with the same length and width as the downsampled image according to a preset depth;
[0120] The optical flow addition unit is used to add the tensor to the optical flow corresponding to the target key point to obtain the initial dense optical flow field of the current key point;
[0121] The affine transformation map generation unit is used to perform affine transformation processing on the downsampled image based on the initial dense optical flow field to obtain the affine transformation map.
[0122] Furthermore, the integrated processing unit includes:
[0123] The segmentation feature generation unit is used to segment animal face features using the UNet network model to obtain segmentation features;
[0124] The convolution processing unit is used to perform convolution processing on the segmentation features to obtain convolutional features;
[0125] The accumulation processing unit is used to perform element-wise multiplication of segmentation features and convolution features to obtain the multiplication result, and then accumulate the multiplication result in a depth-accumulation manner to obtain a dense optical flow field.
[0126] Furthermore, the target animal facial expression animation generation unit 76 includes:
[0127] The feature encoding processing unit is used to obtain the feature encoding result by performing feature decomposition and feature encoding processing on the animal face image;
[0128] The feature transformation processing unit is used to perform feature transformation processing on the feature encoding result through a dense optical flow field to obtain the feature transformation result;
[0129] The feature summarization unit is used to generate target animal facial animation by performing feature decoding and feature summarization on the feature transformation results.
[0130] In this embodiment, a video of a human facial expression and an image of an animal face are acquired. Keypoint detection is performed on both the video and the image to obtain facial keypoints and animal keypoints. Based on these keypoints, a sparse optical flow field for the human face is constructed. A trained transformation network model is used to convert the sparse optical flow field for the human face into a sparse optical flow field for the animal face. Based on the animal face image, the keypoints, and the sparse optical flow field, the sparse optical flow field for the animal face is converted into a dense optical flow field. Feature transformation and feature summarization are performed on the animal face image and the dense optical flow field to generate a target animal facial expression animation. This embodiment of the invention can generate a target animal facial expression animation based on a video of a human facial expression and an image of an animal face, thus solving the problem of the limited number of animal facial expression videos available.
[0131] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 9 , Figure 9 This is a basic structural block diagram of the computer device in this embodiment.
[0132] Computer device 8 includes a memory 81, a processor 82, and a network interface 83 that are interconnected via a system bus. It should be noted that only a computer device 8 with three components—memory 81, processor 82, and network interface 83—is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0133] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.
[0134] The memory 81 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 81 may be an internal storage unit of the computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the memory 81 may also be an external storage device of the computer device 8, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 8. Of course, the memory 81 may also include both internal storage units and external storage devices of the computer device 8. In this embodiment, the memory 81 is typically used to store the operating system and various application software installed on the computer device 8, such as program code for a method of generating animal expression animation based on human facial expressions. In addition, the memory 81 can also be used to temporarily store various types of data that have been output or will be output.
[0135] In some embodiments, processor 82 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. This processor 82 is typically used to control the overall operation of the computer device 8. In this embodiment, processor 82 is used to run program code stored in memory 81 or process data, for example, to run the program code of the above-described method for generating animal expression animation based on facial expressions, to implement various embodiments of the method for generating animal expression animation based on facial expressions.
[0136] The network interface 83 may include a wireless network interface or a wired network interface, which is typically used to establish a communication connection between the computer device 8 and other electronic devices.
[0137] This application also provides another embodiment, namely, providing a computer-readable storage medium storing a computer program that can be executed by at least one processor to cause the at least one processor to perform the steps of the above-described method for generating animal facial expression animation based on human facial expressions.
[0138] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0139] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
Claims
1. A method for generating animal facial expression animations based on human facial expressions, characterized in that, include: Acquire videos of human facial expressions and images of animal faces; Key point detection is performed on the human facial expression video and the animal face image respectively to obtain human facial key points and animal facial key points; Based on the aforementioned facial key points, a sparse optical flow field for the face is constructed. Acquire multiple animal expression videos and crop the animal expression videos to obtain the animal face regions; Extract the key points of the animal face region to obtain the initial animal face key points; Acquire an initial human face video with the same actions and expressions as the animal face expression video, and crop the face from the initial human face video to obtain the face region; Align the human face region with the animal face region, and extract the key points in the human face region to obtain the initial human face key points; Training data is constructed based on the initial animal facial key points and the initial human facial key points, and a transformation network model is trained using the training data to obtain a trained transformation network model. The trained conversion network model is used to convert the sparse optical flow field of the human face into a sparse optical flow field of the animal. Based on the animal face image, the animal face key points, and the animal sparse optical flow field, the animal sparse optical flow field is converted into a dense optical flow field; By performing feature transformation and feature summarization on the animal face image and the dense optical flow field, an animated expression of the target animal is generated.
2. The method for generating animal facial expression animation based on human facial expressions according to claim 1, characterized in that, The construction of a sparse optical flow field for the face based on the facial key points includes: Based on each key point in the face, obtain the initial frame key points in the face key points; Based on the alignment method of the same key points, the displacement of each of the facial key points and the key points of the initial frame is calculated to obtain the optical flow of the key points; The sparse optical flow field of the face is constructed based on the optical flow of the key points.
3. The method for generating animal facial expression animation based on human facial expressions according to claim 1, characterized in that, The step of converting the sparse optical flow field of the animal into a dense optical flow field based on the animal face image, the animal face key points, and the animal sparse optical flow field includes: The animal face image is downsampled to obtain a downsampled image; Based on the downsampled image, the key points of the animal face, and the sparse optical flow field of the animal, a key point heatmap and an affine transformation map are generated. By concatenating the key point heatmap and the affine transformation map, animal face features are obtained; The dense optical flow field is obtained by segmenting and integrating the animal face features.
4. The method for generating animal facial animation based on facial expressions according to claim 3, characterized in that, The process of generating a keypoint heatmap and an affine transformation map based on the downsampled image, the animal facial keypoints, and the animal's sparse optical flow field includes: Based on the animal face key points and the animal sparse optical flow field, target key points are constructed; The target key points are converted into a Gaussian distribution heatmap to obtain the key point heatmap; Generate a tensor with the same length and width as the downsampled image according to a preset depth; The tensor is added to the optical flow corresponding to the target key point to obtain the initial dense optical flow field of the current key point; The downsampled image is subjected to affine transformation based on the initial dense optical flow field to obtain the affine transformation image.
5. The method for generating animal facial expression animation based on facial expressions according to claim 3, characterized in that, The process of segmenting and integrating the animal face features to obtain the dense optical flow field includes: The animal face features are segmented using the UNet network model to obtain segmentation features; The segmentation features are convolved to obtain convolutional features; The segmentation features and the convolution features are multiplied element-wise to obtain the multiplication result. The multiplication result is then accumulated in a depthwise manner to obtain the dense optical flow field.
6. The method for generating animal facial expression animation based on human facial expressions according to any one of claims 1 to 5, characterized in that, The step of generating a target animal facial animation by performing feature transformation and feature summarization on the animal face image and the dense optical flow field includes: By performing feature decomposition and feature encoding on the animal face image, the feature encoding result is obtained; The feature encoding result is subjected to feature transformation processing through the dense optical flow field to obtain the feature transformation result; The target animal's facial expression animation is generated by decoding and summarizing the feature transformation results.
7. A device for generating animal expression animation based on human facial expressions, characterized in that, include: The facial expression video acquisition unit is used to acquire facial expression videos and animal face images; A key point detection unit is used to perform key point detection on the human facial expression video and the animal face image respectively, to obtain human facial key points and animal facial key points; A sparse optical flow field construction unit is used to construct a sparse optical flow field of the face based on the facial key points; An animal facial expression video acquisition unit is used to acquire multiple animal facial expression videos and crop the animal facial expression videos to obtain animal face regions; An initial animal face key point generation unit is used to extract key points of the animal face region to obtain initial animal face key points; A face region generation unit is used to acquire an initial face video with the same actions and expressions as the animal face expression video, and to crop the face from the initial face video to obtain a face region. An initial facial key point generation unit is used to align the face region with the animal face region and extract key points in the face region to obtain initial facial key points. The transformation network model training unit is used to construct training data based on the initial animal face key points and the initial human face key points, and to train the transformation network model through the training data to obtain the trained transformation network model. A sparse optical flow field conversion unit is used to convert the sparse optical flow field of the human face into the sparse optical flow field of the animal through the trained conversion network model. A dense optical flow field generation unit is used to convert the animal sparse optical flow field into a dense optical flow field based on the animal face image, the animal face key points, and the animal sparse optical flow field. The target animal facial expression animation generation unit is used to generate target animal facial expression animation by performing feature transformation and feature summarization on the animal face image and the dense optical flow field.
8. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method for generating animal expression animation based on human facial expressions as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method for generating animal facial expression animation based on human facial expressions as described in any one of claims 1 to 6.