Facial Expression Transfer Method, Device, Equipment, Storage Medium and Program Product
The implicit 3D modeling approach for facial expression transfer improves realism and efficiency by extracting 3D features from source and driver images, addressing the limitations of traditional 3D reconstruction methods.
Patent Information
- Application Number
- CN202111388795.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-22
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-11-22
AI Technical Summary
The existing expression migration method based on three-dimensional modeling has the problem of unreal reconstruction, lack of personalized characteristics, and inefficient efficiency.
Implicit 3D modeling is adopted to extract the 3D feature information of the source map and the driver map, generate 3D dense optical flow sets and masks, and perform expression migration, avoiding the reconstruction of the 3D model and directly generating the driver result map.
It improves the effect and efficiency of expression migration, maintains the identity characteristics of the face objects in the source map, and improves the migration accuracy under complex postures and expressions.
Smart Images

Figure CN116152876B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and particularly to an expression transfer method, apparatus, device, storage medium and program product. Background Art
[0002] Expression transfer refers to transferring the pose and expression of the facial object in the driving image to the facial object in the source image to generate a driving result image. The driving result image is an image with the identity information of the facial object in the source image and the pose and expression information of the facial object in the driving image.
[0003] Related technologies provide an expression transfer method based on 3D modeling. First, a 3D (3 Dimension) model of the facial object in the source image is constructed, and then, based on the pose and expression of the facial object in the driving image, the pose and expression of the above 3D model are fitted so that the 3D model has the pose and expression of the facial object in the driving image.
[0004] In this way, since it is necessary to perform 3D reconstruction on the facial object in the source image, on the one hand, the reconstructed 3D model is often not realistic enough and lacks the personalized features of the facial object in the source image, resulting in an unsatisfactory expression transfer effect; on the other hand, since 3D reconstruction takes a long time, the efficiency of expression transfer is low. Summary of the Invention
[0005] Embodiments of the present application provide an expression transfer method, apparatus, device, storage medium and program product. The technical solutions are as follows:
[0006] According to one aspect of the embodiments of the present application, an expression transfer method is provided. The method includes:
[0007] Generating a first set of 3D key points and a second set of 3D key points according to the source image and the driving image; wherein, the first set of 3D key points is a set of key points used to represent the pose and expression information of the facial object in the source image; the second set of 3D key points is a set of key points used to represent the pose and expression information of the facial object in the source image after being driven by the driving image;
[0008] Generating a 3D dense optical flow set according to the first set of 3D key points and the second set of 3D key points; wherein, the 3D dense optical flow set is used to represent the spatial position change of the corresponding key points in the first set of 3D key points and the second set of 3D key points;
[0009] Processing the 3D feature representation of the source image based on the 3D dense optical flow set to generate a 3D optical flow mask and a 2D occlusion mask; wherein, the 3D optical flow mask is used to linearly combine each 3D dense optical flow in the 3D dense optical flow set, and the 2D occlusion mask is used to selectively retain the feature information of the source image;
[0010] Generating a driving result image corresponding to the source image according to the 3D dense optical flow set, the 3D optical flow mask, the 3D feature representation of the source image, and the 2D occlusion mask; wherein, the driving result image is an image with the identity information of the face object in the source image and the pose and expression information of the face object in the driving image.
[0011] According to one aspect of the embodiments of the present application, a method for training an expression transfer model is provided, and the method includes:
[0012] Obtaining training data for the expression transfer model, where the training data includes source image samples, driving image samples, and target driving result images corresponding to the source image samples;
[0013] Generating a first set of 3D key points and a second set of 3D key points through the expression transfer model according to the source image sample and the driving image sample; wherein, the first set of 3D key points is a set of key points used to represent the pose and expression information of the face object in the source image sample; the second set of 3D key points is a set of key points used to represent the pose and expression information of the face object in the source image sample after being driven by the driving image sample;
[0014] Generating a 3D dense optical flow set according to the first set of 3D key points and the second set of 3D key points; wherein, the 3D dense optical flow set is used to represent the spatial position change of the corresponding key points in the first set of 3D key points and the second set of 3D key points;
[0015] Processing the 3D feature representation of the source image sample through the expression transfer model based on the 3D dense optical flow set to generate a 3D optical flow mask and a 2D occlusion mask; wherein, the 3D optical flow mask is used to linearly combine each 3D dense optical flow in the 3D dense optical flow set, and the 2D occlusion mask is used to selectively retain the feature information of the source image sample;
[0016] Generating an output driving result image corresponding to the source image sample through the expression transfer model according to the 3D dense optical flow set, the 3D optical flow mask, the 3D feature representation of the source image sample, and the 2D occlusion mask;
[0017] Calculate the training loss of the expression transfer model according to the output driving result map and the target driving result map, and adjust the parameters of the expression transfer model based on the training loss.
[0018] According to one aspect of the embodiments of the present application, an expression transfer device is provided, and the device includes:
[0019] A key point generation module, configured to generate a first 3D key point set and a second 3D key point set according to a source map and a driving map; wherein, the first 3D key point set is a key point set used to characterize the posture and expression information of the face object in the source map; the second 3D key point set is a key point set used to characterize the posture and expression information of the face object in the source map after being driven by the driving map;
[0020] An optical flow generation module, configured to generate a 3D dense optical flow set according to the first 3D key point set and the second 3D key point set; wherein, the 3D dense optical flow set is used to characterize the spatial position change of the corresponding key points in the first 3D key point set and the second 3D key point set;
[0021] A mask generation module, configured to process the 3D feature representation of the source map based on the 3D dense optical flow set to generate a 3D optical flow mask and a 2D occlusion mask; wherein, the 3D optical flow mask is used to perform a linear combination of each 3D dense optical flow in the 3D dense optical flow set, and the 2D occlusion mask is used to selectively retain the feature information of the source map;
[0022] A result map generation module, configured to generate a driving result map corresponding to the source map according to the 3D dense optical flow set, the 3D optical flow mask, the 3D feature representation of the source map, and the 2D occlusion mask; wherein, the driving result map is an image having the identity information of the face object in the source map and the posture and expression information of the face object in the driving map.
[0023] According to one aspect of the embodiments of the present application, a training device for an expression transfer model is provided, and the device includes:
[0024] A training data acquisition module, configured to acquire training data of the expression transfer model, where the training data includes a source map sample, a driving map sample, and a target driving result map corresponding to the source map sample;
[0025] The key point generation module is used to generate a first set of 3D key points and a second set of 3D key points through the expression migration model according to the source map sample and the driving map sample; wherein, the first set of 3D key points is a set of key points used to represent the pose and expression information of the face object in the source map sample; the second set of 3D key points is a set of key points used to represent the pose and expression information of the face object in the source map sample after being driven by the driving map sample;
[0026] The optical flow generation module is used to generate a set of 3D dense optical flows according to the first set of 3D key points and the second set of 3D key points; wherein, the set of 3D dense optical flows is used to represent the spatial position change of the corresponding key points in the first set of 3D key points and the second set of 3D key points;
[0027] The mask generation module is used to process the 3D feature representation of the source map sample based on the set of 3D dense optical flows through the expression migration model to generate a 3D optical flow mask and a 2D occlusion mask; wherein, the 3D optical flow mask is used to perform a linear combination of each 3D dense optical flow in the set of 3D dense optical flows, and the 2D occlusion mask is used to selectively retain the feature information of the source map sample;
[0028] The result map generation module is used to generate the output driving result map corresponding to the source map sample through the expression migration model according to the set of 3D dense optical flows, the 3D optical flow mask, the 3D feature representation of the source map sample, and the 2D occlusion mask;
[0029] The parameter adjustment module is used to calculate the training loss of the expression migration model according to the output driving result map and the target driving result map, and adjust the parameters of the expression migration model based on the training loss.
[0030] According to one aspect of the embodiments of the present application, a computer device is provided. The computer device includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the above-mentioned expression migration method or the training method of the above-mentioned expression migration model.
[0031] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided. At least one instruction, at least one program, a code set, or an instruction set is stored in the storage medium. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the above-mentioned expression migration method or the training method of the above-mentioned expression migration model.
[0032] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor reads and executes the computer instructions from the computer-readable storage medium of the operating room to implement the above-mentioned expression migration method or the training method of the above-mentioned expression migration model.
[0033] The technical solution provided by the embodiments of the present application can bring the following beneficial effects:
[0034] By adopting the method of implicit 3D modeling, the 3D feature information in the source image and the driving image is extracted. Based on this 3D feature information, the pose and expression features of the face object in the driving image are migrated to the face object in the source image, and the identity features of the face object in the source image are maintained. Finally, a driving result image after expression migration is generated. The entire process does not require reconstructing the 3D model of the face object in the source image, avoiding the inherent defects of poor authenticity and low efficiency existing in reconstructing the 3D model. While improving the expression migration effect, it also improves the efficiency of expression migration. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a schematic diagram of the implementation environment of the solution provided by an embodiment of the present application;
[0036] Figure 2 is a schematic diagram of the expression migration effect provided by an embodiment of the present application;
[0037] Figure 3 is a flowchart of the expression migration method provided by an embodiment of the present application;
[0038] Figure 4 is a flowchart of the expression migration method provided by another embodiment of the present application;
[0039] Figure 5 is a flowchart of the training method of the expression migration model provided by an embodiment of the present application;
[0040] Figure 6 is a block diagram of the expression migration device provided by an embodiment of the present application;
[0041] Figure 7 is a block diagram of the training device of the expression migration model provided by an embodiment of the present application;
[0042] Figure 8 is a block diagram of the structure of the computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0044] AI (Artificial Intelligence) is to simulate, extend, and expand human intelligence using digital computers or machines controlled by digital computers, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0045] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0046] CV (Computer Vision) is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for machine vision such as target recognition, tracking, and measurement, and further performing graphics processing to make the images processed by the computer more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0047] ML (Machine Learning) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0048] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0049] The solution provided in the embodiments of this application relates to the computer vision technology and machine learning technology of artificial intelligence, and provides an expression transfer method. By adopting the method of implicit 3D modeling, the 3D feature information in the source image and the driving image is extracted, and based on this 3D feature information, the pose and expression features of the face object in the driving image are transferred to the face object in the source image, while maintaining the identity features of the face object in the source image. Finally, a driving result image after expression transfer is generated. The entire process does not require reconstructing the 3D model of the face object in the source image, avoiding the inherent defects of poor authenticity and low efficiency in reconstructing the 3D model, and improving the expression transfer effect while also improving the efficiency of expression transfer.
[0050] Please refer to Figure 1 , which shows a schematic diagram of the solution implementation environment provided by an embodiment of this application. The solution implementation environment may include a model training device 10 and a model using device 20.
[0051] The model training device 10 may be an electronic device such as a computer, a server, a smart robot, etc., or some other electronic device with strong computing power. The model training device 10 is used to train the expression transfer model 30. In the embodiments of this application, the expression transfer model 30 is a neural network model for expression transfer, and the model training device 10 can use machine learning methods to train the expression transfer model 30 to make it have better performance.
[0052] The above-mentioned trained expression transfer model 30 can be deployed in the model usage device 20 for use to provide the function of expression transfer. The model usage device 20 can be a terminal device such as a mobile phone, a computer, a smart TV, a multimedia playback device, a wearable device, a vehicle-mounted terminal, a PC (Personal Computer), etc., or it can be a server, and the present application does not make any limitation thereto.
[0053] As Figure 1 shown, the expression transfer model 30 is used to transfer the pose and expression of the face object in the driving map D to the source map S to generate a driving result map Y. The driving result map Y is an image with the identity information of the face object in the source map S and the pose and expression information of the face object in the driving map D.
[0054] Optionally, the driving map D is an image frame in a driving video, and multiple driving result maps corresponding to the source map S are integrated to generate a driving result video corresponding to the source map S; wherein, the multiple driving result maps are generated according to the source map S and multiple image frames in the driving video. That is, each image frame in the driving video is used as the driving map D in sequence to perform expression driving on the source map S to generate multiple driving result maps, and these driving result maps can be arranged in order to serve as the driving result video. In short, it is to collect the pose and expression information of the face object in each frame of the driving video, and transplant it to the source map S according to the algorithm to generate a new dynamic video (that is, the "driving result video" mentioned above).
[0055] Exemplarily, as Figure 2 shown, the source map S is a static image, and the driving video includes multiple image frames such as image frame 21a, image frame 21b, and image frame 21c. Taking the image frame 21a as the driving map, an image 22a can be generated based on the image frame 21a and the source map S. It can be clearly seen from Figure 2 that the image 22a has the identity information of the source map S (that is, the appearance of the face in the image 22a is the same as the appearance of the face in the source map S), and the image 22a has the pose and expression information of the image frame 21a (that is, the pose and expression of the face in the image 22a are the same as the pose and expression of the face in the image frame 21a). Similarly, taking the image frame 21b as the driving map, an image 22b can be generated based on the image frame 21b and the source map S. Taking the image frame 21c as the driving map, an image 22c can be generated based on the image frame 21c and the source map S. Arranging the above images 22a, 22b, 22c, etc. as image frames as needed can generate the driving result video.
[0056] In some possible application scenarios, the user can upload the source image S through the client and select a driving video. The client or the server generates a driving result video based on the source image S and the driving video, and plays the driving result video through the client. The above driving video can be a video selected from a preset material library or a video uploaded by the user himself / herself. This application does not make any limitation in this regard.
[0057] Next, several method embodiments will be used to introduce and illustrate the technical solution of this application.
[0058] Please refer to Figure 3 , which shows a flowchart of an expression migration method provided by an embodiment of this application. This method can be executed by the model using device 20 in the solution implementation environment shown in Figure 1 . This method can include at least one of the following steps (310 to 340):
[0059] Step 310, generate a first 3D key point set and a second 3D key point set according to the source image and the driving image; wherein, the first 3D key point set is a key point set used to characterize the posture and expression information of the face object in the source image; the second 3D key point set is a key point set used to characterize the posture and expression information of the face object in the source image after being driven by the driving image.
[0060] In the embodiment of this application, the source image can be a static image, and the source image contains a face object. For example, the face object can be a human face. Of course, in some other possible embodiments, the face object can also be the face of a cartoon character, the face of an anime character, the face of an animal, or the face of other types. This application does not make any limitation in this regard.
[0061] In the embodiment of this application, the driving image can also be a static image, and the driving image also contains a face object. For example, the face object can also be a human face. Of course, in some other possible embodiments, the face object can also be the face of a cartoon character, the face of an anime character, the face of an animal, or the face of other types. This application does not make any limitation in this regard.
[0062] The types of the face objects in the source image and the driving image can be of the same type. For example, both the face object in the source image and the face object in the driving image are human faces. In addition, when both the face object in the source image and the face object in the driving image are human faces, the face object in the source image and the face object in the driving image can be the face of the same person or the faces of two different people. For example, the face object in the source image can be the human face of user "Zhang San", and the face object in the driving image can be the human face of a certain star "Li Mou".
[0063] To ensure the expression transfer effect, it is best that the types of facial objects in the source image and the driving image are the same. For example, both are human faces. In this way, the number and distribution of key points of the facial objects in the source image and the driving image will be closer, which helps to improve the expression transfer effect. Of course, in some other possible embodiments, the types of facial objects in the source image and the driving image can also be different. For example, the type of the facial object in the source image is a human face, and the type of the facial object in the driving image is the face of a cartoon character or the face of an animal. Another example is that the type of the facial object in the source image is the face of a cartoon character, and the type of the facial object in the driving image is a human face. The technical solution of this application is also applicable to such scenarios.
[0064] In the embodiment of this application, based on the source image and the driving image, a first 3D key point set and a second 3D key point set can be obtained. Among them, the first 3D key point set is a key point set used to represent the pose and expression information of the facial object in the source image. The first 3D key point set may include K first key points, where K is a positive integer. The second 3D key point set is a key point set used to represent the pose and expression information of the facial object in the source image after being driven by the driving image. The second 3D key point set may include K second key points, where K is a positive integer. The number of key points included in the above first 3D key point set and second 3D key point set is the same. For example, both are K. Exemplarily, K is 30, 40, 50, 60, or 70, etc. This application does not limit this.
[0065] In some embodiments, the above step 310 may include the following sub-steps:
[0066] 1. Extract K normalized 3D key points from the source image, where K is a positive integer;
[0067] The so-called normalized 3D key points refer to key points that are independent of the pose and expression of the facial object in the image and are only related to the identity of the facial object in the image, and these key points are three-dimensional key points, that is, the position of each key point is represented by three-dimensional spatial coordinates. For two images, if the facial objects in these two images are the same facial object, such as the face of the same person, then even if the poses and expressions of the facial objects in these two images are different, the positions of the K normalized 3D key points extracted from these two images are basically the same or similar. For two images, if the facial objects in these two images are two different facial objects, such as the faces of two different people, then even if the poses and expressions of the facial objects in these two images are the same, the positions of the K normalized 3D key points extracted from these two images are different or have a large difference.
[0068] Optionally, K normalized 3D key points are extracted from the source image through a 3D key point extraction network, which can be a neural network, such as a deep convolutional neural network. In the embodiments of the present application, the specific network structure of the 3D key point extraction network is not limited.
[0069] 2. Obtain the image information corresponding to the source image and the driving image respectively. The image information includes the pose angle, translation amount, and offset amount of K 3D key points of the face object.
[0070] The image information corresponding to the source image is used to represent the pose and expression information of the face object in the source image. The image information corresponding to the source image may include the pose angle, translation amount, and offset amount of K 3D key points of the face object in the source image. The image information corresponding to the driving image is used to represent the pose and expression information of the face object in the driving image. The image information corresponding to the driving image may include the pose angle, translation amount, and offset amount of K 3D key points of the face object in the driving image.
[0071] The above pose angle can also be called Euler angle, including pitch angle, yaw angle, and roll angle, which describe the pose of the face object. The translation amount describes the spatial position of the face object, which refers to the translation amount of the position of the face object in the image compared to a certain set position. The translation amount is also a three-dimensional quantity, which includes the translation amounts in 3 spatial dimensions. The offset amount of each 3D key point describes the spatial position of the 3D key point, which refers to the offset amount of the position of the 3D key point on the face object in the image compared to its normalized position. The offset amount is also a three-dimensional quantity, which includes the offset amounts in 3 spatial dimensions.
[0072] Through the image information corresponding to the source image, the pose and expression of the face object in the source image can be described; similarly, through the image information corresponding to the driving image, the pose and expression of the face object in the driving image can be described.
[0073] 2-1. Extract the feature information of the source image and the driving image respectively.
[0074] 2-2. Based on the feature information of the source image, predict the image information corresponding to the source image.
[0075] 2-3. Based on the feature information of the driving image, predict the image information corresponding to the driving image.
[0076] Optionally, the image information corresponding to the source image is extracted from the source image and the image information corresponding to the driving image is extracted from the driving image through an image information extraction network.
[0077] Optionally, the image information extraction network may include a feature extraction network and three branch prediction networks, denoted as the first branch prediction network, the second branch prediction network, and the third branch prediction network. Among them, the feature extraction network is used to extract the feature information of the image, and the first branch prediction network, the second branch prediction network, and the third branch prediction network are used to extract the pose angle, translation amount, and offset amount of K 3D key points of the face object, respectively.
[0078] For example, input the source image into the feature extraction network, output the feature information of the source image through the feature extraction network, and then input the feature information of the source image into the first branch prediction network, the second branch prediction network, and the third branch prediction network respectively. Output the pose angle of the face object in the source image through the first branch prediction network, output the translation amount of the face object in the source image through the second branch prediction network, and output the offset amount of K 3D key points of the face object in the source image through the third branch prediction network.
[0079] Similarly, input the driving image into the feature extraction network, output the feature information of the driving image through the feature extraction network, and then input the feature information of the driving image into the first branch prediction network, the second branch prediction network, and the third branch prediction network respectively. Output the pose angle of the face object in the driving image through the first branch prediction network, output the translation amount of the face object in the driving image through the second branch prediction network, and output the offset amount of K 3D key points of the face object in the driving image through the third branch prediction network.
[0080] The above image information extraction network may be a neural network, such as a deep neural network. In the embodiments of the present application, the specific network structure of the image information extraction network is not limited.
[0081] 3. Perform 3D transformation processing on the K normalized 3D key points using the image information corresponding to the source image to generate a first set of 3D key points;
[0082] After obtaining the pose angle, translation amount, and offset amount of K 3D key points of the face object in the source image, these data can be used to perform 3D transformation processing on the K normalized 3D key points. The 3D transformation processing includes rotation, translation, and deformation, and generates a first set of 3D key points for characterizing the pose and expression information of the face object in the source image.
[0083] Optionally, the k-th key point x in the first set of 3D key points s,k is as follows:
[0084] x s,k = T(x c,k , R s , t s , δ s,k ) = R s x c,k + ts +δ s,k
[0085] where x c,k represents the k-th normalized 3D key point in the source image, R s represents the pose angle of the face object in the source image, t s represents the translation amount of the face object in the source image, δ s,k represents the offset amount of the k-th 3D key point of the face object in the source image, T represents 3D transformation processing, and k is a positive integer. For example, k is a positive integer within the range of [1, K].
[0086] 4. Perform 3D transformation processing on the K normalized 3D key points using the image information corresponding to the driving image to generate a second set of 3D key points.
[0087] After obtaining the pose angle, translation amount, and offset amount of the K 3D key points of the face object in the driving image, these data can be used to perform 3D transformation processing on the K normalized 3D key points. This 3D transformation processing includes rotation, translation, and deformation, and generates a second set of 3D key points for characterizing the pose and expression information of the face object in the source image after being driven by the driving image.
[0088] Optionally, the k-th key point x in the second set of 3D key points d,k is as follows:
[0089] x d,k = T(x c,k , R d , t d , δ d,k) = R d x c,k + t d + δ d,k
[0090] where x c,k represents the k-th normalized 3D key point in the source image, R d represents the pose angle of the face object in the driving image, t d represents the translation amount of the face object in the driving image, δ d,k represents the offset amount of the k-th 3D key point of the face object in the driving image, T represents 3D transformation processing, and k is a positive integer. For example, k is a positive integer within the range of [1, K].
[0091] Step 320, generate a 3D dense optical flow set according to the first set of 3D key points and the second set of 3D key points; wherein, the 3D dense optical flow set is used to characterize the spatial position change of the corresponding key points in the first set of 3D key points and the second set of 3D key points.
[0092] As introduced above, the first 3D key-point set includes K first key points, and the second 3D key-point set includes K second key points. Two key points used to mark the same facial position in the first 3D key-point set and the second 3D key-point set are denoted as a group of corresponding key points. For example, one key point in the first 3D key-point set used to mark the tip of the nose and one key point in the second 3D key-point set used to mark the tip of the nose form a group of corresponding key points. Another example is that one key point in the first 3D key-point set used to mark the right eye corner and one key point in the second 3D key-point set used to mark the right eye corner form another group of corresponding key points. In this way, the key points in the first 3D key-point set and the second 3D key-point set are combined in pairs, and K groups of corresponding key points can be formed.
[0093] In some embodiments, step 320 above includes the following sub-steps:
[0094] 1. For each group of corresponding key points in the first 3D key-point set and the second 3D key-point set, determine the 3D dense optical flow corresponding to the corresponding key points according to the position information of the two 3D key points in the corresponding key points;
[0095] 2. Obtain a 3D dense optical flow set according to the 3D dense optical flows respectively corresponding to each group of corresponding key points.
[0096] For each group of corresponding key points in the above K groups of corresponding key points, determine the 3D dense optical flow corresponding to the corresponding key points according to the position information of the two 3D key points in the corresponding key points; then, synthesize the 3D dense optical flows respectively corresponding to the K groups of corresponding key points to obtain a 3D dense optical flow set.
[0097] Optionally, according to the K groups of corresponding key points (x s,k , x d,k ), based on the approximation method of first-order motion, K 3D dense optical flows can be obtained. Among them, the kth 3D dense optical flow w k is expressed as follows:
[0098]
[0099] Among them, p d is the 3D coordinate of the driving graph feature space, p s is the 3D coordinate of the source graph feature space, and R s and R d are as explained above.
[0100] Step 330: Process the 3D feature representation of the source image based on the 3D dense optical flow set to generate a 3D optical flow mask and a 2D occlusion mask. The 3D optical flow mask is used to linearly combine each 3D dense optical flow in the 3D dense optical flow set, and the 2D occlusion mask is used to selectively retain the feature information of the source image.
[0101] The 3D feature representation of the source image is used to represent the three-dimensional feature information of the face object in the source image. Optionally, the 3D feature representation of the source image is extracted through a 3D feature extraction network. For example, the source image is input into the 3D feature extraction network, and the 3D feature representation of the source image is output through the 3D feature extraction network. In some embodiments, the 3D feature extraction network includes a 2D convolutional network and a 3D convolutional network. First, the source image is input into the 2D convolutional network, and the source image is mapped to a 2D feature space through the 2D convolutional network to obtain the 2D feature representation of the source image. Then, the 2D feature representation of the source image is dimension-changed (such as through a Reshape operation) to the 3D feature space to obtain the initial 3D feature representation of the source image. Next, the initial 3D feature representation of the source image is input into the 3D convolutional network, and the 3D feature representation of the source image is output through the 3D convolutional network. Compared with the 2D feature representation, the 3D feature representation has an additional dimension of depth information.
[0102] After obtaining the 3D feature representation of the source image, the 3D feature representation of the source image can be processed based on the 3D dense optical flow set obtained in step 320 to generate a 3D optical flow mask and a 2D occlusion mask. The 3D optical flow mask can be regarded as a three-dimensional matrix, and the value of each element in the matrix is between [0, 1]. The 2D occlusion mask can be regarded as a two-dimensional matrix, and the value of each element in the matrix is also between [0, 1].
[0103] In some embodiments, step 330 above may include the following sub-steps:
[0104] 1. Based on each 3D dense optical flow in the 3D dense optical flow set, perform transformation processing on the 3D feature representation of the source image respectively to obtain multiple transformed 3D feature representations;
[0105] 2. Concatenate the multiple transformed 3D feature representations to obtain a concatenated 3D feature representation;
[0106] 3. Process the concatenated 3D feature representation to generate a 3D optical flow mask and a 2D occlusion mask.
[0107] The 3D dense optical flow set includes K 3D dense optical flows. For each of the K 3D dense optical flows, the 3D feature representation of the source image is transformed (such as warped) using the 3D dense optical flow to obtain a transformed 3D feature representation. By separately transforming the 3D feature representation of the source image with the K 3D dense optical flows, K transformed 3D feature representations can be obtained. Then, the K transformed 3D feature representations are concatenated to obtain a concatenated 3D feature representation. Subsequently, the concatenated 3D feature representation is used as the input to the motion field estimation network, and the network processes it to output a 3D optical flow mask and a 2D occlusion mask. The motion field estimation network can also be a neural network, and the present application does not limit its network structure either.
[0108] Step 340: Generate a driving result image corresponding to the source image according to the 3D dense optical flow set, the 3D optical flow mask, the 3D feature representation of the source image, and the 2D occlusion mask; wherein, the driving result image is an image with the identity information of the face object in the source image, as well as the pose and expression information of the face object in the driving image.
[0109] In some embodiments, the above step 340 may include the following sub-steps:
[0110] 1. Linearly combine each 3D dense optical flow in the 3D dense optical flow set based on the 3D optical flow mask to obtain a combined 3D dense optical flow;
[0111] Use the 3D optical flow mask obtained in step 330 to linearly combine the K 3D dense optical flows to obtain a combined 3D dense optical flow.
[0112] Optionally, the combined 3D dense optical flow w is expressed as follows:
[0113]
[0114] wherein, w k (p d ) represents the k-th 3D dense optical flow, and m k (p d ) represents the mask corresponding to the k-th 3D dense optical flow obtained from the 3D optical flow mask.
[0115] 2. Based on the combined 3D dense optical flow, transform the 3D feature representation of the source image to obtain a transformed 3D feature representation;
[0116] Use the combined 3D dense optical flow to transform the 3D feature representation of the source image (such as warping) to obtain a transformed 3D feature representation.
[0117] 3. Perform dimensionality reduction on the transformed 3D feature representation to obtain the transformed 2D feature representation;
[0118] Perform dimensionality reduction on the transformed 3D feature representation through a dimensionality change operation (such as a Reshape operation), reducing the dimension from the 3D feature space to the 2D feature space to obtain the transformed 2D feature representation.
[0119] 4. Generate a driving result map corresponding to the source map based on the 2D occlusion mask and the transformed 2D feature representation.
[0120] Optionally, process the transformed 2D feature representation using the 2D occlusion mask to obtain the processed 2D feature representation. For example, calculate the Hadamard product of the 2D occlusion mask and the transformed 2D feature representation to obtain the processed 2D feature representation. Then, decode the processed 2D feature representation to generate a driving result map corresponding to the source map. For example, input the processed 2D feature representation into a decoding network, and output a driving result map corresponding to the source map through this decoding network. The decoding network can also be a neural network, and the present application does not limit its network structure either.
[0121] In the embodiments of the present application, by performing transformation processing on the 3D feature representation of the source map in the 3D space based on the combined 3D dense optical flow, compared with only processing in the 2D space, one-dimensional depth information is considered more, making the migration transformation of the pose and expression more accurate, thereby improving the expression migration effect of the finally generated driving result map.
[0122] In some embodiments, the driving map is an image frame in a driving video. Use multiple image frames in the driving video as driving maps respectively, and generate corresponding driving result maps with the source map. Integrate the multiple driving result maps corresponding to the source map to generate a driving result video corresponding to the source map. The identity information of the face object in the driving result video is consistent with the identity information of the face object in the source map, and the action and expression information of the face object in the driving result video is consistent with the action and expression information of the face object in the driving video. This effect can be as Figure 2 shown. Drive the source map through the driving video to generate a driving result video in which the actions and expressions of a person change with the driving video, improving the interestingness of expression migration in product applications.
[0123] The technical solution provided by the embodiments of this application adopts the method of implicit 3D modeling. By extracting the 3D feature information in the source image and the driving image, based on this 3D feature information, the pose and expression features of the face object in the driving image are transferred to the face object in the source image, and the identity features of the face object in the source image are maintained. Finally, a driving result image after expression transfer is generated. The whole process does not require reconstructing the 3D model of the face object in the source image, avoiding the inherent defects of poor authenticity and low efficiency existing in reconstructing the 3D model. While improving the expression transfer effect, it also improves the efficiency of expression transfer.
[0124] In addition, through the expression transfer model, implicit 3D modeling of the source image and the driving image is achieved, that is, the 3D feature information of the image is implicitly learned in the neural network space. Compared with some expression transfer schemes based on 2D feature information, this application takes more depth information into consideration, which can improve the accuracy and precision of the model's characterization of the feature information of the face object, and improve the generation quality of the driving result image.
[0125] In addition, by extracting multiple normalized 3D key points from the source image, the image information corresponding to the source image and the driving image is obtained. This image information includes the pose angle, translation amount of the face object, and the offset amount of multiple 3D key points. Then, using the image information corresponding to the source image and the driving image respectively, 3D transformation processing is performed on the above-mentioned multiple normalized 3D key points respectively, obtaining a key point set for characterizing the pose and expression information of the face object in the source image, and a key point set for characterizing the pose and expression information of the face object in the source image after being driven by the driving image, realizing the decoupled characterization of the identity information, spatial pose, spatial position, and key point position of the face object in the image. After such decoupling, the model can be more clear about each control quantity. For an image with a complex pose or expression, each different control quantity can be better split, realizing flexible and precise control (for example, if the face object in the driving image only makes a blinking expression, then only the offset amount of some key points in the eyes needs to be changed, and the offset amounts of other key points do not need to be changed, and the pose angle and translation amount of the whole face do not need to be changed). In this way, even in the case of large deflection or distortion of the face object in the image, a high-quality driving result image can be generated, improving the expression transfer effect in complex and large-scale motion scenarios.
[0126] In some embodiments, the above expression transfer method is implemented through an expression transfer model. Exemplarily, such as Figure 4As shown, the expression transfer model may include: a 3D key point extraction network 41, an image information extraction network 42, a 3D feature extraction network 43, a motion field estimation network 44, and a decoding network 45. Among them, the 3D key point extraction network 41 is used to extract K normalized 3D key points from the source image S. The image information extraction network 42 is used to obtain the image information corresponding to the source image S and the driving image D respectively. Among them, the image information corresponding to the source image S is used to perform 3D transformation processing on the K normalized 3D key points to generate a first set of 3D key points. The image information corresponding to the driving image D is used to perform 3D transformation processing on the K normalized 3D key points to generate a second set of 3D key points. The 3D feature extraction network 43 is used to extract the 3D feature representation of the source image S. The motion field estimation network 44 is used to generate a 3D optical flow mask and a 2D occlusion mask. The decoding network 45 is used to generate a driving result image Y corresponding to the source image S.
[0127] See Figure 4, the source image S is input into the 3D key point extraction network 41, and K normalized 3D key points of the source image S are output through the 3D key point extraction network 41. The source image S is input into the image information extraction network 42, and the image information corresponding to the source image S is extracted through the image information extraction network 42, including the pose angle, translation amount of the face object in the source image S, and the offset amounts of the K 3D key points. The driving image D is input into the image information extraction network 42, and the image information corresponding to the driving image D is extracted through the image information extraction network 42, including the pose angle, translation amount of the face object in the driving image D, and the offset amounts of the K 3D key points. Then, the above K normalized 3D key points are subjected to 3D transformation processing using the image information corresponding to the source image S to generate a first set of 3D key points, which is used to represent the pose and expression information of the face object in the source image S; the above K normalized 3D key points are subjected to 3D transformation processing using the image information corresponding to the driving image D to generate a second set of 3D key points, which is used to represent the pose and expression information of the face object in the source image S after being driven by the driving image D. Then, according to the first set of 3D key points and the second set of 3D key points, a 3D dense optical flow set is generated, and the 3D dense optical flow set includes 3D dense optical flows corresponding to K groups of corresponding key points respectively. In addition, the source image S passes through the 3D feature extraction network 43 to output the 3D feature representation of the source image S. Based on the K 3D dense optical flows in the 3D dense optical flow set, the 3D feature representation of the source image S is respectively subjected to transformation processing to obtain K transformed 3D feature representations. The K transformed 3D feature representations are concatenated to obtain a concatenated 3D feature representation, and the concatenated 3D feature representation is processed to generate a 3D optical flow mask and a 2D occlusion mask. Among them, the 3D optical flow mask is used to perform a linear combination of each 3D dense optical flow in the 3D dense optical flow set, and the 2D occlusion mask is used to selectively retain the feature information of the source image S. Then, based on the 3D optical flow mask, a linear combination of the K 3D dense optical flows in the 3D dense optical flow set is performed to obtain a combined 3D dense optical flow. Based on the combined 3D dense optical flow, the 3D feature representation of the source image S is subjected to transformation processing to obtain a transformed 3D feature representation. The transformed 3D feature representation is subjected to dimensionality reduction processing to obtain a transformed 2D feature representation. The Hadamard product of the 2D occlusion mask and the transformed 2D feature representation is calculated to obtain a processed 2D feature representation. Finally, the processed 2D feature representation is input into the decoding network 45, and the driving result image Y corresponding to the source image S is output through the decoding network 45.
[0128] The above introduced the usage process of the expression transfer model. Next, the training process of the expression transfer model will be described through embodiments. It should be noted that the content involved in the usage process of this model and the content involved in the training process correspond to each other and are interconnected. For places not described in detail on one side, the description on the other side can be referred to.
[0129] Please refer to Figure 5 , which shows a flowchart of a method for training an expression transfer model provided by an embodiment of the present application. This method can be executed by the model training device 10 in the implementation environment shown by Figure 1 . This method may include at least one of the following steps (510-560):
[0130] Step 510, obtain training data for the expression transfer model, where the training data includes source graph samples, driving graph samples, and target driving result graphs corresponding to the source graph samples.
[0131] During the model training process, the source graph samples and the driving graph samples are used as the source graph and the driving graph respectively. The target driving result graph corresponding to the source graph sample is used as a label to evaluate the quality of the output driving result graph corresponding to the source graph sample generated by the expression transfer model.
[0132] In some embodiments, step 510 may include the following sub-steps:
[0133] 1. Obtain video samples for generating training data;
[0134] 2. Extract a first image frame and a second image frame from the video sample, where the second image frame is the next image frame of the first image frame;
[0135] 3. Determine the first image frame as the source graph sample;
[0136] 4. Determine the second image frame as the driving graph sample and the target driving result graph corresponding to the source graph sample.
[0137] The video sample may be a video containing a specific facial object, such as a video containing the face of a specific person. From the video sample, multiple groups of image frame pairs are obtained, and each group of image frame pairs may include two image frames of the vector. That is, each group of image frame pairs includes a first image frame and a second image frame, and the second image frame is the next image frame of the first image frame. In the present application, the first image frame is used as the source graph sample, the second image frame is used as the driving graph sample, and at the same time, the second image frame is used as the target driving result graph corresponding to the source graph sample, so that the training data required for model training can be constructed simply and efficiently.
[0138] Of course, image frames can be selected from multiple different video samples to construct training data, thereby enhancing the richness of the training data. For example, the training data can include the faces of different people, which helps to improve the robustness of the finally trained expression transfer model so that it can process source images and driving images of different people.
[0139] Step 520: Generate a first set of 3D key points and a second set of 3D key points according to the source image sample and the driving image sample through the expression transfer model; wherein, the first set of 3D key points is a set of key points used to represent the pose and expression information of the facial object in the source image sample; the second set of 3D key points is a set of key points used to represent the pose and expression information of the facial object in the source image sample after being driven by the driving image sample.
[0140] In some embodiments, the above step 520 may include the following sub-steps:
[0141] 1. Extract K normalized 3D key points from the source image sample, where K is a positive integer;
[0142] Optionally, extract K normalized 3D key points from the source image sample through a 3D key point extraction network.
[0143] 2. Obtain the image information corresponding to the source image sample and the driving image sample respectively, where the image information includes the pose angle, translation amount of the facial object, and the offset amount of K 3D key points;
[0144] Optionally, extract the feature information of the source image sample and the driving image sample respectively; based on the feature information of the source image sample, predict the image information corresponding to the source image sample; based on the feature information of the driving image sample, predict the image information corresponding to the driving image sample.
[0145] Optionally, extract the image information corresponding to the source image sample from the source image sample and the image information corresponding to the driving image sample from the driving image sample through an image information extraction network.
[0146] 3. Perform 3D transformation processing on the K normalized 3D key points by using the image information corresponding to the source image sample to generate a first set of 3D key points;
[0147] 4. Perform 3D transformation processing on the K normalized 3D key points by using the image information corresponding to the driving image sample to generate a second set of 3D key points.
[0148] Step 530: Generate a 3D dense optical flow set according to the first set of 3D key points and the second set of 3D key points; wherein, the 3D dense optical flow set is used to represent the spatial position change of the corresponding key points in the first set of 3D key points and the second set of 3D key points.
[0149] In some embodiments, step 530 may include the following sub-steps:
[0150] 1. For each pair of corresponding key points in the first 3D key point set and the second 3D key point set, determine the 3D dense optical flow corresponding to the corresponding key point according to the position information of the two 3D key points in the corresponding key point;
[0151] 2. Obtain a 3D dense optical flow set according to the 3D dense optical flows respectively corresponding to each group of corresponding key points.
[0152] Step 540, process the 3D feature representation of the source graph sample based on the 3D dense optical flow set through an expression transfer model to generate a 3D optical flow mask and a 2D occlusion mask; wherein, the 3D optical flow mask is used to perform a linear combination of the individual 3D dense optical flows in the 3D dense optical flow set, and the 2D occlusion mask is used to selectively retain the feature information of the source graph sample.
[0153] Optionally, extract the 3D feature representation of the source graph sample through a 3D feature extraction network.
[0154] In some embodiments, step 540 may include the following sub-steps:
[0155] 1. Based on each 3D dense optical flow in the 3D dense optical flow set, perform transformation processing on the 3D feature representation of the source graph sample respectively to obtain multiple transformed 3D feature representations;
[0156] 2. Concatenate the multiple transformed 3D feature representations to obtain a concatenated 3D feature representation;
[0157] 3. Process the concatenated 3D feature representation to generate a 3D optical flow mask and a 2D occlusion mask.
[0158] Optionally, use the concatenated 3D feature representation as the input of a motion field estimation network, and process it through the motion field estimation network to output a 3D optical flow mask and a 2D occlusion mask.
[0159] Step 550, generate an output driving result graph corresponding to the source graph sample through an expression transfer model according to the 3D dense optical flow set, the 3D optical flow mask, the 3D feature representation of the source graph sample, and the 2D occlusion mask.
[0160] In some embodiments, step 550 may include the following sub-steps:
[0161] 1. Based on the 3D optical flow mask, perform a linear combination of the individual 3D dense optical flows in the 3D dense optical flow set to obtain a combined 3D dense optical flow;
[0162] 2. Transform the 3D feature representation of the source image sample based on the combined 3D dense optical flow to obtain the transformed 3D feature representation;
[0163] 3. Perform dimensionality reduction on the transformed 3D feature representation to obtain the transformed 2D feature representation;
[0164] 4. Generate the output driving result map corresponding to the source image sample according to the 2D occlusion mask and the transformed 2D feature representation.
[0165] Optionally, process the transformed 2D feature representation with the 2D occlusion mask to obtain the processed 2D feature representation. Then, input the processed 2D feature representation into the decoding network, and output the output driving result map corresponding to the source image sample through this decoding network.
[0166] Step 560, calculate the training loss of the expression transfer model according to the output driving result map and the target driving result map, and adjust the parameters of the expression transfer model based on the training loss.
[0167] The training loss of the expression transfer model is used to measure the transfer effect of the expression transfer model. For example, when constructing the loss function of this expression transfer model, the proximity between the output driving result map and the target driving result can be considered. During the model training process, aiming at the convergence of the value of the loss function of this expression transfer model (i.e., the training loss), continuously adjust the parameters of each network in the expression transfer model to make the expression transfer model have a better transfer effect.
[0168] In some embodiments, calculate the training loss of the expression transfer model through the following steps:
[0169] 1. Calculate the visual perception loss according to the image features of the output driving result map and the image features of the target driving result map. The visual perception loss is used to measure the difference degree between the image features of the output driving result map and the image features of the target driving result map;
[0170] Optionally, the image features of the driving result map and the image features of the target driving result map can be extracted through an image information extraction network. For example, the image features can be the feature information extracted by the feature extraction network in the image information extraction network.
[0171] 2. Calculate the generative adversarial loss according to the output driving result map and the target driving result map. The generative adversarial loss is used to measure the difference degree between the output driving result map and the target driving result map;
[0172] Optionally, the loss function of the generative adversarial loss can adopt the hinge loss function to improve the authenticity of the output result.
[0173] 3. Calculate the geometric equivariance loss according to the position information of each second key point in the second 3D key point set. The geometric equivariance loss is used to measure the geometric equivariance of the second key points during the thin plate spline interpolation transformation;
[0174] Optionally, the geometric equivariance loss L E = ||x d - T -1 (x T(d) )||1, where x d represents the position information of the second key points, T represents the thin plate spline interpolation transformation, and x T(d) is the result after performing the thin plate spline interpolation transformation on x d , and || ||1 represents the L1 distance.
[0175] Of course, in some other possible embodiments, the geometric equivariance loss can also be calculated according to the position information of each key point in the first 3D key point set and / or the second 3D key point set. The geometric equivariance loss is used to measure the geometric equivariance of the first key points and / or the second key points during the thin plate spline interpolation transformation.
[0176] 4. Calculate the 3D key point prior loss according to the position information of each second key point in the second 3D key point set. The 3D key point prior loss is used to measure the dispersion and depth of each second key point;
[0177] Optionally, the 3D key point prior loss where xd,i and xd,j represent any two second key points in the second 3D key point set, D t is the first threshold, || ||2 represents the L2 distance, Z(x d ) represents the depth value of the second key point, and z t is the second threshold. The values of the above first threshold and second threshold can be preset in advance.
[0178] Of course, in some other possible embodiments, the 3D key point prior loss can also be calculated according to the position information of each key point in the first 3D key point set and / or the second 3D key point set. The 3D key point prior loss is used to measure the dispersion and depth of each first key point and / or the dispersion and depth of each second key point.
[0179] 5. Calculate the face pose loss according to the pose angle of the face object in the driving image. The face pose loss is used to measure the accuracy of face pose estimation;
[0180] Optionally, the face pose loss where R d is the pose angle of the face object in the driving image, is the pose angle of the face object in the driving map obtained by using a pre-trained face pose prediction model, and || ||1 represents the L1 distance.
[0181] Of course, in some other possible embodiments, the face pose loss can also be calculated according to the pose angles of the face object in the source map and / or the driving map.
[0182] 6. Calculate the bias prior loss according to the offset amount of the 3D key points of the face object in the driving map. The bias prior loss is used to measure the deviation amount between the position information of the 3D key points of the face object in the driving map and the reference position information.
[0183] Optionally, the bias prior loss L Δ = ||δ d,k ||1, where δ d,k represents the offset amount of the k-th 3D key point of the face object in the driving map, and || ||1 represents the L1 distance.
[0184] Of course, in some other possible embodiments, the bias prior loss can also be calculated according to the offset amounts of the 3D key points of the face object in the source map and / or the driving map.
[0185] 7. Calculate the training loss of the expression transfer model according to the visual perception loss, the adversarial loss, the geometric equivariance loss, the 3D key point prior loss, the face pose loss, and the bias prior loss.
[0186] Optionally, the calculation formula of the training loss L of the expression transfer model is as follows:
[0187]
[0188] where L P (d,y) is the visual perception loss, L G (d,y) is the generative adversarial loss, L E ({x d,k}) is the geometric equivariance loss, L L ({x d,k}) is the 3D key point prior loss, is the face pose loss, L Δ ({δ d,k}) is the bias prior loss, and λ P , λ G , λ E , λ L , λ H and λ Δ are the weight values corresponding to the above respective losses, and the weight values can be preset.
[0189] In summary, the technical solution provided by the embodiments of the present application trains an expression transfer model. Using this expression transfer model in an implicit 3D modeling manner, by extracting 3D feature information from the source image and the driving image, based on this 3D feature information, the pose and expression features of the face object in the driving image are transferred to the face object in the source image, while maintaining the identity features of the face object in the source image. Finally, a driving result image after expression transfer is generated. The entire process does not require reconstructing the 3D model of the face object in the source image, avoiding the inherent defects of poor authenticity and low efficiency in reconstructing the 3D model. While improving the expression transfer effect, it also improves the efficiency of expression transfer.
[0190] Moreover, when calculating the training loss of the model, the present application considers various losses such as visual perception loss, adversarial loss, geometric equivariance loss, 3D key point prior loss, face pose loss, and bias prior loss, fully ensuring the training effects of each network included in the model.
[0191] In addition, for the details not described in detail in the embodiments of the model training method, reference can be made to the introduction in the above embodiments of the expression transfer method. Some identical or similar content will not be repeated.
[0192] The following is an embodiment of the device of the present application, which can be used to execute the method embodiments of the present application. For the details not disclosed in the embodiments of the device of the present application, please refer to the method embodiments of the present application.
[0193] Please refer to Figure 6 , which shows a block diagram of an expression transfer device provided by an embodiment of the present application. This device has the function of implementing the above expression transfer method. This function can be implemented by hardware or by hardware executing corresponding software. This device can be a model usage device or can be set in a model usage device. The device 600 may include: a key point generation module 610, an optical flow generation module 620, a mask generation module 630, and a result image generation module 640.
[0194] The key point generation module 610 is configured to generate a first set of 3D key points and a second set of 3D key points according to the source image and the driving image; wherein, the first set of 3D key points is a set of key points used to represent the pose and expression information of the face object in the source image; the second set of 3D key points is a set of key points used to represent the pose and expression information of the face object in the source image after being driven by the driving image.
[0195] The optical flow generation module 620 is configured to generate a 3D dense optical flow set according to the first set of 3D key points and the second set of 3D key points; wherein, the 3D dense optical flow set is used to represent the spatial position change of the corresponding key points in the first set of 3D key points and the second set of 3D key points.
[0196] A mask generation module 630 is configured to process the 3D feature representation of the source image based on the 3D dense optical flow set to generate a 3D optical flow mask and a 2D occlusion mask; wherein, the 3D optical flow mask is used to linearly combine each 3D dense optical flow in the 3D dense optical flow set, and the 2D occlusion mask is used to selectively retain the feature information of the source image.
[0197] A result image generation module 640 is configured to generate a driving result image corresponding to the source image according to the 3D dense optical flow set, the 3D optical flow mask, the 3D feature representation of the source image, and the 2D occlusion mask; wherein, the driving result image is an image with the identity information of the face object in the source image and the pose and expression information of the face object in the driving image.
[0198] In some embodiments, the result image generation module 640 is configured to:
[0199] Linearly combine each 3D dense optical flow in the 3D dense optical flow set based on the 3D optical flow mask to obtain a combined 3D dense optical flow;
[0200] Perform a transformation process on the 3D feature representation of the source image based on the combined 3D dense optical flow to obtain a transformed 3D feature representation;
[0201] Perform a dimensionality reduction process on the transformed 3D feature representation to obtain a transformed 2D feature representation;
[0202] Generate a driving result image corresponding to the source image according to the 2D occlusion mask and the transformed 2D feature representation.
[0203] In some embodiments, when the result image generation module 640 generates a driving result image corresponding to the source image according to the 2D occlusion mask and the transformed 2D feature representation, it is specifically configured to:
[0204] Process the transformed 2D feature representation with the 2D occlusion mask to obtain a processed 2D feature representation;
[0205] Decode the processed 2D feature representation to generate a driving result image corresponding to the source image.
[0206] In some embodiments, the mask generation module 630 is configured to:
[0207] Perform a transformation process on the 3D feature representation of the source image respectively based on each 3D dense optical flow in the 3D dense optical flow set to obtain a plurality of transformed 3D feature representations;
[0208] Stitch the multiple transformed 3D feature representations to obtain a stitched 3D feature representation;
[0209] Process the stitched 3D feature representation to generate the 3D optical flow mask and the 2D occlusion mask.
[0210] In some embodiments, the key point generation module 610 is configured to:
[0211] Extract K normalized 3D key points from the source image, where K is a positive integer;
[0212] Obtain the image information corresponding to the source image and the driving image respectively, where the image information includes the pose angle, translation amount of the face object, and the offset amount of K 3D key points;
[0213] Perform 3D transformation processing on the K normalized 3D key points using the image information corresponding to the source image to generate the first set of 3D key points;
[0214] Perform 3D transformation processing on the K normalized 3D key points using the image information corresponding to the driving image to generate the second set of 3D key points.
[0215] In some embodiments, when the key point generation module 610 obtains the image information corresponding to the source image and the driving image respectively, it is specifically configured to:
[0216] Extract the feature information of the source image and the driving image respectively;
[0217] Based on the feature information of the source image, predict the image information corresponding to the source image;
[0218] Based on the feature information of the driving image, predict the image information corresponding to the driving image.
[0219] In some embodiments, the optical flow generation module 620 is configured to:
[0220] For each pair of corresponding key points in the first set of 3D key points and the second set of 3D key points, determine the 3D dense optical flow corresponding to the corresponding key points according to the position information of the two 3D key points in the corresponding key points;
[0221] Obtain the 3D dense optical flow set according to the 3D dense optical flows corresponding to each pair of corresponding key points respectively.
[0222] In some embodiments, the method is implemented through an expression transfer model, and the expression transfer model includes: a 3D key point extraction network, an image information extraction network, a 3D feature extraction network, a motion field estimation network, and a decoding network; where
[0223] The 3D key point extraction network is used to extract K normalized 3D key points from the source image;
[0224] The image information extraction network is used to obtain the image information corresponding to the source image and the driving image respectively. Among them, the image information corresponding to the source image is used to perform 3D transformation processing on the K normalized 3D key points to generate the first set of 3D key points, and the image information corresponding to the driving image is used to perform 3D transformation processing on the K normalized 3D key points to generate the second set of 3D key points;
[0225] The 3D feature extraction network is used to extract the 3D feature representation of the source image;
[0226] The motion field estimation network is used to generate the 3D optical flow mask and the 2D occlusion mask;
[0227] The decoding network is used to generate the driving result image corresponding to the source image.
[0228] In some embodiments, the driving image is an image frame in a driving video, and the device 600 further includes: a video generation module ( Figure 6 not shown in the figure).
[0229] The video generation module is used to integrate multiple driving result images corresponding to the source image to generate a driving result video corresponding to the source image; wherein, the multiple driving result images are generated according to the source image and multiple image frames in the driving video.
[0230] Please refer to Figure 7 , which shows a block diagram of a training device for an expression migration model provided by an embodiment of the present application. This device has the function of implementing the training method of the above-mentioned expression migration model, and this function can be implemented by hardware or by hardware executing corresponding software. This device can be a model training device or can be set in a model training device. The device 700 can include: a training data acquisition module 710, a key point generation module 720, an optical flow generation module 730, a mask generation module 740, a result image generation module 750, and a parameter adjustment module 760.
[0231] The training data acquisition module 710 is used to acquire the training data of the expression migration model, and the training data includes a source image sample, a driving image sample, and a target driving result image corresponding to the source image sample.
[0232] A key point generation module 720, configured to generate a first set of 3D key points and a second set of 3D key points according to the source graph sample and the driving graph sample through the expression transfer model; wherein, the first set of 3D key points is a set of key points used to represent the posture and expression information of the face object in the source graph sample; the second set of 3D key points is a set of key points used to represent the posture and expression information of the face object in the source graph sample after being driven by the driving graph sample.
[0233] An optical flow generation module 730, configured to generate a set of 3D dense optical flows according to the first set of 3D key points and the second set of 3D key points; wherein, the set of 3D dense optical flows is used to represent the spatial position changes of the corresponding key points in the first set of 3D key points and the second set of 3D key points.
[0234] A mask generation module 740, configured to process the 3D feature representation of the source graph sample based on the set of 3D dense optical flows through the expression transfer model to generate a 3D optical flow mask and a 2D occlusion mask; wherein, the 3D optical flow mask is used to perform a linear combination of each 3D dense optical flow in the set of 3D dense optical flows, and the 2D occlusion mask is used to selectively retain the feature information of the source graph sample.
[0235] A result graph generation module 750, configured to generate an output driving result graph corresponding to the source graph sample through the expression transfer model according to the set of 3D dense optical flows, the 3D optical flow mask, the 3D feature representation of the source graph sample, and the 2D occlusion mask.
[0236] A parameter adjustment module 760, configured to calculate the training loss of the expression transfer model according to the output driving result graph and the target driving result graph, and adjust the parameters of the expression transfer model based on the training loss.
[0237] In some embodiments, the expression transfer model includes: a 3D key point extraction network, an image information extraction network, a 3D feature extraction network, a motion field estimation network, and a decoding network; wherein,
[0238] The 3D key point extraction network is used to extract K normalized 3D key points from the source graph sample;
[0239] The image information extraction network is used to obtain the image information corresponding to the source graph sample and the driving graph sample respectively. Among them, the image information corresponding to the source graph sample is used to perform 3D transformation processing on the K normalized 3D key points to generate the first set of 3D key points, and the image information corresponding to the driving graph sample is used to perform 3D transformation processing on the K normalized 3D key points to generate the second set of 3D key points;
[0240] The 3D feature extraction network is used to extract the 3D feature representation of the source map sample;
[0241] The motion field estimation network is used to generate the 3D optical flow mask and the 2D occlusion mask;
[0242] The decoding network is used to generate the output driving result map corresponding to the source map sample.
[0243] In some embodiments, the parameter adjustment module 760 is used for:
[0244] Calculate a visual perception loss according to the image features of the output driving result map and the image features of the target driving result map, where the visual perception loss is used to measure the difference degree between the image features of the output driving result map and the image features of the target driving result map;
[0245] Calculate a generative adversarial loss according to the output driving result map and the target driving result map, where the generative adversarial loss is used to measure the difference degree between the output driving result map and the target driving result map;
[0246] Calculate a geometric equivariance loss according to the position information of each second key point in the second 3D key point set, where the geometric equivariance loss is used to measure the geometric equivariance of the second key point during the thin plate spline interpolation transformation;
[0247] Calculate a 3D key point prior loss according to the position information of each second key point in the second 3D key point set, where the 3D key point prior loss is used to measure the dispersion and depth of each second key point;
[0248] Calculate a face pose loss according to the pose angle of the face object in the driving map, where the face pose loss is used to measure the accuracy of face pose estimation;
[0249] Calculate a bias prior loss according to the offset amount of the 3D key points of the face object in the driving map, where the bias prior loss is used to measure the deviation amount between the position information of the 3D key points of the face object in the driving map and the reference position information;
[0250] Calculate the training loss of the expression transfer model according to the visual perception loss, the adversarial loss, the geometric equivariance loss, the 3D key point prior loss, the face pose loss and the bias prior loss.
[0251] In some embodiments, the training data acquisition module 710 is used for:
[0252] Obtain video samples for generating the training data;
[0253] Extract a first image frame and a second image frame from the video sample, where the second image frame is the next image frame of the first image frame;
[0254] Determine the first image frame as the source graph sample;
[0255] Determine the second image frame as the target driving result graph corresponding to the driving graph sample and the source graph sample.
[0256] It should be noted that for the device provided in the above embodiment, when implementing its functions, only the division of the above function modules is used for illustration. In actual applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the method embodiment belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be repeated here.
[0257] Please refer to Figure 8 , which shows a schematic structural diagram of a computer device provided in an embodiment of the present application. The computer device can be any electronic device with data calculation, processing, and storage functions, and the computer device can be implemented as Figure 1 the model training device 10 and / or the model using device 20 in the implementation environment of the solution shown. When the computer device is implemented as Figure 1 the model training device 10 in the implementation environment of the solution shown, the computer device can be used to implement the training method of the expression transfer model provided in the above embodiment. When the computer device is implemented as Figure 1 the model using device 20 in the implementation environment of the solution shown, the computer device can be used to implement the expression transfer method provided in the above embodiment. Specifically:
[0258] The computer device 800 includes a central processing unit (such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array), etc.) 801, a system memory 804 including a RAM (Random-Access Memory) 802 and a ROM (Read-Only Memory) 803, and a system bus 805 connecting the system memory 804 and the central processing unit 801. The computer device 800 also includes a basic input / output system (Input Output System, I / O system) 806 for facilitating the transfer of information between various components within the server, and a mass storage device 807 for storing an operating system 813, application programs 814, and other program modules 815.
[0259] In some embodiments, the basic input / output system 806 includes a display 808 for displaying information and input devices 809 such as a mouse and a keyboard for user input of information. Among them, both the display 808 and the input devices 809 are connected to the central processing unit 801 through an input / output controller 810 connected to the system bus 805. The basic input / output system 806 may also include an input / output controller 810 for receiving and processing inputs from multiple other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 810 also provides outputs to a display screen, a printer, or other types of output devices.
[0260] The mass storage device 807 is connected to the central processing unit 801 through a mass storage controller (not shown) connected to the system bus 805. The mass storage device 807 and its associated computer-readable medium provide non-volatile storage for the computer device 800. That is to say, the mass storage device 807 may include a computer-readable medium (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.
[0261] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, magnetic tape cartridges, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will know that the computer storage media is not limited to the above several types. The above system memory 804 and mass storage device 807 can be collectively referred to as memory.
[0262] According to an embodiment of the present application, the computer device 800 can also run on a remote computer on the network through a network such as the Internet. That is, the computer device 800 can be connected to the network 812 through the network interface unit 811 connected to the system bus 805. Or rather, the network interface unit 811 can also be used to connect to other types of networks or remote computer systems (not shown).
[0263] The memory further includes at least one instruction, at least one program, a code set or an instruction set, which is stored in the memory and is configured to be executed by one or more processors to implement the above-mentioned expression migration method or the training method of the expression migration model.
[0264] In an exemplary embodiment, a computer-readable storage medium is further provided. At least one instruction, at least one program, a code set or an instruction set is stored in the storage medium. When the at least one instruction, the at least one program, the code set or the instruction set is executed by the processor of the computer device, the above-mentioned expression migration method or the training method of the expression migration model is implemented.
[0265] Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random-Access Memory), SSD (Solid State Drives), or optical discs, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0266] In an exemplary embodiment, a computer program product or a computer program is further provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned expression migration method or the training method of the expression migration model.
[0267] It should be understood that "a plurality of" mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. In addition, the step numbers described in this article only exemplarily show a possible execution sequence between steps. In some other embodiments, the above steps may not be executed in the order of the numbers. For example, two steps with different numbers are executed simultaneously, or two steps with different numbers are executed in the reverse order of the illustration. The embodiments of the present application do not limit this.
[0268] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. An expression transfer method, characterized in that, The method includes: Generating a first set of 3D key points and a second set of 3D key points according to a source map and a driving map; wherein, the first set of 3D key points is a set of key points used to characterize the pose and expression information of a face object in the source map; the second set of 3D key points is a set of key points used to characterize the pose and expression information of the face object in the source map after being driven by the driving map; Generating a set of 3D dense optical flows according to the first set of 3D key points and the second set of 3D key points; wherein, the set of 3D dense optical flows is used to characterize the spatial position changes of corresponding key points in the first set of 3D key points and the second set of 3D key points; Processing the 3D feature representation of the source map based on the set of 3D dense optical flows to generate a 3D optical flow mask and a 2D occlusion mask; wherein, the 3D optical flow mask is used to perform a linear combination of each 3D dense optical flow in the set of 3D dense optical flows, and the 2D occlusion mask is used to selectively retain the feature information of the source map; Generating a driving result map corresponding to the source map according to the set of 3D dense optical flows, the 3D optical flow mask, the 3D feature representation of the source map, and the 2D occlusion mask; wherein, the driving result map is an image having the identity information of the face object in the source map and the pose and expression information of the face object in the driving map.
2. The method according to claim 1, characterized in that, The generating the driving result map corresponding to the source map according to the set of 3D dense optical flows, the 3D optical flow mask, the 3D feature representation of the source map, and the 2D occlusion mask includes: Performing a linear combination of each 3D dense optical flow in the set of 3D dense optical flows based on the 3D optical flow mask to obtain a combined 3D dense optical flow; Performing a transformation process on the 3D feature representation of the source map based on the combined 3D dense optical flow to obtain a transformed 3D feature representation; Performing a dimensionality reduction process on the transformed 3D feature representation to obtain a transformed 2D feature representation; Generating the driving result map corresponding to the source map according to the 2D occlusion mask and the transformed 2D feature representation.
3. The method according to claim 2, wherein The generating the driving result map corresponding to the source map according to the 2D occlusion mask and the transformed 2D feature representation includes: Processing the transformed 2D feature representation using the 2D occlusion mask to obtain a processed 2D feature representation; Decoding the processed 2D feature representation to generate the driving result map corresponding to the source map.
4. The method according to claim 1, wherein The processing the 3D feature representation of the source map based on the set of 3D dense optical flows to generate a 3D optical flow mask and a 2D occlusion mask includes: Performing a transformation process on the 3D feature representation of the source map respectively based on each 3D dense optical flow in the set of 3D dense optical flows to obtain a plurality of transformed 3D feature representations; Stitching the plurality of transformed 3D feature representations to obtain a stitched 3D feature representation; Processing the stitched 3D feature representation to generate the 3D optical flow mask and the 2D occlusion mask.
5. The method according to claim 1, wherein Generating a first set of 3D key points and a second set of 3D key points according to the source image and the driving image, includes: Extracting K normalized 3D key points from the source image, where K is a positive integer; Obtaining the image information corresponding to the source image and the driving image respectively, where the image information includes the pose angle, translation amount of the face object, and the offset amounts of K 3D key points; Performing 3D transformation processing on the K normalized 3D key points using the image information corresponding to the source image to generate the first set of 3D key points; Performing 3D transformation processing on the K normalized 3D key points using the image information corresponding to the driving image to generate the second set of 3D key points.
6. The method according to claim 5, wherein The obtaining the image information corresponding to the source image and the driving image respectively, includes: Extracting the feature information of the source image and the feature information of the driving image respectively; Based on the feature information of the source image, predicting the image information corresponding to the source image; Based on the feature information of the driving image, predicting the image information corresponding to the driving image.
7. The method according to claim 1, wherein Generating a 3D dense optical flow set according to the first set of 3D key points and the second set of 3D key points, includes: For each pair of corresponding key points in the first set of 3D key points and the second set of 3D key points, determining the 3D dense optical flow corresponding to the corresponding key points according to the position information of the two 3D key points in the corresponding key points; Obtaining the 3D dense optical flow set according to the 3D dense optical flows corresponding to each group of the corresponding key points.
8. The method according to any one of claims 1 to 7, characterized in that, Implementing the method through an expression transfer model, where the expression transfer model includes: a 3D key point extraction network, an image information extraction network, a 3D feature extraction network, a motion field estimation network, and a decoding network; wherein, The 3D key point extraction network is used to extract K normalized 3D key points from the source image; The image information extraction network is used to obtain the image information corresponding to the source image and the driving image respectively, where the image information corresponding to the source image is used to perform 3D transformation processing on the K normalized 3D key points to generate the first set of 3D key points, and the image information corresponding to the driving image is used to perform 3D transformation processing on the K normalized 3D key points to generate the second set of 3D key points; The 3D feature extraction network is used to extract the 3D feature representation of the source image; The motion field estimation network is used to generate the 3D optical flow mask and the 2D occlusion mask; The decoding network is used to generate the driving result image corresponding to the source image.
9. The method according to any one of claims 1 to 7, characterized in that, The driving image is an image frame in a driving video, and the method further includes: Integrating the multiple driving result images corresponding to the source image to generate the driving result video corresponding to the source image; Wherein, the multiple driving result images are generated according to the source image and multiple image frames in the driving video.
10. A training method for an expression transfer model, characterized in that The method includes: Obtaining the training data of the expression transfer model, where the training data includes source image samples, driving image samples, and the target driving result images corresponding to the source image samples; Generate a first set of 3D key points and a second set of 3D key points according to the source map sample and the driving map sample by means of the expression transfer model; wherein, the first set of 3D key points is a set of key points used to represent the posture and expression information of the face object in the source map sample; the second set of 3D key points is a set of key points used to represent the posture and expression information of the face object in the source map sample after being driven by the driving map sample; Generate a 3D dense optical flow set according to the first set of 3D key points and the second set of 3D key points; wherein, the 3D dense optical flow set is used to represent the spatial position change of the corresponding key points in the first set of 3D key points and the second set of 3D key points; Process the 3D feature representation of the source map sample by means of the expression transfer model based on the 3D dense optical flow set to generate a 3D optical flow mask and a 2D occlusion mask; wherein, the 3D optical flow mask is used to perform a linear combination of each 3D dense optical flow in the 3D dense optical flow set, and the 2D occlusion mask is used to selectively retain the feature information of the source map sample; Generate the output driving result map corresponding to the source map sample by means of the expression transfer model according to the 3D dense optical flow set, the 3D optical flow mask, the 3D feature representation of the source map sample, and the 2D occlusion mask; Calculate the training loss of the expression transfer model according to the output driving result map and the target driving result map, and adjust the parameters of the expression transfer model based on the training loss.
11. The method according to claim 10, wherein The expression transfer model includes: a 3D key point extraction network, an image information extraction network, a 3D feature extraction network, a motion field estimation network, and a decoding network; wherein, The 3D key point extraction network is used to extract K normalized 3D key points from the source map sample; The image information extraction network is used to obtain the image information corresponding to the source map sample and the driving map sample respectively, wherein the image information corresponding to the source map sample is used to perform 3D transformation processing on the K normalized 3D key points to generate the first set of 3D key points, and the image information corresponding to the driving map sample is used to perform 3D transformation processing on the K normalized 3D key points to generate the second set of 3D key points; The 3D feature extraction network is used to extract the 3D feature representation of the source map sample; The motion field estimation network is used to generate the 3D optical flow mask and the 2D occlusion mask; The decoding network is used to generate the output driving result map corresponding to the source map sample.
12. The method according to claim 10, wherein The calculating the training loss of the expression transfer model according to the output driving result map and the target driving result map includes: Calculate a visual perception loss according to the image features of the output driving result map and the image features of the target driving result map, and the visual perception loss is used to measure the difference degree between the image features of the output driving result map and the image features of the target driving result map; Calculate a generative adversarial loss based on the output driving result map and the target driving result map, where the generative adversarial loss is used to measure the difference between the output driving result map and the target driving result map; Calculate a geometric equivariance loss based on the position information of each second key point in the second 3D key point set, where the geometric equivariance loss is used to measure the geometric equivariance of the second key points during the thin plate spline interpolation transformation; Calculate a 3D key point prior loss based on the position information of each second key point in the second 3D key point set, where the 3D key point prior loss is used to measure the dispersion and depth of each second key point; Calculate a face pose loss based on the pose angle of the face object in the driving map, where the face pose loss is used to measure the accuracy of face pose estimation; Calculate a bias prior loss based on the offset amount of the 3D key points of the face object in the driving map, where the bias prior loss is used to measure the deviation between the position information of the 3D key points of the face object in the driving map and the reference position information; Calculate the training loss of the expression transfer model based on the visual perception loss, the adversarial loss, the geometric equivariance loss, the 3D key point prior loss, the face pose loss, and the bias prior loss.
13. The method according to any one of claims 10 to 12, characterized in that, The obtaining of the training data for the expression transfer model includes: Obtain video samples for generating the training data; Extract a first image frame and a second image frame from the video samples, where the second image frame is the next image frame of the first image frame; Determine the first image frame as the source map sample; Determine the second image frame as the driving map sample and the target driving result map corresponding to the source map sample.
14. An expression transfer device, characterized in that The device includes: A key point generation module, configured to generate a first 3D key point set and a second 3D key point set according to a source map and a driving map; wherein, the first 3D key point set is a key point set for characterizing the pose and expression information of the face object in the source map; the second 3D key point set is a key point set for characterizing the pose and expression information of the face object in the source map after being driven by the driving map; An optical flow generation module, configured to generate a 3D dense optical flow set according to the first 3D key point set and the second 3D key point set; wherein, the 3D dense optical flow set is used to characterize the spatial position change of the corresponding key points in the first 3D key point set and the second 3D key point set; A mask generation module, configured to process the 3D feature representation of the source map based on the 3D dense optical flow set to generate a 3D optical flow mask and a 2D occlusion mask; wherein, the 3D optical flow mask is used to linearly combine each 3D dense optical flow in the 3D dense optical flow set, and the 2D occlusion mask is used to selectively retain the feature information of the source map; A result map generation module, configured to generate a driving result map corresponding to the source map according to the 3D dense optical flow set, the 3D optical flow mask, the 3D feature representation of the source map, and the 2D occlusion mask; wherein, the driving result map is an image with the identity information of the face object in the source map, as well as the pose and expression information of the face object in the driving map.
15. A training device for an expression transfer model, characterized in that, The device includes: A training data acquisition module, configured to acquire training data for the expression transfer model, where the training data includes source map samples, driving map samples, and target driving result maps corresponding to the source map samples; A key point generation module, configured to generate a first 3D key point set and a second 3D key point set through the expression transfer model according to the source map samples and the driving map samples; wherein, the first 3D key point set is a key point set used to represent the pose and expression information of the face object in the source map samples; the second 3D key point set is a key point set used to represent the pose and expression information of the face object in the source map samples after being driven by the driving map samples; An optical flow generation module, configured to generate a 3D dense optical flow set according to the first 3D key point set and the second 3D key point set; wherein, the 3D dense optical flow set is used to represent the spatial position change of the corresponding key points in the first 3D key point set and the second 3D key point set; A mask generation module, configured to process the 3D feature representation of the source map samples through the expression transfer model based on the 3D dense optical flow set to generate a 3D optical flow mask and a 2D occlusion mask; wherein, the 3D optical flow mask is used to perform a linear combination of each 3D dense optical flow in the 3D dense optical flow set, and the 2D occlusion mask is used to selectively retain the feature information of the source map samples; A result map generation module, configured to generate an output driving result map corresponding to the source map samples through the expression transfer model according to the 3D dense optical flow set, the 3D optical flow mask, the 3D feature representation of the source map samples, and the 2D occlusion mask; A parameter adjustment module, configured to calculate the training loss of the expression transfer model according to the output driving result map and the target driving result map, and adjust the parameters of the expression transfer model based on the training loss.
16. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the expression transfer method according to any one of claims 1 to 9, or to implement the training method of the expression transfer model according to any one of claims 10 to 13.
17. A computer-readable storage medium, characterized in that, At least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the expression migration method according to any one of claims 1 to 9, or to implement the training method of the expression migration model according to any one of claims 10 to 13.
18. A computer program product or a computer program, characterized in that, The computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium, and the processor reads and executes the computer instructions from the computer-readable storage medium to implement the expression migration method according to any one of claims 1 to 9, or to implement the training method of the expression migration model according to any one of claims 10 to 13.
Citation Information
Patent Citations
Face key point tracking method and device and electronic device
CN111563490A
Image facial expression migration method and device, electronic equipment and readable storage medium
CN112800869A