Facial expression transfer method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202210981919.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-16
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-08-16
AI Technical Summary
[0004]上述使用稠密运动场来实现人脸表情迁移的方式,在大的头部姿态下,稠密运动场会变得非常复杂(因为稠密运动场描述的是像素的空间位移,当头部姿态变大时,整幅图像中需要移位的像素就增多,例如,人脸微笑只有嘴部的像素有位移,但是甩头发就会整个人脸的像素有位移),这种复杂性导致神经网络的输出产生较大的误差,合成的人脸面部会产生形变,影响了表情迁移效果
[0020]本公开实施例提供一种表情迁移方法、装置、电子设备及存储介质,通过使用源人脸图像的源人脸外貌标识系数、驱动人脸图像的驱动表情系数和驱动姿态系数对平均三维人脸模型进行变形处理,得到目标三维人脸模型,使得目标三维人脸模型具有源人脸图像的外貌信息和驱动人脸图像的表情与姿态信息,通过采样该模型上的关键点,并将关键点的位置信息体现在二维热力图中,再通过该二维热力图和源人脸图像各特征层对应的仿射变换系数对源人脸图像进行仿射变换,得到人脸表情迁移图像,这种方式得到的人脸表情迁移图像的人脸表情和人脸姿态为驱动人脸图像包含的人脸表情和人脸姿态,外貌为源人脸图像包含的人脸外貌,缓解了基于稠密运动场的表情迁移技术中的人脸形变问题,提升了表情迁移效果。
Smart Images

Figure CN115330980B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image transfer technology, and in particular to an expression transfer method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the development of image transfer technology, facial expression transfer has found wide application in face forgery, facial expression database generation, face editing, and art design fields such as film and animation. Facial expression transfer technology refers to a process where, given a driving face video or image and a source face image, the driving face ID and the source face ID are different (i.e., the person in the driving face image and the person in the source face image are not the same person), the facial expression and head pose (also known as facial pose) from the driving face image are transferred to the source face image to synthesize a new source face image or video. In the synthesized image or video, the expression and head pose of the source face are consistent with the driving face.
[0003] In existing facial expression transfer technologies, for all frames of the driving face video, 3D face reconstruction technology is used to reconstruct the driving face's facial expression coefficients, facial pose coefficients, and face ID coefficients (also known as facial appearance identifier coefficients, i.e., data reflecting facial appearance deformation). Then, the reconstructed facial expression coefficients, facial pose coefficients, and face ID coefficients are input into a neural network to calculate a dense motion field. This dense motion field is applied to the source face image, and then passed through another neural network to synthesize the final facial expression transfer image or video. The aforementioned dense motion field describes the displacement of each pixel in the image.
[0004] The above-mentioned method of using dense motion fields to achieve facial expression transfer becomes very complex under large head poses (because dense motion fields describe the spatial displacement of pixels; when the head pose increases, the number of pixels that need to be shifted in the entire image increases; for example, only the pixels around the mouth are shifted when a face smiles, but the entire face is shifted when the hair is tossed). This complexity leads to large errors in the output of the neural network, and the synthesized face will be deformed, affecting the expression transfer effect. Summary of the Invention
[0005] The purpose of this disclosure is to provide a method, apparatus, electronic device, and storage medium for facial expression transfer, so as to improve the effect of facial expression transfer.
[0006] In a first aspect, embodiments of this disclosure provide an expression transfer method, which provides a driving face image and a source face image via an electronic device. The method includes: performing three-dimensional face reconstruction on the driving face image to obtain driving adjustment coefficients; wherein the driving adjustment coefficients include driving expression coefficients for driving facial expression adjustment and driving pose coefficients for driving facial pose adjustment; performing the three-dimensional face reconstruction on the source face image to obtain source face appearance identification coefficients for characterizing facial appearance features in the source face image; applying the source face appearance identification coefficients and the driving adjustment coefficients to deform a preset average three-dimensional face model to obtain a target three-dimensional face model; sampling a preset number of key points from the target three-dimensional face model to obtain a two-dimensional heatmap corresponding to the key points; wherein the two-dimensional heatmap includes the position of the key points in the target three-dimensional face model and the expression information, pose information, and appearance information corresponding to the position; determining affine transformation coefficients based on the source face image and the two-dimensional heatmap, and performing an affine transformation on the source face image based on the affine transformation coefficients to obtain a facial expression transfer image.
[0007] In some implementations, applying the source face appearance identification coefficient and the driving adjustment coefficient to deform a preset average 3D face model includes: applying the source face appearance identification coefficient to deform the appearance features of the preset average 3D face model; applying the driving expression coefficient in the driving adjustment coefficient to deform the expression features of the average 3D face model; and applying the driving posture coefficient in the driving adjustment coefficient to deform the head posture features of the average 3D face model.
[0008] In some implementations, the source face appearance identifier coefficient and the driving adjustment coefficient are used to deform a preset average 3D face model, including deformation processing using the following formula:
[0009] M=R d .(M0+sp s .V sp +ep d .V ep )+t d Where M represents the target 3D face model, M0 represents the preset average 3D face model, and V sp This represents the preset face shape base, sp s V represents the source face appearance identification coefficient. ep This represents the preset facial expression base, ep d R represents the driving expression coefficient. d and t d This represents the rotation and translation parameters in the driving attitude coefficients.
[0010] In some implementations, the two-dimensional heat map corresponding to the key point is obtained using the following formula:
[0011] in, This represents the orthogonal projection matrix, used to project the key points onto a two-dimensional plane. s G(.) represents the projection scaling factor of the source face image, G(.) represents the Gaussian blur function, and M represents the projection scaling factor of the source face image. n This indicates that n key points are sampled from the target 3D face model, H is the generated 2D heat map, and n is a preset value.
[0012] In some implementations, the affine transformation coefficients include scaling parameters, rotation parameters, and translation parameters.
[0013] In some implementations, determining the affine transformation coefficients based on the source face image and the two-dimensional heatmap includes: performing convolution processing on the source face image and the two-dimensional heatmap using a texture encoder to obtain feature maps corresponding to a first number of feature layers; wherein the feature maps of different feature layers correspond to different region features; and determining the affine transformation coefficients of the feature maps corresponding to each feature layer using a transform encoder.
[0014] In some implementations, the step of performing an affine transformation on the source face image according to the affine transformation coefficients to obtain a facial expression transfer image includes: for each feature map corresponding to a feature layer, performing an affine transformation on the feature map according to the affine transformation coefficients of the feature map using an affine transformer to obtain an affine transformation map of the feature map; and decoding the affine transformation map corresponding to each feature layer using a texture decoder to obtain a facial expression transfer image.
[0015] In some embodiments, performing an affine transformation on the feature map using an affine transformer based on the affine transformation coefficients of the feature map includes: inputting the affine transformation coefficients of the feature map and the feature map into an affine transformer; and performing the affine transformation using the following formula within the affine transformer:
[0016] Where, x c y c Represents the coordinates of pixels on the feature map. This represents the coordinates of the pixel after the affine transformation. Let p represent the affine transformation matrix consisting of the affine transformation coefficients. s p is the scaling parameter in the affine transformation coefficients. θ The rotation parameter in the affine transformation coefficients, The x-axis translation parameter in the affine transformation coefficients. y-axis translation parameter in the affine transformation coefficients.
[0017] Secondly, embodiments of this disclosure also provide an expression transfer device, which provides a driving face image and a source face image via an electronic device. The device includes: a first coefficient acquisition module, configured to acquire driving adjustment coefficients of the driving face image; wherein the driving adjustment coefficients include driving expression coefficients for driving facial expression adjustment and driving posture coefficients for driving facial pose adjustment; a second coefficient acquisition module, configured to acquire source face appearance identification coefficients of the source face image, wherein the source face appearance identification coefficients are used to characterize the facial appearance features in the source face image; and a processing module, configured to apply the source face appearance identification coefficients and The driving adjustment coefficients deform a preset average 3D face model to obtain a target 3D face model. A heatmap acquisition module samples a preset number of key points from the target 3D face model to obtain a 2D heatmap corresponding to each key point. The 2D heatmap includes the position of each key point in the target 3D face model and the corresponding expression, posture, and appearance information. An affine transformation module determines affine transformation coefficients based on the source face image and the 2D heatmap, and performs an affine transformation on the source face image based on the affine transformation coefficients to obtain a facial expression transfer image.
[0018] Thirdly, embodiments of this disclosure also provide an electronic device, including a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above-described expression transfer method.
[0019] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the aforementioned facial expression transfer method.
[0020] This disclosure provides an expression transfer method, apparatus, electronic device, and storage medium. It involves deforming an average 3D face model using source face appearance identification coefficients from a source face image, driving expression coefficients from a driving face image, and driving pose coefficients from a driving face image to obtain a target 3D face model. This target 3D face model possesses the appearance information of the source face image and the expression and pose information of the driving face image. By sampling key points on the model and representing their location information in a 2D heatmap, the source face image is then subjected to an affine transformation using the 2D heatmap and the affine transformation coefficients corresponding to each feature layer of the source face image. This yields a transferred expression image. The facial expression and pose of the transferred image obtained in this way are the same as those contained in the driving face image, and the appearance is the same as that contained in the source face image. This alleviates the facial deformation problem in expression transfer techniques based on dense motion fields and improves the expression transfer effect. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the specific embodiments of this disclosure or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of an implementation environment according to an embodiment of the present disclosure;
[0023] Figure 2 This is a flowchart illustrating an expression transfer method according to an embodiment of the present disclosure;
[0024] Figure 3 This is an example diagram of another facial expression transfer method in the embodiments of this disclosure;
[0025] Figure 4 This is an example diagram of the adaptive affine transformation module in an embodiment of this disclosure;
[0026] Figure 5 This is a schematic diagram of the structure of an expression transfer device according to an embodiment of the present disclosure;
[0027] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0028] The technical solutions of this disclosure will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0029] In 3D face reconstruction technology, a dense motion field is typically obtained by reconstructing facial expression coefficients, facial pose coefficients, and facial ID coefficients (also known as facial appearance identifier coefficients, which reflect data on facial appearance deformation) to drive the face. This dense motion field is then applied to the source face image to synthesize the final image or video with transferred facial expressions. However, this method of using a dense motion field to achieve facial expression transfer suffers from several drawbacks. Because the dense motion field describes the spatial displacement of pixels, as the head pose increases, the number of pixels that need to be shifted in the entire image increases, causing deformation of the synthesized face and affecting the expression transfer effect. To address these issues, this disclosure provides an expression transfer method, apparatus, electronic device, and storage medium that can improve the expression transfer effect.
[0030] In one embodiment of this disclosure, the expression transfer method can be run on an electronic device (such as a terminal device or a server). The terminal device can be a local terminal device. When the expression transfer method runs on a server, it can be implemented and executed based on a cloud interaction system, which includes a server and client devices.
[0031] In an optional implementation, various cloud applications, such as cloud gaming, can run under the cloud interaction system. Taking cloud gaming as an example, cloud gaming refers to a gaming method based on cloud computing. In the cloud gaming operating mode, the game program and the game screen presentation are separated. The storage and execution of the facial expression transfer method are completed on the cloud gaming server. The client device is used for data reception, transmission, and game screen presentation. For example, the client device can be a display device with data transmission capabilities located close to the user, such as a mobile terminal, television, computer, PDA, or game console; however, the terminal device for information processing is the cloud gaming server in the cloud. When playing the game, the player operates the client device to send operation commands to the cloud gaming server. The cloud gaming server runs the game according to the operation commands, encodes and compresses the game screen and other data, returns it to the client device via the network, and finally, the client device decodes and outputs the game screen.
[0032] In an alternative implementation, the terminal device can be a local terminal device. Taking a game as an example, the local terminal device stores the game program and is used to display the game screen. The local terminal device is used to interact with the player through a graphical user interface, that is, conventionally downloading, installing, and running the game program via an electronic device. The local terminal device can provide the graphical user interface to the player in various ways, such as rendering it on the terminal's display screen, or providing it to the player through holographic projection. For example, the local terminal device can include a display screen for displaying the graphical user interface, which includes game screens, and a processor for running the game, generating the graphical user interface, and controlling the display of the graphical user interface on the display screen.
[0033] In one possible implementation, embodiments of this disclosure provide an expression transfer method that provides a graphical user interface (GUI) via a terminal device. The terminal device can be either the aforementioned local terminal device or a client device within the aforementioned cloud interaction system. The GUI provided by the terminal device can include a driving face image (or a video frame containing the driving face image) and a source face image.
[0034] This disclosure provides embodiments that Figure 1 The diagram illustrates the implementation environment. This environment may include a first terminal device, a server, and a second terminal device. The first terminal device and the second terminal device communicate with the server to achieve data communication. In this embodiment, the first terminal device and the second terminal device are each equipped with a client that executes the facial expression transfer method provided by this disclosure, and the server is a server-side application that executes the facial expression transfer method provided by this disclosure. Through the client, the first terminal device and the second terminal device can communicate with the server respectively.
[0035] This disclosure provides an expression transfer method applied to electronic devices (such as mobile phones, computers, tablets, etc.), which provides a driving face image and a source face image through the electronic device.
[0036] Applying the aforementioned facial expression transfer method to games involves capturing real-time video frames of the player's face using a camera. The captured facial images from these frames serve as the driving face image, while the player character's face image in the game is used as the source face image. By applying the method described above, the player character's facial expressions can be transferred, ensuring that the player character's expressions and facial postures match those of a real-world player. Here, "player character" refers to a virtual object that can be controlled by the player and moves within the game environment; in some video games, it may also be called a shikigami (spirit) or hero character. A player character can be at least one of several different forms, such as a virtual character or a virtual anime character.
[0037] The aforementioned electronic devices can be touch-enabled devices or non-touch-enabled devices. Touch-enabled devices can be operated via a touchscreen display, while non-touch-enabled devices can be operated via external devices such as a mouse, keyboard, or gamepad. See also Figure 2 The diagram illustrates a method for transferring facial expressions. The electronic device described above is the executing entity of this method, which includes the following steps:
[0038] Step S202: Obtain the driving adjustment coefficients for driving the face image; wherein, the driving adjustment coefficients include driving expression coefficients for driving face expression adjustment and driving pose coefficients for driving face pose adjustment.
[0039] The aforementioned driving face image can be a face image directly acquired by an image acquisition device (such as a camera), a face image in a video frame sequence acquired in real time by a video acquisition device (such as a camera), or a face image crawled from the network; there is no limitation on this.
[0040] Three-dimensional face reconstruction can be performed on driven face images based on 3DMM (3D Morphable Model) technology, yielding the aforementioned driven expression coefficients (also known as driven face expression coefficients) and driven pose coefficients (also known as driven face pose coefficients). In addition, driven face appearance identification coefficients, or simply driven face ID coefficients, can also be obtained. The expressions in the driven expression coefficients can include coefficients corresponding to emotionally charged expressions such as joy, anger, sorrow, and happiness, as well as coefficients for expressions such as contemplation, blankness, and drowsiness. The driven pose coefficients typically refer to head posture information, such as coefficients for head tilting, turning left, turning right, or tilting the head.
[0041] Step S204: Obtain the source face appearance identification coefficient of the source face image. The source face appearance identification coefficient is used to characterize the face appearance features in the source face image.
[0042] The aforementioned source face image can be a face image directly acquired by an image acquisition device (such as a camera), a face image from a video frame sequence captured in real time by a video acquisition device (such as a webcam), a face image crawled from the internet, or a face image of the object to be processed in the work; there are no limitations on this. The source face image can be reconstructed using the same 3DMM technology as the aforementioned driving face image 3D face reconstruction to obtain the aforementioned source face appearance identification coefficient, which can also be simply referred to as the source face ID coefficient. In addition to this coefficient, the source face expression coefficient and the aforementioned source face pose coefficient can also be obtained during the 3D face reconstruction process.
[0043] In addition to being obtained using 3D face reconstruction technology, the aforementioned driving adjustment coefficients and source face appearance identification coefficients can also be determined using a neural network model. For example, the aforementioned driving face image and the aforementioned source face image can be input into a pre-trained neural network model, which outputs the driving adjustment coefficients of the driving face image and the source face appearance identification coefficients of the source face image.
[0044] Step S206: Apply the source face appearance identification coefficient and driving adjustment coefficient to deform the preset average three-dimensional face model to obtain the target three-dimensional face model.
[0045] The source face appearance identifier coefficient and the driving adjustment coefficient can be used to deform the preset average 3D face model in the above 3DMM to obtain the target 3D face model. Specifically, the source face appearance identifier coefficient is used to deform the appearance features of the average 3D face model, the driving expression coefficient in the driving adjustment coefficient is used to deform the expression features of the average 3D face model, and the driving pose coefficient in the driving adjustment coefficient is used to deform the head pose features of the average 3D face model. The order of deformation processing of the appearance features, expression features, and head pose features of the average 3D face model is not fixed; they can be performed sequentially or simultaneously.
[0046] Step S208: Sample a preset number of key points from the target 3D face model to obtain a 2D heat map corresponding to the key points; wherein, the 2D heat map contains the position of the key points in the target 3D face model and the corresponding expression information, posture information and appearance information.
[0047] The aforementioned target 3D face model is a mesh containing a large number of vertices. Since each vertex has its own vertex index, a preset number of vertex indices can be selected in advance and used as sampling parameters for a sampling function. Then, this sampling function samples a preset number of key points from the target 3D face model. During expression transfer, vertex indices corresponding to the facial features are typically selected. Sampling these vertex indices allows the sampled key points to represent the expression and appearance information of the target 3D face model. Therefore, the 3D information corresponding to the sampled key points includes both facial expression and appearance information.
[0048] The steps for obtaining the two-dimensional heatmap corresponding to the key points described above can include the following operations: projecting the key points onto a two-dimensional plane to obtain a two-dimensional image of the key points; wherein, the pixel value corresponding to the key points in the two-dimensional image of the key points is 255, and the value of the remaining pixels is 0; performing convolution processing on the two-dimensional image of the key points using a preset Gaussian convolution kernel to obtain the two-dimensional heatmap. When projecting the key points onto the two-dimensional plane, the key points can be rotated and translated according to the pose information of the face. Therefore, the points corresponding to the key points in the above two-dimensional heatmap include facial expression information, appearance information, and pose information. Here, the facial expression information specifically refers to the information corresponding to the driving expression coefficients, the facial pose information specifically refers to the information corresponding to the driving pose coefficients, and the facial appearance information specifically refers to the information corresponding to the source facial appearance identifier coefficients.
[0049] Step S210: Determine the affine transformation coefficients based on the source face image and the two-dimensional heatmap, and perform an affine transformation on the source face image based on the affine transformation coefficients to obtain a face expression transfer image.
[0050] In practice, the corresponding affine transformation coefficients can be determined based on the source face image and the facial expression, appearance, and pose information contained in the two-dimensional heatmap. Then, the source face image is subjected to an affine transformation using these coefficients to obtain a facial expression transfer image. Since the facial expression, appearance, and pose information contained in the two-dimensional heatmap originates from the driving expression coefficients, the source face appearance identification coefficients, and the driving pose coefficients, respectively, the facial expression and pose of the transferred image are identical to those of the driving face image, and the appearance of the transferred image is identical to that of the source face image.
[0051] This disclosure provides an expression transfer method. It involves deforming an average 3D face model using source face appearance identification coefficients from a source face image, driving expression coefficients from a driving face image, and driving pose coefficients from a driving face image to obtain a target 3D face model. This target 3D face model possesses the appearance information of the source face image and the expression and pose information of the driving face image. By sampling key points on the model and representing their location information in a 2D heatmap, the source face image is then subjected to an affine transformation using the 2D heatmap and the affine transformation coefficients corresponding to each feature layer of the source face image. This yields a transferred expression image. The facial expression and pose of the transferred image obtained in this way are the same as those contained in the driving face image, and the appearance is the same as that contained in the source face image. This alleviates the facial deformation problem in expression transfer techniques based on dense motion fields and improves the expression transfer effect.
[0052] In the aforementioned affine transformation process, the source face image and the aforementioned two-dimensional heatmap can be convolved to obtain feature maps corresponding to different feature layers. In this embodiment, different feature layers encode different regional features. Different affine transformations are performed on different feature layers during the affine transformation process; that is, the affine transformation coefficients corresponding to different feature layers are used to perform affine transformation processing on the feature maps corresponding to those feature layers. After all feature maps have undergone affine transformation processing, the face expression transfer image can be obtained through decoding. This method can achieve more complex spatial deformations, and therefore, even in application scenarios with large variations in face pose, it can ensure that the face appearance in the face expression transfer image is consistent with the appearance of the source face.
[0053] As one possible implementation, the step of applying the source face appearance identifier coefficient and the driving adjustment coefficient to deform the preset average 3D face model can include the following operation: performing deformation processing using the following formula: M = R d .(M0+sp s .V sp +ep d .V ep )+t d Where M represents the target 3D face model, M0 represents the preset average 3D face model, and V sp This represents the preset face shape base, sp s V represents the source face appearance identification coefficient. ep This represents the preset facial expression base, ep d R represents the driving expression coefficient. d and t d This represents the rotation and translation parameters in the driving attitude coefficients.
[0054] The face shape basis can be a reference face shape image, such as an oval face, a heart-shaped face, or other facial features. This reference face shape image can be determined from the statistical results of multiple face images, or a face shape image can be manually specified as the face shape basis.
[0055] The aforementioned facial expression base can refer to typical facial expressions with salient features, such as blinking, frowning, tugging at the corners of the mouth, squinting, etc. The aforementioned driving expression coefficients can be the coefficients used in the linear combination with each facial expression base when reconstructing facial expressions. In the above formula, the preset facial shape base V... sp In the source facial appearance identifier coefficient sp s The face shape is deformed under the action of [something], resulting in the deformed face shape sp. s .V sp Preset facial expression base V ep In the above-mentioned driving expression coefficient ep dThe face undergoes deformation under the influence of the action, resulting in the deformed facial expression (ep). d .V ep ; the deformed face shape sp s .V sp And the deformed facial expressions ep d .V ep The effect is applied to a preset average 3D face model M0 to obtain a face model with altered shape and appearance (M0+sp). s .V sp +ep d .V ep ); Face model after shape and appearance changes (M0+sp) s .V sp +ep d .V ep After rotation parameter R d Perform rotation and translation parameter t d After translation, a deformed 3D face model is obtained (i.e., the target 3D face model mentioned above). As a possible implementation, the step of deforming the preset average 3D face model using the source face appearance identifier coefficient and driving adjustment coefficient can include the following operation: performing deformation processing using the following formula: M = R d .M0+t d +sp s .R d V sp +ep d .R d V ep Where M represents the target 3D face model, M0 represents the preset average 3D face model, and V sp This represents the preset face shape base, sp s V represents the source face appearance identification coefficient. ep This represents the preset facial expression base, ep d R represents the driving expression coefficient. d and t d This represents the rotation and translation parameters in the driving attitude coefficients.
[0056] Using the above formula, the average 3D face model can be deformed using the source face appearance identification coefficient and the driving adjustment coefficient to obtain a target 3D face model that has the appearance in the source face image and the expression and posture in the driving face image. The model achieves the expected results well.
[0057] As one possible implementation method, the two-dimensional heat map corresponding to the above key points can be obtained using the following formula: in, This represents the orthogonal projection matrix, used to project the aforementioned key points onto a two-dimensional plane. s G(.) represents the projection scaling factor of the source face image, G(.) represents the Gaussian blur function, and M represents the projection scaling factor of the source face image. n This indicates that n key points are sampled from the above target 3D face model, H is the generated 2D heat map, and n is a preset value, such as 68.
[0058] The two-dimensional heatmap obtained by the above formula is relatively simple. The influence range of key points can be expanded by using the Gaussian blur function, making the expression transfer effect more accurate. The specific formula form of H in the above embodiments is not limited, and can be flexibly determined according to actual application.
[0059] As one possible implementation, the step of determining the affine transformation coefficients based on the source face image and the two-dimensional heatmap can include the following operations: performing convolution processing on the source face image and the two-dimensional heatmap using a texture encoder to obtain feature maps corresponding to a first number of feature layers; wherein the feature maps of different feature layers correspond to different region features; determining the affine transformation coefficients of the feature maps corresponding to each feature layer using a transform encoder; wherein the affine transformation coefficients include scaling parameters, rotation parameters, and translation parameters.
[0060] For example, if the source face image is an RGB three-channel color image and the heatmap is a single-channel grayscale image, the source face image and the heatmap can be stacked together, that is, the RGB three channels of the source face image and the grayscale channel of the heatmap are stacked together to form a four-channel image. Then, the four-channel image is input into the texture encoder, which performs convolution processing on the four-channel image and outputs a feature map F of size C×H×W (that is, the feature maps corresponding to the first number of feature layers mentioned above), where C represents the number of feature maps (that is, the first number mentioned above, for example, 512), and H and W represent the height and width of the feature map, respectively. Then, each feature map F is input into the transform encoder, which calculates the corresponding affine transform coefficients for the feature map of each feature layer and outputs the affine transform coefficients for the feature map corresponding to each feature layer.
[0061] By employing the above-mentioned method of calculating affine transformation coefficients for each feature map of each feature layer, the computational load can be minimized, thereby making the computation results more accurate and effectively improving the accuracy of expression transfer.
[0062] As one possible implementation, the step of performing an affine transformation on the source face image based on the affine transformation coefficients to obtain a facial expression transfer image may include the following operation: for each feature map corresponding to a feature layer, an affine transformer is used to perform an affine transformation on the feature map based on the affine transformation coefficients of the feature map to obtain an affine transformation map of the feature map; and a texture decoder is used to decode the affine transformation map corresponding to each feature layer to obtain a facial expression transfer image.
[0063] Continuing from the previous example, after the transform encoder outputs the affine transform coefficients of the feature maps corresponding to each feature layer, for each feature map corresponding to a feature layer, the affine transform coefficients of that feature map are obtained. For each feature map corresponding to a feature layer, the affine transformer performs affine transform processing on the feature map based on the affine transform coefficients of that feature map, and outputs the affine transform map of that feature map. For each affine transform map corresponding to a feature layer, the affine transform map corresponding to each feature layer is input into the texture decoder. The texture decoder decodes the affine transform map corresponding to each feature layer, and synthesizes all the decoded affine transform maps into a facial expression transfer image, and then outputs the facial expression transfer image.
[0064] The above-described operation method of performing affine transformation processing on the feature map of each feature layer based on the affine transformation coefficients of each feature layer is adopted. The affine transformation coefficients of different feature layers are usually different, and the feature maps of different feature layers are also different. Therefore, the affine transformations performed on the feature maps corresponding to different feature layers are also different. Thus, this method can handle complex spatial deformations. By performing affine transformations on the feature maps corresponding to each feature layer separately, the amount of computation is small each time, thereby making the calculation results more accurate and effectively improving the accuracy of expression transfer. Even when performing expression transfer under large head poses, facial appearance deformation will not occur.
[0065] As one possible implementation, the above-mentioned affine transformation processing of the feature map based on the affine transformation coefficients of the feature map by the affine transformer may include the following operation: inputting the affine transformation coefficients of the feature map and the feature map into the affine transformer;
[0066] Affine transformation is performed using the following formula within the affine transformer:
[0067] Where, x c y c Represents the coordinates of pixels on the feature map. Represents the coordinates of a pixel after affine transformation. Let p represent the affine transformation matrix consisting of the affine transformation coefficients. s p is the scaling parameter in the affine transformation coefficients mentioned above. θ The rotation parameter in the affine transformation coefficients mentioned above, The x-axis translation parameter in the above affine transformation coefficients, y-axis translation parameter in the above affine transformation coefficients.
[0068] Using the above formula can ensure the accuracy and rationality of the calculation. In addition to the above formula, affine transformation can also be performed in a similar way. Alternatively, depending on the needs of the actual application scenario, at least one of the above scaling parameters, rotation parameters, or translation parameters can be added with a preset constant to reduce errors in the image acquisition or processing process and improve the affine transformation effect.
[0069] As one possible implementation, the steps of performing 3D face reconstruction on the driving face image and performing 3D face reconstruction on the source face image described above can include the following operation methods:
[0070] (1) Using the same face key point detection algorithm, the first key point in the above-mentioned driving face image and the same number of second key points in the above-mentioned source face image are detected respectively; wherein, the number of the first key point, the second key point and the key points sampled from the target three-dimensional face model are equal.
[0071] The facial landmark detection algorithm mentioned above can be determined according to actual needs, such as using OpenFace, dlib, etc., without any restrictions.
[0072] (2) Using a pre-set 3D face model, the target is optimized. Calculate the driving face ID coefficients (i.e., source face appearance identifier coefficients that characterize the facial features in the driving face image) of the driving face image. d The driving facial expression coefficient (i.e., the driving expression coefficient mentioned above) ep d and the driving face pose coefficient (i.e., the driving pose coefficient mentioned above, including the projection scaling coefficient s) d Rotation parameter R d Translation parameter t d ), and calculate the source face ID coefficient of the source face image (i.e., the source face appearance identifier coefficient that characterizes the facial features in the source face image) sp s , source facial expression coefficient ep s Heyuan face pose coefficient (including projection scaling coefficient s) s Rotation parameter R s Translation parameter t s ).
[0073] In the above formula, sp represents the face ID coefficient of the 2D image, ep represents the face expression coefficient of the 2D image, s, R, and t represent the projection scale coefficient, rotation parameter, and translation parameter in the face pose coefficient (i.e., head pose coefficient) of the 2D image, respectively, n represents the number of keypoints (first keypoint or second keypoint) detected in the 2D image (driving face image or source face image), l represents the index of the keypoints detected in the 2D image, and p l This represents the l-th keypoint detected in the two-dimensional image. This indicates that the l-th keypoint detected in the 2D image corresponds to the keypoint on the 3D face model.
[0074] The smaller the value of the above optimization objective, the more accurate the 3D face reconstruction result. The face ID coefficient sp refers to a typical facial shape with significant features, such as eyes, eyebrows, mouth, nose, ears, chin, forehead, and face shape. The face expression coefficient ep refers to a typical facial expression with significant features, such as blinking, frowning, tugging at the corners of the mouth, and squinting. The projection scale coefficient s is used to characterize the distance between the face and the camera; its value is related to the distance between the face and the camera, as well as the size of the 2D image. Since the 3D face model is constructed relative to a virtual camera at a certain orientation, and the projection scale coefficient s is the corresponding value in the 3D face model, once the 3D face model is fixed, the corresponding projection scale coefficient s is a fixed value.
[0075] For ease of understanding, here we will use Figure 3 and Figure 4 The above expression transfer method is described as an example below:
[0076] Given a source face image and a driving face video, the goal is to synthesize a face video where the appearance of the face in the synthesized video is the same as that in the source face image, and the facial expression and head pose of the face in the synthesized video are the same as those in the driving face video. To achieve this, the expression transfer method described above can be performed in three stages: 3D face reconstruction, facial landmark projection, and face image generation.
[0077] (I) Three-dimensional face reconstruction stage
[0078] In the 3D face reconstruction stage, a general 3D face reconstruction algorithm is used to calculate the facial expression coefficients, facial pose coefficients, and face ID coefficients in the source face image and the driving face image. See also... Figure 3 As shown, the specific process of 3D face reconstruction includes the following parts:
[0079] (1) Extract video frame images containing faces in the driving face video B1, and use each extracted video frame image as a driving face image; use the same face key point detection algorithm to detect 68 first key points in the face contained in each driving face image and 68 second key points in the face contained in the source face image A1.
[0080] (2) See Figure 3 As shown, 3DMM, or three-dimensional face reconstruction, is used to optimize the target... Calculate the driving face ID coefficient sp for each driving face image. d , driving facial expression coefficient ep d and driving face pose coefficients (including projection scaling coefficients s) d Rotation parameter R d Translation parameter t d ), and calculate the source face ID coefficients sp of the source face image A1. s , source facial expression coefficient ep s Heyuan face pose coefficient (including projection scaling coefficient s) s Rotation parameter R s Translation parameter t s The meanings of the symbols in the above formulas are the same as those in the aforementioned related content, and will not be repeated here.
[0081] (II) Facial Key Point Projection Stage
[0082] See Figure 3 As shown, the main tasks in the facial landmark projection stage include the following parts:
[0083] (1) The source face ID coefficients sp s , driving facial expression coefficient ep d The face pose coefficients are applied to the preset average 3D face model in 3DMM, and the deformation is performed using the following formula: M = R d .(M0+sp s .V sp +ep d .V ep )+t d This yields the deformed 3D face M, which is the aforementioned target 3D face model. The meanings of the symbols in the above formulas are the same as those in the previous related content, and will not be repeated here.
[0084] (2) Key points were sampled from the deformed 3D face C1, and 68 key points were obtained by sampling on the deformed 3D face M using the following formula: M 68 =S(M), where M 68This indicates that 68 key points are sampled on the deformed 3D face M, and S(.) represents the sampling function.
[0085] (3) Projecting from 3D to 2D: The 68 key points sampled from the deformed 3D face are projected onto a 2D plane using the following formula, and a corresponding heatmap H is generated: H = Where H represents the heat map, Let s represent the orthogonal projection matrix. s G(.) represents the projection scaling factor of the source face, and G(.) represents Gaussian blur.
[0086] (III) Face Image Generation Stage
[0087] See Figure 3 As shown, the main goal of the face image generation stage is to input the source face image A1 and the heatmap H, and synthesize the final face expression transfer result C1 through the adaptive affine transformation module T affine transformation.
[0088] See Figure 4 As shown, the basic structure of the adaptive affine transformation module includes a texture encoder E. app A transform encoder E trans An affine transformer T trans A texture decoder D app First, the source face image and heatmap are stacked together and input into the texture encoder E. app In the process, a feature map F of size C×H×W is calculated; then the feature map F is input into the transform encoder E. trans In this process, for each feature map, a transformation encoder E is used. trans Calculate the scaling parameter p corresponding to the feature map. s Rotation parameter p θ x-axis translation parameters and y-axis translation parameters Specifically, the transform encoder can be a convolutional neural network.
[0089] For each feature map, the scaling parameter p corresponding to that feature map is... s Rotation parameter p θ x-axis translation parameters and y-axis translation parameters The feature map is then input into an affine transformer, which performs an affine transformation on the feature map using the following formula: The affine transformation map corresponding to the feature map is obtained; then, the affine transformation map corresponding to each feature map is further input into the texture decoder D. app In the middle, through the texture decoder D appThe affine transformation map corresponding to each feature map is decoded, and all the decoded affine transformation maps are synthesized into a facial expression transfer image. The final synthesized result (i.e., the facial expression transfer result) is then output. This synthesized result consists of multiple synthesized images, and the time of each synthesized image in the synthesized face video corresponds to the time of each driving face image in the driving face video. In this way, a face video segment is synthesized.
[0090] Based on the above method embodiments, this disclosure also provides an expression transfer device, see [link to relevant documentation]. Figure 5 As shown, the device may include the following modules:
[0091] The first coefficient acquisition module 502 is used to acquire the driving adjustment coefficients of the driving face image; wherein the driving adjustment coefficients include driving expression coefficients for driving face expression adjustment and driving posture coefficients for driving face posture adjustment.
[0092] The second coefficient acquisition module 504 is used to acquire the source face appearance identification coefficient of the source face image, and the source face appearance identification coefficient is used to characterize the face appearance features in the source face image.
[0093] The processing module 506 is used to apply the source face appearance identification coefficient and the driving adjustment coefficient to deform the preset average three-dimensional face model to obtain the target three-dimensional face model.
[0094] The heatmap acquisition module 508 is used to sample a preset number of key points from the target three-dimensional face model and acquire a two-dimensional heatmap corresponding to the key points; wherein, the two-dimensional heatmap includes the position of the key points in the target three-dimensional face model and the expression information, posture information and appearance information corresponding to the position.
[0095] The affine transformation module 510 is used to determine affine transformation coefficients based on the source face image and the two-dimensional heatmap, and to perform an affine transformation on the source face image based on the affine transformation coefficients to obtain a face expression transfer image; wherein the face expression and face pose of the face expression transfer image are the same as those of the face expression and face pose contained in the driving face image, and the appearance of the face expression transfer image is the same as the appearance of the face contained in the source face image.
[0096] This disclosure provides an expression transfer device that deforms an average 3D face model by using the source face appearance identification coefficient of the source face image, the driving expression coefficient of the driving face image, and the driving pose coefficient to obtain a target 3D face model. This target 3D face model possesses the appearance information of the source face image and the expression and pose information of the driving face image. By sampling key points on the model and representing the location information of these key points in a 2D heatmap, and then performing an affine transformation on the source face image using the 2D heatmap and the affine transformation coefficients corresponding to each feature layer of the source face image, an expression transfer image is obtained. The expression and pose of the expression transfer image obtained in this way are the same as those contained in the driving face image, and the appearance is the same as that contained in the source face image. This alleviates the face deformation problem in expression transfer techniques based on dense motion fields and improves the expression transfer effect.
[0097] The aforementioned processing module 506 can also be used to: apply the source face appearance identification coefficient to deform the appearance features of a preset average three-dimensional face model; apply the driving expression coefficient in the driving adjustment coefficient to deform the expression features of the average three-dimensional face model; and apply the driving pose coefficient in the driving adjustment coefficient to deform the head pose features of the average three-dimensional face model. For example, deformation processing can be performed using the following formula: M = R d .(M0+sp s .V sp +ep d .V ep )+t d Where M represents the target 3D face model, M0 represents the preset average 3D face model, and V sp This represents the preset face shape base, sp s V represents the source face appearance identification coefficient. ep This represents the preset facial expression base, ep d R represents the driving expression coefficient. d and t d This represents the rotation and translation parameters in the driving attitude coefficients.
[0098] The heatmap acquisition module 508 described above can also be used to: obtain the two-dimensional heatmap corresponding to the key point using the following formula: in, This represents the orthogonal projection matrix, used to project the key points onto a two-dimensional plane. s G(.) represents the projection scaling factor of the source face image, G(.) represents the Gaussian blur function, and M represents the projection scaling factor of the source face image. n This indicates that n key points are sampled from the target 3D face model, H is the generated 2D heat map, and n is a preset value.
[0099] The affine transformation coefficients mentioned above include scaling parameters, rotation parameters, and translation parameters.
[0100] The aforementioned affine transformation module 510 can also be used to: perform convolution processing on the source face image and the two-dimensional heatmap using a texture encoder to obtain feature maps corresponding to a first number of feature layers; wherein the feature maps of different feature layers correspond to different region features; and determine the affine transformation coefficients of the feature maps corresponding to each feature layer using a transformation encoder.
[0101] The aforementioned affine transformation module 510 can also be used to: for each feature map corresponding to a feature layer, perform affine transformation processing on the feature map according to the affine transformation coefficients of the feature map through an affine transformer to obtain the affine transformation map of the feature map; and decode the affine transformation map corresponding to each feature layer through a texture decoder to obtain a facial expression transfer image.
[0102] The above-mentioned affine transformation processing of the feature map based on the affine transformation coefficients of the feature map by the affine transformer includes: inputting the affine transformation coefficients of the feature map and the feature map itself into the affine transformer; and performing the affine transformation processing using the following formula within the affine transformer: Where, x c y c Represents the coordinates of pixels on the feature map. This represents the coordinates of the pixel after the affine transformation. Let p represent the affine transformation matrix consisting of the affine transformation coefficients. s p is the scaling parameter in the affine transformation coefficients. θ The rotation parameter in the affine transformation coefficients, The x-axis translation parameter in the affine transformation coefficients. y-axis translation parameter in the affine transformation coefficients.
[0103] The first coefficient acquisition module 502 is also used to perform three-dimensional face reconstruction on the driving face image to obtain driving adjustment coefficients;
[0104] The second coefficient acquisition module 504 is also used to perform the three-dimensional face reconstruction on the source face image to obtain the source face appearance identification coefficient.
[0105] The facial expression transfer device provided in this disclosure has the same implementation principle and technical effect as the aforementioned facial expression transfer method embodiment. For the sake of brevity, any parts of the facial expression transfer device embodiment not mentioned can be referred to the corresponding content in the aforementioned facial expression transfer method embodiment.
[0106] This disclosure also provides an electronic device, such as... Figure 6The diagram shows the structure of the electronic device, which includes a processor 61 and a memory 60. The memory 60 stores computer-executable instructions that can be executed by the processor 61. The processor 61 executes the computer-executable instructions to implement the above-mentioned expression transfer method.
[0107] exist Figure 6 In the illustrated embodiment, the electronic device further includes a bus 62 and a communication interface 63, wherein the processor 61, the communication interface 63, and the memory 60 are connected via the bus 62.
[0108] The memory 60 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 63 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 62 may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 62 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0109] Processor 61 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 61 or by instructions in software form. Processor 61 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this disclosure can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the memory. The processor 61 reads the information in the memory and, in conjunction with its hardware, completes the steps of the expression transfer method described in the foregoing embodiment.
[0110] This disclosure also provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are invoked and executed by a processor, they cause the processor to implement the aforementioned facial expression transfer method. For specific implementation details, please refer to the foregoing method embodiments, which will not be repeated here.
[0111] The computer program products of the facial expression transfer method, apparatus, electronic device, and storage medium provided in this disclosure include a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.
[0112] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0113] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0114] In the description of this disclosure, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this disclosure and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this disclosure. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0115] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. A facial expression transfer method, characterized in that, The method of providing a driving face image and a source face image via an electronic device includes: Obtain the driving adjustment coefficients of the driving face image; wherein, the driving adjustment coefficients include driving expression coefficients for driving facial expression adjustment and driving pose coefficients for driving facial pose adjustment. Obtain the source face appearance identification coefficient of the source face image, the source face appearance identification coefficient is used to characterize the face appearance features in the source face image; The source face appearance identification coefficient and the driving adjustment coefficient are used to deform the preset average three-dimensional face model to obtain the target three-dimensional face model; A preset number of key points are sampled from the target 3D face model to obtain a 2D heat map corresponding to the key points; wherein, the 2D heat map includes the position of the key points in the target 3D face model and the expression information, posture information and appearance information corresponding to the position; The affine transformation coefficients are determined based on the source face image and the two-dimensional heatmap, and the source face image is subjected to affine transformation based on the affine transformation coefficients to obtain a facial expression transfer image. The step of obtaining the two-dimensional heatmap corresponding to the key point includes: projecting the key point onto a two-dimensional plane to obtain a two-dimensional image of the key point; and performing convolution processing on the two-dimensional image of the key point using a preset Gaussian convolution kernel to obtain a two-dimensional heatmap. The process of applying the source face appearance identification coefficient and the driving adjustment coefficient to deform a preset average 3D face model includes: applying the source face appearance identification coefficient to deform the appearance features of the preset average 3D face model; applying the driving expression coefficient in the driving adjustment coefficient to deform the expression features of the average 3D face model; and applying the driving pose coefficient in the driving adjustment coefficient to deform the head pose features of the average 3D face model.
2. The method according to claim 1, characterized in that, The preset average 3D face model is deformed using the source face appearance identification coefficient and the driving adjustment coefficient, including: The deformation process is performed using the following formula: ,in, This represents a 3D human face model. This represents a preset average 3D face model. This represents the preset face shape base. This represents the source face appearance identifier coefficient. This represents the preset facial expression base. Indicates the driving expression coefficient; and This represents the rotation and translation parameters in the driving attitude coefficients.
3. The method according to claim 1, characterized in that, The two-dimensional heat map corresponding to the key point is obtained using the following formula: ,in, This represents the orthogonal projection matrix, used to project the key points onto a two-dimensional plane. This represents the projection scaling factor of the source face image. Represents the Gaussian blur function. This indicates that n key points are sampled from the target 3D face model. The generated two-dimensional heatmap is n, which is a preset value.
4. The method according to claim 1, characterized in that, The affine transformation coefficients include scaling parameters, rotation parameters, and translation parameters.
5. The method according to claim 1, characterized in that, Determining the affine transformation coefficients based on the source face image and the two-dimensional heatmap includes: The source face image and the two-dimensional heatmap are convolved by a texture encoder to obtain feature maps corresponding to a first number of feature layers; wherein the feature maps of different feature layers correspond to different region features. The affine transformation coefficients of the feature map corresponding to each feature layer are determined by the transform encoder.
6. The method according to claim 5, characterized in that, The step of performing an affine transformation on the source face image based on the affine transformation coefficients to obtain a facial expression transfer image includes: For each feature map corresponding to a feature layer, an affine transform is performed on the feature map according to the affine transform coefficients of the feature map to obtain the affine transform map of the feature map. The facial expression transfer image is obtained by decoding the affine transformation map corresponding to each feature layer using a texture decoder.
7. The method according to claim 6, characterized in that, The affine transformation of the feature map using an affine transformer based on the affine transformation coefficients includes: The affine transformation coefficients of the feature map and the feature map itself are input into the affine transformer; Affine transformation is performed using the following formula within the affine transformer: ,in, , Represents the coordinates of pixels on the feature map. , This represents the coordinates of the pixel after the affine transformation. Let represent the affine transformation matrix consisting of affine transformation coefficients, where The scaling parameter in the affine transformation coefficients. The rotation parameter in the affine transformation coefficients, The x-axis translation parameter in the affine transformation coefficients. y-axis translation parameter in the affine transformation coefficients.
8. The method according to claim 1, characterized in that, Obtaining the driving adjustment coefficients of the driving face image includes: performing three-dimensional face reconstruction on the driving face image to obtain the driving adjustment coefficients; Obtaining the source face appearance identification coefficient of the source face image includes: performing the three-dimensional face reconstruction on the source face image to obtain the source face appearance identification coefficient.
9. An expression transfer device, characterized in that, The device provides a driving face image and a source face image via an electronic device, the device comprising: The first coefficient acquisition module is used to acquire the driving adjustment coefficients of the driving face image; wherein, the driving adjustment coefficients include driving expression coefficients for driving face expression adjustment and driving posture coefficients for driving face posture adjustment. The second coefficient acquisition module is used to acquire the source face appearance identification coefficient of the source face image, and the source face appearance identification coefficient is used to characterize the face appearance features in the source face image; The processing module is used to apply the source face appearance identification coefficient and the driving adjustment coefficient to deform the preset average three-dimensional face model to obtain the target three-dimensional face model; A heatmap acquisition module is used to sample a preset number of key points from the target 3D face model and acquire a two-dimensional heatmap corresponding to the key points; wherein, the two-dimensional heatmap includes the position of the key points in the target 3D face model and the expression information, posture information and appearance information corresponding to the position; wherein, acquiring the two-dimensional heatmap corresponding to the key points includes: projecting the key points onto a two-dimensional plane to obtain a two-dimensional image of the key points; performing convolution processing on the two-dimensional image of the key points using a preset Gaussian convolution kernel to obtain a two-dimensional heatmap; An affine transformation module is used to determine affine transformation coefficients based on the source face image and the two-dimensional heatmap, and to perform an affine transformation on the source face image based on the affine transformation coefficients to obtain a facial expression transfer image. The process of applying the source face appearance identification coefficient and the driving adjustment coefficient to deform a preset average 3D face model includes: applying the source face appearance identification coefficient to deform the appearance features of the preset average 3D face model; applying the driving expression coefficient in the driving adjustment coefficient to deform the expression features of the average 3D face model; and applying the driving pose coefficient in the driving adjustment coefficient to deform the head pose features of the average 3D face model.
10. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor to implement the method of any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when invoked and executed by a processor, cause the processor to perform the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Face 3D key point detection method and system based on joint thermodynamic diagram
CN110516643A
Facial expression migration method and device, electronic equipment and storage medium
CN113762147A