Fusion method and device for face images

By obtaining the attribute characteristics of the template face image and the identity characteristics of the user face image, using the attention map to distinguish areas and perform feature fusion, the problem of face swap effect limitation caused by relying on the face segmentation model in the existing technology is solved, and a more efficient face swap effect is achieved.

CN113762022BActive Publication Date: 2025-07-18BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110178117.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-09
Publication Date
2025-07-18
Estimated Expiration
2041-02-09

AI Technical Summary

Technical Problem

The existing intelligent face swap technology relies on face segmentation models, resulting in limited face swap effect and time-consuming training, especially when face segmentation effect is not good.

Method used

Using a method that does not rely on the face segmentation model, by obtaining the attribute characteristics of the template face image and the identity characteristics of the user face image, the attention map is used to distinguish the attribute stable area and the identity sensitive area, and perform identity migration and attribute recovery processing, and combine adaptive instance normalization operations to achieve feature fusion.

Benefits of technology

The face change effect is improved, the dependence on the face segmentation model is reduced, the training process is simplified, and the efficiency and effect of the face change process is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113762022B_ABST
    Figure CN113762022B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for fusing face images, relating to the field of image processing. The method includes obtaining the attribute features of a template face image; obtaining the identity features of a user face image; determining an attention map of the template face image based on the attribute features of the template face image, where the attention map is used to distinguish the attribute stable regions and identity sensitive regions in the template face image; taking the attribute features of the template face image in the attribute stable regions as the features of the attribute stable regions; performing identity migration and attribute restoration processing according to the attribute features of the template face image in the identity sensitive regions and the identity features of the user face image in the identity sensitive regions to obtain the features of the identity sensitive regions; and fusing the features of the attribute stable regions and the features of the identity sensitive regions by using the attention map to obtain a fused face image. The face swapping process does not require the assistance of a face segmentation model, improving the face swapping effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing, and particularly to a method and apparatus for fusing face images. Background Art

[0002] Using artificial intelligence technology, the face in image A is replaced with the face in image B, and the face in the synthesized image C is consistent with the face in image A, while the background, hairstyle, etc. are consistent with image B. Intelligent face replacement technology has a wide range of applications, such as movie post-production, privacy protection, etc. For example, a stuntman first completes difficult actions, and then the face of the protagonist is replaced with the face of the stuntman. Another example is that on some public websites, a virtual face is used to replace the real face of the user.

[0003] The face replacement effect of some related intelligent face replacement technologies highly depends on the face segmentation model. A face segmentation model with poor face segmentation effect will seriously affect the face replacement effect, while a face segmentation model with good face segmentation effect requires a large amount of labeled data for training, which is time-consuming and laborious. Summary of the Invention

[0004] Embodiments of the present disclosure propose an intelligent face replacement solution that does not require the assistance of a face segmentation model, improving the face replacement effect.

[0005] Some embodiments of the present disclosure propose a method for fusing face images, including:

[0006] Obtaining the attribute features of the template face image;

[0007] Obtaining the identity features of the user face image;

[0008] Based on the attribute features of the template face image, determining the attention map of the template face image, where the attention map is used to distinguish the attribute stable region and the identity sensitive region in the template face image;

[0009] Taking the attribute features of the template face image in the attribute stable region as the features of the attribute stable region;

[0010] Performing identity migration and attribute restoration processing according to the attribute features of the template face image in the identity sensitive region and the identity features of the user face image in the identity sensitive region to obtain the features of the identity sensitive region;

[0011] Using the attention map to fuse the features of the attribute stable region and the features of the identity sensitive region to obtain a fused face image.

[0012] In some embodiments, performing identity migration and attribute restoration processing according to the attribute features of the template face image in the identity sensitive region and the identity features of the user face image in the identity sensitive region to obtain the features of the identity sensitive region includes:

[0013] Calculate the mean and variance related to the identity features based on the identity features of the user's face image in the identity-sensitive region;

[0014] Calculate the mean and variance related to the attribute features based on the attribute features of the template face image in the identity-sensitive region;

[0015] Perform the first adaptive instance normalization operation according to the attribute features of the template face image in the identity-sensitive region and the mean and variance related to the identity features, and perform identity transfer;

[0016] Perform the second adaptive instance normalization operation according to the result of the first adaptive instance normalization operation and the mean and variance related to the attribute features, and perform attribute restoration to obtain the features of the identity-sensitive region.

[0017] In some embodiments, the features of the identity-sensitive region are represented as follows:

[0018]

[0019] Among them, F [Is-areas] represents the features of the identity-sensitive region, Conv represents the convolution operation, represents the normalization operation on Conv( ), F att represents the attribute features of the template face image in the identity-sensitive region, β id , γ id respectively represent the calculated mean and variance related to the identity features, β att , γ att respectively represent the calculated mean and variance related to the attribute features, represents the Hadamard product, represents the Hadamard sum.

[0020] In some embodiments, the features of the attribute-stable region and the features of the identity-sensitive region are fused using the attention map to obtain the fused face image, including:

[0021]

[0022] Among them, O represents the fused face image, M represents the attention map, F [As-areas] represents the features of the attribute-stable region, F [Is-areas] represents the features of the identity-sensitive region, represents the Hadamard product.

[0023] In some embodiments, the method recited in claim 1 is performed using a fusion network, which includes an attribute network and a face swapping network. The attribute network performs the step of obtaining the attribute features of the template face image, and the face swapping network performs the steps of identity feature acquisition, attention map determination, feature determination of the attribute stable region, feature determination of the identity sensitive region, and fused face image determination.

[0024] The attribute network includes one or more sequentially cascaded encoders for convolution and one or more sequentially cascaded decoders for deconvolution after the last encoder. The last encoder and each decoder output the attribute features of the template face image at their respective levels.

[0025] The face swapping network includes an identity feature acquisition module and one or more face swapping modules sequentially cascaded after the identity feature acquisition module. The two input terminals of the first face swapping module are respectively connected to the identity feature acquisition module and the last encoder, and the three input terminals of the other face swapping modules are respectively connected to the output terminals of the corresponding decoders, the output terminal of the identity feature acquisition module, and the output terminal of the previous face swapping module. The identity feature acquisition module is configured to obtain and output the identity features of the user's face image based on face recognition technology, and each face swapping module is configured to perform the steps of attention map determination, feature determination of the attribute stable region, feature determination of the identity sensitive region, and fused face image determination based on the input information of the current face swapping module.

[0026] In some embodiments, when the face swapping module performs the fused face image determination step, it includes:

[0027]

[0028] where O (l) represents the fused face image determined by the l-th face swapping module in the cascading order of the face swapping modules, O (l-1) represents the fused face image determined by the (l - 1)-th face swapping module in the cascading order of the face swapping modules, M (l) represents the attention map determined by the l-th face swapping module in the cascading order of the face swapping modules, represents the feature of the attribute stable region determined by the l-th face swapping module in the cascading order of the face swapping modules, represents the feature of the identity sensitive region determined by the l-th face swapping module in the cascading order of the face swapping modules, represents the Hadamard product.

[0029] In some embodiments, it further includes: jointly training the attribute network and the face swapping network using face image training data until the training loss meets the requirements.

[0030] Among them, the face image training data includes multiple groups of second template face images and second user face images, and the training loss includes the identity loss between the second user face image and the second fused face image, where the second fused face image is an image output by the face swapping network based on the second template face image and the second user face image.

[0031] In some embodiments, the training loss further includes at least one of the following:

[0032] The attribute consistency loss between the second fused face image and the second template face image;

[0033] The reconstruction loss of the second fused face image relative to the second template face image;

[0034] The adversarial loss between the second fused face image and a preset real face image;

[0035] The regularization constraint of the attention map.

[0036] In some embodiments, the attribute stable regions include the background and hair, the attribute features include the background, hairstyle, skin color, expression, and pose; the identity sensitive regions include the facial feature regions and contour regions of the face.

[0037] Some embodiments of the present disclosure propose a face image fusion device, including:

[0038] A memory; and

[0039] A processor coupled to the memory, the processor being configured to execute the face image fusion method according to any of the embodiments based on instructions stored in the memory.

[0040] Some embodiments of the present disclosure propose a face image fusion device, including:

[0041] An attribute network for obtaining the attribute features of the template face image;

[0042] A face swapping network for:

[0043] Obtaining the identity features of the user face image;

[0044] Determining the attention map of the template face image based on the attribute features of the template face image, where the attention map is used to distinguish the attribute stable regions and identity sensitive regions in the template face image;

[0045] Taking the attribute features of the template face image in the attribute stable regions as the features of the attribute stable regions;

[0046] Perform identity migration and attribute restoration processing based on the attribute features of the template face image in the identity-sensitive area and the identity features of the user face image in the identity-sensitive area to obtain the features of the identity-sensitive area;

[0047] Use the attention map to fuse the features of the attribute-stable area and the features of the identity-sensitive area to obtain a fused face image.

[0048] In some embodiments, the attribute network includes one or more encoders cascaded in sequence for performing convolution and one or more decoders cascaded in sequence after the last encoder for performing deconvolution. The last encoder and each decoder respectively output the attribute features of the template face image at their respective levels;

[0049] The face-swapping network includes an identity feature acquisition module and one or more face-swapping modules cascaded in sequence after the identity feature acquisition module. The two input terminals of the first face-swapping module are respectively connected to the identity feature acquisition module and the last encoder, and the three input terminals of the other face-swapping modules are respectively connected to the output terminals of the corresponding decoders, the output terminal of the identity feature acquisition module, and the output terminal of the previous face-swapping module. The identity feature acquisition module is configured to acquire and output the identity features of the user face image based on face recognition technology, and each face-swapping module is configured to perform an attention map determination step, an attribute-stable area feature determination step, an identity-sensitive area feature determination step, and a fused face image determination step based on its input information.

[0050] In some embodiments, each face-swapping module includes:

[0051] An attention map determination unit for determining the attention map of the template face image based on the attribute features of the template face image input to this face-swapping module. The attention map is used to distinguish the attribute-stable area and the identity-sensitive area in the template face image;

[0052] An attribute-stable area feature determination unit for using the attribute features of the template face image in the attribute-stable area input to this face-swapping module as the features of the attribute-stable area;

[0053] An identity-sensitive area feature determination unit for performing identity migration and attribute restoration processing based on the attribute features of the template face image in the identity-sensitive area input to this face-swapping module and the identity features of the user face image in the identity-sensitive area to obtain the features of the identity-sensitive area;

[0054] A fused face image determination unit for using the attention map to fuse the features of the attribute-stable area and the features of the identity-sensitive area to obtain a fused face image;

[0055] Among them, the identity-sensitive area feature determination unit includes:

[0056] The first adaptive instance normalization unit is configured to calculate the mean and variance related to the identity features based on the identity features of the user's face image in the identity-sensitive region; perform the first adaptive instance normalization operation according to the attribute features of the template face image in the identity-sensitive region and the mean and variance related to the identity features, and perform identity migration.

[0057] The second adaptive instance normalization unit is configured to calculate the mean and variance related to the attribute features based on the attribute features of the template face image in the identity-sensitive region; perform the second adaptive instance normalization operation according to the result of the first adaptive instance normalization operation and the mean and variance related to the attribute features, and perform attribute restoration to obtain the features of the identity-sensitive region.

[0058] Some embodiments of the present disclosure propose a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the face image fusion method described in any embodiment are implemented.

[0059] In the embodiments of the present disclosure, the face image is softly partitioned through the attention mechanism to distinguish the attribute-stable region and the identity-sensitive region therein. The attribute features of the template face image in the attribute-stable region are used as the features of the attribute-stable region. Identity migration and attribute restoration processing are performed according to the attribute features of the template face image in the identity-sensitive region and the identity features of the user's face image in the identity-sensitive region to obtain the features of the identity-sensitive region. The attention map is used to fuse the features of the attribute-stable region and the features of the identity-sensitive region to obtain the fused face image. The face swapping process does not require the assistance of a face segmentation model, improving the face swapping effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. The present disclosure can be more clearly understood according to the following detailed description with reference to the drawings.

[0061] Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0062] Figure 1a A schematic diagram of a fusion network (also referred to as a fusion device) for fusing face images according to some embodiments of the present disclosure is shown.

[0063] Figure 1b A schematic diagram of a face swapping module according to some embodiments of the present disclosure is shown.

[0064] Figure 2 A flowchart of the face image fusion method according to some embodiments of the present disclosure is shown.

[0065] Figure 3 Schematic structural diagram of a face image fusion device according to some embodiments of the present disclosure. Detailed implementation manners

[0066] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure.

[0067] Unless otherwise specified, the descriptions such as "first", "second", etc. in the present disclosure are used to distinguish different objects and do not represent meanings such as size or time sequence.

[0068] Some embodiments of the present disclosure propose an intelligent face swapping solution for fusing face images based on a fusion network. The face swapping process does not require the assistance of a face segmentation model, which improves the face swapping effect.

[0069] Figure 1a Schematic diagram of a fusion network (also referred to as a fusion device) for fusing face images according to some embodiments of the present disclosure. That is, the fusion network can be used as a fusion device for fusing face images.

[0070] As Figure 1a shown, the fusion network (or, the fusion device) of this embodiment includes: an attribute network 110 (denoted as AttNet) and a face swapping network 120 (denoted as SwapNet).

[0071] The attribute network 110 is used to obtain the attribute features of a template face image (denoted as Ir). The attribute features include background, hairstyle, skin color, expression, posture, etc. The attribute network 110 includes one or more sequentially cascaded encoders 111 for performing convolution and one or more sequentially cascaded decoders 112 for performing deconvolution after the last-level encoder 111. The last-level encoder 111 and each decoder 112 respectively output the attribute features of the template face image at their respective levels. According to the cascading order, the sizes of the image features output by each encoder 111 become smaller and the number of channels becomes larger, and the sizes of the image features output by each decoder 112 become larger and the number of channels becomes smaller.

[0072] The face swapping network 120 is used to obtain the identity features of the user's face image (set as Is) by using face recognition tools such as InsightFace and DeepFace. The identity features include facial feature features, facial contour features, etc. Based on the attribute features of the template face image, the attention map of the template face image is determined by using convolutional operations. The attention map (also known as the swapping attention map) is used to distinguish the attribute stable area and the identity sensitive area in the template face image. The attribute features of the template face image in the attribute stable area are used as the features of the attribute stable area. Identity migration and attribute restoration processing are performed according to the attribute features of the template face image in the identity sensitive area and the identity features of the user's face image in the identity sensitive area to obtain the features of the identity sensitive area. The attention map is used to fuse the features of the attribute stable area and the features of the identity sensitive area to obtain the fused face image.

[0073] The attribute stable area includes the background and hair. The identity sensitive area includes the facial feature area and the contour area of the face. "Facial features" refer to the five facial features that affect appearance, namely "eyebrows, eyes, ears, nose, and mouth".

[0074] Fusing the features of the attribute stable area and the features of the identity sensitive area by using the attention map to obtain the fused face image includes:

[0075]

[0076] Among them, O represents the fused face image, M represents the attention map, F [As-areas] represents the features of the attribute stable area, F [Is-areas] represents the features of the identity sensitive area, represents the Hadamard product. So that the identity information in the swapped fused face image is consistent with the user's face image, and the attribute information in the swapped fused face image is consistent with the template face image.

[0077] The face swapping network 120 includes an identity feature acquisition module 121 and one or more face swapping modules 122 cascaded in sequence behind the identity feature acquisition module 121. The two input ends of the first face swapping module 122 are respectively connected to the identity feature acquisition module 121 and the last-stage encoder 111. The three input ends of other face swapping modules 122 are respectively connected to the output end of the corresponding decoder 112, the output end of the identity feature acquisition module 121, and the output end of the previous-stage face swapping module 122. The identity feature acquisition module 121 is configured to acquire and output the identity features of the user's face image based on face recognition technology. Each face swapping module 122 is configured to perform an attention map determination step, an attribute stable area feature determination step, an identity sensitive area feature determination step, and a fused face image determination step based on the input information of this face swapping module 122.

[0078] The identity feature acquisition module 121 uses face recognition tools such as InsightFace and DeepFace to acquire and output the identity features of the user's face image.

[0079] A face swapping module 122 performs the steps of determining the fused face image, including:

[0080]

[0081] Among them, O (l) represents the fused face image determined by the l-th face swapping module 122 in the cascading order of each face swapping module 122, and O (l-1) represents the fused face image determined by the (l - 1)-th face swapping module 122 in the cascading order of each face swapping module 122, and M (l) represents the attention map determined by the l-th face swapping module 122 in the cascading order of each face swapping module 122, represents the features of the attribute stable region determined by the l-th face swapping module 122 in the cascading order of each face swapping module 122, represents the features of the identity sensitive region determined by the l-th face swapping module 122 in the cascading order of each face swapping module 122, represents the Hadamard product.

[0082] It should be noted that the identity features of the user's face images input to each face swapping module 122 are the same, and the attribute features of the template face images input to each face swapping module 122 are different, and are different in terms of feature size and number of channels, and respectively come from the attribute features of the template face images output by the last encoder 111 or decoder 112 connected to it.

[0083] In Figure 1a exemplarily shows 4 encoders 111, 3 decoders 112, and 4 face swapping modules 122. However, the present disclosure is not limited to the above examples, and the number of encoders 111, the number of decoders 112, and the number of face swapping modules 122 can also be set to other values.

[0084] Figure 1b The schematic diagram of a face swapping module 122 showing some embodiments of the present disclosure. The structures and functions of each face swapping module 122 are the same, Figure 1b The schematic diagram of one of the face swapping modules 122 is shown. The structures and functions of the other face swapping modules 122 are the same as those of this face swapping module 122, and will not be described in detail hereinafter.

[0085] As Figure 1b shown, the face swapping module 122 of this embodiment includes:

[0086] An attention map determination unit 1221, configured to determine an attention map of a template face image based on the attribute features of the template face image input to the face swapping module 122, where the attention map is used to distinguish the attribute stable region and the identity sensitive region in the template face image;

[0087] An attribute stable region feature determination unit 1222, configured to use the attribute features of the template face image input to the face swapping module 122 in the attribute stable region as the features of the attribute stable region;

[0088] An identity sensitive region feature determination unit 1223, configured to perform identity migration and attribute restoration processing based on the attribute features of the template face image input to the face swapping module 122 in the identity sensitive region and the identity features of the user face image in the identity sensitive region, to obtain the features of the identity sensitive region;

[0089] A fused face image determination unit 1224, configured to fuse the features of the attribute stable region and the features of the identity sensitive region by using the attention map to obtain a fused face image.

[0090] Among them, the identity sensitive region feature determination unit 1223 includes:

[0091] A first Adaptive Instance Normalization (AdaIN) unit 1223a, configured to calculate the mean and variance related to the identity features based on the identity features of the user face image in the identity sensitive region; perform a first adaptive instance normalization operation according to the attribute features of the template face image in the identity sensitive region and the mean and variance related to the identity features, for identity migration.

[0092] A second adaptive instance normalization unit 1223b, configured to calculate the mean and variance related to the attribute features based on the attribute features of the template face image in the identity sensitive region; perform a second adaptive instance normalization operation according to the result of the first adaptive instance normalization operation and the mean and variance related to the attribute features, for attribute restoration, to obtain the features of the identity sensitive region.

[0093] The features of the identity sensitive region are represented as follows:

[0094]

[0095] Among them, F [Is-areas] represents the features of the identity sensitive region, that is (the features of the identity sensitive region of the l-th face swapping module), the company omits the annotation of the face swapping module, Conv represents the convolution operation, represents performing a normalization operation on Conv( ), F attRepresents the attribute features of the template face image in the identity-sensitive region, β id and γ id respectively represent the mean and variance related to the identity features calculated, β att and γ att respectively represent the mean and variance related to the attribute features calculated, represents the Hadamard product, represents the Hadamard sum. represents the first adaptive instance normalization operation. represents the second adaptive instance normalization operation.

[0096] After constructing the fusion network, it is also necessary to train the fusion network, that is, jointly train the attribute network 110 and the face swapping network 120 using the face image training data until the training loss meets the requirements. For example, the training loss is less than a preset value, so that the trained fusion network can fuse face images. During the training process, the parameters of the fusion network are continuously updated, such as the parameters of the encoder 111, the parameters of the decoder 112, the convolutional parameters for generating the attention map, etc.

[0097] Among them, the face image training data includes multiple groups of second template face images and second user face images. The training loss includes the identity loss between the second user face image and the second fused face image, and the second fused face image is the image output by the face swapping network 120 based on the second template face image and the second user face image. The training loss also includes at least one of the following: the attribute consistency loss between the second fused face image and the second template face image; the reconstruction loss of the second fused face image relative to the second template face image; the adversarial loss between the second fused face image and a preset real face image (which can be the second template face image, the second user face image, or other real face images obtained by shooting); the regularization constraint of the attention map.

[0098] The training loss consists of 5 parts, expressed as:

[0099] L total = λ1L id + λ2L att + λ3L rec + λ4L adv + λ5L reg

[0100] Among them: represents the identity loss between the second user face image and the second fused face image, calculated using the cosine distance, is the identity feature of the second user face image, is the identity feature of the generated second fused face image.

[0101] represents the attribute consistency loss between the second fused face image and the second template face image. is the attribute feature of the second template face image input to the l-th face swapping module 122. is the attribute feature of the second fused face image output by the l-th face swapping module 122, and L is the total number of face swapping modules 122.

[0102] The reconstruction loss L rec is:

[0103]

[0104] where M L represents the attention map learned by the last face swapping module 122, I s is the second user face image used during training, I r is the second template face image used during training, and O is the second fused face image output during training.

[0105] L adv represents the adversarial loss between the second fused face image and a preset real face image (which can be the second template face image, the second user face image, or other real face images obtained by shooting), and the adversarial loss can refer to the prior art.

[0106] represents the regularization constraint on the attention map, which can make the face swapping network 120 put more effort on identity-sensitive regions. M l represents the attention map learned by the l-th face swapping module 122.

[0107] λ1 to λ5 represent the weighting coefficients of each loss and can be set.

[0108] From the training process, it can be seen that although there is no labeled information, the self-supervised training method is used to adaptively learn the attention map, and the attention map can effectively distinguish the identity-sensitive regions (mostly concentrated in the facial feature regions and contour regions) and the attribute-stable regions (background, hair, etc. regions). The identity information in the fused face image after face swapping is consistent with the user face image, and the attribute information in the fused face image after face swapping is consistent with the template face image.

[0109] According to the cascading sequence, the value of the attention map will gradually increase. The value of the previous attention map is relatively small, indicating that the face swapping network 120 pays more attention to learning the overall attribute features, such as background, character attributes, etc. in the early stage. And for the later attention maps, as the resolution increases, the saliency of the identity-sensitive regions will increase accordingly.

[0110] After constructing the fusion network for face images and training it using self-supervised learning methods, the fusion network can be used to fuse face images.

[0111] Figure 2 It is a schematic flowchart of the method for fusing face images according to some embodiments of the present disclosure.

[0112] As Figure 2 shown, the method for fusing face images in this embodiment includes steps 210-260. Among them, step 210 is executed by the attribute network 110, steps 220-260 are executed by the face swapping network 120, and steps 230-260 are executed by each face swapping module 122 in the face swapping network 120.

[0113] In step 210, obtain the attribute features of the template face image.

[0114] Each encoder 111 in the attribute network 110 performs convolutional processing on the template face image, and each decoder 112 performs deconvolutional processing on the template face image. The last-level encoder 111 and each decoder 112 respectively output the attribute features of the template face image at their respective levels. According to the cascading order, the size of the image features output by each encoder 111 becomes smaller and the number of channels becomes larger, and the size of the image features output by each decoder 112 becomes larger and the number of channels becomes smaller.

[0115] In step 220, obtain the identity features of the user's face image.

[0116] The identity feature acquisition module 121 in the face swapping network 120 uses face recognition tools such as InsightFace and DeepFace to obtain and output the identity features of the user's face image.

[0117] In step 230, based on the attribute features of the template face image, determine the attention map of the template face image, and the attention map is used to distinguish the attribute stable region and the identity sensitive region in the template face image.

[0118] Perform a convolutional operation on the attribute features of the template face image to obtain the attention map of the template face image.

[0119] In step 240, use the attribute features of the template face image in the attribute stable region as the features of the attribute stable region.

[0120] In step 250, perform identity transfer and attribute restoration processing according to the attribute features of the template face image in the identity sensitive region and the identity features of the user's face image in the identity sensitive region to obtain the features of the identity sensitive region.

[0121] Determining the features of the identity sensitive region includes:

[0122] Calculate the mean and variance related to the identity features based on the identity features of the user's face image in the identity-sensitive region;

[0123] Calculate the mean and variance related to the attribute features based on the attribute features of the template face image in the identity-sensitive region;

[0124] Perform the first adaptive instance normalization operation according to the attribute features of the template face image in the identity-sensitive region and the mean and variance related to the identity features, and perform identity transfer;

[0125] Perform the second adaptive instance normalization operation according to the result of the first adaptive instance normalization operation and the mean and variance related to the attribute features, and perform attribute restoration to obtain the features of the identity-sensitive region.

[0126] The features of the identity-sensitive region are represented as follows:

[0127]

[0128] where F [Is-areas] represents the features of the identity-sensitive region, Conv represents the convolution operation, represents the normalization operation on Conv( ), F att represents the attribute features of the template face image in the identity-sensitive region, β id , γ id respectively represent the calculated mean and variance related to the identity features, β att , γ att respectively represent the calculated mean and variance related to the attribute features, represents the Hadamard product, represents the Hadamard sum.

[0129] represents the first adaptive instance normalization operation. represents the second adaptive instance normalization operation.

[0130] In step 260, fuse the features of the attribute-stable region and the features of the identity-sensitive region using the attention map to obtain a fused face image.

[0131] Fusing the features of the attribute-stable region and the features of the identity-sensitive region using the attention map to obtain a fused face image includes:

[0132]

[0133] where O represents the fused face image, M represents the attention map, F [As-areas] represents the features of the attribute-stable region, F [Is-areas] represents the features of the identity-sensitive region, Denotes the Hadamard product.

[0134] When there are multiple face swapping modules 122, each face swapping module 122 fuses the features of the attribute stable region and the identity sensitive region using the attention map, and the obtained fused face image includes:

[0135]

[0136] Where, O (l) Denotes the fused face image determined by the l-th face swapping module 122 in the cascading order of each face swapping module 122, and O (l-1) Denotes the fused face image determined by the (l - 1)-th face swapping module 122 in the cascading order of each face swapping module 122, and M (l) Denotes the attention map determined by the l-th face swapping module 122 in the cascading order of each face swapping module 122, Denotes the features of the attribute stable region determined by the l-th face swapping module 122 in the cascading order of each face swapping module 122, Denotes the features of the identity sensitive region determined by the l-th face swapping module 122 in the cascading order of each face swapping module 122, Denotes the Hadamard product.

[0137] Through the attention mechanism, the face image is softly partitioned to distinguish the attribute stable region and the identity sensitive region therein. The attribute features of the template face image in the attribute stable region are used as the features of the attribute stable region. Identity transfer and attribute restoration processing are performed based on the attribute features of the template face image in the identity sensitive region and the identity features of the user face image in the identity sensitive region to obtain the features of the identity sensitive region. The features of the attribute stable region and the features of the identity sensitive region are fused using the attention map to obtain the fused face image. The face swapping process does not require the assistance of a face segmentation model, improving the face swapping effect.

[0138] Figure 3 It is a schematic structural diagram of a fusion device for face images according to some embodiments of the present disclosure.

[0139] As Figure 3 shown, the device 300 of this embodiment includes: a memory 310 and a processor 320 coupled to the memory 310. The processor 320 is configured to execute the methods in any of the foregoing embodiments based on the instructions stored in the memory 310.

[0140] For example, obtain the attribute features of a template face image; obtain the identity features of a user face image; based on the attribute features of the template face image, determine an attention map of the template face image, where the attention map is used to distinguish the attribute stable region and the identity sensitive region in the template face image; use the attribute features of the template face image in the attribute stable region as the features of the attribute stable region; perform identity migration and attribute restoration processing according to the attribute features of the template face image in the identity sensitive region and the identity features of the user face image in the identity sensitive region to obtain the features of the identity sensitive region; use the attention map to fuse the features of the attribute stable region and the features of the identity sensitive region to obtain a fused face image.

[0141] Among them, the memory 310 can include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a boot loader, and other programs.

[0142] The device 300 may further include an input / output interface 330, a network interface 340, a storage interface 350, etc. These interfaces 330, 340, 350 and the memory 310 and the processor 320 can be connected through a bus 360, for example. Among them, the input / output interface 330 provides a connection interface for input / output devices such as a display, a mouse, a keyboard, and a touch screen. The network interface 340 provides a connection interface for various networking devices. The storage interface 350 provides a connection interface for external storage devices such as an SD card and a USB flash drive.

[0143] For example, obtain the attribute features of a template face image; obtain the identity features of a user face image; based on the attribute features of the template face image, determine an attention map of the template face image, where the attention map is used to distinguish the attribute stable region and the identity sensitive region in the template face image; use the attribute features of the template face image in the attribute stable region as the features of the attribute stable region; perform identity migration and attribute restoration processing according to the attribute features of the template face image in the identity sensitive region and the identity features of the user face image in the identity sensitive region to obtain the features of the identity sensitive region; use the attention map to fuse the features of the attribute stable region and the features of the identity sensitive region to obtain a fused face image.

[0144] Some embodiments of the present disclosure propose a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method for fusing face images are implemented.

[0145] For example, obtain the attribute features of the template face image; obtain the identity features of the user face image; based on the attribute features of the template face image, determine the attention map of the template face image, where the attention map is used to distinguish the attribute stable region and the identity sensitive region in the template face image; use the attribute features of the template face image in the attribute stable region as the features of the attribute stable region; perform identity migration and attribute restoration processing according to the attribute features of the template face image in the identity sensitive region and the identity features of the user face image in the identity sensitive region to obtain the features of the identity sensitive region; use the attention map to fuse the features of the attribute stable region and the features of the identity sensitive region to obtain the fused face image.

[0146] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more non-transitory computer-readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer program code.

[0147] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0148] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufacture including an instruction device that realizes the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0149] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0150] The above are only the preferred embodiments of the present disclosure, and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for fusing face images, characterized in that, Including: Obtain the attribute features of the template face image, where the attribute features include at least one of background, hairstyle, skin color, expression, and pose; Obtain the identity features of the user's face image; Based on the attribute features of the template face image, use convolution operations to determine the attention map of the template face image. The attention map is used to perform soft partitioning on the template face image to distinguish the attribute stable region and the identity sensitive region in the template face image. The attribute stable region includes at least one of background and hair; Use the attribute features of the template face image in the attribute stable region as the features of the attribute stable region; Perform identity transfer and attribute restoration processing according to the attribute features of the template face image in the identity sensitive region and the identity features of the user's face image in the identity sensitive region to obtain the features of the identity sensitive region; Use the attribute stable region and the identity sensitive region distinguished by the attention map to fuse the features of the attribute stable region and the features of the identity sensitive region to obtain a fused face image; The fusion network includes an attribute network and a face swapping network. The attribute network performs the step of obtaining the attribute features of the template face image, and the face swapping network performs the steps of identity feature acquisition, attention map determination, determination of the features of the attribute stable region, determination of the features of the identity sensitive region, and determination of the fused face image; Jointly train the attribute network and the face swapping network using face image training data. The training loss includes the regularization constraint of the attention map, so that the face swapping network devotes more effort to the identity sensitive region.

2. The method according to claim 1, wherein The features of the identity sensitive region are represented as follows: Among them, F [Is-areas] represents the feature of the identity-sensitive region, Conv represents the convolution operation, represents the operation of for normalization, F att represents the attribute feature of the template face image in the identity-sensitive region, β id , γ id respectively represent the mean and variance related to the calculated identity feature, β att , γ att respectively represent the mean and variance related to the calculated attribute feature, represents the Hadamard product, represents the Hadamard sum.

3. The method according to claim 1, wherein Using the attention map to fuse the features of the attribute stable region and the features of the identity sensitive region to obtain a fused face image includes: Among them, represents the fused face image, M represents the attention map, represents the features of the attribute stable region, represents the features of the identity sensitive region, represents the Hadamard product.

4. The method according to claim 1, wherein: The attribute network includes one or more sequentially cascaded encoders for performing convolution and one or more decoders for performing deconvolution sequentially cascaded after the last encoder. The last encoder and each decoder respectively output the attribute features of the template face image at their respective levels; Wherein, the face swapping network includes an identity feature acquisition module and one or more face swapping modules sequentially cascaded after the identity feature acquisition module. The two input ends of the first face swapping module are respectively connected to the identity feature acquisition module and the last encoder, and the three input ends of other face swapping modules are respectively connected to the output end of the corresponding decoder, the output end of the identity feature acquisition module, and the output end of the previous face swapping module. The identity feature acquisition module is configured to obtain and output the identity features of the user's face image based on face recognition technology, and each face swapping module is configured to perform the steps of attention map determination, determination of the features of the attribute stable region, determination of the features of the identity sensitive region, and determination of the fused face image based on the input information of the current face swapping module.

5. The method according to claim 4, wherein The face swapping module performing the step of determining the fused face image includes: Among them, O (l) represents the fused face image determined by the l-th face-swapping module in the cascading order of each face-swapping module, O (l -1) represents the fused face image determined by the (l - 1)-th face-swapping module in the cascading order of each face-swapping module, M (l) represents the attention map determined by the l-th face-swapping module in the cascading order of each face-swapping module, represents the feature of the attribute stable region determined by the l-th face-swapping module in the cascading order of each face-swapping module, represents the feature of the identity sensitive region determined by the l-th face-swapping module in the cascading order of each face-swapping module, represents the Hadamard product.

6. The method according to claim 4, wherein The face image training data includes multiple groups of second template face images and second user face images, and the training loss includes the identity loss between the second user face image and the second fused face image, where the second fused face image is an image output by the face swapping network based on the second template face image and the second user face image.

7. The method according to claim 6, wherein The training loss further includes at least one of the following: The attribute consistency loss between the second fused face image and the second template face image; The reconstruction loss of the second fused face image relative to the second template face image; The adversarial loss between the second fused face image and a preset real face image.

8. The method according to any one of claims 1-7, wherein The identity sensitive region includes the facial feature regions and the contour region of the face.

9. The method according to any one of claims 1-7, characterized in that, Performing identity transfer and attribute restoration processing based on the attribute features of the template face image in the identity sensitive region and the identity features of the user face image in the identity sensitive region to obtain the features of the identity sensitive region, including: Calculating the mean and variance related to the identity features based on the identity features of the user face image in the identity sensitive region; calculating the mean and variance related to the attribute features based on the attribute features of the template face image in the identity sensitive region; performing a first adaptive instance normalization operation according to the attribute features of the template face image in the identity sensitive region and the mean and variance related to the identity features to perform identity transfer; and performing a second adaptive instance normalization operation according to the result of the first adaptive instance normalization operation and the mean and variance related to the attribute features to perform attribute restoration, so as to obtain the features of the identity sensitive region.

10. A face image fusion device, comprising: A memory; And A processor coupled to the memory, the processor being configured to execute the face image fusion method according to any one of claims 1-9 based on instructions stored in the memory.

11. A face image fusion device, comprising: An attribute network for obtaining the attribute features of the template face image, where the attribute features include at least one of background, hairstyle, skin color, expression, and pose; A face swapping network for: Obtaining the identity features of the user face image; Based on the attribute features of the template face image, using a convolution operation to determine an attention map of the template face image, where the attention map is used for soft partitioning of the template face image to distinguish the attribute stable region and the identity sensitive region in the template face image; Regarding the attribute features of the template face image in the attribute stable region as the features of the attribute stable region, where the attribute stable region includes at least one of background and hair; Performing identity transfer and attribute restoration processing according to the attribute features of the template face image in the identity sensitive region and the identity features of the user face image in the identity sensitive region to obtain the features of the identity sensitive region; Using the attribute stable region and the identity sensitive region distinguished by the attention map to fuse the features of the attribute stable region and the features of the identity sensitive region to obtain a fused face image. The fusion network includes an attribute network and a face-swapping network. The attribute network performs the step of obtaining the attribute features of the template face image, and the face-swapping network performs the steps of identity feature acquisition, attention map determination, feature determination of the attribute stable region, feature determination of the identity-sensitive region, and fused face image determination. The attribute network and the face-swapping network are jointly trained using face image training data. The training loss includes the regularization constraint of the attention map, so that the face-swapping network devotes more effort to the identity-sensitive region.

12. The device according to claim 11, wherein The face-swapping network includes an identity feature acquisition module and one or more face-swapping modules cascaded in sequence behind the identity feature acquisition module. The two input ends of the first face-swapping module are respectively connected to the identity feature acquisition module and the last-stage encoder, and the three input ends of the other face-swapping modules are respectively connected to the output end of the corresponding decoder, the output end of the identity feature acquisition module, and the output end of the previous face-swapping module. The identity feature acquisition module is configured to acquire and output the identity features of the user's face image based on face recognition technology, and each face-swapping module is configured to perform the steps of attention map determination, feature determination of the attribute stable region, feature determination of the identity-sensitive region, and fused face image determination based on its own input information.

13. The device according to claim 12, wherein Each face-swapping module includes: An attention map determination unit, configured to determine the attention map of the template face image based on the attribute features of the template face image input to this face-swapping module. The attention map is used to distinguish the attribute stable region and the identity-sensitive region in the template face image; A feature determination unit for the attribute stable region, configured to use the attribute features of the template face image in the attribute stable region input to this face-swapping module as the features of the attribute stable region; A feature determination unit for the identity-sensitive region, configured to perform identity transfer and attribute restoration processing according to the attribute features of the template face image in the identity-sensitive region and the identity features of the user's face image in the identity-sensitive region to obtain the features of the identity-sensitive region; A fused face image determination unit, configured to fuse the features of the attribute stable region and the features of the identity-sensitive region using the attention map to obtain a fused face image; Among them, the feature determination unit for the identity-sensitive region includes: A first adaptive instance normalization unit, configured to calculate the mean and variance related to the identity features based on the identity features of the user's face image in the identity-sensitive region; perform the first adaptive instance normalization operation according to the attribute features of the template face image in the identity-sensitive region and the mean and variance related to the identity features to perform identity transfer; A second adaptive instance normalization unit, configured to calculate the mean and variance related to the attribute features based on the attribute features of the template face image in the identity-sensitive region; perform the second adaptive instance normalization operation according to the result of the first adaptive instance normalization operation and the mean and variance related to the attribute features to perform attribute restoration and obtain the features of the identity-sensitive region.

14. A non-transitory computer-readable storage medium storing a computer program which, when executed by a processor, implements the steps of the method for fusing face images according to any one of claims 1-9.

Citation Information

Patent Citations

  • Generative adversarial network training method, image face changing and video face changing method and device

    CN111783603A