Facial attribute editing method and apparatus, device, storage medium, and product

By extracting region images from the image to be processed and generating masks for fusion processing, the problem of unrealistic and unstable facial attribute editing effects in the existing technology is solved, achieving more realistic and stable facial attribute editing effects, especially reducing inter-frame jitter in videos.

WO2026103741A1PCT designated stage Publication Date: 2026-05-21BEIJINGLUOTA INFORMATION TECHNOLOGYCO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJINGLUOTA INFORMATION TECHNOLOGYCO LTD
Filing Date
2025-11-12
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing technologies have poor facial attribute editing effects, resulting in unrealistic and noisy images. In particular, facial attribute editing is unstable in videos and prone to inter-frame jitter.

Method used

By extracting the region image corresponding to the preset face attributes from the image to be processed, generating the generated image with the preset face attribute editing effect and the face attribute mask, and using the mask as the fusion weight to fuse the region image and the generated image, a region fusion image is obtained, which finally replaces the corresponding region of the image to be processed.

Benefits of technology

It effectively reduces noise in generated images, improves the realism and stability of facial attribute editing, and maintains good temporal stability in videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025134411_21052026_PF_FP_ABST
    Figure CN2025134411_21052026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a facial attribute editing method and apparatus, a device, a storage medium, and a product. In a technical solution provided in the embodiments of the present application, a regional image corresponding to a preset facial attribute is extracted from an image to be processed, a generated image and a facial attribute mask corresponding to a preset facial attribute editing effect are generated on the basis of the regional image, the facial attribute mask is used as a fusion weight for performing fusion processing on the regional image and the generated image, to obtain a regional fusion image, and a corresponding region in the image to be processed is replaced with the regional fusion image, to obtain a target image. This can effectively reduce noise in a generated image, making the effect of facial attribute editing on the image realistic, and improving the effect of facial attribute editing.
Need to check novelty before this filing date? Find Prior Art

Description

A method, apparatus, device, storage medium, and product for editing facial attributes.

[0001] This application claims priority to Chinese Patent Application No. 202411642851.6, filed on November 18, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of image processing technology, and in particular to a method, apparatus, device, storage medium, and product for editing facial attributes. Background Technology

[0003] With the development of internet technology and image processing technology, people are using digital entertainment products more and more widely, and digital entertainment products are offering more and more functions and ways to play.

[0004] For example, video applications often offer facial attribute editing features, such as transforming facial attributes in videos (e.g., hair color, expression, beard, age, etc.). Traditional facial attribute editing schemes typically input the entire face image into a Generative Adversarial Network (GAN), which then outputs the transformed face image. Because GANs usually predict images by learning pixel-to-pixel mappings, they are prone to producing unrealistic and noisy images, resulting in poor facial attribute editing performance. Summary of the Invention

[0005] This application provides a method, apparatus, device, storage medium, and product for editing facial attributes, in order to solve the technical problem of poor facial attribute editing effect in related technologies, reduce noise in generated images, and make the facial attribute editing effect of images more realistic, thereby effectively improving the facial attribute editing effect.

[0006] In a first aspect, embodiments of this application provide a method for editing facial attributes, including:

[0007] Acquire an image to be processed, and extract the region image corresponding to the preset face attributes from the image to be processed;

[0008] Generate a generated image and a facial attribute mask corresponding to the preset facial attribute editing effect based on the region image;

[0009] The facial attribute mask is used as a fusion weight to fuse the region image and the generated image to obtain a region fused image;

[0010] The target image is obtained by replacing the corresponding region in the image to be processed with the fused image of the region.

[0011] In a second aspect, embodiments of this application provide a face attribute editing device, including an image extraction module, an image processing module, an image fusion module, and an image generation module, wherein:

[0012] The image extraction module is configured to acquire an image to be processed and extract a region image from the image to be processed.

[0013] The image processing module is configured to generate a generated image and a face attribute mask corresponding to a preset face attribute editing effect based on the region image.

[0014] The image fusion module is configured to use the face attribute mask as a fusion weight to perform fusion processing on the region image and the generated image to obtain a region fusion image.

[0015] The image generation module is configured to replace the corresponding region in the image to be processed with the region fusion image to obtain the target image.

[0016] In a third aspect, embodiments of this application provide a face attribute editing device, including: a memory and one or more processors;

[0017] The memory is used to store one or more programs;

[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the face attribute editing method as described in the first aspect.

[0019] In a fourth aspect, embodiments of this application provide a non-volatile storage medium for storing computer-executable instructions, which, when executed by a computer processor, are used to perform the face attribute editing method as described in the first aspect.

[0020] In a fifth aspect, embodiments of this application provide a computer program product comprising a computer program stored in a computer-readable storage medium, wherein at least one processor of the device reads from the computer-readable storage medium and executes the computer program, causing the device to perform the face attribute editing method as described in the first aspect.

[0021] This application embodiment extracts a region image corresponding to a preset face attribute from the image to be processed, generates a generated image and a face attribute mask corresponding to the preset face attribute editing effect based on the region image, and uses the face attribute mask as a fusion weight to fuse the region image and the generated image to obtain a region fusion image. The region fusion image is then used to replace the corresponding region in the image to be processed to obtain the target image. This can effectively reduce the noise of the generated image, make the face attribute editing effect of the image more realistic, and improve the face attribute editing effect. Attached Figure Description

[0022] Figure 1 is a flowchart of a face attribute editing method provided in an embodiment of this application;

[0023] Figure 2 is a flowchart of another face attribute editing method provided in an embodiment of this application;

[0024] Figure 3 is a schematic diagram of a face attribute editing device provided in an embodiment of this application;

[0025] Figure 4 is a structural schematic diagram of a face attribute editing device provided in an embodiment of this application. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this application clearer, specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely for explaining this application and not for limiting it. It should also be noted that, for ease of description, only the parts relevant to this application are shown in the drawings, not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but additional steps not included in the drawings may also be present. The above processes can correspond to methods, functions, procedures, subroutines, subroutines, etc.

[0027] The facial attribute editing method provided in this application can be applied to facial editing of images, such as in mobile live streaming, video editing, and image editing scenarios. It enables the editing of facial attributes, such as modifying hair color, expression, beard, and age of a person in an image. The method aims to extract region images corresponding to preset facial attributes from the image to be processed, generate a generated image corresponding to the preset facial attribute editing effect based on the region images, and a facial attribute mask. The facial attribute mask, region images, and generated images are then fused to obtain a region fusion image. Finally, the region fusion image replaces the corresponding region in the image to be processed to obtain the target image.

[0028] Existing facial attribute editing schemes typically employ either face mesh methods or generative methods. Face mesh methods primarily involve labeling the image with corresponding feature points based on facial attributes, dividing the entire image into several triangular mesh regions, and then modifying the specified facial attribute within those mesh regions. Generative methods generally predefine multiple facial attributes and then train a generative adversarial network (GAN) to extract the latent encoding of the face. Editing facial attributes is achieved by modifying this latent encoding. However, these methods are computationally intensive. In mobile scenarios, where real-time operation is required, pixel-to-pixel GAN ​​generation is commonly used. However, pixel-to-pixel GAN ​​generation requires constraints on each pixel in the image to learn the mapping relationship from the source image to the target image. The GAN needs to learn pixel-to-pixel mappings, which can lead to unreliable facial attribute generation results (such as inaccuracy or high noise levels). Furthermore, in continuous video, slight facial movement can cause significant differences between consecutive frames, resulting in video frame jitter and poor facial attribute editing performance. Based on this, an embodiment of the present application provides a face attribute editing method to solve the technical problem of poor existing face attribute editing effects, provide reliable and stable face attribute editing effects, and at the same time ensure good attribute editing image and video temporal stability.

[0029] Figure 1 shows a flowchart of a face attribute editing method provided in an embodiment of this application. The face attribute editing method provided in this embodiment of the application can be executed by a face attribute editing device, which can be implemented by hardware and / or software and integrated into a face attribute editing device.

[0030] The following description uses a face attribute editing device to perform a face attribute editing method as an example. Referring to Figure 1, the face attribute editing method includes:

[0031] S110: Obtain the image to be processed and extract the region image corresponding to the preset face attributes from the image to be processed.

[0032] The image to be processed provided in this solution can be a single video frame or multiple video frames from a video that requires facial attribute editing. For example, the image to be processed that requires facial attribute editing is acquired, and the region image corresponding to the preset facial attributes is extracted from the image to be processed.

[0033] For example, facial landmark detection can be performed on the image to be processed to identify multiple facial landmarks in the image. Facial landmarks are important feature points of various parts of the face in the image, typically contour points and corner points, including eyebrows, eyes, nose, and mouth. Facial landmarks can be represented using 68, 106, or 240 points. For example, to balance the accuracy and time efficiency of landmark detection, 106 points can be used to label facial landmarks. Optionally, a pre-trained face detection model can be used to detect facial landmarks in the image to be processed.

[0034] After obtaining facial key points, affine transformations can be used to rotate, translate, and crop the key points corresponding to preset facial attributes to obtain a facial attribute image. Taking a beard attribute as an example, the key points corresponding to the beard attribute include multiple key points of the nose, mouth, and contour. The center point of the beard attribute and the preset size to be cropped can be calculated based on the multiple key points of the nose, mouth, and contour. The beard image is then obtained from the image to be processed through affine transformation alignment. This beard image is the region image corresponding to the beard attribute.

[0035] S120: Generate a generated image and a face attribute mask corresponding to the preset face attribute editing effect based on the region image.

[0036] For example, after obtaining the region image, facial attribute editing processing can be performed on the region image based on preset facial attribute editing effects to obtain a generated image and a facial attribute mask corresponding to the region image. The generated image and the facial attribute mask have the same dimensions as the region image. The facial attribute mask reflects the distribution range of pixels corresponding to preset facial attributes in the region image. The value range of each pixel in the facial attribute mask is [0, 1]. The higher the value of the pixel corresponding to the facial attribute mask, the higher the probability that the pixel corresponds to the preset facial attribute.

[0037] The generated image provided by this solution, compared to the region image, reflects changes in facial attributes corresponding to preset facial attribute editing effects. For example, if the preset facial attribute editing effect is to remove facial whiskers, whiskers exist in the region image, but are removed in the generated image. Since there may be significant color differences in skin tone between the generated and unprocessed images, directly replacing the corresponding region in the unprocessed image with the generated image could lead to unreliable whisker removal effects (e.g., noticeable color shifts, obvious seams, unrealistic appearance, high noise levels). This solution, in addition to editing facial attributes in the region image to obtain the generated image, also generates a facial attribute mask. The facial attribute mask, region image, and generated image are then fused, resulting in a fused image that more closely resembles facial skin tone, and a more natural transition between the preset facial attribute editing effect and the unedited facial region.

[0038] S130: Use the face attribute mask as a fusion weight to fuse the region image and the generated image to obtain a region fused image.

[0039] For example, after obtaining the generated image and the facial attribute mask, the facial attribute mask can be used as a fusion weight to fuse the region image and the generated image, and the fusion result can be used as the region fusion image. For instance, the pixel values ​​of each point in the facial attribute mask can be used as fusion weights to perform a weighted summation of the pixel values ​​at corresponding pixel positions in the region image and the generated image, thus completing the fusion processing of each pixel position in the region image and the generated image, and obtaining the region fusion image. Optionally, the region fusion image provided by this solution can be determined by the following formula: I out =M pred ×I pred +(aM pred )×I in

[0040] Among them, M pred For facial attribute masking, I pred To generate the image, 'a' is a preset maximum pixel value, for example, a = 255, I in For the region image, the facial attribute mask is used as the fusion weight to fuse the region image and the generated image. The skin color of the face in the generated fused region image is closer to that in the image to be processed, reducing the color cast in the facial attribute editing effect, and the transition between the preset facial attributes and other positions is more natural.

[0041] S140: Replace the corresponding region in the image to be processed with the region fusion image to obtain the target image.

[0042] For example, after obtaining the region fusion image, the region fusion image is used to replace the region in the image to be processed that corresponds to the region image to obtain the target image. At this time, the target image achieves the preset face attribute editing effect in the region corresponding to the region image.

[0043] The above method extracts the region image corresponding to the preset facial attributes from the image to be processed, generates the generated image corresponding to the preset facial attribute editing effect and the facial attribute mask based on the region image, and uses the facial attribute mask as the fusion weight to fuse the region image and the generated image to obtain the region fusion image. The region fusion image is then used to replace the corresponding region in the image to be processed to obtain the target image. This method can effectively reduce the noise of the generated image, make the facial attribute editing effect of the image more realistic, and improve the facial attribute editing effect.

[0044] Based on the above embodiments, Figure 2 shows a flowchart of another face attribute editing method provided by the embodiments of this application. This face attribute editing method is a concretization of the above-described face attribute editing method. Referring to Figure 2, the face attribute editing method includes:

[0045] S210: Obtain the image to be processed, and extract the region image corresponding to the preset face attributes from the image to be processed.

[0046] S220: Input the region image into the trained image processing model. The image processing model generates a generated image and a face attribute mask corresponding to the preset face attribute editing effect based on the region image. The objective function of training the image processing model is determined based on the generative adversarial network loss function, absolute value error loss function, perceptual loss function and L1 regularization loss function.

[0047] This solution uses a trained image processing model to obtain the generated image and facial attribute mask corresponding to the region image. After training the image processing model, it can be configured in a facial attribute editing device. The output of the image processing model includes the generated image corresponding to the preset facial attribute editing effect, as well as the facial attribute mask corresponding to the region image. The training of the image processing model can be based on sample region images and corresponding label region images (control samples). The objective function for training the image processing model is determined based on the Generative Adversarial Network Loss (GAN Loss), the L1 Loss, the Perceptual Loss, and the L1 Regularization Loss.

[0048] For example, after obtaining the region image, the region image is input into the image processing model, and the image processing model processes the region image to generate a generated image and a face attribute mask corresponding to the preset face attribute editing effect.

[0049] Optionally, the objective function for training the image processing model can be obtained by summing the generative adversarial network loss function, absolute value error loss function, perceptual loss function, and L1 regularization loss function, or it can be obtained based on a weighted sum of the generative adversarial network loss function, absolute value error loss function, perceptual loss function, and L1 regularization loss function, using preset loss weights. For example, the objective function can be expressed by the following formula: L = λ GAN L GAN (x,y)+λ pixel L pixel (x,y)+λ perc L perc (x,y)+λ mask L mask (x)

[0050] Where x is the sample region image, y is the label region image, and L GAN (x,y) is the loss function of the generative adversarial network, λ GAN To generate the loss weights corresponding to the adversarial network loss function, L pixel (x,y) is the absolute value error loss function, λ pixel L represents the loss weights corresponding to the absolute value error loss function. perc (x,y) is the perceptual loss function, λ perc To perceive the loss weights corresponding to the loss function, L mask (x) is the L1 regularization loss function, λ mask These are the loss weights corresponding to the L1 regularization loss function.

[0051] This scheme accurately determines the generated image and facial attribute mask corresponding to the region image through an image processing model. The image processing model is trained using an objective function determined by the generative adversarial network loss function, absolute value error loss function, perceptual loss function, and L1 regularization loss function. The adversarial network loss function and absolute value error loss function make the output of the image processing model closer to the control sample. Furthermore, the perceptual loss function, based on calculating low-level feature losses (pixel color, edges, etc.), calculates the loss by comparing the convolutional outputs of the original image and the generated image, thus leveraging the ability of convolutional layers to abstract high-level features and perceiving the image from a global, high-dimensional perspective. Using perceptual loss preserves the overall information of the image, effectively ensuring the realism of the generated image and improving the quality of facial attribute editing. The L1 regularization loss function can be used to constrain the facial attribute mask generated by the image processing model, improving the accuracy of the generated facial attribute mask.

[0052] In one embodiment, the purpose of facial attribute editing is to use a GAN model to infer and generate an edited facial attribute image (i.e., a generated image) from an input source facial attribute image (i.e., a region image). This scheme takes an image processing model built on a Generative Adversarial Network (GAN) as an example, providing [sample region image x, label region image y (GroundTruth, GT)] as training data pairs to train the generator G and discriminator D in the image processing model. Optionally, the GAN loss function can be determined by the following formula: L GAN (x,y)=E x,y [logD(x,y)]+E x,z [log(1-D(x,G(x)))]

[0053] Where D(·)∈[0,1] is the output of the discriminator D, representing the probability that the predicted image is true, G(x) is the generated predicted image, and Ex,y [logD(x,y)] represents the discriminator D's output value of 1 when it receives [x,y] data pairs as input, i.e., maximizing D(x,y). x,z [log(1-D(x,G(x)))] means that when the discriminator D is input with the data pair [x,G(x)], the discriminator D should be input as 0, that is, D(x,G(x)) should be as small as possible.

[0054] In learning pixel-to-pixel mapping, the generated image is typically constrained pixel-by-pixel to make it closer to the label region image; this involves calculating the loss between the generated image and the label region image. This loss can be represented by different absolute value error loss functions such as L1 loss, L2 loss, and smooth L1 loss. In one embodiment, taking L1 loss as the absolute value error loss function as an example, the absolute value error loss function can be expressed as: L pixel (x,y)=||yG(x)||1

[0055] The perceptual loss function, based on the calculation of low-level feature losses (such as pixel color, edges, etc.), perceives the image from a global high-dimensional level, effectively preserving the global information of the image. Taking the VGG-19 network as the perceptual loss benchmark as an example, the L1 distance of multiple convolutions is used as the perceptual loss. The perceptual loss function can be determined by the following formula:

[0056] Among them, F i (·) represents the feature map of the i-th convolutional layer, λ i The preset loss weights are the feature maps of the i-th convolutional layer.

[0057] Using the aforementioned baseline network, by learning local pixel information and global information of an image, it is possible to generate edited facial attribute images that basically meet expectations. However, calculating only pixel-wise loss and perceptual loss on the image often fails to balance the differences between global and local information in the generated image, potentially leading to color casts that do not match the original image. In this case, further constraints can be imposed on facial attributes, editing only the attribute region, which can effectively ensure the consistency between the regions outside the attribute in the image and the source image.

[0058] To edit facial attributes only, the model needs to accurately identify the pixel locations of specified facial attributes, i.e., a facial attribute mask. However, supervised learning of facial attribute masks requires manual annotation of each image, leading to ambiguity in the details of different attributes and a massive workload. Therefore, this solution proposes an unsupervised approach for learning facial attribute masks. In traditional point-to-point image mapping methods, the input image directly generates the generated image through the model, i.e., I... pred=G(x). This scheme can generate a facial attribute mask, i.e., I, based on the image generated by the image processing model output. pred M pred =G(x), this step does not require the direct participation of the label region image; it only requires constraint by adding a regularization term to the input sample region image. Correspondingly, the L1 regularization loss function provided by this scheme can be determined by the following formula: L mask (x)=||x||1

[0059] In one possible embodiment, the absolute value error loss function provided by this solution can be determined based on the pixel-wise loss of the sample-generated image and the label region image generated by the image processing model from the sample region image, and the pixel-wise loss of the sample region fusion image and the label region image. The sample region fusion image is obtained by fusing the sample face attribute mask, the sample-generated image, and the sample region image generated by the image processing model (e.g., using the sample face attribute mask as a fusion weight to fuse the sample region image and the sample-generated image to obtain the sample region fusion image). For example, the absolute value error loss function can be the sum of the pixel-wise losses of the sample-generated image and the label region image generated by the image processing model from the sample region image, and the pixel-wise losses of the sample region fusion image and the label region image, or a weighted sum of both. In one possible embodiment, the absolute value error loss function provided by this solution is determined by the following formula: L′ pixel (x,y)=||yI pred ||1+||yI blend ||1

[0060] Where x is the sample region image, y is the label region image, and I pred Generate images for the samples, I blend For the sample region fused image, where ||yI pred ||1 represents the pixel-wise loss of the image processing model, which generates the sample region image and the label region image based on the sample region image.||yI blend ||1 represents the pixel-wise loss between the fused sample region image and the label region image. Based on this, the objective function can be expressed as: L′=λ GAN L GAN (x,y)+λ pixel L′ pixel (x,y)+λ perc L perc (x,y)+λ mask L mask (x)

[0061] This scheme determines the absolute value error loss function based on the pixel-wise loss of the sample generated image and the label region image, as well as the pixel-wise loss of the sample region fusion image and the label region image. This effectively constrains the generation of the face attribute mask and ensures that the fused region image has better consistency with the image to be processed, thereby improving the face attribute editing effect.

[0062] In one possible embodiment, the image processing model provided by this solution can also be fine-tuned based on inter-frame motion constraint regularization. It should be explained that while generative adversarial networks (GANs) can effectively learn point-to-point image mappings, they are highly sensitive to small changes in the input. Typical GANs perform well in generating images, but when applied to video sequences, they can lead to temporal instability, causing flickering, artifacts, and other inconsistencies over time. Convolutional neural networks (CNNs) typically rely on estimating dense frame-to-frame motion information (optical flow) during training or inference, or on exploring recurrent learning structures to adapt the model to video sequences. However, applying such methods to GANs for face attribute editing incurs additional and unnecessary computational overhead. This solution addresses this problem by using temporal stability as a regularization function for the loss function. The regularization is configured to consider different types of motion that may occur between frames, enabling the training of temporally stable GANs without requiring video material or expensive motion estimation.

[0063] To ensure temporal consistency in videos, CNNs typically consider calculating the loss between two consecutive frames, E = ||y||. t -W(y t-1 )||2

[0064] Here, W describes the warping transformation operation from frame to frame using the optical flow field between two frames. If there are inter-frame variations that cannot be explained by the flow field motion, these variations will be recorded as inconsistencies. To regularize this metric without requiring video data or optical flow information, this scheme introduces intra-frame warping and geometric transformations W(x) = T(x). The perturbation function T(·) in all introduced regularization terms depends on the transformation of the input image and can effectively simulate the motion that may occur between frames in a video sequence. This scheme uses simple geometric transformations to simulate the motion that may occur between frames in a video sequence. These transformations include translation, rotation, scaling, and shearing, which are applied as transformation matrices to the input image. This matrix is ​​randomly assigned for each image, and its transformation parameters are selected from the range in the table below through a uniform distribution.

[0065] By simulating the motion that may occur between frames in a video sequence through the above transformation operations, we can use x and T(x) to simulate two consecutive frames. Furthermore, the generator G of the image processing model is used to infer the output results G(x) and G(T(x)) of the two frames. If G(x) and G(T(x)) are temporally consistent, then the two frames should produce the same result. Based on this, the inter-frame motion constraint regularization term provided by this scheme can be determined by the following formula: L stability (x)=||G(T(x))-T(G(x))||2

[0066] Where x is the sample region image, G(·) is the image output by the generator in the image processing model, and T(·) is the result of inter-array motion prediction of the image.

[0067] Based on this, fine-tuning of the inter-frame motion constraint regularization term can be performed using the following objective function: L″=λ GAN L GAN (x,y)+λ pixel L pixel (x,y)+λ perc L perc (x,y)+λ mask L mask (x)+λ stability L stability (x)

[0068] Or L″=λ GAN L GAN (x,y)+λ pixel L′ pixel (x,y)+λ perc L perc (x,y)+λ mask L mask (x)+λ stability L stability (x)

[0069] Where, λ stability The weights of the inter-frame motion constraint regularization terms are the corresponding inter-frame motion constraint regularization terms.

[0070] For example, a training set containing sample pairs [sample region image x, label region image y] can be used, based on an objective function L or L ’ The image processing model is trained, and then fine-tuned based on the objective function L″. Specifically, the training is based on the objective function L″ or L″. ’When training the image processing model, the weights corresponding to the adversarial network loss function, absolute value error loss function, perceptual loss function, and L1 regularization loss function can be set to a combination of relatively large first preset values. However, when fine-tuning the image processing model based on the objective function L″, the weights corresponding to these functions can be set to a combination of relatively small second preset values, while the weights of the inter-frame motion constraint regularization term are set to a relatively large value (greater than the second preset value). This scheme fine-tunes the image processing model based on the inter-frame motion constraint regularization term, ensuring good temporal stability of the video sequence, improving the face attribute editing effect, and, when constraining inter-frame motion, fine-tuning only occurs during the training phase of the image processing model, without additional computational overhead for inference. This makes it computationally friendly for mobile devices, reducing the difficulty of adapting the image processing model to mobile platforms.

[0071] S230: Use the face attribute mask as a fusion weight to fuse the region image and the generated image to obtain a region fused image.

[0072] S240: Replace the corresponding region in the image to be processed with the region fusion image to obtain the target image.

[0073] In one possible embodiment, before fusing the region image and the generated image using the face attribute mask as a fusion weight to obtain the region fused image, the face attribute editing method provided in this solution can further transform the face attribute mask according to preset transformation coefficients. Based on this, the face attribute editing method provided in this solution, which uses the face attribute mask as a fusion weight to fuse the region image and the generated image to obtain the region fused image, can specifically use the transformed face attribute mask as a fusion weight to fuse the region image and the generated image to obtain the region fused image.

[0074] For example, after obtaining the facial attribute mask corresponding to the region image, the facial attribute mask can be transformed according to preset transformation coefficients to transform the fusion weights. After completing the transformation of the facial attribute mask, the transformed facial attribute mask can be used as the fusion weights to fuse the region image and the generated image to obtain a region fused image. Optionally, the transformation of the facial attribute mask can be a power-law transformation (gamma transformation). For example, the transformation of the facial attribute mask can be expressed by the following formula:

[0075] Among them, M′ pred M is the face attribute mask after transformation. predThis is the face attribute mask before transformation, where 'a' is the preset maximum pixel value (e.g., a = 255), and γ is the preset transformation coefficient (gamma transformation coefficient). This scheme uses the transformed face attribute mask to fuse the region image and the generated image. The preset transformation coefficients allow for control over the intensity of face attribute editing, improving the flexibility of face attribute editing.

[0076] The above describes a method that extracts region images corresponding to preset facial attributes from the image to be processed. Based on these region images, it generates a generated image and a facial attribute mask corresponding to the preset facial attribute editing effect. The facial attribute mask is then used as a fusion weight to fuse the region images and the generated images, resulting in a region fusion image. This fusion image is then used to replace the corresponding region in the image to be processed to obtain the target image. This method effectively reduces noise in the generated image, resulting in a more realistic facial attribute editing effect and improving the overall facial attribute editing performance. Furthermore, by accurately determining the generated image and facial attribute mask corresponding to the region images through an image processing model, and by training the image processing model using a target function determined by the generative adversarial network loss function, absolute value error loss function, perceptual loss function, and L1 regularization loss function, the accuracy of the image processing model in generating facial attribute masks and generated images can be effectively improved.

[0077] Figure 3 is a schematic diagram of a face attribute editing device provided in an embodiment of this application. Referring to Figure 3, the face attribute editing device includes an image extraction module 31, an image processing module 32, an image fusion module 33, and an image generation module 34.

[0078] The image extraction module 31 is configured to acquire the image to be processed and extract the region image from the image to be processed; the image processing module 32 is configured to generate a generated image corresponding to the preset face attribute editing effect and a face attribute mask based on the region image; the image fusion module 33 is configured to use the face attribute mask as a fusion weight to perform fusion processing on the region image and the generated image to obtain a region fusion image; and the image generation module 34 is configured to replace the corresponding region in the image to be processed with the region fusion image to obtain the target image.

[0079] The above method extracts the region image corresponding to the preset facial attributes from the image to be processed, generates the generated image corresponding to the preset facial attribute editing effect and the facial attribute mask based on the region image, and uses the facial attribute mask as the fusion weight to fuse the region image and the generated image to obtain the region fusion image. The region fusion image is then used to replace the corresponding region in the image to be processed to obtain the target image. This method can effectively reduce the noise of the generated image, make the facial attribute editing effect of the image more realistic, and improve the facial attribute editing effect.

[0080] In one possible embodiment, the image processing module 32 generates a generated image and a face attribute mask corresponding to a preset face attribute editing effect based on the region image. The configuration is as follows: the region image is input into the trained image processing model, and the image processing model generates a generated image and a face attribute mask corresponding to the preset face attribute editing effect based on the region image. The objective function of training the image processing model is determined based on the generative adversarial network loss function, the absolute value error loss function, the perceptual loss function, and the L1 regularization loss function.

[0081] In one possible embodiment, the absolute value error loss function is determined based on the pixel-wise loss of the sample generated image and the label region image generated by the image processing model from the sample region image, as well as the pixel-wise loss of the sample region fusion image and the label region image. The sample region fusion image is obtained by fusing the sample face attribute mask, the sample generated image, and the sample region image generated by the image processing model.

[0082] In one possible embodiment, the absolute value error loss function is determined by the following formula: L′ pixel (x,y)=||yI pred ||1+||yI blend ||1

[0083] Where x is the sample region image, y is the label region image, and I pred Generate images for the samples, I blend This is a fused image of the sample regions.

[0084] In one possible embodiment, the image processing model is fine-tuned based on an inter-frame motion constraint regularization term, which is determined by the following formula: L stability (x)=||G(T(x))-T(G(x))||2

[0085] Where x is the sample region image, G(·) is the image output by the generator in the image processing model, and T(·) is the result of inter-array motion prediction of the image.

[0086] In one possible embodiment, the region fusion image is determined by the following formula: I out =M pred ×I pred +(255-M pred )×I in

[0087] Among them, M pred For facial attribute masking, I pred To generate the image, 'a' is the preset maximum pixel value, and 'I' is the maximum pixel value. in This is a region image.

[0088] In one possible embodiment, the face attribute editing device further includes a mask transformation module, which is configured to transform the face attribute mask according to a preset transformation coefficient.

[0089] The image fusion module 33 uses the face attribute mask as the fusion weight to fuse the region image and the generated image to obtain a region fusion image. The configuration is as follows: the transformed face attribute mask is used as the fusion weight to fuse the region image and the generated image to obtain a region fusion image.

[0090] It is worth noting that in the above embodiments of the face attribute editing device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this application.

[0091] This application also provides a face attribute editing device, which can integrate the face attribute editing apparatus provided in this application. Figure 4 is a schematic diagram of the structure of a face attribute editing device provided in this application. Referring to Figure 4, the face attribute editing device includes: an input device 43, an output device 44, a memory 42, and one or more processors 41; the memory 42 is used to store one or more programs; when one or more programs are executed by one or more processors 41, the one or more processors 41 implement the face attribute editing method provided in the above embodiments. The face attribute editing apparatus, device, and computer provided above can be used to execute the face attribute editing method provided in any of the above embodiments, and have corresponding functions and beneficial effects.

[0092] This application also provides a non-volatile storage medium storing computer-executable instructions, which, when executed by a computer processor, are used to perform the face attribute editing method provided in the above embodiments. Of course, the computer-executable instructions provided in this application are not limited to the face attribute editing method provided above; they can also perform related operations in the face attribute editing method provided in any embodiment of this application. The face attribute editing apparatus, device, and storage medium provided in the above embodiments can execute the face attribute editing method provided in any embodiment of this application. Technical details not described in detail in the above embodiments can be found in the face attribute editing method provided in any embodiment of this application.

[0093] Based on the above embodiments, this application also provides a computer program product. The technical solution of this application, in essence or in other words, the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer program product is stored in a storage medium and includes several instructions to cause a computer device, mobile terminal, or processor therein to execute all or part of the steps of the face attribute editing method provided in the various embodiments of this application.

Claims

1. A face attribute editing method, wherein, include: Acquire an image to be processed, and extract the region image corresponding to the preset face attributes from the image to be processed; Generate a generated image and a facial attribute mask corresponding to the preset facial attribute editing effect based on the region image; The facial attribute mask is used as a fusion weight to fuse the region image and the generated image to obtain a region fused image; The target image is obtained by replacing the corresponding region in the image to be processed with the fused image of the region. 2.The face attribute editing method of claim 1, wherein, The step of generating a generated image and a facial attribute mask corresponding to a preset facial attribute editing effect based on the region image includes: The region image is input into the trained image processing model. The image processing model generates a generated image and a face attribute mask corresponding to the preset face attribute editing effect based on the region image. The objective function of the image processing model is determined based on the generative adversarial network loss function, the absolute value error loss function, the perceptual loss function, and the L1 regularization loss function.

3. The face attribute editing method of claim 2, wherein, The absolute value error loss function is determined based on the pixel-wise loss of the sample generated image and the label region image generated by the image processing model from the sample region image, as well as the pixel-wise loss of the sample region fusion image and the label region image. The sample region fusion image is obtained by fusing the sample face attribute mask generated by the image processing model, the sample generated image, and the sample region image.

4. The face attribute editing method of claim 3, wherein, The absolute value error loss function is determined by the following equation: L' pixel (x,y) = ||y - I pred ||1 + ||y - I blend ||1 wherein x is a sample region image, y is a label region image, I pred is a sample region image, I blend is a sample region image, I 5.The face attribute editing method of claim 2, wherein, The image processing model is fine-tuned based on an inter-frame motion constraint regularization term determined by the following formula: L stability (x) = ||G(T(x))-T(G(x))||2 Where x is the sample region image, G(·) is the image output by the generator in the image processing model, and T(·) is the result of inter-array motion prediction of the image. 6.The face attribute editing method of claim 1, wherein, The region fusion image is determined by the following equation: I out = M pred × I pred + (255 - M pred ) × I in Wherein, M pred is a face attribute mask, I pred is a generated image, a is a preset maximum pixel value, I in is a region image.

7. The face attribute editing method of claim 1, wherein, Before fusing the region image and the generated image using the facial attribute mask as fusion weights to obtain the region fused image, the method further includes: The facial attribute mask is transformed according to preset transformation coefficients; The step of fusing the region image and the generated image using the facial attribute mask as fusion weights to obtain a region fused image includes: The transformed facial attribute mask is used as a fusion weight to fuse the region image and the generated image, resulting in a region fused image.

8. A face attribute editing apparatus, comprising: It includes an image extraction module, an image processing module, an image fusion module, and an image generation module, among which: The image extraction module is configured to acquire an image to be processed and extract a region image from the image to be processed. The image processing module is configured to generate a generated image and a face attribute mask corresponding to a preset face attribute editing effect based on the region image. The image fusion module is configured to use the face attribute mask as a fusion weight to perform fusion processing on the region image and the generated image to obtain a region fusion image. The image generation module is configured to replace the corresponding region in the image to be processed with the region fusion image to obtain the target image.

9. A face attribute editing apparatus, comprising: include: Memory and one or more processors; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the face attribute editing method according to any one of claims 1-7.

10. A non-transitory storage medium storing computer-executable instructions that, when executed by a processor of a device, cause the device to perform steps comprising: The computer executable instructions, when executed by a computer processor, perform the face attribute editing method according to any one of claims 1-7.

11. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the face attribute editing method according to any one of claims 1-7.