Model Training Method, Image Processing Method, Apparatus, Device, and Storage Medium
By combining the generation of antagonistic network and the defect segmentation network, a flawless face image processing model is trained to generate a flawless face image processing model, solving the problems of loss of details and low processing efficiency caused by defect removal in the existing beauty treatment solutions, and achieving efficient, real and natural defect removal effect, which is suitable for real-time video image processing.
Patent Information
- Application Number
- CN202111497159.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-12-09
AI Technical Summary
When the existing beauty treatment solutions remove facial blemishes, they will lead to loss of skin details. The higher the degree of beautification, the more serious the loss of details, the overall texture is unreal, and the deep learning-based solution processing process will take a long time and be inefficient.
By combining the generative adversarial network with the defect segmentation network, we generate defect-free face images by generating the adversarial network, and segmenting them using the defect segmentation network. Combined with the constraint generation process of suppressing the defect loss function, the target generation network is trained as a face image processing model.
Generating real and natural flawless face images improves the efficiency of flaw removal and enhances real-time performance. It is suitable for application scenarios such as live broadcast and video calls for real-time video image processing.
Smart Images

Figure CN114187201B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of image processing, and in particular to a model training method, an image processing method, an apparatus, a device, and a storage medium. Background Art
[0002] The beauty function has become one of the important functions in many applications. Using the beauty function, the face image can be beautified, and then the defects in the facial skin can be removed, such as pimples, moles, skin spots, and thick pores, etc.
[0003] Currently, common beauty processing solutions include implementing based on edge-preserving filters, such as bilateral filters, guided filters, surface blur filters, and local mean filters, or a combination and superposition of multiple filters, retaining parts of edges with large pixel gradient values such as eyebrows, hair, and background, and smoothing parts with relatively uniform colors and small pixel gradients such as skin. However, this solution will cause some loss of skin details. The higher the degree of beautification, the more serious the detail loss, and the less realistic the overall texture. The whole image becomes illusory, deviating from the texture of real skin, with a false face feeling, the beauty effect is not ideal, and it affects the comfort of viewers. In addition, there is also a beauty processing solution based on deep learning. This solution needs to first detect the defects in the face through a target detection model, and then use methods such as local skin smoothing or filling to remove the defects. However, the timeliness of the target detection model is poor, and combined with the subsequent processing process, it will cause the entire image processing process to be very time-consuming and the image processing efficiency is low. Summary of the Invention
[0004] Embodiments of the present invention provide a model training method, an image processing method, an apparatus, a device, and a storage medium, which can optimize the existing solutions for beautifying face images.
[0005] In a first aspect, an embodiment of the present invention provides a model training method, and the method includes:
[0006] Obtain a first training sample set including training sample pairs, where each training sample pair includes an original face sample image containing defects and a target face sample image obtained by removing the defects on the basis of the original face sample image;
[0007] Train a preset original network model based on the first training sample set using a preset loss function to obtain a target network model. Among them, the preset original network model includes a generative adversarial network with parameters to be adjusted and a pre-trained flaw segmentation network with fixed parameters. The generative adversarial network includes a generator network and a discriminator network. The generator network is used to generate a flawless face image in the target domain. The flaw segmentation network is used to segment the generated image output by the generator network to obtain a segmentation result based on the flaw region and the non-flaw region. The preset loss function includes a flaw suppression loss function, which is transformed from the segmentation result and is used to constrain the generation of flaws in the generated image during the generation process;
[0008] Determine a face image processing model according to the target generator network included in the target network model. Among them, the face image processing model is used to process the face image to be processed to remove the flaws included in the face image to be processed.
[0009] In a second aspect, an embodiment of the present invention provides an image processing method, and the method includes:
[0010] Obtain a face image to be processed;
[0011] Input the face image to be processed into the face image processing model to output a target face image corresponding to the face image to be processed after flaw removal processing, where the face image processing model is obtained by the model training method provided by the embodiment of the present invention.
[0012] In a third aspect, an embodiment of the present invention provides a model training device, and the device includes:
[0013] A sample set acquisition module, configured to acquire a first training sample set including training sample pairs, where the training sample pairs include an original face sample image with flaws and a target face sample image after removing the flaws on the basis of the original face sample image;
[0014] A model training module, configured to train a preset original network model based on the first training sample set using a preset loss function to obtain a target network model. Among them, the preset original network model includes a generative adversarial network with parameters to be adjusted and a pre-trained flaw segmentation network with fixed parameters. The generative adversarial network includes a generator network and a discriminator network. The generator network is used to generate a flawless face image in the target domain. The flaw segmentation network is used to segment the generated image output by the generator network to obtain a segmentation result based on the flaw region and the non-flaw region. The preset loss function includes a flaw suppression loss function, which is transformed from the segmentation result and is used to constrain the generation of flaws in the generated image during the generation process;
[0015] A model determination module, configured to determine a face image processing model according to a target generation network included in the target network model, where the face image processing model is used to process a to-be-processed face image to remove defects included in the to-be-processed face image.
[0016] In a fourth aspect, an embodiment of the present invention provides a model training apparatus, which includes:
[0017] A to-be-processed image acquisition module, configured to acquire a to-be-processed face image;
[0018] An image processing module, configured to input the to-be-processed face image into the face image processing model to output a target face image corresponding to the to-be-processed face image after defect removal processing, where the face image processing model is obtained by the model training method provided in the embodiment of the present invention.
[0019] In a fifth aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the model training method and / or the image processing method provided in the embodiment of the present invention are implemented.
[0020] In a sixth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the model training method and / or the image processing method provided in the embodiment of the present invention are implemented.
[0021] The model training scheme provided in the embodiment of the present invention obtains a first training sample set including training sample pairs, wherein the training sample pairs include an original face sample image including defects and a target face sample image after the defects are removed based on the original face sample image, and uses the first training sample set to train a preset original network model based on a preset loss function to obtain a target network model, wherein the preset original network model includes a generative adversarial network with parameters to be adjusted and a pre-trained defect segmentation network with fixed parameters, the generative adversarial network includes a generative network and a discriminant network, the generative network is used to generate a defect-free face image in the target domain, the defect segmentation network is used to segment the generated image output by the generative network to obtain a segmentation result based on defect areas and non-defect areas, the preset loss function includes a defect suppression loss function, the defect suppression loss function is obtained by converting the segmentation result, and is used to constrain the generation of defects in the generated image during the generation process, and the face image processing model is determined according to the target generative network included in the target network model, wherein the face image processing model is used to process the face image to be processed to remove the defects contained in the face image to be processed. By adopting the above technical solution, a generative adversarial network is combined with a defect segmentation network, and the generative adversarial network is responsible for generating flawless faces, which is conducive to maintaining the consistency of other features except skin texture. The defect segmentation network can promote the generation process of the generative adversarial network as an auxiliary network. Combined with the defect suppression loss function, the generative network can learn to eliminate defects and then output flawless images. After the training is completed, the trained target generative network is used as a facial image processing model for defect elimination. Using this model, more realistic and natural flawless facial images can be obtained, and there is no need for multi-stage processing. Flawless facial images can be output at one time, which can improve the efficiency of defect removal processing and enhance real-time performance. It can be well applied to application scenarios with high timeliness requirements such as live broadcasts or video calls that process real-time video images. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A flowchart of a model training method provided by an embodiment of the present invention;
[0023] Figure 2 A flowchart of another model training method provided by an embodiment of the present invention;
[0024] Figure 3 A schematic diagram of a defect segmentation network training process provided by an embodiment of the present invention;
[0025] Figure 4 A schematic diagram of a preset original network model training process provided by an embodiment of the present invention;
[0026] Figure 5Schematic flowchart of an image processing method provided by an embodiment of the present invention;
[0027] Figure 6 Block diagram of a model training device provided by an embodiment of the present invention;
[0028] Figure 7 Block diagram of an image processing device provided by an embodiment of the present invention;
[0029] Figure 8 Block diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners
[0030] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the drawings. Furthermore, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0031] Figure 1 Schematic flowchart of a model training method provided by an embodiment of the present invention. This method can be executed by a model training device, where the device can be implemented by software and / or hardware and is generally integrated in a computer device. As Figure 1 shown, this method includes:
[0032] Step 101: Obtain a first training sample set containing training sample pairs, where each training sample pair includes an original face sample image with defects and a target face sample image after removing the defects based on the original face sample image.
[0033] Exemplarily, the defects on a human face can include objects that affect the facial beauty, such as pimples, acne marks, moles, blood streaks, skin spots, and enlarged pores. Many real face images contain defects. By processing the defects, a more beautiful face image can be obtained to meet the aesthetic needs of users. The defects in a face image can be understood as the defect image regions obtained after the defects on a real human face are collected by an image acquisition device, which are simply referred to as defects in this article.
[0034] In an embodiment of the present invention, a face image containing defects can be collected as an original face sample image. For the original face sample image, professional image processing software can be used for processing to eliminate the defects contained therein. The processing process can be completed by a professional operating the image processing software to obtain a more standard beautified image as the target face sample image. An original face sample image and the corresponding target face sample image form a training sample pair. For example, the original face sample image is A, and after processing A, a defect-free target face sample image a is obtained, and A and a form a training sample pair.
[0035] Exemplarily, the first training sample set may include a first preset number of training sample pairs. The first preset number can be set according to the actual situation. To ensure the training effect, the first preset number can be set larger.
[0036] Step 102: Train a preset original network model based on the first training sample set using a preset loss function to obtain a target network model. The preset original network model includes a generative adversarial network with parameters to be adjusted and a pre-trained defect segmentation network with fixed parameters. The generative adversarial network includes a generator network and a discriminator network. The generator network is used to generate a defect-free face image in the target domain, and the defect segmentation network is used to segment the generated image output by the generator network to obtain a segmentation result based on the defect area and the non-defect area. The preset loss function includes a defect suppression loss function, and the defect suppression loss function is transformed from the segmentation result and is used to constrain the generation of defects in the generated image during the generation process.
[0037] Exemplarily, a preset original network model can be constructed, and the preset original network model includes a generative adversarial network and a defect segmentation network. The specific internal structures of the generative adversarial network and the defect segmentation network are not limited.
[0038] Among them, the generative adversarial network (GAN), also known as the generative discriminative network, has extremely strong learning and generation capabilities for image information. The generative adversarial network includes two parts, a generator network (also known as generator G) and a discriminator network (also known as discriminator D). The generator network can be used to learn to remove defects in the image and generate a defect-free face image in the target domain, and can maintain the consistency of features except for skin texture; the discriminator network can be used to discriminate whether the face image output by the generator network is real and whether it is close to the target face sample image. During the training process, the parameters in the generative adversarial network can be continuously optimized to achieve the effect of enhancing the model's ability.
[0039] Optionally, during the process of training the preset original network model based on the first training sample set using the preset loss function, the generation network takes the original face sample image as input and outputs a generated image, and the discrimination network takes the corresponding target face sample image and the generated image as input and outputs whether the target face sample image and the generated image are the same.
[0040] The defect segmentation network is used to assist in the training of the generative adversarial network and can learn the defect segmentation signal in a supervised manner. Before this step, the preset original defect segmentation network can be trained in advance to obtain a trained defect segmentation network, which can segment the defect area and non-defect area in the image. In the preset original network model, the defect segmentation network is used to segment the generated image output by the generation network to obtain a segmentation result based on the defect area and non-defect area. During the training process of the preset original network model, the parameters in the defect segmentation network remain unchanged.
[0041] In the embodiment of the present invention, the role of the preset loss function is to provide guidance for the training process of the preset original network model. During training, the parameters of the generative adversarial network are continuously optimized to make the loss function smaller, thereby achieving the effect of enhancing the model's ability. The preset loss function includes a defect suppression loss function, which is transformed from the segmentation result and is used to constrain the generation of defects in the generated image during the generation process. Specifically, it can be used to constrain the gap between the generated image and the target image of the defect-free area. Optionally, the defect suppression loss function can specifically be to constrain the number of defect areas included in the generated image. Exemplarily, the defect suppression loss function is represented by the total number of defect areas included in the generated image output by the defect segmentation network.
[0042] Exemplarily, the preset loss function may also include other loss functions, such as a reconstruction loss function, an adversarial loss function, a perceptual loss function, and a structural information loss function, etc., which are not specifically limited.
[0043] Exemplarily, after the training process is completed, a target network model is obtained. The target network model may include a trained generative adversarial network (which can be called a target generative adversarial network) and an unchanged defect segmentation network. The target generative adversarial network includes a trained generation network (which can be called a target generation network) and a trained discrimination network (which can be called a target discrimination network).
[0044] Step 103: Determine a face image processing model according to the target generation network included in the target network model, where the face image processing model is used to process the to-be-processed face image to remove the defects included in the to-be-processed face image.
[0045] Exemplarily, in actual applications, it is necessary to process face images containing defects to remove the defects therein. In fact, the defect segmentation network and the target discrimination network in the target network model do not participate in the processing of face images. Therefore, the defect segmentation network and the discrimination network can be understood as networks for assisting the training of the model and are not required in the inference stage (i.e., the application stage). The main network to be used is the target generation network. Therefore, the target generation network can be directly used as the face image processing model, or further optimized and adjusted based on the target generation network to obtain the face image processing model. When applying, the face image to be processed can be input into the face image processing model, and the face image processing model can be used to perform defect removal processing on the face image to be processed, and a processed, real, natural, and defect-free face image can be quickly output.
[0046] In the model training method provided in the embodiments of the present invention, a method combining a generative adversarial network and a defect segmentation network is adopted. The generative adversarial network is responsible for generating defect-free faces, which is beneficial to maintaining the consistency of other features except skin texture. The defect segmentation network, as an auxiliary network, can promote the generation process of the generative adversarial network. Combining with the defect suppression loss function enables the generative network to learn to eliminate defects and then output defect-free images. After training, the trained target generation network is used as the face image processing model for defect elimination. Using this model, a more real and natural defect-free face image can be obtained, and it is not necessary to go through multiple stages of processing. A defect-free face image is output at one time, which can improve the efficiency of defect removal processing and enhance real-time performance, and can be well applied to application scenarios with high timeliness requirements such as live broadcast or video call for processing real-time video images.
[0047] In some embodiments, the preset loss function further includes a reconstruction loss function and an adversarial loss function. Among them, the reconstruction loss function is used to constrain the gap between the generated image and the corresponding target face sample image. The advantage of such a setting is that it can enable the generative network to better learn to remove defects.
[0048] Exemplarily, the reconstruction loss function can make the generated image (denoted as gen_img) close to the target face sample image (denoted as target_img), which helps the generated image maintain the characteristics of the attribute image (i.e., the original face sample image). The reconstruction loss function can be, for example, the L1 norm loss function, and can be expressed by the following expression:
[0049] recon_loss = ||gen_img - target_img||2
[0050] Among them, recon_loss represents the reconstruction loss function.
[0051] Exemplarily, the adversarial loss function makes the generated images more realistic and natural, and closer to the data distribution of beauty-enhanced human faces, which can be achieved by a conventional adversarial loss and denoted as gan_loss.
[0052] In some embodiments, before training the preset original network model based on the preset loss function using the training sample set, it further includes: obtaining a second training sample set containing the training sample pairs; segmenting and labeling the original human face sample images according to the differences between the original human face sample images and the target human face sample images in each training sample pair in the second training sample set to obtain target segmentation images containing defective regions and non-defective regions; using the original human face sample images as the input of a preset original defect segmentation network and the corresponding target segmentation images as the expectations, and training the preset original defect segmentation network based on a preset segmentation loss function to obtain a defect segmentation network. The advantage of this setting is that a defect segmentation network that can accurately divide defective regions and non-defective regions can be trained.
[0053] Exemplarily, the second training sample set may contain a second preset number of training sample pairs, and the second preset number may be the same as or different from the first preset number. The training sample pairs contained in the second training sample set also include original human face sample images with defects and target human face sample images after removing the defects on the basis of the original human face sample images. Each image in the training sample pairs contained in the second training sample set may be the same as or different from each image in the training sample pairs contained in the first training sample set, and no specific limitation is made.
[0054] Exemplarily, the target human face sample image can be considered as a defect-free human face image. If the gap between the first pixel region in the original human face sample image and the corresponding first target pixel region in the target human face sample image is large, it can be considered that the first pixel region contains defects, and thus the original human face sample image is segmented and labeled. The defect segmentation network is used to identify the defective regions and non-defective regions in the image. The original human face sample image is input into the preset original defect segmentation network, and a segmentation image with the defective regions and non-defective regions segmented is output. If the preset original defect segmentation network can accurately segment, the output segmentation image should be consistent with the corresponding target segmentation image. Accordingly, the preset original defect segmentation network is trained using the preset segmentation loss function to continuously adjust the parameters in the preset original defect segmentation network to obtain a defect segmentation network that can accurately perform defect segmentation.
[0055] In some embodiments, segmenting and labeling the original face sample image according to the difference between the original face sample image and the target face sample image in each training sample pair in the second training sample set includes: for each training sample pair in the second training sample set, calculating the absolute difference between the first pixel and the second pixel corresponding to each pixel position, where the first pixel is from the original face sample image and the second pixel is from the target face sample image; determining a target segmentation threshold according to each of the absolute differences; using the target segmentation threshold to segment and label the original face sample image, where the part corresponding to the absolute difference greater than or equal to the target segmentation threshold is labeled as a defective area, and the part corresponding to the absolute difference less than the target segmentation threshold is labeled as a non-defective area. The advantage of such a setting is that the target segmentation threshold can be reasonably determined according to the overall difference situation, and then the training samples can be labeled quickly and accurately based on the target segmentation threshold.
[0056] Exemplarily, the original face sample image and the target face sample image have the same size and the same number of pixels. A coordinate system can be constructed with a certain vertex of the image as the coordinate origin, and the pixel positions are represented by coordinates. The absolute difference is the absolute value of the difference. Calculate the absolute value of the difference between the pixel values of the first pixel from the original face sample image and the second pixel from the target face sample image at the same pixel position. For example, if the pixel value of the first pixel is m and the pixel value of the second pixel is n, the absolute difference is the absolute value of m - n. Optionally, analyze the absolute differences and determine the target segmentation threshold according to the distribution or concentration degree of each absolute difference, etc. For example, the target segmentation threshold can be determined according to the median, average value or percentile, etc. Further, the target segmentation threshold can also be determined by the product obtained by multiplying the median, average value or percentile, etc. by a preset coefficient.
[0057] Exemplarily, when segmenting and labeling the original face sample image, a binary mask image (which can be understood as Ground-Truth) can be obtained. The defective area is marked as 1 and the non-defective area is marked as 0. After the original face sample image passes through the defective segmentation network, a binary segmentation mask image of the same size as the original face sample image can be obtained. During the process of training the preset original defective segmentation network, a preset segmentation loss function can be used to constrain the gap between the output binary segmentation mask image and the binary mask image as the target segmentation image, that is, to make the binary segmentation mask image approach the binary mask image. Through gradient descent and parameter update, an accurate defective segmentation network can be obtained.
[0058] The segmentation in the embodiments of the present invention can be regarded as a pixel classification problem. Conventional classification problems usually use cross-entropy loss. However, for face images, according to the characteristics of skin blemishes, the proportion of skin blemishes in the human face skin is usually small, resulting in a serious imbalance in the proportion of positive and negative samples. The number of pixels of negative samples (positions without blemishes, such as large areas of skin and background) is significantly more than the number of pixels of positive samples (blemish positions such as pimples or age spots). The commonly used cross-entropy loss cannot achieve a better training effect. In the embodiments of the present invention, in view of the above characteristics, an unconventional loss function can be adopted to weaken the influence of the imbalance in the proportion of positive and negative samples on the training effect.
[0059] Exemplarily, the preset segmentation loss function may include, for example, a weighted cross-entropy loss function (Weighted crossentropy Loss), a focal loss function (Focal Loss), and a dice loss function (Dice Loss), etc. Among them, DiceLoss is currently only used for medical image segmentation, and in the present invention, it is applied to the segmentation of blemished face images. It is found through experiments that good training effects can be obtained, and the problem of imbalance between positive and negative samples in blemished face images can be well solved.
[0060] Figure 2 FIG. is a schematic flowchart of another model training method provided by an embodiment of the present invention. This method is optimized based on the above-mentioned various alternative embodiments, such as Figure 2 As shown, the method may include:
[0061] Step 201, obtain a second training sample set including training sample pairs.
[0062] Among them, each training sample pair includes an original face sample image containing blemishes and a target face sample image obtained by removing the blemishes based on the original face sample image.
[0063] Step 202, according to the difference between the original face sample image and the target face sample image in each training sample pair in the second training sample set, perform segmentation annotation on the original face sample image to obtain a target segmentation image including a blemish area and a non-blemish area.
[0064] Exemplarily, for each training sample pair, calculate the absolute value of the difference between the pixel values of the first pixel from the original face sample image and the second pixel from the target face sample image corresponding to each pixel position, and determine a target segmentation threshold according to each absolute value. Use the target segmentation threshold to annotate the original face sample image to obtain a target segmentation image including a blemish area and a non-blemish area, which may specifically be a binary mask image, where the blemish area is marked as 1 and the non-blemish area is marked as 0.
[0065] Step 203: Use the original face sample image as the input of the preset original defect segmentation network, use the corresponding target segmentation image as the expectation, and train the preset original defect segmentation network based on the preset segmentation loss function to obtain the defect segmentation network.
[0066] Among them, the preset segmentation loss function is Dice Loss.
[0067] Figure 3 It is a schematic diagram of the training process of a defect segmentation network provided by an embodiment of the present invention. As Figure 3 shown, use the original face sample image as the input image and input it into the preset original defect segmentation network to output a segmentation image, and calculate the preset segmentation loss function based on the segmentation image and the target segmentation image. For example, use Dice Loss to constrain the binary segmentation mask image output from the preset original defect segmentation network to approach the binary mask image, and continuously adjust the weight parameters in the preset original defect segmentation network until the trained defect segmentation network is obtained. In the binary mask image, black represents 0 and white represents 1, that is, white represents the defect area.
[0068] Step 204: Construct a preset original network model according to the defect segmentation network and the generative adversarial network.
[0069] Figure 4 It is a schematic diagram of the training process of a preset original network model provided by an embodiment of the present invention. As Figure 4 shown, the preset original network model includes a generative adversarial network composed of a generative network and a discriminative network, and also includes a defect segmentation network. The generated image output by the generative network is used as the input of the defect segmentation network, and the target face sample image and the generated image output by the generative network are used as the inputs of the discriminative network.
[0070] It should be noted that Figure 3 and Figure 4 in each image containing a face, the eyes are mosaicked, and this process is generally not performed during the actual training process and application process.
[0071] Step 205: Obtain a first training sample set containing training sample pairs.
[0072] Step 206: Use the first training sample set to train the preset original network model based on the preset loss function to obtain the target network model.
[0073] Among them, the preset loss function includes a reconstruction loss function, an adversarial loss function, and a defect suppression loss function.
[0074] As Figure 4As shown, the generation network takes the original face sample image as input and outputs a generated image. With the constraints of the reconstruction loss function and the adversarial loss function, the generated image approaches the target face sample image, thereby learning to generate the target face sample image after removing defects. The discriminative network takes the generated image and the target face sample image as input and learns to distinguish between them. The defect segmentation network uses the defect suppression loss function to penalize the generation network that fails to successfully remove defects. The defect segmentation network takes the generated image as input and outputs a binary segmentation mask image. The defect suppression loss function is calculated based on the binary segmentation mask image. The defect suppression loss function constrains that the generated image has no defects, which is equivalent to constraining that all the pixel values corresponding in the binary segmentation mask image are 0 (such as Figure 4 the all-black image in
[0075] . Optionally, the preset loss function can be expressed as the weighted sum of the reconstruction loss function, the adversarial loss function, and the defect suppression loss function. For example, it can be expressed by the following expression:
[0076] Loss = a * recon_loss + b * gan_loss + c * res_loss
[0077] where Loss represents the preset loss function, recon_loss represents the reconstruction loss function, gan_loss represents the adversarial loss function, res_loss represents the defect suppression loss function, a represents the first weight coefficient, b represents the second weight coefficient, and c represents the third weight coefficient. The values of a, b, and c can be set according to actual needs.
[0078] During the training process, the parameters of the defect segmentation network are fixed, and the parameters in the generation network and the discriminative network can be fixed separately and optimized alternately. The optimization is aimed at minimizing the preset loss function, and the respective parameters are updated through gradient descent.
[0079] Step 207: Determine a face image processing model according to the target generation network included in the target network model.
[0080] wherein, the face image processing model is used to process the face image to be processed to remove the defects included in the face image to be processed.
[0081] The model training method provided by the embodiments of the present invention results in a face image processing model with a wide range of applications. It can be applied to any skin type, has a strong sense of reality, and can handle large areas of dense blemishes well without leaving mottled marks. For skin that is originally flawless, the generated image is basically the same as the original image, without causing damage such as quality or color difference. In practical applications, only through one network, namely the generation network, it is superior to the two-step blemish removal solution of detection plus filling or local skin smoothing, and also has an advantage in terms of time. It can support video-level acne removal and is well applicable to application scenarios with high timeliness requirements for processing real-time video images, such as live broadcasts or video calls.
[0082] Figure 5 As shown in the flowchart of an image processing method provided by the embodiments of the present invention, this method can be executed by an image processing device, which can be implemented by software and / or hardware and is generally integrated in a computer device. Figure 5 As shown, the method includes:
[0083] Step 501, obtain a face image to be processed.
[0084] Exemplarily, the specific source of the face image to be processed is not limited. It can be a local face image of a computer device, a face image from the network, or a real-time captured face image, etc. Optionally, the face image to be processed can be a real-time video image containing a face in a video call, or a video frame containing a face in a live stream, etc.
[0085] Step 502, input the face image to be processed into the face image processing model to output a target face image corresponding to the face image to be processed after blemish removal.
[0086] Among them, the face image processing model is obtained through the model training method provided by the embodiments of the present invention.
[0087] Inputting the face image to be processed into the face image processing model provided by the embodiments of the present invention can output a target face image after blemish removal, thereby achieving a beauty effect.
[0088] The image processing method provided by the embodiments of the present invention can obtain a more realistic and natural flawless face image, and does not require multi-stage processing. It outputs a flawless face image at one time, which can improve the efficiency of blemish removal processing, enhance real-time performance, and is well applicable to application scenarios with high timeliness requirements for processing real-time video images, such as live broadcasts or video calls.
[0089] Figure 6The following is a structural block diagram of a model training device provided by an embodiment of the present invention. The device can be implemented by software and / or hardware, and is generally integrated in a computer device. Model training can be performed by executing a model training method. As Figure 6 shown, the device includes:
[0090] A sample set acquisition module 601, configured to acquire a first training sample set including training sample pairs, where each training sample pair includes an original face sample image with defects and a target face sample image obtained by removing the defects based on the original face sample image;
[0091] A model training module 602, configured to train a preset original network model based on the first training sample set using a preset loss function to obtain a target network model. The preset original network model includes a generative adversarial network with parameters to be adjusted and a pre-trained defect segmentation network with fixed parameters. The generative adversarial network includes a generator network and a discriminator network. The generator network is configured to generate a defect-free face image in the target domain, and the defect segmentation network is configured to segment the generated image output by the generator network to obtain a segmentation result based on defect regions and non-defect regions. The preset loss function includes a defect suppression loss function, which is obtained by transforming the segmentation result and is used to constrain the generation of defects in the generation process of the generated image;
[0092] A model determination module 603, configured to determine a face image processing model according to the target generator network included in the target network model. The face image processing model is configured to process a face image to be processed to remove the defects included in the face image to be processed.
[0093] In the model training device provided by the embodiment of the present invention, a combination of a generative adversarial network and a defect segmentation network is adopted. The generative adversarial network is responsible for generating defect-free faces, which is beneficial to maintaining the consistency of other features except skin texture. The defect segmentation network, as an auxiliary network, can promote the generation process of the generative adversarial network. Combining the defect suppression loss function enables the generator network to learn to eliminate defects and then output defect-free images. After training, the trained target generator network is used as the face image processing model for defect elimination. Using this model, a more realistic and natural defect-free face image can be obtained, and it is not necessary to go through multiple stages of processing. The defect-free face image is output at one time, which can improve the efficiency of defect removal processing and enhance real-time performance, and can be well applied to application scenarios with high timeliness requirements such as live broadcast or video call for processing real-time video images.
[0094] Figure 7The following is a structural block diagram of an image processing device provided by an embodiment of the present invention. The device can be implemented by software and / or hardware, and is generally integrated in a computer device, and can perform image processing by executing an image processing method. As Figure 7 shown, the device includes:
[0095] An image to be processed acquisition module 701, configured to acquire a face image to be processed;
[0096] An image processing module 702, configured to input the face image to be processed into a face image processing model to output a target face image corresponding to the face image to be processed after defect removal processing, where the face image processing model is obtained by the model training method provided by the embodiment of the present invention.
[0097] The image processing device provided by the embodiment of the present invention can obtain a more realistic and natural flawless face image, and does not require multi-stage processing, and outputs a flawless face image at one time, which can improve the efficiency of defect removal processing, enhance real-time performance, and can be well applied to application scenarios with high timeliness requirements such as live broadcast or video call for processing real-time video images.
[0098] The embodiment of the present invention provides a computer device, and the model training device provided by the embodiment of the present invention can be integrated in the computer device. Figure 8 The following is a structural block diagram of a computer device provided by an embodiment of the present invention. The computer device 800 includes a memory 801, a processor 802, and a computer program stored on the memory 801 and executable on the processor 802. When the processor 802 executes the computer program, the model training method and / or the image processing method provided by the embodiment of the present invention are implemented.
[0099] The embodiment of the present invention further provides a storage medium including computer-executable instructions, and the computer-executable instructions are used to execute the model training method and / or the image processing method provided by the embodiment of the present invention when executed by a computer processor.
[0100] The model training device, the image processing device, the device, and the storage medium provided in the above embodiments can execute the corresponding methods provided by any embodiment of the present invention, and have the corresponding functional modules and beneficial effects of the execution methods. For technical details not described in detail in the above embodiments, reference may be made to the model training method and the image processing method provided by any embodiment of the present invention.
[0101] Note that the above is only a preferred embodiment of the present invention. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments only. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the claims.
Claims
1. A model training method, characterized in that, Including: Obtain a first training sample set including training sample pairs, where each training sample pair includes an original face sample image with defects and a target face sample image obtained by removing the defects based on the original face sample image; Use the first training sample set to train a preset original network model based on a preset loss function to obtain a target network model. The preset original network model includes a generative adversarial network with parameters to be adjusted and a pre-trained defect segmentation network with fixed parameters. The generative adversarial network includes a generator network and a discriminator network. The generator network is used to generate a defect-free face image in the target domain. The defect segmentation network is used to segment the generated image output by the generator network to obtain a segmentation result based on the defect region and the non-defect region. The preset loss function includes a defect suppression loss function, which is obtained by transforming the segmentation result and is used to constrain the generation of defects in the generated image during the generation process; Determine a face image processing model according to the target generator network included in the target network model, where the face image processing model is used to process a face image to be processed to remove the defects included in the face image to be processed.
2. The method according to claim 1, wherein During the process of using the first training sample set to train the preset original network model based on the preset loss function, the generator network takes the original face sample image as input and outputs a generated image. The discriminator network takes the corresponding target face sample image and the generated image as input and outputs whether the target face sample image and the generated image are the same.
3. The method according to claim 2, characterized in that, The preset loss function further includes a reconstruction loss function and an adversarial loss function, where the reconstruction loss function is used to constrain the gap between the generated image and the corresponding target face sample image.
4. The method according to claim 1, wherein Before using the training sample set to train the preset original network model based on the preset loss function, it further includes: Obtain a second training sample set including the training sample pairs; According to the differences between the original face sample images and the target face sample images in each training sample pair in the second training sample set, segment and label the original face sample images to obtain target segmentation images including defect regions and non-defect regions; Use the original face sample images as the input of a preset original defect segmentation network, and use the corresponding target segmentation images as the expectation, and train the preset original defect segmentation network based on a preset segmentation loss function to obtain a defect segmentation network.
5. The method according to claim 4, characterized in that, The segmenting and labeling the original face sample images according to the differences between the original face sample images and the target face sample images in each training sample pair in the second training sample set includes: For each training sample pair in the second training sample set, calculate the absolute difference between the first pixel and the second pixel corresponding to each pixel position, where the first pixel is from the original face sample image and the second pixel is from the target face sample image; Determine a target segmentation threshold according to each of the absolute differences; Segment and label the original face sample image by using the target segmentation threshold. Among them, the part corresponding to the absolute difference greater than or equal to the target segmentation threshold is labeled as the defect area, and the part corresponding to the absolute difference less than the target segmentation threshold is labeled as the non-defect area.
6. The method according to claim 4, wherein The preset segmentation loss function includes the Dice loss function.
7. According to the method according to any one of claims 1-6, characterized in that, The defect suppression loss function is represented by the total number of defect areas included in the generated image output by the defect segmentation network.
8. An image processing method, characterized in that, It includes: Obtain the face image to be processed; Input the face image to be processed into the face image processing model to output the target face image corresponding to the face image to be processed after defect removal, where the face image processing model is obtained by the model training method according to any one of claims 1-7.
9. A model training device, characterized in that, It includes: A sample set acquisition module, configured to acquire a first training sample set including training sample pairs, where each training sample pair includes an original face sample image with defects and a target face sample image after defect removal based on the original face sample image; A model training module, configured to train a preset original network model based on a preset loss function by using the first training sample set to obtain a target network model, where the preset original network model includes a generative adversarial network with parameters to be adjusted and a defect segmentation network with pre-trained fixed parameters. The generative adversarial network includes a generator network and a discriminator network. The generator network is used to generate a defect-free face image in the target domain, and the defect segmentation network is used to segment the generated image output by the generator network to obtain a segmentation result based on the defect area and the non-defect area. The preset loss function includes a defect suppression loss function, and the defect suppression loss function is transformed from the segmentation result and is used to constrain the generation of defects in the process of generating the generated image; A model determination module, configured to determine a face image processing model according to the target generator network included in the target network model, where the face image processing model is used to process the face image to be processed to remove the defects included in the face image to be processed.
10. An image processing apparatus, characterized in that, It includes: An image to be processed acquisition module, configured to acquire a face image to be processed; An image processing module, configured to input the face image to be processed into the face image processing model to output the target face image corresponding to the face image to be processed after defect removal, where the face image processing model is obtained by the model training method according to any one of claims 1-7.
11. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1-8 is implemented.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method according to any one of claims 1-8 is implemented.
Citation Information
Patent Citations
Glass flaw detection method based on variational auto-encoder
CN113344903A
Defective Pixel Correction Using Adversarial Networks
US20200065945A1