Image enhancement model processing method and device, storage medium and program product

By degrading the face image and optimizing the image enhancement model, the problem of low-quality face images is solved and more efficient image quality improvement is achieved.

CN120125459APending Publication Date: 2025-06-10XIAMEN MEITUEVE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510177953.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

During the process of facial image acquisition, the quality of facial images acquired by low-cost equipment or non-standardized lighting conditions is often low, resulting in difficulty in subsequent analysis and poor enhancement effect of existing image processing algorithms.

Method used

A variety of low-quality sample images are generated by degrading the original face image, and the initial image enhancement model is used to enhance these sample images. The training loss value is calculated based on the enhanced image and the original image, and the model is optimized until the convergence conditions are reached, and a trained image enhancement model is obtained for improving the quality of the face image.

Benefits of technology

This method can stably improve image quality when facing various low-quality facial images, and significantly improve the quality enhancement effect of facial images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125459A_ABST
    Figure CN120125459A_ABST
Patent Text Reader

Abstract

The invention relates to an image enhancement model processing method and device, a storage medium and a program product. The method comprises the following steps: performing degradation processing on an original face image to obtain a sample face image; performing enhancement processing on the sample face image through an initial image enhancement model to obtain an enhanced face image; determining a training loss value based on the enhanced face image and the original face image; optimizing the initial image enhancement model based on the training loss value until a convergence condition is reached, and obtaining a trained image enhancement model; the image enhancement model is used for enhancing the face image. By adopting the method, the quality enhancement effect of the face image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and particularly to a method and apparatus for processing an image enhancement model, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] During the process of face image acquisition, the image quality is easily affected by the performance of the acquisition device and the environmental lighting conditions. Especially under low-cost devices or non-standard lighting conditions, the acquired face images often lack sufficient clarity, and the low quality of the face images makes it impossible to accurately analyze the face images subsequently.

[0003] Currently, the enhancement of the image quality of face images mainly relies on traditional image processing algorithms, such as contrast stretching, histogram equalization, sharpening, etc. However, the enhancement effects of these algorithms are poor. Summary of the Invention

[0004] Based on this, this application provides a method and apparatus for processing an image enhancement model, a computer device, a computer-readable storage medium, and a computer program product, which can improve the enhancement effect of the quality of face images.

[0005] On the one hand, this application provides a method for processing an image enhancement model, including:

[0006] Performing degradation processing on an original face image to obtain a sample face image;

[0007] Performing enhancement processing on the sample face image through an initial image enhancement model to obtain an enhanced face image;

[0008] Determining a training loss value based on the enhanced face image and the original face image;

[0009] Optimizing the initial image enhancement model based on the training loss value until a convergence condition is reached to obtain a trained image enhancement model; the image enhancement model is used to enhance face images.

[0010] In one embodiment, the degradation processing includes skin problem removal processing; the performing degradation processing on an original face image to obtain a sample face image includes:

[0011] Performing skin problem removal processing on the original face image according to the maximum removal intensity to obtain a face image after problem removal;

[0012] Performing image fusion on the face image after problem removal and the original face image according to a target removal intensity to obtain a partially face image after problem removal;

[0013] Generate a sample face image based on the post-removal face image with the target removal intensity.

[0014] In one embodiment, the degradation process further includes noise addition processing; the generating a sample face image based on the post-removal face image with the target removal intensity includes:

[0015] Perform noise addition processing on the post-removal face image with the target removal intensity according to a preset noise addition algorithm to obtain a sample face image.

[0016] In one embodiment, the training loss value includes a first training loss value and a second training loss value; the initial image enhancement model includes a denoising sub-model and an enhancement sub-model; the enhancing the sample face image through the initial image enhancement model to obtain an enhanced face image includes:

[0017] Perform denoising processing on the sample face image through the denoising sub-model to obtain a denoised sample face image;

[0018] Perform skin problem enhancement processing on the denoised sample face image through the enhancement sub-model to obtain an enhanced face image;

[0019] The determining the training loss value based on the enhanced face image and the original face image includes:

[0020] Determine a first training loss value based on the enhanced face image and the original face image;

[0021] Determine a second training loss value based on the denoised sample face image and the post-removal face image with the target removal intensity.

[0022] In one embodiment, the method further includes determining a first training loss value based on the enhanced face image and the original face image, including:

[0023] Determine a skin problem mask based on the original face image and the post-removal face image with partial problems removed;

[0024] Determine a basic loss value based on the enhanced face image and the original face image;

[0025] Adjust the basic loss value based on the skin problem mask to obtain a first training loss value.

[0026] In one embodiment, the determining the second training loss value based on the denoised sample face image and the post-removal face image with the target removal intensity includes:

[0027] Determine an image difference value based on the denoised sample face image and the target removal intensity for the post-removal face image;

[0028] Determine a second training loss value according to the image difference value.

[0029] In one embodiment, the method further includes:

[0030] Obtain a face image to be processed;

[0031] Perform enhancement processing on the face image to be processed through the image enhancement model to obtain a processed face image.

[0032] On the one hand, the present application also provides a processing device for an image enhancement model, including:

[0033] A sample face image acquisition module, configured to perform degradation processing on an original face image to obtain a sample face image;

[0034] A face image enhancement module, configured to perform enhancement processing on the sample face image through an initial image enhancement model to obtain an enhanced face image;

[0035] A loss value determination module, configured to determine a training loss value based on the enhanced face image and the original face image;

[0036] A model optimization module, configured to optimize the initial image enhancement model based on the training loss value until a convergence condition is reached to obtain a trained image enhancement model.

[0037] On the one hand, the present application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of any of the above methods are implemented.

[0038] On the one hand, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above methods are implemented.

[0039] On the one hand, the present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above methods are implemented.

[0040] The processing method, device, computer equipment, computer-readable storage medium and computer program product of the above image enhancement model can obtain a variety of low-quality sample face images by degrading the original face image, which can better simulate the image quality problems that occur in the actual acquisition process, so that the subsequent model training is more targeted. The sample face images are enhanced by the initial image enhancement model to obtain enhanced face images; the training loss value is determined based on the enhanced face images and the original face images; the initial image enhancement model is optimized based on the training loss value until the convergence condition is reached, so that the trained image enhancement model can stably improve the image quality when facing various low-quality facial images, thereby improving the effect of quality enhancement for face images. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.

[0042] Figure 1 It is an application environment diagram of the processing method of the image enhancement model in an embodiment;

[0043] Figure 2 It is a flowchart of the processing method of the image enhancement model in an embodiment;

[0044] Figure 3 It is a schematic diagram of the image enhancement model in an embodiment;

[0045] Figure 4 It is a schematic diagram of a face image in an embodiment;

[0046] Figure 5 It is a flowchart of the processing method of the image enhancement model in another embodiment;

[0047] Figure 6 It is a flowchart of the processing method of the image enhancement model in another embodiment;

[0048] Figure 7 It is a structural block diagram of the processing device of the image enhancement model in an embodiment;

[0049] Figure 8 It is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] In order to make the objectives, technical solutions, and beneficial effects of this application more clearly understood, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not used to limit this application.

[0051] The processing method of the image enhancement model provided by the embodiments of this application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed in the cloud or on other network servers. The processing method of this image enhancement model can be executed independently by the terminal 102 or the server 104, or jointly executed by the terminal 102 and the server 104. In some embodiments, the processing method of this image enhancement model is executed by the server 104. The server 104 performs a degradation process on the original face image to obtain a sample face image; performs an enhancement process on the sample face image through an initial image enhancement model to obtain an enhanced face image; determines a training loss value based on the enhanced face image and the original face image; optimizes the initial image enhancement model based on the training loss value until the convergence condition is reached to obtain a trained image enhancement model; the image enhancement model is used to enhance the face image.

[0052] Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, skin detection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server 104 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0053] In an exemplary embodiment, the processing method of this image enhancement model is executed by a computer device, and the computer device can be, for example, Figure 1 the terminal 102 or the server 104 shown. As Figure 2 shown, the method can include the following steps:

[0054] S202, perform a degradation process on the original face image to obtain a sample face image.

[0055] Among them, the original face image is a high-quality face image, specifically, it can be a high-definition face image; the sample face image is a low-quality face image, specifically, it can be an image obtained by degrading the original face image.

[0056] The degradation process refers to performing a series of operations on the high-quality original face image to reduce its quality to simulate low-quality images in actual acquisition. Specifically, it can be at least one of noise addition processing and skin problem removal processing. Noise addition processing means adding random noise to the original high-quality image to simulate the degradation of image quality caused by sensor problems, signal interference, or information loss during compression. Skin problem removal processing means eliminating skin problems (such as pimples, spots, wrinkles, pores, etc.) on the skin surface through certain algorithms, that is, eliminating the details on the skin surface to make the image look more blurred or distorted.

[0057] Specifically, the computer device can obtain the original face image collected by the high-definition image acquisition device and perform degradation processing on the original face image according to a preset degradation processing algorithm to obtain a sample image.

[0058] Among them, the high-definition image acquisition device refers to an image acquisition device with a high image resolution, a low noise level, and a strong detail capture ability, which can capture clear and detailed images. Specifically, it can be devices such as a high-resolution camera, a high-definition camera, a depth camera, and a professional face recognition camera.

[0059] S204, perform enhancement processing on the sample face image through the initial image enhancement model to obtain an enhanced face image.

[0060] Among them, the initial image enhancement model is an image enhancement model to be trained. The initial image enhancement model can specifically be a neural network model, specifically a neural network model based on deep learning.

[0061] In one embodiment, the initial image enhancement model includes an enhancement sub-model. After the computer device obtains the post-removal face image with the target removal intensity, it can directly use the post-removal face image with the target removal intensity as the sample face image and input it into the enhancement sub-model of the initial image enhancement model. The enhancement sub-model performs skin problem enhancement processing on the sample face image through each network layer and outputs an enhanced face image.

[0062] Among them, the enhancement sub-model can specifically adopt a U-shaped network architecture and perform multi-scale processing on the image in combination with the Laplacian pyramid.

[0063] The enhancement processing process of the enhancement sub-model is illustrated by the following example, such as Figure 3Schematic diagram of the enhancer model shown on the right side in the middle. The enhancer model has a U-shaped network architecture. The input image is first decomposed into three high-frequency layers at different scales ( ) and a low-frequency layer . The low-frequency layer is enhanced through a BGN module to obtain the enhanced low-frequency layer . The enhanced low-frequency layer is gradually upsampled and fused with the high-frequency layer to generate the enhanced high-frequency layer ( ). Finally, the enhanced high-frequency layer is re-fused with the enhanced low-frequency layer through Laplacian reconstruction to reconstruct the enhanced face image .

[0064] In one embodiment, the initial image enhancement model includes an enhancer model and a denoising model. After the computer device performs noise addition processing on the face image after removing the target removal intensity according to a preset noise addition algorithm to obtain a sample face image, the sample face image can be input into the initial image enhancement model, and denoising processing and skin problem enhancement processing are performed through the denoising model and the enhancer model in turn to obtain the enhanced face image.

[0065] Among them, the denoising model can specifically be a denoising network, such as Figure 3 the schematic diagram of the denoising model shown on the left side in the middle on the right. The denoising model is SCUNet, and it can also be replaced with other denoising models according to specific applications. The input of the denoising model is the original image in RGB format, and the output after processing is an image with reduced noise.

[0066] S206, determining the training loss value based on the enhanced face image and the original face image.

[0067] Among them, the training loss value is a quantitative index used in machine learning and deep learning models during the training process to measure the difference between the model output (i.e., the prediction result) and the true label or target value. The smaller the training loss value, the better the performance of the model on the training data, and the closer the output prediction result is to the true label.

[0068] In one embodiment, when the initial image enhancement model includes an enhancer model, the computer device can directly determine the first training loss value based on the difference between the enhanced face image and the original face image, and use this first training loss value as the training loss value.

[0069] S208, optimizing the initial image enhancement model based on the training loss value until the convergence condition is reached to obtain the trained image enhancement model.

[0070] Among them, the trained image enhancement model is used to enhance the face image. The convergence condition refers to the stopping criterion of the training process, which means that after the parameters of the model are updated multiple times, they reach a stable state, no longer change significantly or the change is small enough. Specifically, the convergence condition can be that the training loss no longer decreases significantly or the maximum number of iterations is reached.

[0071] Specifically, after obtaining the training loss value, the computer device can determine the gradient of the network parameters of the initial image enhancement model based on the training loss value, update the network parameters of the initial image enhancement model according to the gradient to obtain an updated image enhancement model, and check whether the updated image enhancement model has reached the convergence condition. If not, continue to train the updated image enhancement model using other sample face images until the convergence condition is reached. When the convergence condition is reached, stop training to obtain the trained image enhancement model.

[0072] In the above processing method of the image enhancement model, various low-quality sample face images can be obtained by degrading the original face image, which can better simulate the image quality problems that occur in the actual acquisition process, making the subsequent model training more targeted. The sample face images are enhanced by the initial image enhancement model to obtain enhanced face images; the training loss value is determined based on the enhanced face images and the original face images; the initial image enhancement model is optimized based on the training loss value until the convergence condition is reached, so that the trained image enhancement model can stably improve the image quality when facing various low-quality facial images, thereby improving the effect of face image quality enhancement.

[0073] In one embodiment, the degradation process includes skin problem removal processing; the process by which the computer device degrades the original face image to obtain sample face images includes the following steps: perform skin problem removal processing on the original face image according to the maximum removal intensity, perform skin problem removal processing on the original face image to obtain a face image after problem removal; perform image fusion on the face image after problem removal and the original face image according to the target removal intensity to obtain a face image after removal with the target removal intensity; generate sample face images based on the face image after removal with the target removal intensity.

[0074] Among them, the removal intensity refers to a parameter that controls the degree or intensity of removing at least one skin problem in the original face image during the skin problem removal process. The skin problems can specifically be acne, spots, wrinkles, pores, etc. The maximum removal intensity refers to maximizing the removal of at least one skin problem in the skin problem removal process, that is, performing the strongest removal or repair process on the image. The target removal intensity is less than the maximum removal intensity, and the target removal intensity can be set according to actual needs. For example, if the maximum removal intensity is characterized as 1, the target removal intensity can be set to 0.2, 0.5, 0.8, etc.

[0075] The skin problem removal process can be specifically implemented through manual retouching, traditional image algorithms, neural network models, etc.

[0076] Specifically, for a certain skin problem, the computer device can use the removal method corresponding to the skin problem to perform skin problem removal processing on the original face image, so as to achieve skin problem removal processing of the original face image for the skin problem according to the maximum removal intensity, obtain the face image after problem removal, and then determine the first weight corresponding to the face image after problem removal and the second weight corresponding to the original face image based on the target removal intensity, and fuse the face image after problem removal and the original face image based on the first weight and the second weight to obtain the face image after removal with the target removal intensity, and generate a sample face image according to the face image after removal with the target removal intensity.

[0077] Taking skin problems such as freckles and acne problems (which can be denoted as freckle-acne problems) as an example, the above embodiments are described. Assume that the original face image is a high-definition image , with a size of (3xHxW), and perform freckle and acne removal on the high-definition image through any one of manual retouching, traditional image algorithms, and existing freckle and acne removal network models to obtain the face image after problem removal (3xHxW), and use the fusion method shown in the following formula to fuse the high-definition image and the face image after removal to obtain the face image after removal with the target removal intensity . .

[0078]

[0079] Among them, is the first weight corresponding to the face image after problem removal , which can also be called the fusion coefficient, and is specifically determined according to the target removal intensity. is the second weight corresponding to the original face image . The value range of is between 0 and 1, and can control the face image after problem removal and the original face image

[0080] to be fused at different ratios. Taking skin problem as an example of wrinkle problem, the above embodiments are described. Assume that the original face image is a high-definition image , with a size of (3xHxW). The high-definition image is subjected to wrinkle removal by any one of manual image retouching, traditional image algorithms, and existing anti-wrinkle network models, and the face image after problem removal (3xHxW) is obtained. The high-definition image and the face image after removal are fused to obtain the face image after removal with the target removal intensity .

[0081] Taking skin problem as an example of pore problem, the above embodiments are described. Assume that the original face image is a high-definition image , with a size of (3xHxW). The high-definition image is subjected to pore removal by any one of manual image retouching, traditional image algorithms, and existing anti-pore network models, and the face image after problem removal (3xHxW) is obtained. The high-definition image and the face image after removal are fused to obtain the face image after removal with the target removal intensity .

[0082] In addition, since pores are generally small and widely distributed in the face image, in addition to obtaining the face image after removal with the target removal intensity by the method in the above example , it can also be obtained by the following method: After the high-definition image is bilinearly filtered, the pore area will be smoothed, achieving the effect of pore removal or pore lightening. By adjusting the parameters of the bilinear filter, different degrees of pore removal can be controlled. Different from the method in the above example, the face image after removal with the target removal intensity is directly obtained after bilinear filtering, and there is no need to fuse it with the high-definition image any more.

[0083] It can be understood that when the degradation process only includes skin problem removal processing, the process of the computer device generating the sample face image according to the face image after removal with the target removal intensity can specifically be: directly using the face image after removal with the target removal intensity as the sample face image.

[0084] In the above embodiments, the computer device performs skin problem removal processing on the original face image according to the maximum removal intensity, and obtains the face image after problem removal; performs image fusion on the face image after problem removal and the original face image according to the target removal intensity to obtain the face image after removal with the target removal intensity; generates a sample face image based on the face image after removal with the target removal intensity, so as to simulate a low-quality face image for training the model, enabling the trained image enhancement model to stably improve the image quality when facing various low-quality facial images, thereby improving the effect of quality enhancement for face images.

[0085] In one embodiment, the degradation processing further includes noise addition processing. The process by which the computer device generates a sample face image based on the face image after removal with the target removal intensity includes: performing noise addition processing on the face image after removal with the target removal intensity according to a preset noise addition algorithm to obtain the sample face image.

[0086] Among them, noise addition processing refers to adding a certain amount of noise to the image. The preset noise addition algorithm can specifically be at least one of a picture quality compression algorithm, a Gaussian noise addition algorithm, a Poisson noise addition algorithm, and an image blurring processing algorithm. The picture quality compression algorithm is used to simulate the distortion during image compression, and can specifically be a JPEG format compression algorithm. Gaussian noise is a type of random noise based on the Gaussian distribution, and the Gaussian noise addition algorithm is used to introduce random pixel value changes in the image; Poisson noise is a type of noise based on the Poisson distribution, and the Poisson noise addition algorithm is used to introduce changes based on uneven brightness changes in the image; image blurring processing makes the image smoother by reducing the details of the image. The image blurring processing algorithm can specifically be a mean blurring processing algorithm or a Gaussian blurring processing algorithm. The mean blurring processing algorithm replaces the central pixel with the average value of pixel values within a local area, thereby making the image more blurred. The Gaussian blurring processing algorithm is a weighted blurring method that blurs the image using a Gaussian function. Different from the mean blurring processing algorithm, the Gaussian blurring processing algorithm performs smoothing by assigning different weights to the neighborhood around the pixel, and the weights are usually larger in the center and gradually decrease towards the periphery.

[0087] Specifically, after the computer device obtains the face image after removal with the target removal intensity, it can determine the preset noise addition algorithm to be used, and then perform noise addition processing on the face image after removal with the target removal intensity according to the preset noise addition algorithm to obtain the sample face image.

[0088] Assume that for any image to be subjected to noise addition processing, it is called the target image. Then the computer device can use at least one of the following methods to perform noise addition processing on the face image after removal with the target removal intensity to obtain the sample face image:

[0089] 1) Randomly compress the quality of the target image using JPEG compression, where the value of the random ratio can be between 60 and 95, so that when the model is trained using the compressed image obtained later, the model can learn to eliminate the noise caused by JPEG compression at different ratios;

[0090] 2) Add Gaussian noise with different intensities to the target image, where the value of the random intensity can be between 2 and 20. It should be noted that there is a 50% chance of adding noise independently and randomly on the three RGB channels, and a 50% chance of adding the same noise to the three RGB channels.

[0091] 3) Add Poisson noise with different intensities to the target image, where the value of the random intensity can be between 1 and 6. It should be noted that there is a 50% chance of adding noise independently and randomly on the three RGB channels, and a 50% chance of adding the same noise to the three RGB channels.

[0092] 4) Add mean blur with different kernel sizes to the target image, where the value of the kernel size can be randomly selected from 3, 5, and 7.

[0093] 5) Add Gaussian blur with different intensities to the target image, where the standard deviation sigma corresponding to the Gaussian function is randomly selected between 0 and 1.

[0094] In the above embodiments, the computer device performs noise addition processing on the post-removal human face image with the target removal intensity according to a preset noise addition algorithm to obtain a sample human face image, so as to simulate different low-quality human face images for training the model, enabling the trained image enhancement model to stably improve the image quality when facing various low-quality facial images, thereby improving the effect of quality enhancement for human face images.

[0095] In one embodiment, the training loss value includes a first training loss value and a second training loss value; the initial image enhancement model includes a denoising sub-model and an enhancement sub-model; the process of the computer device enhancing the sample human face image through the initial image enhancement model to obtain an enhanced human face image includes the following steps: performing denoising processing on the sample human face image through the denoising sub-model to obtain a denoised sample human face image; performing skin problem enhancement processing on the denoised sample human face image through the enhancement sub-model to obtain an enhanced human face image; the process of the computer device determining the training loss value based on the enhanced human face image and the original human face image includes the following steps: determining the first training loss value based on the enhanced human face image and the original human face image; determining the second training loss value based on the denoised sample human face image and the post-removal human face image with the target removal intensity.

[0096] Specifically, after the computer device performs noise addition processing on the post-removal human face image with the target removal intensity according to the preset noise addition algorithm to obtain a sample human face image, the sample human face image can be input into the initial image enhancement model. The sample human face image is processed through each network layer of the denoising sub-model to obtain a denoised sample human face image. The denoised sample human face image is input into the enhancement sub-model, and the enhancement sub-model performs skin problem enhancement processing on the denoised sample human face image to obtain an enhanced human face image. After that, the computer device can determine a first training loss value based on the enhanced human face image and the original human face image; determine a second training loss value based on the denoised sample human face image and the post-removal human face image with the target removal intensity; adjust the network parameters of the enhancement sub-model based on the first training loss value, and adjust the network parameters of the denoising sub-model based on the second training loss value until the convergence condition is reached, and obtain a trained image enhancement model.

[0097] As Figure 3 shown, the image enhancement model includes a denoising sub-model and an enhancement sub-model, where the denoising sub-model is SCUNet and the enhancement sub-model is a U-shaped network architecture.

[0098] In the above embodiment, the computer device performs denoising processing on the sample human face image through the denoising sub-model to obtain a denoised sample human face image; performs skin problem enhancement processing on the denoised sample human face image through the enhancement sub-model to obtain an enhanced human face image; determines a first training loss value based on the enhanced human face image and the original human face image; determines a second training loss value based on the denoised sample human face image and the post-removal human face image with the target removal intensity, so that the network parameters of the enhancement sub-model can be adjusted based on the first training loss value, and the network parameters of the denoising sub-model can be adjusted based on the second training loss value until the convergence condition is reached, and a trained image enhancement model is obtained, enabling the trained image enhancement model to stably improve the image quality when facing various low-quality facial images, thereby improving the effect of quality enhancement for human face images.

[0099] In one embodiment, the process of the computer device determining the first training loss value based on the enhanced human face image and the original human face image includes the following steps: determining a skin problem mask based on the original human face image and the post-removal human face image with the target removal intensity; determining a basic loss value based on the enhanced human face image and the original human face image; adjusting the basic loss value based on the skin problem mask to obtain the first training loss value.

[0100] Among them, the basic training loss value reflects the deviation between the output of the enhancer model and the original face image. When determining the basic training loss value, the L1 loss function or the L2 loss function can be used. In addition, perceptual loss (such as VGGLoss) and structural similarity metrics such as SSIM can also be added to enhance the model's perception ability of details; the L1 loss function calculates the average of the absolute differences of each pixel between the enhanced image and the original image; the L2 loss function calculates the average of the squares of the differences of each pixel between the enhanced image and the original image; the perceptual loss is usually calculated by extracting the features of the intermediate layer in a pre-trained convolutional neural network (such as the VGG network) to calculate the perceptual difference between the enhanced image and the original image, rather than directly calculating the difference at the pixel level; SSIM is an index used to measure the structural similarity of two images, considering brightness, contrast and structural information.

[0101] Specifically, after the computer device performs skin problem removal processing on the original face image to obtain the face image after removal with the target removal intensity, the difference value between the face image after removal with the target removal intensity and the original face image can be determined, and this difference value is used as the skin problem mask corresponding to the face image after removal with the target removal intensity. After obtaining the enhanced face image, the difference between the enhanced face image and the original face image can be determined, and the basic loss value can be determined according to this difference, and the basic loss value can be adjusted based on the skin problem mask to obtain the first training loss value.

[0102] For example, for the face image after removal with the target removal intensity and the original face image , the corresponding skin problem mask can be characterized as follows:

[0103]

[0104] For example, for the face image after removal with the target removal intensity and the original face image , the corresponding skin problem mask can be characterized as follows:

[0105]

[0106] In one embodiment, the process of adjusting the basic loss value based on the skin problem mask to obtain the first training loss value includes the following steps: performing a dot product operation on the basic loss value and the skin problem mask, and using the result of the dot product operation as the first training loss value.

[0107] In the above embodiments, the computer device determines a skin problem mask based on the post-removal face image obtained from the original face image and the target removal intensity; determines a basic loss value based on the enhanced face image and the original face image; adjusts the basic loss value based on the skin problem mask to obtain a first training loss value, so that the network parameters of the enhancer sub-model can be adjusted based on the first training loss value, enabling the model to perform more precise optimization on the skin problem areas, highlighting the detailed performance in areas such as age spots and wrinkles, and ensuring a significant improvement in the image quality of these areas.

[0108] In one embodiment, the process by which the computer device determines a second training loss value based on the post-removal face image obtained from the denoised sample face image and the target removal intensity includes the following steps: determining an image difference value based on the denoised sample face image and the post-removal face image obtained from the target removal intensity; determining the second training loss value according to the image difference value.

[0109] Specifically, after the computer device determines the image difference value based on the denoised sample face image and the post-removal face image obtained from the target removal intensity, it can use L1, L2 loss, or other similarity measurement methods to convert the image difference value into a second training loss value. The second training loss value is mainly used to guide the denoising sub-model to optimize its denoising ability and reduce the impact of noise on the image quality.

[0110] In the above embodiments, the computer device determines an image difference value based on the denoised sample face image and the post-removal face image obtained from the target removal intensity; determines the second training loss value according to the image difference value, so that the network parameters of the denoising sub-model can be adjusted based on the second training loss value, enabling the model to accurately remove the noise in the image and further improve the image quality.

[0111] In one embodiment, the processing method of the above image enhancement model further includes the following steps: obtaining a face image to be processed; performing enhancement processing on the face image to be processed through the image enhancement model to obtain a processed face image.

[0112] Among them, the face image to be processed is an image that needs to enhance the image quality, and specifically can be a face image collected by an image acquisition device with poor image quality.

[0113] Specifically, after the computer device obtains the face image to be processed, it can input the face image to be processed into the trained image enhancement model, perform denoising processing on the face image to be processed through the denoising sub-model of the image enhancement model to obtain a denoised face image, and input the denoised face image into the enhancer sub-model of the image enhancement model, and perform skin problem enhancement processing on the denoised face image through the enhancer sub-model to obtain a processed face image.

[0114] In one embodiment, the computer device may also obtain the enhancement intensity of the target skin problem, and input the denoised face image and the enhancement intensity of the target skin problem into the enhancement sub-model of the image enhancement model. The enhancement sub-model performs skin problem enhancement processing on the denoised face image for the target skin problem according to the skin problem enhancement intensity, and obtains the processed face image.

[0115] Among them, the target skin problem may be at least one of color acne problem, wrinkle problem, and pore problem, and the skin problem enhancement intensity may specifically be a value between 0 and 1.

[0116] As Figure 4 shown, Figure 4 in (A) is the face image collected by an image acquisition device with poor image quality, Figure 4 in (B) is the processed face image obtained by enhancing it through the image enhancement model. It can be seen from the figure that for the processed face image obtained by the enhancement processing, compared with the original low-quality face image, the details in areas such as skin spots and wrinkles are enhanced.

[0117] In one embodiment, a processing method for an image enhancement model is also provided. The processing method of the image enhancement model is executed by a computer device, and the computer device may be, for example, Figure 1 shown in the terminal 102 or the server 104. As Figure 5 shown, the method may include the following steps:

[0118] S502, perform skin problem removal processing on the original face image according to the maximum removal intensity to obtain the face image after problem removal; perform image fusion on the face image after problem removal and the original face image according to the target removal intensity to obtain the partially face image after problem removal; perform noise addition processing on the face image after removal of the target removal intensity according to the preset noise addition algorithm to obtain the sample face image.

[0119] S504, perform denoising processing on the sample face image through the denoising sub-model of the initial image enhancement model to obtain the denoised sample face image.

[0120] S506, perform skin problem enhancement processing on the denoised sample face image through the enhancement sub-model of the initial image enhancement model to obtain the enhanced face image.

[0121] S508, determine the first training loss value based on the enhanced face image and the original face image.

[0122] S510, determine the second training loss value based on the denoised sample face image and the face image after removal of the target removal intensity.

[0123] S512. Optimize the network parameters of the enhancer model based on the first training loss value, and optimize the network parameters of the denoising sub-model based on the second training loss value until the convergence condition is reached, to obtain a trained image enhancement model.

[0124] S514. Obtain the face image to be processed; perform enhancement processing on the face image to be processed through the image enhancement model to obtain the processed face image.

[0125] This application also provides an application scenario, which applies the processing method of the above image enhancement model. Refer to Figure 6 the flowchart shown, the process of applying the processing method of the above image enhancement model in this application scenario includes the following steps:

[0126] Step 1. Acquisition and degradation processing of high-definition images.

[0127] Specifically, obtain the high-definition face image taken by a single-lens reflex camera, denoted as , since the image taken by a single-lens reflex camera has a high resolution and quality, with good detail levels, and can be used as the target image for training; it is necessary to perform degradation processing on for specific skin areas. Through the above processing methods for areas such as skin spots, wrinkles, and pores, degrade the original target image to obtain the image after removing these areas, denoted as .

[0128] Step 2. Addition of noise gain.

[0129] After obtaining the image, in order to simulate various noises introduced by shooting with low-quality image devices, during training, perform random noise gain on to generate a noisy image .

[0130] Step 3. Processing of the denoising sub-model.

[0131] During training, input the image after noise gain into the denoising sub-model. The task of the denoising sub-model is to perform denoising processing on the input image, reduce various noise interferences introduced by shooting with low-quality devices, and generate an image closer to the quality of the target image. After denoising processing, the output image is denoted as .

[0132] Step 3. Processing of the enhancer model.

[0133] Next, input the denoised image Continue to be used as input and fed into the enhancer model. The enhancer model is responsible for enhancing specific skin problem areas. Through mechanisms such as Laplacian pyramid decomposition and reconstruction, the enhancer model can process the high-frequency and low-frequency information of the image separately, thereby enhancing the skin detail performance.

[0134] Finally, the enhancer model outputs the enhanced image, denoted as . At this time, should have skin detail effects similar to those of the high-definition image , and at the same time, the noise has been effectively removed.

[0135] Step Four: Loss Calculation and Model Optimization.

[0136] During model training, and the target high-definition image are used for loss calculation. By comparing the difference between and , the loss value is obtained, and this loss value reflects the deviation between the output of the enhancer model and the target high-definition image.

[0137] At the same time, and the image after degradation processing are used for loss calculation. By comparing the difference between and , the loss value is obtained, and this loss value reflects the deviation between the output of the denoising sub-model and the low-noise image after degradation processing.

[0138] In the design of the loss function, in addition to common losses (such as L1 or L2 losses), perceptual losses (such as VGGLoss) and structural similarity metrics such as SSIM are also added to enhance the model's perception ability of details.

[0139] When calculating the loss value of the enhancer model, a region-weighted loss is specially designed. This loss is based on the masks of problem areas such as skin spots and wrinkles generated during the training data pair. The specific implementation method of the region-weighted loss is as follows: First, calculate the L1 or L2 loss value of the entire image, and then perform a dot product operation with the corresponding mask . In this way, the model can be more precisely optimized on the skin problem areas, focusing on enhancing the details of areas such as skin spots and wrinkles to ensure a significant improvement in the image quality of these areas.

[0140] Finally, the loss values and are comprehensively used to optimize the model's parameters through the backpropagation algorithm. Specifically, the loss value It is mainly used to guide the noise reduction sub - model to optimize its denoising ability and reduce the impact of noise on image quality; the loss value is mainly used to guide the enhancement sub - model to improve its ability to enhance skin details, enabling the model to gradually improve the image clarity and detail performance of the skin problem areas. Through multiple iterative trainings, the weights of the model are continuously updated, gradually enhancing its comprehensive performance in noise removal and detail enhancement, achieving the improvement of the image quality of low - quality images, and finally outputting an enhancement effect close to the target high - definition image.

[0141] It should be understood that although the steps in the flowcharts involved in the above - mentioned embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above - mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0142] Based on the same inventive concept, an embodiment of the present application also provides an image enhancement model processing device for implementing the above - mentioned image enhancement model processing method. The solution provided by this device to solve problems is similar to the solution described in the above - mentioned method. Therefore, the specific limitations in one or more embodiments of the following image enhancement model processing device can refer to the limitations on the image enhancement model processing method in the above text, and will not be repeated here.

[0143] In an exemplary embodiment, as Figure 7 shown, an image enhancement model processing device is provided, including: a sample face image acquisition module 702, a face image enhancement module 704, a loss value determination module 706, and a model optimization module 708, where:

[0144] The sample face image acquisition module 702 is used to perform degradation processing on the original face image to obtain a sample face image;

[0145] The face image enhancement module 704 is used to perform enhancement processing on the sample face image through an initial image enhancement model to obtain an enhanced face image;

[0146] The loss value determination module 706 is used to determine a training loss value based on the enhanced face image and the original face image;

[0147] The model optimization module 708 is used to optimize the initial image enhancement model based on the training loss value until the convergence condition is reached, and a trained image enhancement model is obtained.

[0148] In the above embodiments, a variety of low-quality sample face images can be obtained by degrading the original face image, which can better simulate the image quality problems that occur in the actual acquisition process, making the subsequent model training more targeted. The sample face images are enhanced by the initial image enhancement model to obtain enhanced face images; the training loss value is determined based on the enhanced face images and the original face images; the initial image enhancement model is optimized based on the training loss value until the convergence condition is reached, so that the trained image enhancement model can stably improve the image quality when facing various low-quality facial images, thereby improving the effect of quality enhancement for face images.

[0149] In one embodiment, the degradation process includes skin problem removal processing; the sample face image acquisition module 702 is further configured to: perform skin problem removal processing on the original face image according to the maximum removal intensity to obtain a face image after problem removal; perform image fusion on the face image after problem removal and the original face image according to the target removal intensity to obtain a partially face image after problem removal; generate a sample face image based on the face image after removal with the target removal intensity.

[0150] In one embodiment, the degradation process further includes noise addition processing; the sample face image acquisition module 702 is further configured to: perform noise addition processing on the face image after removal with the target removal intensity according to a preset noise addition algorithm to obtain a sample face image.

[0151] In one embodiment, the training loss value includes a first training loss value and a second training loss value; the initial image enhancement model includes a denoising sub-model and an enhancement sub-model; the face image enhancement module 704 is further configured to: perform denoising processing on the sample face image through the denoising sub-model to obtain a denoised sample face image; perform skin problem enhancement processing on the denoised sample face image through the enhancement sub-model to obtain an enhanced face image; the loss value determination module 706 is further configured to: determine the first training loss value based on the enhanced face image and the original face image; determine the second training loss value based on the denoised sample face image and the face image after removal with the target removal intensity.

[0152] In one embodiment, the loss value determination module 706 is further configured to: determine a skin problem mask based on the original face image and the partially face image after problem removal; determine a basic loss value based on the enhanced face image and the original face image; adjust the basic loss value based on the skin problem mask to obtain the first training loss value.

[0153] In one embodiment, the loss value determination module 706 is further configured to: determine an image difference value based on the denoised sample face image and the face image after removal with the target removal intensity; and determine a second training loss value according to the image difference value.

[0154] In one embodiment, the face image enhancement module 704 is further configured to: obtain a face image to be processed; and perform enhancement processing on the face image to be processed through an image enhancement model to obtain a processed face image.

[0155] Each module in the processing device of the above image enhancement model can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.

[0156] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 8 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store image data. The input / output interface of the computer device is used for the processor to exchange information with external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a processing method of an image enhancement model.

[0157] Those skilled in the art can understand that Figure 8 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0158] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0159] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.

[0160] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.

[0161] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0162] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0163] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0164] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A processing method for an image enhancement model, characterized in that: The method comprises: Performing degradation processing on the original face image to obtain a sample face image; Performing enhancement processing on the sample face image by using an initial image enhancement model to obtain an enhanced face image; Determining a training loss value based on the enhanced face image and the original face image; The initial image enhancement model is optimized based on the training loss value until a convergence condition is reached to obtain a trained image enhancement model; the image enhancement model is used to enhance facial images.

2. The method according to claim 1, characterized in that The degeneration treatment includes skin problem removal treatment; The step of performing degradation processing on the original face image to obtain a sample face image includes: According to the maximum removal strength, the original face image is subjected to skin problem removal processing to obtain a face image after problem removal; According to the target removal strength, the face image after the problem is removed and the original face image are fused to obtain the face image after some of the problems are removed; A sample face image is generated based on the face image after removal of the target removal intensity.

3. The method according to claim 2, characterized in that The degradation processing also includes noise adding processing; the generation of a sample face image based on the face image after removal of the target removal strength includes: According to a preset denoising algorithm, the face image after the target removal intensity is denoised to obtain a sample face image.

4. The method according to claim 3, characterized in that The training loss value includes a first training loss value and a second training loss value; the initial image enhancement model includes a denoising sub-model and an enhancement sub-model; The step of performing enhancement processing on the sample face image by using the initial image enhancement model to obtain an enhanced face image includes: Performing denoising processing on the sample face image by using the denoising sub-model to obtain a denoised sample face image; Performing skin problem enhancement processing on the denoised sample face image by using the enhancer model to obtain an enhanced face image; The determining of the training loss value based on the enhanced face image and the original face image comprises: Determine a first training loss value based on the enhanced face image and the original face image; A second training loss value is determined based on the denoised sample face image and the face image removed with the target removal strength.

5. The method according to claim 4, characterized in that The method further includes determining a first training loss value based on the enhanced face image and the original face image, including: Determine a skin problem mask based on the original face image and the face image after the partial problem is removed; Determine a basic loss value based on the enhanced face image and the original face image; The basic loss value is adjusted based on the skin problem mask to obtain a first training loss value.

6. The method according to claim 4, characterized in that The determining of a second training loss value based on the denoised sample face image and the face image removed with the target removal strength comprises: Determining an image difference value based on the denoised sample face image and the face image after removal of the target removal intensity; A second training loss value is determined according to the image difference value.

7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: Obtain the face image to be processed; The face image to be processed is enhanced by the image enhancement model to obtain a processed face image.

8. A processing device for an image enhancement model, characterized in that: The device comprises: A sample face image acquisition module is used to perform degradation processing on the original face image to obtain a sample face image; A face image enhancement module, used for performing enhancement processing on the sample face image through an initial image enhancement model to obtain an enhanced face image; A loss value determination module, used to determine a training loss value based on the enhanced face image and the original face image; The model optimization module is used to optimize the initial image enhancement model based on the training loss value until a convergence condition is reached to obtain a trained image enhancement model.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • A method, electronic device, system, and storage medium for enhancing human portrait images.

    CN122573739A