Image Restoration Method, Device, Electronic Device and Storage Medium
By using the diffusion module and noise-added noise reduction mode in the image restoration technology, the masked image is processed multiple times and the problem of image restoration distortion in the prior art is solved, and high-quality image restoration is achieved, especially when restoring the obscured face image, the effect is significant.
Patent Information
- Application Number
- CN202210908887.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-07-29
AI Technical Summary
The existing image restoration technology has distortion problems when restoring images, especially when dealing with occluded face images, the restored image quality is not high.
The image restoration method of diffusion module combined with the noise reduction mode is adopted. By performing multiple additions of noise and denoising on the masked image, the image restoration model is trained until the loss function value is optimal, thereby achieving high-quality image restoration.
This method can quickly and accurately restore the obstructed face image, and the restored image has high authenticity, avoiding the confrontation training process and improving training stability.
Smart Images

Figure CN115239593B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to an image restoration method, apparatus, electronic device, and storage medium. Background Art
[0002] Image restoration technology is a type of image processing technology that can restore the pixel features of damaged parts in an image, reconstruct and generate damaged areas that approximate the semantics of a high-quality original image, thereby improving the image quality. Currently, image restoration is usually performed using an image restoration model based on a generative adversarial network or an autoencoder, training a specific mask distribution using an adversarial training method, and patching the semantically damaged areas of the image. However, this restoration method only performs a simple texture extension of the missing area using the surrounding pixel area, and there is a distortion problem in the restored image. Summary of the Invention
[0003] An object of the present application is to solve at least to some extent one of the technical problems existing in the related art.
[0004] To this end, an object of an embodiment of the present application is to provide an image restoration method, apparatus, electronic device, and storage medium that can quickly and accurately restore an occluded face image.
[0005] To achieve the above object, a first aspect of an embodiment of the present application provides an image restoration method, including:
[0006] Obtain a first face image to be trained, and input the first face image into an image restoration model, where the image restoration model includes a trained diffusion module;
[0007] Perform mask processing on the first face image to obtain a mask image;
[0008] Train the image restoration model according to the mask image until the first loss function value of the image restoration model is optimal, to obtain a trained image restoration model. During the training process, the diffusion module performs multiple noise addition processes on the mask image to obtain a white noise image that follows a standard normal distribution, and forms a noise function according to the noise added in the noise addition process. The image restoration model restores the mask image through a noise addition and noise reduction mode according to the white noise image, the standard normal distribution, and the noise function;
[0009] Obtain a second face image to be restored, and perform image restoration on the second face image through the trained image restoration model to obtain a target restored image.
[0010] In some embodiments, the image restoration method further includes: training the diffusion module until the value of the second loss function of the diffusion module is minimized, to obtain a trained diffusion module;
[0011] Wherein, the training of the diffusion module includes:
[0012] Obtaining training images;
[0013] Based on the forward Markov process, performing multiple noise addition processes on the training images to obtain a second white noise image that follows a standard normal distribution;
[0014] Based on the transition distribution, performing multiple noise removal processes on the second white noise image to restore the white noise image to the training images, where the number of executions of the noise removal process is equal to the number of executions of the noise addition process;
[0015] Calculating the value of the second loss function of the diffusion module, and adjusting the parameters of the diffusion module according to the value of the second loss function.
[0016] In some embodiments, one noise-added image is obtained each time the noise addition process is performed; the value of the second loss function includes the gap loss between the noise-added image and the second white noise image.
[0017] In some embodiments, the restoring the masked image through the noise addition and removal mode according to the white noise image, the standard normal distribution, and the noise function includes:
[0018] Using the white noise image as the initial restored image, and performing multiple iterative restoration processes on the masked image through the noise addition and removal mode using the standard normal distribution and the noise function, and using the restored image obtained from the last restoration process as the predicted restored image, where the number of restoration processes is equal to the total number of noise addition processes performed by the diffusion module.
[0019] In some embodiments, the masked image includes a masked part and an unmasked part; the restoration process includes:
[0020] Based on the noise addition process of the diffusion module, performing a noise addition process on the masked image to obtain a first processed image, where the noise addition process is related to the standard normal distribution;
[0021] Based on the noise removal process of the diffusion module, performing a noise reduction process on the previous restored image to obtain a second processed image, where the noise reduction process is related to the noise function;
[0022] Calculate a first product value of the pixel matrix of the first processed image and the masked part, calculate a second product value of the pixel matrix of the second processed image and the unmasked part, and use the sum of the first product value and the second product value as the current restored image.
[0023] In some embodiments, before the step of performing noise reduction processing on the previous restored image, it further includes:
[0024] Perform multiple resamplings on the previous restored image to update the previous restored image.
[0025] In some embodiments, the first loss function value of the image restoration model is calculated according to the following method:
[0026] Calculate the mean square error according to the first mean of the first face image, the first variance of the first face image, the second mean of the predicted restored image, and the second variance of the predicted restored image;
[0027] Calculate the structural error between the first face image and the predicted restored image;
[0028] Use the sum of the mean square error and the structural error as the first loss function value.
[0029] To achieve the above object, a second aspect of the embodiments of the present application provides an image restoration device, including:
[0030] An input module, configured to obtain a first face image to be trained, and input the first face image into an image restoration model, where the image restoration model includes a trained diffusion module;
[0031] A masking module, configured to perform masking processing on the first face image to obtain a masked image;
[0032] A training module, configured to train the image restoration model according to the masked image until the first loss function value of the image restoration model is optimal, to obtain a trained image restoration model. During the training process, the diffusion module performs multiple noise addition processes on the masked image to obtain a white noise image that follows a standard normal distribution, and forms a noise function according to the noise added in the noise addition process. The image restoration model performs restoration processing on the masked image through a noise addition and noise reduction mode according to the white noise image, the standard normal distribution, and the noise function;
[0033] A restoration module, configured to obtain a second face image to be restored, and perform image restoration on the second face image through the trained image restoration model to obtain a target restored image.
[0034] To achieve the above object, a third aspect of the embodiments of the present application further provides an electronic device, which includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory. When the program is executed by the processor, the above-mentioned image restoration method is realized.
[0035] To achieve the above object, a fourth aspect of the embodiments of the present application further provides a computer-readable storage medium, which stores computer-executable instructions for causing a computer to execute the above-mentioned image restoration method.
[0036] The image restoration method, device, electronic device and storage medium disclosed in the embodiments of the present application obtain a first face image to be trained, input the first face image into an image restoration model, and the image restoration model includes a trained diffusion module; perform masking processing on the first face image to obtain a masked image; train the image restoration model according to the masked image until the first loss function value of the image restoration model is optimal, and obtain a trained image restoration model. During the training process, the diffusion module performs multiple noise addition processes on the masked image to obtain a white noise image that follows a standard normal distribution, and forms a noise function according to the noise added in the noise addition process. The image restoration model restores the masked image through a noise addition and noise reduction mode based on the white noise image, the standard normal distribution, and the noise function; obtain a second face image to be restored, and perform image restoration on the second face image through the trained image restoration model to obtain a target restored image. It uses a noise addition and noise reduction mode to train image data, deeply excavates global information, and can then restore the pixel missing area to more realistic semantic content; avoids the adversarial training process and improves training stability; can quickly restore the occluded face image, and the restored image has high authenticity. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the related technical solution drawings in the embodiments of the present application or the prior art. It should be understood that the drawings in the following introduction are only for conveniently and clearly expressing some embodiments of the technical solutions in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 It is a schematic diagram of the principle of the diffusion module provided by the embodiments of the present application;
[0039] Figure 2 It is a schematic diagram of the principle of the restoration process;
[0040] Figure 3 It is a step diagram of the image restoration method provided by the embodiments of the present application;
[0041] Figure 4 It is a step diagram of training the diffusion module;
[0042] Figure 5 It is a step diagram of the restoration process;
[0043] Figure 6 It is a step diagram of resampling;
[0044] Figure 7 It is a step diagram of calculating the value of the first loss function;
[0045] Figure 8 It is a structural diagram of the image restoration device provided by the embodiments of the present application;
[0046] Figure 9 It is a structural diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners
[0047] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0048] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0050] First, several nouns involved in the present application are analyzed:
[0051] Artificial Intelligence (AI): It is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0052] Image processing technology: It is the technology of processing image information by computer. Image processing technology generally includes image compression, enhancement and restoration, matching, description, and recognition.
[0053] Image restoration technology: It is the technology of using the prior knowledge of the degradation process to restore the original appearance of the degraded image. The reasons for degradation may be optical system aberrations or defocus, relative motion between the imaging system and the object being imaged, noise in the electronic or optical system, and atmospheric turbulence between the imaging system and the object being imaged. First, an appropriate estimation of the entire image degradation process is required. Based on this, an approximate degradation mathematical model is established. After that, the model also needs to be appropriately corrected to compensate for the distortion that occurs during the degradation process, so as to ensure that the image obtained after restoration approaches the original image and realizes the optimization of the image.
[0054] Machine Learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning (deep learning) usually includes technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0055] Neural Networks: It is an algorithmic mathematical model that mimics the behavioral characteristics of animal neural networks and performs distributed parallel information processing for function estimation or approximation. A neural network consists of a large number of artificial neurons connected for computing. This network relies on the complexity of the system and adjusts the relationships between a large number of internal nodes to achieve the purpose of processing information. An artificial neural network is a technical reproduction of a biological neural network in a simplified sense. Its main task is to build a practical artificial neural network model according to the principles of biological neural networks and the needs of practical applications, design corresponding learning algorithms, simulate certain intelligent activities of the human brain, and then implement them technically to solve practical problems.
[0056] To solve the problems in the related art, an object of an embodiment of the present application is to provide an image restoration method, apparatus, electronic device, and storage medium. By obtaining a first face image to be trained and inputting the first face image into an image restoration model, the image restoration model includes a trained diffusion module; performing a masking process on the first face image to obtain a masked image; training the image restoration model according to the masked image until the first loss function value of the image restoration model is optimal to obtain a trained image restoration model. During the training process, the diffusion module performs multiple noise addition processes on the masked image to obtain a white noise image that follows a standard normal distribution, and forms a noise function according to the noise added in the noise addition process. The image restoration model restores the masked image through a noise addition and noise reduction mode according to the white noise image, the standard normal distribution, and the noise function; obtaining a second face image to be restored, and performing image restoration on the second face image through the trained image restoration model to obtain a target restored image; using the noise addition and noise reduction mode to train the image data, deeply mining global information, and then being able to restore the pixel missing area to more real semantic content; avoiding the adversarial training process and improving the training stability; being able to quickly restore the occluded face image, and the restored image has high authenticity.
[0057] The implementation environment of an image restoration method provided by an embodiment of the present application is as follows. The main software and hardware entities of this implementation environment include an operation terminal and a server, and the operation terminal is communicatively connected to the server. Among them, the training method of this retrieval-based dialogue model can be configured to be executed alone on the operation terminal, or can be configured to be executed alone on the server, or can be executed based on the interaction between the operation terminal and the server. Specifically, it can be appropriately selected according to the actual application situation, and this embodiment does not make specific limitations in this regard. In addition, the operation terminal and the server can be nodes in a blockchain, and this embodiment does not make specific limitations in this regard.
[0058] Specifically, the operation terminal in the present application may include, but is not limited to, any one or more of a smart watch, a smart phone, a computer, a personal digital assistant (PDA), a smart voice interaction device, a smart home appliance, or a vehicle-mounted terminal. The server may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. A communication connection may be established between the operation terminal 101 and the server 102 through a wireless network or a wired network. The wireless network or the wired network uses standard communication technologies and / or protocols. The network may be set as the Internet, or any other network, such as including, but not limited to, any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network, or a virtual private network.
[0059] In addition, the present application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0060] Referring to Figure 3 , Figure 3 is a step diagram of an image restoration method. An embodiment of the present application provides an image restoration method, which includes, but is not limited to, the following steps:
[0061] Step S100, obtain a first face image to be trained, and input the first face image into an image restoration model, where the image restoration model includes a trained diffusion module;
[0062] Step S200: Perform masking processing on the first face image to obtain a masked image;
[0063] Step S300: Train the image restoration model according to the masked image until the first loss function value of the image restoration model is optimal, obtaining a trained image restoration model. During the training process, the diffusion module performs multiple noise addition processes on the masked image to obtain a white noise image that follows the standard normal distribution, and forms a noise function based on the noise added in the noise addition process. The image restoration model restores the masked image through a denoising mode based on the white noise image, the standard normal distribution, and the noise function;
[0064] Step S400: Obtain a second face image to be restored, and perform image restoration on the second face image through the trained image restoration model to obtain a target restored image.
[0065] Among them, steps S100 to S300 are the training process of the image restoration model, and step S400 is the application process of the image restoration model.
[0066] Among them, for the image restoration model, the model includes a trained diffusion module, and the model is established based on the diffusion module.
[0067] Refer to Figure 1 , Figure 1 is a schematic diagram of the principle of the diffusion module. The diffusion module is a type of generative network. The diffusion module defines a Markov chain of diffusion steps, gradually adds random noise to the data, and then learns the inverse diffusion process to construct the required data samples from the noise.
[0068] The diffusion module is mainly divided into two processes: the forward process and the backward process. In the forward process, noise is gradually added to the real image of the data set to form a white noise image that follows the standard normal distribution; in the backward process, the white noise image is gradually denoised to restore the original real image.
[0069] For the training of the diffusion module, it mainly includes the following steps: Step S10: Train the diffusion module until the second loss function value of the diffusion module is the smallest, obtaining a trained diffusion module;
[0070] Refer to Figure 4 , Figure 4 is the step diagram for training the diffusion module. Among them, training the diffusion module includes:
[0071] Step S11: Obtain training images;
[0072] Step S12: Based on the forward Markov process, perform multiple noise addition processes on the training images to obtain a second white noise image that follows the standard normal distribution;
[0073] Step S13: Based on the transfer distribution, perform multiple noise removal processes on the second white noise image to restore the white noise image to the training image. The number of executions of the noise removal process is equal to the number of executions of the noise addition process.
[0074] Step S14: Calculate the second loss function value of the diffusion module and adjust the parameters of the diffusion module according to the second loss function value.
[0075] Among them, the second loss function value includes the gap loss between the noisy image and the second white noise image. Based on the goal of minimizing the second loss function value of the diffusion module, use the stochastic gradient descent method to solve the parameters of the diffusion module and then adjust the parameters of the diffusion module. At the same time, use the early stopping method to avoid overfitting of the diffusion module. When the training loss continuously decreases and tends to be stable, and the test loss decreases but does not rebound, stop the training.
[0076] For the initial data distribution x 0 ~q(x), in the forward diffusion process, gradually add Gaussian noise to the data distribution, perform the noise addition process T times, and obtain a white noise image that follows the standard normal distribution; specifically, the standard normal distribution is the Gaussian normal distribution. Each execution of the noise addition process obtains a noisy image, generating a series of noisy images, and the noisy images are represented as x 1 , x 2 ,..., x t ,..., x T . In the process of adding noise from x t to x t-1 , the standard deviation or variance of the noise is determined by a fixed value β t within the interval (0, 1). The mean value is determined by the fixed value β t and the image data x t at the current moment.
[0077] The process q(x t-1 ) of obtaining x t from x t |x t-1 ), satisfies the Gaussian normal distribution The Gaussian normal distribution is with the mean value and β t as the variance. Then the noise addition process can be expressed as Therefore, the noise can be determined by the fixed value β t and the image data x t at the current moment, and it is a fixed value rather than a learnable process. As long as there is x 0 , and the fixed value β 1 for each step is determined in advance, β2 ,...,β t ,...,β T , we can deduce the noise-added image x at any step 1 ,x 2 ,...,x t ,...,x T In general, as t increases, the data is closer to a random Gaussian distribution, and β t The larger the value of is, the more 1 ,β 2 ,...,β t ,...,β T , there is β 1 <β 2 <...<β t <...<β T ; Further, the fixed value sequence is a geometric progression with values between 0 and 1.
[0078] As t increases, the initial data x 0 It will gradually lose its characteristics. Finally, when T→∞, it approaches an independent Gaussian distribution and forms a white noise image.
[0079] In the process of gradually adding noise, it is not necessary to start from x step by step. 0 ,x 1 ,...,x t Go iterate to get x t , can be directly from x 0 and fixed value sequences Directly calculate x t . Define α t =1-β t , Then there is Through parameter renormalization, it becomes a form consisting only of random variables z, z~N(0,I).
[0080] For the backward process, it is actually also a Markov process, which changes the direction of the forward process, that is, from q(x t-1 |x t ), that is, we can reconstruct a real original sample from a random Gaussian distribution N(0,I), that is, we can get a real initial image from a bunch of completely messy noisy pictures. However, since we need to find the data distribution from the complete data set, it is difficult to simply predict q(x t-1 |x t ), so we need to learn a model p θ To approximate the conditional probability, we can run the backward process. For the conditional probability model of the backward process, we have pθ (x t-1 |x t )~N(x t-1 ; μ θ (x t , t), ∑ θ (x t , t))。
[0081] For the second loss function value of the diffusion module, the second loss function value includes the gap loss between the noisy image and the second white noise image. Specifically, the second loss function value can be expressed as: In the formula, Loss 2 represents the second loss function value, ε is a noise vector sampled from the standard normal distribution, ε~N(0, I). When the second loss function value becomes smaller and smaller, it means that the noisy image generated in the forward process is approaching or even overlapping with the white noise image that follows the Gaussian distribution.
[0082] For the training steps of the image restoration model, obtain multiple first face images, and the first face images have no occluded parts. Perform masking processing on the first face images to obtain masked images. Since people usually wear masks, sunglasses or hats for occlusion, the mouth and nose area, eye area or top of the head area of the first face images can be masked. It can be understood that on the one hand, the mouth and nose area, eye area or top of the head area of the first face images can be masked manually; on the other hand, the mouth and nose area, eye area or top of the head area of the first face images can be recognized through image recognition technology, and then the mouth and nose area, eye area or top of the head area of the first face images can be automatically masked by the computer.
[0083] The occluded part masked in the masked image is the masked part, and the unoccluded part not masked in the masked image is the unmasked part.
[0084] Input the first face image and the masked image into the image restoration model.
[0085] The image restoration model restores the masked image through the noise addition and reduction mode according to the white noise image, the standard normal distribution, and the noise function, including:
[0086] Taking the white noise image as the initial restored image, the masked image is restored through the noise addition and reduction mode by using the standard normal distribution and the noise function for multiple iterations. The restored image obtained from the last restoration process is used as the predicted restored image. The number of restoration processes is equal to the total number of noise addition processes performed by the diffusion module. That is, if the total number of noise addition processes performed by the diffusion module is T times, then the number of restoration processes is also T times. Among them, for the multiple-iteration restoration process, the output of the previous restoration process is used as the input of the next restoration process.
[0087] That is, based on the white noise image x T and the masked image it is gradually restored to x t , until it is restored to the unobscured face image x 0 .
[0088] Refer to Figure 5 , Figure 5 which is the step diagram of the restoration process. The restoration process includes:
[0089] Step S311: Based on the noise addition process of the diffusion module, the masked image is subjected to noise addition to obtain the first processed image. The noise addition process is related to the standard normal distribution;
[0090] Step S312: Based on the noise removal process of the diffusion module, the previous restored image is subjected to noise reduction to obtain the second processed image. The noise reduction process is related to the noise function;
[0091] Step S313: Calculate the first product value of the pixel matrix of the first processed image and the masked part, calculate the second product value of the pixel matrix of the second processed image and the unmasked part, and take the sum of the first product value and the second product value as the current restored image.
[0092] Refer to Figure 2 , Figure 2 which is the schematic diagram of the principle of the restoration process. Among them, based on the principle of the noise addition process of the diffusion module, there is The first processed image can be expressed as: In the formula, represents the first processed image, ε is the noise vector sampled from the standard normal distribution, ε~N(0,I).
[0093] Based on the noise removal process of the diffusion module, there is The second processed image can be expressed as: In the formula, represents the second processed image, ε θ (x t ,t) represents the sum of the accumulated noise in the noise addition function up to x t , x tDenote the previous restored image as σ t Denote the variance of the previous restored image, z is a random variable and z ~ N(0, I).
[0094] For the current restored image, it can be expressed as: In the formula, x t-1 is the current restored image, m represents the pixel matrix of the masked part, and (1 - m) represents the pixel matrix of the unmasked part.
[0095] In addition, when applying the above restoration process to restore the masked image, only the content type of the restored part matches that of the unmasked part, but there are semantic errors; although the diffusion module utilizes the context of the unmasked part, it does not well coordinate other parts of the image. The diffusion module is trained to generate an image located in the data distribution and naturally wants to generate a consistent structure. By using this feature to coordinate the input of the restoration process through resampling, the semantics of the restored part can be made to match that of the unmasked part.
[0096] Refer to Figure 6 , before the step of reducing noise for the previous restored image, it further includes:
[0097] Step S3111, perform multiple resamplings on the previous restored image to update the previous restored image. For the resampled previous restored image x t , there is
[0098] Refer to Figure 7 , where the first loss function value of the image restoration model is calculated according to the following method:
[0099] Step S321, calculate the mean square error according to the first mean of the first face image, the first variance of the first face image, the second mean of the predicted restored image, and the second variance of the predicted restored image;
[0100] Step S322, calculate the structural error between the first face image and the predicted restored image;
[0101] Step S323, take the sum of the mean square error and the structural error as the first loss function value.
[0102] For the mean square error, it can be expressed by the following formula: In the formula, Loss structure represents the mean square error, μ X represents the first mean of the first face image, μ X′ represents the second mean of the predicted restored image, σ X represents the first variance of the first face image, σX′ represents the second variance of the predicted restored image, C 1 and C 2 are two parameters for maintaining stability, C 1 and C 2 are constants.
[0103] For the structural error, it can be expressed by the following formula: Loss MSE =‖X - X′‖ 2 , where Loss MSE represents the structural error.
[0104] For the first loss function value, it can be expressed by the following formula: Loss 1 = Loss structure + Loss MSE , where Loss 1 represents the first loss function value.
[0105] Based on the goal of minimizing the first loss function value, the parameters of the image restoration model are solved by using the stochastic gradient descent method to adjust the parameters of the image restoration model. At the same time, the early stopping method is used to avoid the phenomenon of overfitting in the diffusion module. When the training loss continuously decreases and tends to be stable, and the test loss decreases but does not rebound, the training is stopped.
[0106] For the application steps of the image restoration model, multiple second face images are obtained, and the second face images have occluded parts. The second face images are input into the trained image restoration model, and the image restoration model restores the occluded parts according to the unoccluded parts.
[0107] The image restoration model performs mask processing on the occluded parts to obtain a mask image; the mask image includes a masked part and an unmasked part.
[0108] The image restoration model restores the mask image through a noise addition and reduction mode according to the white noise image, the standard normal distribution, and the noise function, including:
[0109] Taking the white noise image as the initial restored image, using the standard normal distribution and the noise function to perform multiple iterative restoration processing on the mask image through the noise addition and reduction mode, and taking the restored image obtained from the last restoration processing as the predicted restored image. The number of restoration processing times is equal to the total number of noise addition processing times of the diffusion module, that is, if the total number of noise addition processing times of the diffusion module is T times, then the number of restoration processing times is also T times. Among them, for the multiple iterative restoration processing, the output of the previous restoration processing is used as the input of the next restoration processing.
[0110] The restoration process includes: adding noise processing based on the diffusion module to add noise to the masked image to obtain a first processed image, where the adding noise processing is related to the standard normal distribution; removing noise processing based on the diffusion module to reduce noise from the previous restored image to obtain a second processed image, where the reducing noise processing is related to the noise function; calculating a first product value of the first processed image and the pixel matrix of the masked part, calculating a second product value of the second processed image and the pixel matrix of the unmasked part, and taking the sum of the first product value and the second product value as the current restored image.
[0111] Among them, for the principle of adding noise processing based on the diffusion module, there is The first processed image can be expressed as: In the formula, represents the first processed image, ε is a noise vector sampled from the standard normal distribution, ε~N(0,I).
[0112] For the removing noise processing based on the diffusion module, there is The second processed image can be expressed as: In the formula, represents the second processed image, ε θ (x t ,t) represents the sum of the accumulated noise up to x t in the noise adding function, x t represents the previous restored image, σ t represents the variance of the previous restored image, z is a random variable and z~N(0,I).
[0113] For the current restored image, it can be expressed as: In the formula, x t-1 is the current restored image, m represents the pixel matrix of the masked part, and (1 - m) represents the pixel matrix of the unmasked part.
[0114] In addition, when applying the above restoration process to restore the masked image, the restored part only matches the content type of the unmasked part, but there are semantic errors; although the diffusion module utilizes the context of the unmasked part, it does not well coordinate the other parts of the image. The diffusion module is trained to generate an image within the data distribution and naturally wants to generate a consistent structure. By using this feature to coordinate the input of the restoration process through resampling, the semantics of the restored part can be made to match the unmasked part. Before the step of reducing noise from the previous restored image, it further includes: resampling the previous restored image multiple times to update the previous restored image. For the resampled previous restored image x t , there is
[0115] That is, based on the white noise image x T and the mask image obtained from the second face image gradually restore it to x t until it is restored to the unobscured face image x corresponding to the second face image 0 .
[0116] To achieve the above object, an embodiment of the present application further provides an image restoration device.
[0117] Referring to Figure 8 , Figure 8 is the structural diagram of the image restoration device. The image restoration device includes an input module 110, a mask module 120, a training module 130, and a restoration module 140.
[0118] Among them, the input module 110 is used to obtain the first face image to be trained and input the first face image into the image restoration model. The image restoration model includes a trained diffusion module; the mask module 120 is used to perform mask processing on the first face image to obtain a mask image; the training module 130 is used to train the image restoration model according to the mask image until the first loss function value of the image restoration model is optimal, and obtain the trained image restoration model. During the training process, the diffusion module performs multiple noise addition processes on the mask image to obtain a white noise image that follows the standard normal distribution, and forms a noise function according to the noise added in the noise addition process. The image restoration model restores the mask image through the noise addition and noise reduction mode according to the white noise image, the standard normal distribution, and the noise function; the restoration module 140 is used to obtain the second face image to be restored and perform image restoration on the second face image through the trained image restoration model to obtain the target restored image. It realizes the ability to quickly restore the occluded face image, and the restored image has high authenticity.
[0119] The image restoration device can train the image restoration model and apply the trained image restoration model. The training of the image restoration model is realized through the input module 110, the mask module 120, and the training module 130, and the application of the image restoration model is realized through the restoration module 140.
[0120] For the training steps of the image restoration model, multiple first face images are obtained through the input module 110, and the first face images do not have occluded parts. The mask module 120 performs mask processing on the first face image to obtain a mask image. The occluded part in the mask image that has been masked is the masked part, and the unmasked part in the mask image that has not been masked is the unmasked part.
[0121] The image restoration model is trained by the training module 130. Using a white noise image as the initial restored image, the masked image is iteratively restored multiple times through the noise addition and reduction mode by using the standard normal distribution and the noise function, and the restored image obtained from the last restoration process is used as the predicted restored image. The number of restoration processes is equal to the total number of noise addition processes performed by the diffusion module. That is, if the total number of noise addition processes performed by the diffusion module is T times, then the number of restoration processes is also T times. Among them, for the multiple iterative restoration processes, the output of the previous restoration process is used as the input of the next restoration process. That is, based on the white noise image x T and the masked image is gradually restored to x t , until it is restored to the unobscured face image x 0 .
[0122] Among them, for the restoration process, based on the principle of the noise addition process of the diffusion module, there is The first processed image can be expressed as: In the formula, represents the first processed image, ε is a noise vector sampled from the standard normal distribution, ε~N(0, I).
[0123] Based on the noise removal process of the diffusion module, there is The second processed image can be expressed as: In the formula, represents the second processed image, ε θ (x t , t) represents the sum of the accumulated noises in the noise addition function up to x t , x t represents the previous restored image, σ t represents the variance of the previous restored image, and z is a random variable and z~N(0, I).
[0124] For the current restored image, it can be expressed as: In the formula, x t-1 is the current restored image, m represents the pixel matrix of the masked part, and (1 - m) represents the pixel matrix of the unmasked part.
[0125] Before the step of reducing the noise of the previous restored image, the previous restored image is resampled multiple times to update the previous restored image. For the resampled previous restored image x t , there is
[0126] According to the first loss function value Loss 1 = Loss structure + Loss MSEContinuously adjust the parameters of the image restoration model.
[0127] Among them, for the mean square error, it can be expressed by the following formula: In the formula, Loss structure represents the mean square error, μ X represents the first mean of the first face image, μ X′ represents the second mean of the predicted restored image, σ X represents the first variance of the first face image, σ X′ represents the second variance of the predicted restored image, C 1 and C 2 are two parameters for maintaining stability, C 1 and C 2 are constants.
[0128] For the structural error, it can be expressed by the following formula: Loss MSE =‖X - X′‖ 2 , in the formula, Loss MSE represents the structural error.
[0129] Until the value of the first loss function is minimized, a trained image restoration model is obtained.
[0130] The restoration module 140 uses the trained image restoration model to perform image restoration on the second face image to be restored, and obtains the target restored image.
[0131] It can be understood that the content in the image restoration method embodiment is applicable to the image restoration device embodiment of the present application. The functions specifically implemented by the image restoration device embodiment of the present application are the same as those of the image restoration method embodiment, and the beneficial effects achieved are also the same as those of the image restoration method embodiment.
[0132] To achieve the above object, an embodiment of the present application also provides an electronic device. Referring to Figure 9 , Figure 9 is the structural diagram of the electronic device. The electronic device includes a memory 220, a processor 210, a program stored on the memory 220 and executable on the processor 210, and a data bus 230 for realizing the connection and communication between the processor 210 and the memory 220. When the program is executed by the processor 210, the above image restoration method is implemented.
[0133] In this embodiment, by obtaining a first face image to be trained, the first face image is input into an image restoration model, and the image restoration model includes a trained diffusion module; the first face image is masked to obtain a masked image; the image restoration model is trained according to the masked image until the first loss function value of the image restoration model is optimal, and a trained image restoration model is obtained. During the training process, the diffusion module performs multiple noise addition processes on the masked image to obtain a white noise image that follows a standard normal distribution, and forms a noise function according to the noise added during the noise addition process. The image restoration model restores the masked image through a noise addition and denoising mode based on the white noise image, the standard normal distribution, and the noise function; a second face image to be restored is obtained, and the trained image restoration model is used to perform image restoration on the second face image to obtain a target restored image; the occluded face image can be restored quickly, and the restored image has high authenticity.
[0134] As a non-transitory computer-readable storage medium, the memory 220 can be used to store non-transitory software programs and non-transitory computer-executable programs, such as the image restoration method in the above embodiments of the present invention. The processor 210 realizes the image restoration method in the above embodiments of the present invention by running the non-transitory software programs and programs stored in the memory 220.
[0135] The memory 220 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data required for executing the image restoration method in the above embodiments of the present invention, etc. In addition, the memory 220 may include a high-speed random access memory 220, and may also include a non-transitory memory 220, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 220 may optionally include a memory 220 that is remotely disposed relative to the processor 210, and these remote memories 220 can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0136] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium, and the computer-readable storage medium stores computer-executable instructions for causing a computer to execute the above image restoration method.
[0137] In this embodiment, by obtaining a first face image to be trained and inputting the first face image into an image restoration model, the image restoration model includes a trained diffusion module; performing a masking process on the first face image to obtain a masked image; training the image restoration model according to the masked image until the first loss function value of the image restoration model is optimal, thereby obtaining a trained image restoration model. During the training process, the diffusion module performs multiple noise addition processes on the masked image to obtain a white noise image that follows a standard normal distribution, and forms a noise function based on the noise added during the noise addition process. The image restoration model restores the masked image through a noise addition and noise reduction mode according to the white noise image, the standard normal distribution, and the noise function; obtaining a second face image to be restored, and performing image restoration on the second face image through the trained image restoration model to obtain a target restored image; it can quickly restore an occluded face image, and the restored image has high authenticity.
[0138] Those of ordinary skill in the art will understand that all or some of the steps and systems disclosed above in the methods can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVDs) or other optical disk storage, magnetic cassettes, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium. In the foregoing description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0139] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application. The scope of the present application is defined by the claims and their equivalents.
[0140] The above has specifically described the preferred embodiments of the present application, but the present application is not limited to the embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.
[0141] In the description of this specification, the descriptions referring to terms such as "one embodiment", "another embodiment", or "certain embodiments" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
Claims
1. An image restoration method, characterized in that, comprising: Obtain a first face image to be trained, and input the first face image into an image restoration model, where the image restoration model includes a trained diffusion module; Perform masking processing on the first face image to obtain a masked image; Train the image restoration model according to the masked image until the first loss function value of the image restoration model is optimal, to obtain a trained image restoration model. During the training process, the diffusion module performs multiple noise addition processes on the masked image to obtain a white noise image that follows a standard normal distribution, and forms a noise function according to the noise added in the noise addition process. The image restoration model restores the masked image through a noise addition and noise reduction mode according to the white noise image, the standard normal distribution, and the noise function; Obtain a second face image to be restored, and perform image restoration on the second face image through the trained image restoration model to obtain a target restored image; The image restoration method further includes: training the diffusion module until the second loss function value of the diffusion module is minimized, to obtain a trained diffusion module; wherein, the training of the diffusion module includes: Obtain training images; Based on the forward Markov process, perform multiple noise addition processes on the training images to obtain a second white noise image that follows a standard normal distribution; Based on the transition distribution, perform multiple noise removal processes on the second white noise image to restore the white noise image to the training image. The number of executions of the noise removal process is equal to the number of executions of the noise addition process; Calculate the second loss function value of the diffusion module, and adjust the parameters of the diffusion module according to the second loss function value; The restoring the masked image through a noise addition and noise reduction mode according to the white noise image, the standard normal distribution, and the noise function includes: Using the white noise image as the initial restored image, perform multiple iterative restoration processes on the masked image through the noise addition and noise reduction mode using the standard normal distribution and the noise function, and use the restored image obtained from the last restoration process as the predicted restored image. The number of restoration processes is equal to the total number of noise addition processes performed by the diffusion module.
2. The image restoration method according to claim 1, characterized in that, One noise-added image is obtained each time the noise addition process is performed; the second loss function value includes the gap loss between the noise-added image and the second white noise image.
3. The image restoration method according to claim 1, characterized in that, The masked image includes a masked part and an unmasked part; The restoration process includes: Based on the noise addition process of the diffusion module, perform noise addition processing on the masked image to obtain a first processed image, and the noise addition process is related to the standard normal distribution; Based on the noise removal process of the diffusion module, perform noise reduction processing on the previous restored image to obtain a second processed image, and the noise reduction process is related to the noise function; Calculate a first product value of the pixel matrix of the first processed image and the masked part, calculate a second product value of the pixel matrix of the second processed image and the unmasked part, and use the sum of the first product value and the second product value as the current restored image.
4. An image restoration method according to claim 3, wherein, before the step of performing noise reduction processing on the previous restored image, it further includes: Performing multiple resamplings on the previous restored image to update the previous restored image.
5. An image restoration method according to claim 1, wherein, The first loss function value of the image restoration model is calculated as follows: Calculate the mean square error according to the first mean of the first face image, the first variance of the first face image, the second mean of the predicted restored image, and the second variance of the predicted restored image; Calculate the structural error between the first face image and the predicted restored image; Use the sum of the mean square error and the structural error as the first loss function value.
6. An image restoration device, wherein, it includes: An input module for obtaining a first face image to be trained and inputting the first face image into an image restoration model, where the image restoration model includes a trained diffusion module; A masking module for masking the first face image to obtain a masked image; A training module for training the image restoration model according to the masked image until the first loss function value of the image restoration model is optimal to obtain a trained image restoration model. During the training process, the diffusion module performs multiple noise addition processes on the masked image to obtain a white noise image that follows a standard normal distribution, and forms a noise function according to the noise added in the noise addition process. The image restoration model restores the masked image through a noise addition and noise reduction mode according to the white noise image, the standard normal distribution, and the noise function; A restoration module for obtaining a second face image to be restored and performing image restoration on the second face image through the trained image restoration model to obtain a target restored image; Train the diffusion module until the second loss function value of the diffusion module is minimized to obtain a trained diffusion module; wherein, the training of the diffusion module includes: Obtain training images; Based on the forward Markov process, perform multiple noise addition processes on the training images to obtain a second white noise image that follows a standard normal distribution; Based on the transition distribution, perform multiple noise removal processes on the second white noise image to restore the white noise image to the training image. The number of executions of the noise removal process is equal to the number of executions of the noise addition process; Calculate the second loss function value of the diffusion module and adjust the parameters of the diffusion module according to the second loss function value; The restoring the masked image through a noise addition and noise reduction mode according to the white noise image, the standard normal distribution, and the noise function includes: Using the white noise image as the initial restored image, the masked image is subjected to multiple iterative restoration processes through a noise addition and reduction mode using the standard normal distribution and the noise function, and the restored image obtained from the last restoration process is used as the predicted restored image, where the number of times of the restoration process is equal to the total number of times of the noise addition process performed by the diffusion module.
7. An electronic device, characterized in that, the electronic device includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory, and when the program is executed by the processor, it realizes the image restoration method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores computer-executable instructions for causing a computer to execute the image restoration method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Image augmentation processing method and device based on artificial intelligence, equipment and storage medium
CN112132106A