Image restoration method, electronic equipment and storage medium
By using an adversarial generation network and amplitude phase decoupling module in the image restoration method, the problem of poor image restoration effect in the prior art is solved, and a better image quality restoration effect is achieved.
Patent Information
- Application Number
- CN202311558404.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-05-30
AI Technical Summary
Image restoration methods in the prior art perform poorly, and traditional algorithms are difficult to effectively process complex degradation functions and noise, while deep learning-based methods ignore global information, resulting in poor practical application results.
Adversarial generation network (GAN) is used as the basis of the image restoration method, and by initializing the adversarial generation network, including a first generator and a first discriminator, multiple amplitude phase decoupling modules are used to Fourier transform and inverse transform the input features, extract global context information, and iteratively alternately train the generator and discriminator to obtain the image restoration model.
Through the training of the adversarial generation network, noise and blur can be effectively removed, image restoration effect can be improved, and structural information of the original image can be retained while removing deterioration factors, improving the quality of image restoration.
Smart Images

Figure CN120070291A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of image processing, and particularly relates to an image restoration method, an electronic device, and a storage medium. Background Art
[0002] Image degradation refers to the phenomenon that the image quality deteriorates and information is lost during generation and transmission due to the influence of acquisition devices, light, etc., and can generally be expressed as where g represents the degraded image, h represents the degradation function, f represents the original image, and n represents the noise. The sources of common degradation functions may include at least one of blurring, distortion, occlusion, etc. The purpose of image restoration is to improve the quality of the degraded image and restore the original image from the degraded image as much as possible.
[0003] Image restoration is mainly divided into two categories: traditional algorithms and deep learning. Traditional algorithms mainly invert the equation of the degraded image based on the mathematical models of various degradation functions / noise to obtain the original image. However, in actual scenarios, the degradation functions and noise are usually complex and do not match well with the mathematical models, often unable to produce ideal results.
[0004] In the image restoration method based on deep learning, the receptive field of the obtained model is limited, and it pays more attention to the local features of the image while ignoring the global information, so it performs poorly in actual applications. Summary of the Invention
[0005] The embodiments of this application provide an image restoration method, an electronic device, and a storage medium, which can solve the problem that the image restoration method in the related technology performs poorly.
[0006] In a first aspect, the embodiments of this application provide an image restoration method, which includes: initializing an adversarial generative network, where the adversarial generative network includes a first generator and a first discriminator. The first generator is used to convert the input image belonging to the degradation domain into a first image, and the first discriminator is used to identify whether the first image belongs to the clear domain; wherein the first generator includes a plurality of amplitude-phase decoupling modules, and each amplitude-phase decoupling module is used to perform Fourier transform on the input feature to obtain an amplitude component and a phase component, process the amplitude component and the phase component respectively and then perform inverse Fourier transform, and obtain the output feature based on the result of the inverse Fourier transform; input the degraded reference image into the first generator, and input the clear reference image into the first discriminator to iteratively and alternately train the generator and the discriminator; obtain an image restoration model based on the trained first generator, and the image restoration model is used to restore the degraded image.
[0007] Second aspect, an embodiment of the present application provides an image restoration method, which includes: obtaining an image restoration model, the image restoration model is trained based on a generative adversarial network and includes a first generator, the first generator includes a plurality of amplitude-phase decoupling modules, and each amplitude-phase decoupling module is used to perform a Fourier transform on the input feature to obtain an amplitude component and a phase component, process the amplitude component and the phase component respectively and then perform an inverse Fourier transform, and obtain an output feature based on the result of the inverse Fourier transform; inputting the degraded image into the image restoration model to obtain a restored image.
[0008] Third aspect, an embodiment of the present application provides an image restoration device, which includes: an initialization module, configured to initialize a generative adversarial network, the generative adversarial network includes a first generator and a first discriminator, the first generator is used to convert an input image belonging to the degraded domain into a first image, and the first discriminator is used to identify whether the first image belongs to the clear domain; wherein the first generator includes a plurality of amplitude-phase decoupling modules, and each amplitude-phase decoupling module is used to perform a Fourier transform on the input feature to obtain an amplitude component and a phase component, process the amplitude component and the phase component respectively and then perform an inverse Fourier transform, and obtain an output feature based on the result of the inverse Fourier transform; a training module, configured to input a degraded reference image into the first generator and input a clear reference image into the first discriminator to iteratively and alternately train the generator and the discriminator; an obtaining module, configured to obtain an image restoration model based on the trained first generator, and the image restoration model is used to restore the degraded image.
[0009] Fourth aspect, an embodiment of the present application provides an image restoration device, which includes: an obtaining module, configured to obtain an image restoration model, the image restoration model is trained based on a generative adversarial network and includes a first generator, the first generator includes a plurality of amplitude-phase decoupling modules, and each amplitude-phase decoupling module is used to perform a Fourier transform on the input feature to obtain an amplitude component and a phase component, process the amplitude component and the phase component respectively and then perform an inverse Fourier transform, and obtain an output feature based on the result of the inverse Fourier transform; a restoration module, configured to input the degraded image into the image restoration model to obtain a restored image.
[0010] Fifth aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable by the processor. When the processor executes the computer program, the image restoration method described in the first aspect or the second aspect above is implemented.
[0011] Sixth aspect, an embodiment of the present application provides a computer-readable storage medium, storing a computer program, and when the computer program is executed by a processor, the image restoration method described in the first aspect or the second aspect above is implemented.
[0012] In a seventh aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device is caused to execute the image restoration method described in the first aspect or the second aspect above.
[0013] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: The first generator is trained based on an adversarial generative network, and the training objective of the first generator is to convert a degraded image into a clear image. Based on the trained first generator, an image restoration model can be obtained. The first generator includes an amplitude-phase decoupling module, and the amplitude-phase decoupling module uses Fourier transform. Compared with traditional convolution, it can effectively extract global context information and improve the generalization ability of the model. Moreover, the amplitude-phase decoupling module processes the amplitude component and the phase component obtained by Fourier transform separately and then performs inverse Fourier transform. The Fourier transform used in image processing is a two-dimensional discrete Fourier transform. The obtained amplitude component (also called amplitude spectrum, usually simply referred to as spectrum) mainly contains information related to degradation, while the phase component (also called phase spectrum) mainly contains information related to the image structure and spatial layout. The amplitude-phase decoupling module processes the amplitude component and the phase component separately, separating the degradation information and the structure information in the image, enabling the image restoration model to more retain the structure information of the original image while removing factors causing image degradation such as noise and blur, and improving the effect of image restoration. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0015] Figure 1 is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0016] Figure 2 is a schematic flowchart of an image restoration method provided by an embodiment of the present application;
[0017] Figure 3 is a schematic structural diagram of an adversarial generative network provided by an embodiment of the present application;
[0018] Figure 4 is Figure 3 a schematic structural diagram of the DOAP module in
[0019] Figure 5 is Figure 3 a schematic structural diagram of the image content information extraction module in
[0020] Figure 6 It is a schematic flowchart of an image restoration method provided by another embodiment of the present application;
[0021] Figure 7 It is a schematic structural diagram of an image restoration device provided by an embodiment of the present application;
[0022] Figure 8 It is a schematic structural diagram of an image restoration device provided by another embodiment of the present application. Detailed implementation manners
[0023] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0024] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0025] It should also be understood that the term "and / or" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0026] As used in the specification of the present application and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" depending on the context.
[0027] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0028] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that specific features, structures or characteristics described in connection with that embodiment are included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0029] The image restoration method provided by the embodiments of this application can be applied to an electronic device, and the electronic device includes, but is not limited to, electronic devices with computing functions such as servers, server clusters, mobile phones, tablet computers, laptop computers, desktop computers, personal digital assistants and wearable devices, etc. The embodiments of this application do not impose any restrictions on the specific type of the electronic device.
[0030] Figure 1 Shown is a block diagram of a part of the structure of the electronic device provided by the embodiments of this application. Refer to Figure 1 , the electronic device includes: a processor 10, a memory 20, a bus 30, an input device 40, an output device 50, and a communication device 60. The processor 10 and the memory 20 are connected to each other through the bus 30, and the input device 40, the output device 50, and the communication device 60 are also connected to the bus 30. Those skilled in the art can understand that Figure 1 the structure of the electronic device shown in
[0031] does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. Figure 1 The following specifically introduces each component of the electronic device in combination with
[0032] The processor 10 is the control center of the electronic device and can execute various functions and process data by running programs stored in the memory 20. The processor 10 can be a Central Processing Unit (CPU), and this processor 10 can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or this processor can also be any conventional processor, etc. In some embodiments, the processor 10 can include an AI (Artificial Intelligence) processor, and this AI processor is used to process computational operations related to machine learning.
[0033] The memory 20 is used to store the operating system, application programs, BootLoader, data, and other programs, such as the program code of computer programs, etc. The memory 20 can also be used to temporarily store the data required for executing programs and the data generated. The memory 20 can include high-speed random access memory, and can also include non-volatile memory, such as flash memory, hard disks, multimedia cards, card-type memories, etc. The memory 20 can include storage units provided inside the electronic device, such as the hard disk of the electronic device, and / or removable external storage units, such as external hard disks, USB flash drives, Smart Media Cards (SMCs), Secure Digital (SD) cards, etc.
[0034] The input device 40 can include at least one of a keyboard, a mouse, a touch panel, a joystick, etc., and is used to collect the input operations of the user to generate corresponding input signals.
[0035] The output device 50 is used to output the information to be provided to the user. The output device 50 generally includes a display, and optionally, a Liquid Crystal Display (LCD), an Organic Light-Emitting Diode (OLED), etc. can be adopted. In addition, the output device can further include a speaker.
[0036] The communication device 60 can include a modem, a network card, etc., and is used to establish a network connection with other electronic devices and communicate with each other.
[0037] The image restoration method provided by the embodiments of the present application can be implemented as a computer software program. For example, an embodiment of the present application provides a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 60, and / or installed from a removable external storage unit. When the computer program is executed by a processor 10, various functions defined in the image restoration method provided by the embodiments of the present application are implemented.
[0038] Figure 2 The schematic flowchart of the image restoration method provided by an embodiment of the present application is shown. By way of example and not limitation, this method can be applied to the above-mentioned electronic device.
[0039] S1: Initialize the adversarial generation network.
[0040] Using a matching image pair to train the adversarial generation network (GAN) can achieve domain transfer of images, that is, style conversion. Among them, the content of each image in the matching image pair is the same, but the domains to which they belong are different, that is, the styles are different.
[0041] Specifically, the GAN includes a generator and a discriminator. The generator generates an image that it believes is closer to the real sample (that is, the image belonging to the target domain paired with the input image) based on the input image belonging to the original domain. The goal is to maximize the reduction of the gap between the image it generates and the real sample. The discriminator judges the authenticity of the image generated by the generator and the real sample, that is, whether it belongs to the target domain. The goal is to maximize the determination that the image generated by the generator is fake and the real sample is true. The generator and the discriminator each have their own goals. In each round of training, they alternately maximize their own expectations. In the process of mutual restraint and confrontation, the generation technology of the generator is continuously improved until finally the discriminator can no longer distinguish between true and false.
[0042] In this embodiment, the domain transfer that needs to be performed is from the degraded domain to the clear domain. Since a matching image pair needs to be used as a training sample to enable a one-way GAN (that is, a GAN that only includes a single generator) to achieve domain transfer. According to the expression of the degraded image The degraded image g belongs to the degraded domain, and the image paired with g is the original image f, which belongs to the clear domain. In practical applications, when the acquired image is a degraded image, it means that there are problems in the imaging, transmission, etc. processes, and the original image cannot be directly obtained. Therefore, the actually directly obtained degraded image cannot be used to train a one-way GAN to achieve domain transfer. To solve this problem, the actually acquired clear image can be used as the original image, and the clear image can be degraded according to the expression of the degraded image, such as adding blur, noise, etc., to generate a degraded image. The clear image and the degraded image generated therefrom can form an image pair to be used as the training samples of the one-way GAN.
[0043] The generative adversarial network (GAN) includes a first generator and a first discriminator. The first generator is used to convert an input image belonging to the degraded domain into a first image, and the first discriminator is used to identify whether the first image belongs to the clear domain. In addition, the first discriminator is also used to identify whether an image belonging to the clear domain that matches the input image belongs to the clear domain.
[0044] The first generator may include a plurality of amplitude-phase decoupling (DOAP) modules. In some embodiments, the first generator includes a first encoder and a first decoder. The first encoder and the first decoder each include at least one amplitude-phase decoupling (DOAP) module, and the number of DOAP modules in the first encoder and the first decoder is equal. If the number of DOAP modules included in the first encoder and the first decoder is greater than 1, these DOAP modules are generally cascaded. In the first encoder, the output of the previous DOAP module is generally input to the next DOAP module after downsampling; in the first decoder, the output of the previous DOAP module is generally input to the next DOAP module after upsampling.
[0045] Each DOAP module is used to perform a Fourier transform on the input feature to obtain an amplitude component and a phase component, then process the amplitude component and the phase component separately, then perform an inverse Fourier transform by combining the separately processed amplitude component and phase component, and then obtain the output feature based on the result of the inverse Fourier transform.
[0046] The Fourier transform used in image processing is a two-dimensional discrete Fourier transform. The obtained magnitude component (also known as the magnitude spectrum, usually simply referred to as the spectrum) mainly contains information related to degradation, while the phase component (also known as the phase spectrum) mainly contains information related to the image structure and spatial layout. Optionally, the specific processing of the magnitude component and the phase component may include: performing residual convolution on the magnitude component to extract local spatial features and retaining the original phase component, and then the result of the residual convolution, that is, the local spatial features and the phase component can be combined for inverse Fourier transform. The result of the inverse Fourier transform can be directly used as the output feature of the DOAP module. Or, in the DOAP module, the part that uses the Fourier transform to extract global information is called the frequency branch. In addition, the DOAP module also includes a channel branch, which is used to perform pooling and convolution on the input features to obtain channel features, and then the sum of the channel features and the result of the inverse Fourier transform is used as the output feature.
[0047] Optionally, the first generator further includes an image content information extraction module for extracting the context information of the image based on the attention mechanism. The first feature output by the first encoder and the information feature output by the image content information extraction module are concatenated and then input into the first decoder for decoding.
[0048] The image content information extraction module enables the model to focus on specific regions or features in the input image, and the generator can dynamically adjust the generated output according to different regions or features of the input image, so as to obtain better generation results.
[0049] Optionally, the image content information extraction module includes an attention module and a high-frequency enhancement module connected in sequence. The attention module can be a plug-and-play attention module that combines self-attention and convolution to generate an attention weight map. The high-frequency enhancement module is used to perform high-frequency enhancement on the attention weight map output by the attention module, so as to enhance the expression ability of high-frequency information.
[0050] Specifically, the high-frequency enhancement module is specifically used to perform Fourier transform on the attention weight map to obtain the spectrum map of the attention weight map, use a filter to decompose the spectrum map into a first sub-map and a second sub-map, the frequency retained by the first sub-map is lower than that of the second sub-map, use a learnable weight to weight the second sub-map, and after combining the weighted second sub-map and the first sub-map, perform inverse Fourier transform to obtain the high-frequency enhanced attention weight map.
[0051] Specifically, the mathematical expression of the two-dimensional discrete Fourier transform is as follows:
[0052]
[0053] Among them, F(u, v) is a complex value in the frequency domain, (u, v) represents the coordinates in the frequency domain, f(x, y) is the pixel value in the attention map, and N is the size of the attention weight map. In the frequency domain, that is, the spectrogram of the attention weight map, low-frequency information is located in the center of the image, and high-frequency information is located at the edges of the image. To perform low-pass filtering on the image, a translation operation can be used to move the low-frequency part to the center of the image, and the specific formula for translation is as follows.
[0054]
[0055] Retain the central region of the spectrogram after translation, which contains low-frequency information, and set the remaining parts to zero to obtain the first sub-image Z l , and the second sub-image is the difference between the spectrogram after translation and the first sub-image:
[0056] Z h = F'(u, v) - Z l
[0057] The reweighted spectrogram is obtained as follows:
[0058] Z′ = Z l + WZ h
[0059] where W represents learnable weight parameters, which are initialized to 1 and can be directly optimized through backpropagation. The reweighted spectrogram is transformed back to the spatial domain through inverse Fourier transform to obtain a new attention map, in which the high-frequency information is adaptively adjusted according to the weighting strategy.
[0060] Then, perform inverse Fourier transform on the weighted spectrogram, and the mathematical expression is as follows:
[0061]
[0062] Although the generated degraded image can form an image pair with its original image to train a one-way GAN, the generation process is essentially based on the degraded mathematical model summarized by predecessors and may not necessarily cover all possible degradations that actually occur. Optionally, to further improve the generalization ability of the finally obtained image restoration model, the actually obtained degraded images can be used to train the GAN. In this case, the one-way GAN cannot achieve domain transfer. To solve this problem, a cycle adversarial generation network (cycleGAN) can be introduced to achieve domain transfer using unmatched images.
[0063] In some embodiments, the adversarial generation network is a cyclic adversarial generation network. In addition to the first generator and the first discriminator, it further includes a second generator and a second discriminator. The second generator is used to convert the input image belonging to the clear domain into a second image, and the second discriminator is used to identify whether the second image belongs to the degraded domain.
[0064] Generally speaking, the structure of the second generator is the same as that of the first generator, but the parameters are not shared. Specifically, the second generator includes a second encoder and a second decoder, and the second encoder and the second decoder each include at least one of the amplitude-phase decoupling modules. If the adversarial generation network includes an image content information extraction module, the first generator and the second generator share the same image content information extraction module, and the second features output by the second encoder are concatenated with the information features output by the image content information extraction module and then input into the second decoder for decoding.
[0065] The following specifically describes the specific structure of the adversarial generation network provided by an embodiment of the present application in conjunction with the accompanying drawings.
[0066] Please refer to Figure 3 , the cyclic adversarial generation network provided by an embodiment of the present application includes: Figure 3 The first generator G in the upper left and lower right AB , the second generator G in the lower left and upper right BA , the first discriminator above and the second discriminator below.
[0067] The first generator G AB and the second generator G BA have the same structure, but the parameters are not shared, and no specific distinction will be made in the following description.
[0068] The first discriminator and the second discriminator have the same structure, and their parameters can be shared or not shared.
[0069] The generator includes an encoder and a decoder. The encoder consists of 3 cascaded encoding modules. Each encoding module first applies a 3×3 convolutional layer to extract shallow features from its input, and the shallow features obtain high-level features through a DOAP module. The high-level features output by the previous encoding module are downsampled and then input into the next encoding module. The output of the encoder, that is, the high-level features output by the last encoding module and the information features output by the image content information extraction module are concatenated and then input into the decoder. The decoder includes 3 cascaded DOAP modules. Except for the last DOAP module, the features output by each DOAP module are input into the next DOAP module after passing through a 1×1 convolution and upsampling, and the features output by the last DOAP module are convolved by a 3×3 convolution to obtain the output of the decoder and the generator.
[0070] Please refer to Figure 4, the DOAP module contains a frequency branch and a channel branch. The frequency branch is used to learn global context information. Specifically, the input feature F is subjected to Fourier transform, and the obtained amplitude and phase components are processed separately. Specifically, it includes: using a 3×3 residual group to perform convolution on the amplitude component to obtain local spatial features, combining the local spatial features with the original phase information, and obtaining the feature S through inverse Fourier transform. The above process can be expressed as:
[0071] A, P = FFT(F)
[0072] A = conv_res(A)
[0073] S = IFFT(A, P)
[0074] Among them, FFT represents Fourier transform, conv_res represents residual convolution, A represents amplitude, P represents phase, and IFFT represents inverse Fourier transform.
[0075] The channel branch first performs max pooling and average pooling respectively, fuses the two pooling results through a fully connected layer, then uses convolution operations to learn local texture details, and finally adds them to the feature S obtained by inverse Fourier transform to obtain the final output. The above process can be expressed as:
[0076] out 1 = ReLU(conv3(FC(MP(F)+AP(F)))+b 1 )
[0077] out = S + conv3(out 1 )
[0078] Among them, MP represents max pooling, AP represents average pooling, FC represents fully connected layer, conv3 represents convolution operation with a convolution kernel size of 3×3, ReLU represents activation function, and b 1 represents the bias parameter of the network.
[0079] Please refer to Figure 5 , the image content information extraction module contains multiple residual blocks. The residual blocks are composed of 3×3 kernel size convolutions, and a plug-and-play attention module is inserted into the last residual block. The attention module combines the advantages of self-attention and convolution operators, uses 3×3 convolution, average pooling, 1×1 convolution, and finally generates attention weights through softmax. The generation process of the attention weight map can be expressed as:
[0080] Z = soft max(conv1(AP(conv3(X))))
[0081] Among them, AP represents average pooling, softmax represents the normalized exponential function, conv1 represents the convolution operation with a convolution kernel size of 1×1, conv3 represents the convolution operation with a convolution kernel size of 3×3, and X represents the input feature of the attention module.
[0082] Next, the attention weight map Z is transformed to the frequency domain by Fourier transform and then decomposed by a filter to obtain two subgraphs, namely the low-frequency and high-frequency subgraphs, i.e., the first subgraph and the second subgraph. Then, the high-frequency subgraph is reweighted using learnable parameters to highlight important frequency components. Finally, the weighted high-frequency subgraph is merged with the low-frequency subgraph and an inverse Fourier transform is performed to obtain the attention weight map with enhanced high-frequency information.
[0083] Then, through the 1×1 and 3×3 convolution operations after the inverse Fourier transform, the adjusted attention weight map is applied to the input feature, and then summed with the original input feature to obtain the output feature of the image content information extraction module.
[0084] The discriminator can be a commonly used discriminator in the GAN network. For example, the discriminator can double the number of channels of the input image X through a convolutional layer, and then pass through a LeakyReLU activation function, which is used to introduce non-linearity to help the discriminator better distinguish real images and generated images. Then, the image size is reduced by half through a downsampling layer. Repeat this process, doubling the number of channels of the image and reducing the image size by half each time, where each pixel represents the authenticity of the input image. Among them, each convolutional layer is followed by an instance normalization layer and a LeakyReLU activation function. The above process can be expressed as:
[0085] out 1 =Down 1 (Leaky ReLU(conv3(X)))
[0086] out 2 =Down 2 (Leaky ReLU(IN(conv3(out 1 ))))
[0087] out 3 =Down 3 (Leaky ReLU(IN(conv3(out 2 ))))
[0088] out=conv1(out 3 )
[0089] Among them, conv1 represents a convolution operation with a convolution kernel size of 1×1, conv3 represents a convolution operation with a convolution kernel size of 3×3, In represents instance normalization operation, LeakyReLU represents an activation function, and Down represents a downsampling operation.
[0090] S2: Input the degraded reference image into the first generator, and input the clear reference image into the first discriminator to iteratively and alternately train the generator and the discriminator.
[0091] Specifically, in each round of training, the parameters of the generator can be fixed first to optimize the parameters of the discriminator; then the parameters of the discriminator are fixed to optimize the parameters of the generator. Or, the parameters of the discriminator are fixed first to optimize the parameters of the generator; then the parameters of the generator are fixed to optimize the parameters of the discriminator. Repeat the training process until the stop condition is met to complete the training. The stop condition can be that the number of iterations reaches a threshold or the loss function is less than a preset value, etc.
[0092] The loss function based on which the parameters of the discriminator are optimized includes the adversarial loss function of the discriminator. The loss function based on which the parameters of the generator are optimized includes the adversarial loss function of the generator.
[0093] If the generative adversarial network is cycleGAN, during the training process, the input of the second generator includes the clear reference image, and the input of the second discriminator includes the degraded reference image. During the alternating iterative training process, generally, the two generators are optimized together, and the two discriminators are optimized together. In some embodiments, the loss function based on which the parameters of the generator are optimized can further include a cycle consistency loss to further narrow the space of the mapping function of the generator. Refer to Figure 3 the right half of, convert the first image to the third image through the second generator, and convert the second image to the fourth image through the first generator. Ideally, the degraded reference image should be the same as the third image, and the clear reference image should be the same as the fourth image. Calculate the cycle consistency loss function based on the error between the degraded reference image and the third image, and the error between the clear reference image and the fourth image. The loss function L loss is as follows.
[0094] L loss =L GAN (G AB ,D,X,Y)+L GAN (G BA ,D,X,Y)+λL 2 (G AB ,G BA )
[0095] L GAN (G AB ,D,X,Y)=E y~Y[logD(y)] + E x~X [log(1 - D(G AB (x)))
[0096] L GAN (G BA , D, X, Y) = E y~Y [logD(y)] + E x~X [log(1 - D(G BA (x)))
[0097] L 2 (G AB , G BA ) = E x~pdata(x) [||G AB (G BA (x)) - x|| 1 + E y~pdata(y) [||G BA (G AB (y)) - y|| 1
[0098] Where λ represents the weight magnitude.
[0099] S3: Obtain an image restoration model based on the trained first generator.
[0100] The image restoration model is used to restore the degraded image.
[0101] If the adversarial generation network does not include an image content information extraction module, the trained first generator can be directly used as the image restoration model. If the adversarial generation network includes an image content information extraction module, the trained first generator can be directly used as the image restoration model; or, in order to achieve a better restoration effect, the trained first generator and the image content information extraction module are combined according to the connection method in the adversarial generation network, that is, the features output by the image content information extraction module and the output features of the first encoder in the first generator are concatenated and then input into the first encoder to obtain the image restoration model.
[0102] Through the implementation of this embodiment, the first generator is trained based on the adversarial generation network. The training objective of the first generator is to convert the degraded image into a clear image. An image restoration model can be obtained based on the trained first generator. The first generator includes an amplitude-phase decoupling module, and the amplitude-phase decoupling module uses Fourier transform. Compared with traditional convolution, it can effectively extract global context information and improve the generalization ability of the model.
[0103] The amplitude-phase decoupling module processes the amplitude component and the phase component obtained by Fourier transform separately and then performs inverse Fourier transform. The Fourier transform used in image processing is two-dimensional discrete Fourier transform. The obtained amplitude component (also called amplitude spectrum, usually simply referred to as spectrum) mainly contains information related to degradation, while the phase component (also called phase spectrum) mainly contains information related to image structure and spatial layout. The amplitude-phase decoupling module processes the amplitude component and the phase component separately, separating the degradation information and the structure information in the image, enabling the image restoration model to retain more structure information of the original image while removing factors causing image degradation such as noise and blur, and improving the effect of image restoration.
[0104] In addition, in the case of introducing an image content information extraction module, this module can utilize the correspondence relationship between the two image domains of clear images and degraded images, extract and fuse the content information of the two domains, thereby enhancing the model's ability to reconstruct image details and textures.
[0105] Figure 6 The schematic flowchart of an image restoration method provided by another embodiment of the present application is shown. As an example but not a limitation, this method can be applied to the above-mentioned electronic device.
[0106] S10: Obtain an image restoration model.
[0107] The image restoration model is trained based on a generative adversarial network and includes a first generator. The first generator includes multiple amplitude-phase decoupling modules. Each amplitude-phase decoupling module is used to perform Fourier transform on the input feature to obtain an amplitude component and a phase component, process the amplitude component and the phase component separately and then perform inverse Fourier transform, and obtain an output feature based on the result of the inverse Fourier transform. For specific descriptions, reference can be made to the relevant content in the foregoing embodiments.
[0108] S20: Input the degraded image into the image restoration model to obtain a restored image.
[0109] Figure 7 The structural schematic diagram of an image restoration device provided by an embodiment of the present application is shown. The image restoration device includes an initialization module 11, a training module 12, and an acquisition module 13.
[0110] The initialization module 11 is used to initialize a generative adversarial network. The generative adversarial network includes a first generator and a first discriminator. The first generator is used to convert an input image belonging to the degraded domain into a first image, and the first discriminator is used to identify whether the first image belongs to the clear domain. The first generator includes multiple amplitude-phase decoupling modules. Each amplitude-phase decoupling module is used to perform Fourier transform on the input feature to obtain an amplitude component and a phase component, process the amplitude component and the phase component separately and then perform inverse Fourier transform, and obtain an output feature based on the result of the inverse Fourier transform.
[0111] A training module 12 for inputting a degraded reference image into a first generator and inputting a clear reference image into a first discriminator to iteratively and alternately train the generator and the discriminator.
[0112] An acquisition module 13 for obtaining an image restoration model based on the trained first generator, where the image restoration model is used to restore a degraded image.
[0113] Figure 8 The structural schematic diagram of an image restoration device provided by an embodiment of the present application is shown. The image restoration device includes an acquisition module 21 and a restoration module 22.
[0114] An acquisition module 21 for obtaining an image restoration model, where the image restoration model is trained based on an adversarial generation network and includes a first generator. The first generator includes a plurality of amplitude-phase decoupling modules. Each amplitude-phase decoupling module is used to perform a Fourier transform on the input feature to obtain an amplitude component and a phase component, process the amplitude component and the phase component respectively, then perform an inverse Fourier transform, and obtain an output feature based on the result of the inverse Fourier transform.
[0115] A restoration module 22 for inputting a degraded image into the image restoration model to obtain a restored image.
[0116] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0117] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.
[0118] The embodiments of the present application also provide a computer-readable storage medium storing a computer program, which when executed by a processor can implement the steps in the above-mentioned method embodiments.
[0119] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0120] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0121] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0122] In the embodiments provided in the present application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0123] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0124] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An image restoration method, characterized in that, the method includes: Initializing an adversarial generation network, the adversarial generation network includes a first generator and a first discriminator, the first generator is used to convert an input image belonging to the degradation domain into a first image, and the first discriminator is used to identify whether the first image belongs to the clear domain; Wherein the first generator includes a plurality of amplitude-phase decoupling modules, each amplitude-phase decoupling module is used to perform Fourier transform on the input features to obtain an amplitude component and a phase component, process the amplitude component and the phase component respectively and then perform inverse Fourier transform, and obtain output features based on the result of the inverse Fourier transform; Input the degraded reference image into the first generator, and input the clear reference image into the first discriminator to iteratively and alternately train the generator and the discriminator; Obtain an image restoration model based on the trained first generator, and the image restoration model is used to restore the degraded image.
2. The method according to claim 1, characterized in that, The inverse Fourier transform after separately processing the amplitude component and the phase component includes: Performing residual convolution on the amplitude component to extract local spatial features; Combining the result of the residual convolution with the phase component and performing inverse Fourier transform.
3. The method according to claim 1, characterized in that, The obtaining of the output features based on the result of the inverse Fourier transform includes: Performing pooling and convolution on the input features to obtain channel features; Taking the sum of the channel features and the result of the inverse Fourier transform as the output features.
4. The method according to claim 1, characterized in that, The first generator includes a first encoder, a first decoder and an image content information extraction module, the first encoder and the first decoder each include at least one of the amplitude-phase decoupling modules, the image content information extraction module is used to extract the context information of the input image based on the attention mechanism, and the input features of the first decoder are obtained by splicing the first features output by the first encoder and the information features output by the image content information extraction module.
5. The method according to claim 4, characterized in that, The image content information extraction module includes an attention module and a high-frequency enhancement module connected in sequence, and the high-frequency enhancement module is used to perform high-frequency enhancement on the attention weight map output by the attention module.
6. The method according to claim 5, characterized in that, The high-frequency enhancement module is specifically used to perform Fourier transform on the attention weight map to obtain the frequency spectrum map of the attention weight map, decompose the frequency spectrum map using a filter to obtain a first sub-map and a second sub-map, the frequency retained by the first sub-map is lower than that of the second sub-map, use a learnable weight to weight the second sub-map, and merge the weighted second sub-map with the first sub-map and then perform inverse Fourier transform to obtain the high-frequency enhanced attention weight map.
7. The method according to claim 4, characterized in that, The obtaining of the image restoration model based on the trained first generator includes: Combine the trained first generator and the image content information extraction module according to the connection mode in the adversarial generation network to obtain the image restoration model.
8. The method according to claim 1, wherein, the adversarial generation network further includes a second generator and a second discriminator, the second generator is used to convert an input image belonging to the clear domain into a second image, and the second discriminator is used to identify whether the second image belongs to the degradation domain.
9. The method according to claim 8, wherein, the second generator includes a second encoder and a second decoder, and the second encoder and the second decoder each include at least one of the amplitude-phase decoupling modules.
10. An image restoration method, wherein, the method includes: Obtain an image restoration model, the image restoration model is trained based on an adversarial generation network and includes a first generator, the first generator includes a plurality of amplitude-phase decoupling modules, and each amplitude-phase decoupling module is used to perform a Fourier transform on the input feature to obtain an amplitude component and a phase component, process the amplitude component and the phase component respectively and then perform an inverse Fourier transform, and obtain an output feature based on the result of the inverse Fourier transform; Input the degraded image into the image restoration model to obtain a restored image.
11. The method according to claim 10, wherein, the first generator includes a first encoder, a first decoder and an image content information extraction module, the first encoder and the first decoder each include at least one of the amplitude-phase decoupling modules, the image content information extraction module is used to extract the context information of the input image based on the attention mechanism, and the input feature of the first decoder is obtained by splicing the first feature output by the first encoder and the information feature output by the image content information extraction module.
12. An electronic device, including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein, when the processor executes the computer program, the method according to any one of claims 1 to 11 is implemented.
13. A computer-readable storage medium, the computer-readable storage medium stores a computer program, wherein, when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.