Image super-resolution method with universal robustness
The image super-resolution method achieves universal robustness by using median random smoothing in its inference phase, effectively handling a wide range of disturbances and outperforming traditional methods in terms of computational efficiency and robustness.
Patent Information
- Application Number
- FR2023014472
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing image super-resolution methods struggle with universal robustness against a wide variety of disturbances, such as noise and adversarial attacks, especially when these disturbances were not encountered during training.
The method employs a neural network trained to produce high-resolution images from low-resolution ones, utilizing an inference phase that generates multiple noisy images by applying Gaussian noise. The network predicts images for these noisy inputs, and the high-resolution output is calculated as the median of these predictions, enhancing robustness through median random smoothing.
This approach results in a universally robust image super-resolution method that maintains high performance even against disturbances not seen during training, while also being computationally efficient compared to adversarial training methods.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Image super-resolution method with universal robustness Field of the invention
[0001] The present invention belongs to the field of image super-resolution. In particular, the invention relates to an image processing method and device based on a neural network for improving the resolution of an image. The method has universal robustness in the face of a wide variety of disturbances (noise, adverse attacks, etc.). State of the art
[0002] The objective of super-resolution (SR) is to improve the spatial resolution of a low-resolution (LR) image by producing a clear, artifact-free high-resolution (HR) image. For example, the resolution of the HR image is increased by a factor of four compared to the resolution of the LR image (e.g., going from a 120x120 resolution to a 480x480 resolution). Image super-resolution is used in various fields, such as oceanography, surveillance, or medical imaging.
[0003] In recent years, super-resolution methods based on deep neural networks (DNNs) have made considerable progress and now outperform traditional methods such as linear interpolation or covariance or correlation estimation in a low-resolution image. These traditional methods fail to faithfully capture high-frequency details in images, and therefore often produce results that appear blurry.
[0004] Most super-resolution machine learning methods rely on pairs of high- and low-resolution images used to train a neural network in a supervised manner. However, in the real world, such image pairs are difficult to obtain in practice. Therefore, high-resolution images are typically bicubically downsampled to artificially generate the associated low-resolution images. Unfortunately, this strategy introduces artifacts into the super-resolved images by removing or modifying features of the low-resolution images. As a result, super-resolution neural networks trained in this way often have difficulty generalizing to images that are not bicubically downsampled (especially real-world images).
[0005] In documents [1] and [2], the authors developed methods to recover noise and details lost during bicubic subsampling. These methods are, however, computationally expensive because they require creating several neural network models, on the one hand to generate pairs of high and low resolution images that have the same characteristics, and on the other hand to train a super-resolution model from the data thus generated. In addition, the resulting models perform relatively poorly on images with a type of noise or disturbance that is not present in the training data.
[0006] Document [3] describes a GAN (Generative Adversarial Network) type image super-resolution model composed of two neural networks: a generator and a discriminator. The generator is responsible for creating a high-resolution image from a low-resolution image. The discriminator is responsible for distinguishing real images from generated images. The model named ESRGAN (acronym for "Enhanced Super-Resolution Generative Adversarial Networks") is an image super-resolution model that offers high-quality results. Like most neural networks currently used for super-resolution, ESRGAN loses precision for real-world images or for images affected by slight noise or an adversarial attack.
[0007] Document [4] describes a learning method for building a model that is robust against small perturbations. In this document, the authors use an adversarial learning method: first, an adversarial attack is used to create examples that target the weaknesses of the model; then, the generated adversarial examples are used during training to improve the model's ability to process noisy low-resolution images. In this document, the adversarial attack used is of the PGD type (acronym for "Projected Gradient Descent") and it takes into account both a loss at the pixel level (based for example on a calculation of mean squared error, MSE for "Mean Squared Error" in English) and a loss at the level of perceptual image similarity (based on LPIPS type measures ("Learned Perceptual Image Patch Similarity").
[0008] A first drawback of this method is that applying adversarial learning to image super-resolution is particularly computationally expensive. A second drawback is that the model obtained with specific adversarial learning is not robust against all types of disturbance. For example, the model is not robust against Gaussian noise with a small standard deviation, nor against adversarial attacks of a type different from that considered for the adversarial learning. Also, the accuracy of the model is decreased on real unattacked or unperturbed images.
[0009] In another field, namely object detection, document [5] describes a method for certifying a neural network by a "median randomized smoothing" (MRS) method with Monte Carlo sampling. However, it appears that a large number of samples (of the order of 2000) is necessary for classification or regression procedures in the field of object detection. Statement of the invention
[0010] The present invention aims to remedy all or part of the drawbacks of the prior art, in particular those set out above, by proposing a method for image super-resolution which remains universally robust in the face of a wide variety of disturbances (noise or adverse attacks), and in particular in the face of disturbances which have not necessarily been observed by the model during its training.
[0011] For this purpose, and according to a first aspect, the present invention proposes a method for image super-resolution. The method is implemented by an image processing device storing a neural network previously trained to produce a high-resolution image from a low-resolution image. The method comprises an inference phase comprising the following steps: - generating a plurality of noisy images from a low-resolution image, each noisy image being obtained by a random application to the low-resolution image of Gaussian noise having a predefined standard deviation, - a prediction, by the previously trained neural network, of an image for each of the noisy images thus obtained, - generation of a high-resolution image corresponding to a median of the predicted images thus obtained.
[0012] In particular embodiments, the invention may further comprise one or more of the following characteristics, taken in isolation or in all technically possible combinations.
[0013] In particular embodiments, the prior training of the neural network comprises an adjustment phase based on a training data set comprising a set of image pairs, each image pair comprising a high-resolution version and a low-resolution version of the same image. The adjustment phase comprises, for each low-resolution image of the training data set, and for each of a plurality of different standard deviation values: - a generation of a plurality of noisy images, each noisy image being obtained by a random application to the low-resolution image considered of a Gaussian noise having a standard deviation equal to the standard deviation value considered, - a prediction by the neural network of a high-resolution image for each of the noisy images thus generated, - a generation of a high-resolution image corresponding to a median of the predicted images thus obtained, - a calculation of a loss function from the median thus generated and the high-resolution image associated with the low-resolution image considered, - an update of the parameters of the neural network by a gradient descent based on the loss function.
[0014] In particular embodiments, the loss function comprises a component relating to a loss at the pixel level and a component relating to a loss at the level of a perceptual image similarity.
[0015] In particular embodiments, the value of the standard deviation of the Gaussian noise to be used for the generation of the plurality of noisy images during the inference phase is predefined according to a metric of perceptual image similarity measured for different images obtained at the output of the inference phase and for different candidate values of standard deviation.
[0016] In particular embodiments, the value of the standard deviation of the Gaussian noise used for the generation of the plurality of noisy images is between 0.005 and 0.1.
[0017] In particular modes of implementation, the number of noisy images whose predictions are used for the generation of the median during the inference phase is between 5 and 100, or even between 10 and 50, or even between 15 and 25.
[0018] In particular embodiments, two different standard deviation values are used during the adjustment phase.
[0019] In particular embodiments, the two different standard deviation values used during the adjustment phase comprise a first standard deviation value between 0.01 and 0.06 and a second standard deviation value between 0.1 and 0.6.
[0020] In particular embodiments, the number of noisy images whose predictions are used for generating a median during the adjustment phase is between 2 and 8.
[0021] In particular embodiments, the neural network is a generative adversarial network.
[0022] In particular embodiments, the neural network is based on the ESRGAN model.
[0023] In particular modes of implementation, the prior training of the neural network comprises an adversarial training phase of the BIM, FGSM, PGD or CW type.
[0024] According to a second aspect, the present invention relates to a method for training a neural network for image super-resolution. The training method comprises an adjustment phase based on a training data set comprising a set of image pairs, each image pair comprising a high-resolution version and a low-resolution version of the same image. The adjustment phase comprises, for each low-resolution image of the training data set, and for each of a plurality of different standard deviation values: - a generation of a plurality of noisy images, each noisy image being obtained by a random application to the low-resolution image considered of a Gaussian noise having a standard deviation equal to the standard deviation value considered, - a prediction by the neural network of a high-resolution image for each of the noisy images thus generated, - a generation of a high-resolution image corresponding to a median of the predicted images thus obtained, - a calculation of a loss function from the median thus generated and the high-resolution image associated with the low-resolution image considered, - an update of the parameters of the neural network by a gradient descent based on the loss function.
[0025] According to a third aspect, the present invention relates to an image processing device configured to implement a method according to any one of the preceding embodiments. In particular, the device comprises a memory storing a neural network previously trained to produce a high-resolution image from a low-resolution image. The device comprises a processor configured to implement an inference phase, using the neural network, in order to generate a high-resolution image from a low-resolution image. The image processing device is characterized in that the inference phase comprises: - a generation of a plurality of noisy images from the low-resolution image, each noisy image being obtained by a random application to the low-resolution image of Gaussian noise having a predefined standard deviation, - a prediction, by the previously trained neural network, of an image for each of the noisy images thus obtained, - a generation of the high-resolution image corresponding to a median of the predicted images thus obtained. Presentation of figures
[0026] The invention will be better understood on reading the following description, given by way of non-limiting example, and made with reference to Figures 1 to 15 which represent:
[0027] [Fig-1] a schematic representation of an image super-resolution method according to the invention,
[0028] [Fig.2] a schematic representation of an image processing device configured to implement the image super-resolution method according to the invention,
[0029] [Fig.3] a schematic representation of the main steps of the inference phase of the super-resolution method according to the invention,
[0030] [Fig.4] another schematic representation of the inference phase illustrated in [Fig.l],
[0031] [Fig.5] a schematic representation of a particular mode of implementation of the image super-resolution method according to the invention, with an adjustment phase prior to the inference phase,
[0032] [Fig.6] a schematic representation of the main stages of the adjustment phase,
[0033] [Fig.7] another schematic representation of the adjustment phase illustrated in [Fig.6],
[0034] [Fig.8] a table showing the performance of a first example of implementation of the method according to the invention (CertSR), in comparison with other methods, on raw images, noisy images and blurred images,
[0035] [Fig.9] a table showing the performance of CertSR, in comparison with other methods, on images attacked by different types of adversarial attacks,
[0036] [Fig. 10] a table showing the performance of CertSR, in comparison with other methods, on real-world images with different types of natural noise,
[0037] [Fig. 11] a table comparing the performances of different examples of implementation of the invention (CertSR, MSR-FT and MSR-INF),
[0038] [Fig. 12] a table comparing the performance of an implementation of the invention on different neural networks,
[0039] [Fig. 13] a table comparing the performance of CertSR with a regularization method,
[0040] [Fig. 14] a table illustrating the choice of standard deviation values used in the phase CertSR adjustment,
[0041] [Fig. 15] a table illustrating the choice of the standard deviation value used in the inference phase of CertSR.
[0042] In these figures, identical references from one figure to another designate identical or similar elements. For reasons of clarity, the elements represented are not necessarily on the same scale, unless otherwise stated. Detailed description of the invention
[0043] As indicated previously, the present invention aims to propose an image super-resolution method which remains universally robust in the face of a wide variety of disturbances (noise or adverse attacks), and in particular in the face of disturbances which have not necessarily been observed by the model during its training.
[0044] As illustrated in [Fig.l], the method 100 according to the invention is based on a specific inference phase 300 from a neural network previously trained to produce a high-resolution image from a low-resolution image. This specific inference phase 300 is based on median random smoothing. The prior training 200 is not necessarily part of the method 100 according to the invention. Indeed, the invention is mainly based on the inference phase 300, and the way in which the prior training 200 is implemented is not necessarily important for the invention.
[0045] Neural networks for image super-resolution are generally trained in a supervised manner, from a data set comprising a set of pairs of high-resolution and low-resolution images. However, nothing would prevent considering a neural network trained in an unsupervised manner. Different types of neural networks can be envisaged. As will be seen later, the method 100 according to the invention has been extensively tested based on a generative adversarial network, namely the ESRGAN model described in [3].
[0046] The method 100 according to the invention is implemented by an image processing device. [Fig. 2] schematically illustrates such an image processing device 10. The device 10 comprises a memory 12 in which the previously trained neural network 13 is stored, as well as one or more processors 11 configured to implement the inference phase 300. The memory 12 may in particular comprise a set of lines of code which, when executed by the processor(s) 11, configure the processor(s) 11 to implement the inference phase 300 of the method 100 according to the invention.
[0047] Figures 3 and 4 schematically represent the main steps of the inference phase 300 according to the invention. As illustrated in these figures, the inference phase 300 aims to produce a high resolution image from a low resolution LR image.
[0048] The inference phase 300 firstly comprises a step 310 of generating several noisy images from the low-resolution image. Each noisy image is obtained by applying to the low-resolution image LR a random sample of Gaussian noise N(0, a) having a zero mean and a predefined standard deviation 0. We denote by N the number of noisy images obtained: LRb LR2, ..., LRt„ .... LRn. Each noisy image LRfl therefore corresponds to the addition, pixel by pixel, of the low-resolution image LR with a random draw of the Gaussian noise N(Q, a). The draw is independent and identically distributed for each pixel of each image.
[0049] According to a first example, it is possible, for each pixel, to make an independent and identically distributed random draw and to apply the drawn value to each of the R, G and B components of the pixel (one draw per pixel). According to a second example, it is possible to make an independent and identically distributed random draw for each of the R, G and B components of the pixel (three draws per pixel).
[0050] The inference phase 300 then comprises a prediction step 320, by the previously trained neural network 13, of a high-resolution image HRn for each of the noisy images LRn obtained at the end of step 310.
[0051] Finally, the inference phase 300 comprises a step 330 of generating a high-resolution HR image corresponding to a median of the predicted HRn images obtained at the end of step 320. This median is calculated pixel by pixel. By calculating the median, a smoothed model is obtained which is certified in an interval of percentiles which depends on the disturbance present in the input image. Indeed, a smoothed function by percentiles (lp with a disturbance 5 can be bounded as follows:
[0052] [Math.l] Qp(x} ^P(x + Ô) V II 5 II 2 <e
[0053] such that p = and = | ), with the function standard Gaussian distribution (CDF for “Cumulative Distribution Function”).
[0054] The optimal value of the standard deviation 17 of the Gaussian noise to be used for the generation 310 of the noisy images LRn depends on the type of disturbance which affects the input images. The optimal standard deviation value to be used may in particular be predefined as a function of a perceptual image similarity metric (LPIPS, as defined in [8]) measured for different images obtained at the output of the inference phase 300 and for different candidate standard deviation values. As a non-limiting example, the standard deviation value used for the generation 310 of the noisy images may be between 0.005 and 0.1.
[0055] It is advantageous to use smoothing by the median, and not by the mean as is generally the case in certification methods for the classification task. Indeed, the median is almost unaffected by outliers that may be present in low resolution LR images. Unlike the median, the mean tends to smooth regions where predictions are locally constant, which is disadvantageous for images since they often contain textures.
[0056] On the other hand, it appears that median smoothing is particularly well suited to image super-resolution because the variations at the pixel level in the predicted images remain relatively low, and it is possible to control this instability by this certification method with a low number of samples (significantly lower than the number of samples required in the field of classification or regression in object detection tasks). Thus, the number N of noisy images whose predictions are used for the generation 330 of the median image HR during the inference phase 300 can remain relatively low. As a non-limiting example, N can be between 5 and 100, or even between 10 and 50, or even between 15 and 25.
[0057] As illustrated in [Fig.5], if the neural network 13 used is sensitive to noise Gaussian, it is advantageous to carry out, during the preliminary training 200 of the neural network 13, an adjustment phase 220 ("fine-tuning" in English) which follows an initial training phase 210. The adjustment phase 220 also comprises a median random smoothing. In this case, the adjustment phase 220 is part of the invention, but the initial training phase 210 is not necessarily part of the invention (in particular the way in which this initial training phase 210 is implemented is not essential to the invention). The number of epochs in the initial training 210 is generally much greater than the number of epochs in the adjustment phase 220 (the number of epochs is a hyperparameter defining the number of passes of the learning algorithm on a complete set of training data).In a variant, rather than successively carrying out conventional initial training followed by adjustment by median random smoothing, it is possible to integrate median random smoothing throughout the training.
[0058] The adjustment phase 220 can be implemented by the image processing device 10 described with reference to [Fig. 2]. However, nothing would prevent the adjustment phase 220 and the inference phase 300 from each being implemented by a separate device.
[0059] The adjustment phase 220 is based on a training data set comprising a set of image pairs, each image pair comprising a high-resolution version and a low-resolution version of the same image. It can be the same training data set as that used for the initial training 210 of the neural network 13. However, nothing prevents the adjustment phase 220 from being carried out on a training data set different from that used for the initial training phase 210.
[0060] Figures 6 and 7 schematically represent the main steps of the adjustment phase 220 of the method 100 according to the invention.
[0061] For each low resolution image RR of the training data set of the adjustment phase 220, and for each value ^p among a number P of different standard deviation values, the adjustment phase 220 comprises the following steps.
[0062] The adjustment phase 220 initially comprises a generation 221 of several noisy images. Each noisy image is obtained by applying to the low-resolution image RR a random sample of Gaussian noise 7V(0, (Tp) having a zero mean and a predefined standard deviation aP. We denote by M the number of noisy images obtained b RRp^ RR,^ ..., LR^- Each noisy image RRp / n therefore corresponds to the addition, pixel by pixel, of the low-resolution image RR with a random draw of the Gaussian noise / V(0, ¢7^). Here again, it is possible to make a single independent and identically distributed random draw per pixel and to apply this draw to each of the R, G and B components of the pixel (one draw per pixel), or to make an independent and identically distributed random draw for each of the R, G and B components of the pixel (three draws per pixel).
[0063] The adjustment phase 220 then comprises a prediction step 222, by the neural network 13, of a high-resolution image for each of the noisy RRpm images obtained at the end of step 221.
[0064] The adjustment phase 220 then comprises a step 223 of generating a high-resolution image corresponding to a median of the predicted images. obtained at the end of step 222. The median is calculated pixel by pixel. Kpjn
[0065] The adjustment phase 220 then comprises a calculation 224 of a loss function Rp from the median image and the high-resolution image HR associated with the image low resolution RR.
[0066] The adjustment phase 220 finally comprises an update 225 of the parameters (weights) of the neural network 13 by a gradient descent based on the calculated loss function R'p. A loss function R1 can also be calculated for each RR image from the different values R'p (R1 corresponds for example to the sum of the R1^. As illustrated in FIG. 7, a loss function R calculated for a prediction of the RR image by the neural network 13 can also be taken into account for update 225 of the parameters.
[0067] The loss function may advantageously include a component Lpix relating to a loss at the pixel level and a component Lperc relating to a loss at the level of a perceptual similarity of image. When the model used is based on an adversarial generative network, the loss function may also include a component ^adv relating to an adverse loss. The total loss function Lt0( is for example equal to the sum (or a linear combination) of these three components: ^tot — Lpix + Lperc + Ladv^
[0068] For the adjustment phase 220, relatively low values can be chosen for the number P of different standard deviation values to be used and the number M of noisy images to be considered. For example, P = 2 and M = 2 can be chosen.
[0069] If M = 2 the median is equivalent to the mean. In practice, it may be advantageous to use the median rather than the mean, especially if anomalies may be present in the pixel values of the images. In this case it is preferable to take M > 2.
[0070] It should be noted that there is no requirement that the number of noisy images generated for a given standard deviation value be identical for all the standard deviation values considered. In other words, we can have a number Mp of noisy images from the value aP which varies with the index p.
[0071] As a non-limiting example, it is possible to choose a first standard deviation value between 0.01 and 0.06 and a second standard deviation value ^2 between 0.1 and 0.6. Adversary attacks and adversary learning:
[0072] In the following, we will consider different types of adversarial attacks to verify the robustness of the image super-resolution method 100 according to the invention. These adversarial attacks take into account both a loss at the pixel level and a loss at the level of perceptual image similarity. Different types of noise present in real-world images will also be considered. More particularly, four types of adversarial attacks (FGSM, BIM, PGD and CW) and two particular types of noise (a Gaussian noise similar to sensor noise, and a blurring modeled by convolution with a Gaussian kernel) will be considered.
[0073] Neural network-based prediction models have been shown to be susceptible to adversarial example attacks (also known as adversarial example attacks).
[0074] An adversarial attack can be represented as follows: from an input data x, an adversarial example of A takes the form v' = x + 5, with <5 a small value perturbation in front of x. This means that x' is slightly different from x, but a prediction of A' by a given model can differ significantly from the prediction of x. The elements x and 5 belong, for example, to the set p) [ j”. In the example considered, x is an image, that is, a set of pixels, and each pixel can take a normalized value between 0 and 1.
[0075] Several methods have been proposed to generate adversarial examples in order to deceive a prediction model. These methods are, however, relatively little studied in the field of image resolution.
[0076] To create adversarial examples, it is necessary to search for the most effective perturbation to be applied to an input data x to deceive the model considered. The perturbation 5 is determined so as not to modify a pixel of the input image by a value greater than a parameter s. The perturbation is therefore sought in a ball B^X, s) centered on x and of radius e. The ball Bp is defined relative to a norm lp. For adversarial attacks of the FGSM, BIM and PGD type, the norm is considered. For the adversarial attack of the CW type, the norm / 2 is considered. These adversarial attacks are more effective on these norms.
[0077] For the FGSM (Fast Gradient Sign Method) adversarial attack, the sign of the gradient of the loss function with respect to the input data is used to determine how each pixel of the input data should be modified to obtain the most effective perturbation. Mathematically, the adversarial example x' can be written:
[0078] [Math.2] x' = x+e • sig}{^L(f^ >'))
[0079] In this expression, L is the loss function (including both pixel-level loss and perceptual similarity loss), x is the input data (low-resolution image), 3' is the high-resolution image associated with the low-resolution image (ground truth), f 0 1]” [0 1^” is the function corresponding to the neural network considered, therefore corresponds to the prediction of x by the neural network, and Vx is the gradient function with respect to the input data x. The higher the value of the parameter s, the easier it becomes to degrade the performance of the prediction model.
[0080] The BIM adversarial attack (acronym for "Basic Iterative Method") is an extension of the FGSM adversarial attack. The BIM attack consists of iteratively executing the FGSM attack for a number T of iterations, with a step size a = y:
[0081] [Math.3] xT = xTA + a ■ sigrA VI f (x L y \ \ ~ T-\J /
[0082] In this expression, xt represents an example obtained after T iterations; xn ~ x is the input data.
[0083] The PGD (Projected Gradient Descent) adversarial attack can be considered as a generalization of the BIM adversarial attack, for which the condition a = y is not required. Furthermore, the initialization of xo is based on a perturbation of x according to a uniform distribution in the interval [ - g, g] (uniform distribution U( - s, g)). Mathematically, this can be translated as:
[0084] [Math.4] x0 = x + uu~V( -e, g)
[0085] r / r / \ \\\ xT = chpj^! + a- sigr^JxT^( fx J y
[0086] In this expression, the clipxe function corresponds to a clipping of the adverse example values so that they are located within a neighborhood e of the original data x.
[0087] In an adversarial attack of type CW (named after the initials of its authors, Carlini and Wagner) the perturbation is not constrained by the ball Bp(x, e) according to the infinite norm l„, but it aims to be minimal for the norm G- The objective of this attack is to maximize the loss function by attacking the images with the optimal perturbation. This is an optimization problem defined by the following expression:
[0088] [Math.5] min^ || d || 9-c • L^f^x), y such that x + Ô^ [0,1]”
[0089] In this expression, || 5 || 9 represents the / 2 norm of the perturbation 5, etc is a compromise hyperparameter relative to the norm of the perturbation. To ensure that y + 5 g [0 1]”' we can introduce a parameter w and determine 5 as follows:
[0090] [Math.6] Ô = 4(tanh(H') +1) -x
[0091] There are mainly two solutions to strengthen neural networks against adversarial attacks. A first solution is based on the use of regularization methods to avoid noise expansion through the neural network. A second solution, the most used, is based on adversarial learning. Learning Adversarial training consists of generating adversarial examples from the training set data and training the model on the generated adversarial examples to improve the robustness of the model.
[0092] The objective of a classical learning method is to find, among the set 0 of all possible parameters 0 of the model considered, the optimal parameters 0* which minimize a total loss function corresponding to the average of the individual loss functions y for each of the N pairs (x, y) of the set training D. This can be translated by the following expression:
[0093] [Math.7] 0* = (ï) ) te®
[0094] Adversarial learning consists of learning on adverse examples generated from a perturbation 5. Adversarial learning is generally presented in the form of a robust min-max optimization problem of the type:
[0095] [Math. 8] = ar gmini +ô), j
[0096] The training of the model on the adversarial examples is generally handled using an optimization algorithm based on gradient descent on mini-batches. At each iteration of the optimization process, the parameters of the neural network are updated and it is necessary to calculate the adversarial perturbations with respect to these new parameters for each new iteration. This step therefore represents a significant additional computation time compared to conventional learning. In addition, and as indicated previously, adversarial learning has the disadvantage of protecting a model effectively only against the type of adversarial attack used to generate the adversarial examples exploited during adversarial learning. Experimental results:
[0097] To illustrate the performance of the image super-resolution method 100 according to the invention, different image super-resolution methods were tested and compared. These methods are based on the ESRGAN model described in document [3]. As previously indicated, the ESRGAN model is a generative adversarial neural network used for image super-resolution. The resolution of an image generated by the model is increased by a factor of four compared to the image provided as input to the model. The ESRGAN model is previously trained on a training data set comprising the 2650 2K resolution images from the Flickr2K image set and the associated low-resolution images generated by sub- bicubic sampling.
[0098] The different methods that were tested are each based on a different adjustment phase. The adjustment is based on a set of images from the DIV2K set (the DIV2K image set comprises 800 images of 2K resolution). For the method 100 according to the invention, the adjustment phase 220 previously described with reference to Figures 6 and 7 and the inference phase 300 previously described with reference to Figures 3 and 4 were used. This first example of implementation of the method according to the invention is named “CertSR” in the tables of Figures 8 to 10. The other methods tested simply correspond to an image prediction using respectively: - of the ES RG AN model optimized by an adjustment phase on the set of images from the DIV2K set (method named “ESRGAN” in the tables of figures 8 to 10), - the ESRGAN model optimized by an adverse adjustment phase of the PGD type (this is the model described in document [4], the associated method is called “AD-L-PGD” in the tables of figures 8 to 10), - the ESRGAN model optimized by an adverse adjustment phase of the FGSM type (method named “AD-L-FGSM” in the tables of figures 8 to 10), - of the ESRGAN model optimized by an adversarial adjustment phase of BIM type (method named "AD-L-BIM" in the tables of figures 8 to 10), - of the ESRGAN model optimized by an adversarial adjustment phase of CW type (method named "AD-L-CW" in the tables of figures 8 to 10).
[0099] For each of the adversarial adjustment phases FGSM, BIM and CW, adversarial examples are created from the images of the DIV2K set and adversarial training is carried out on the attacked images. It should be noted that, to the inventors' knowledge, before the work described in the present application, there were no image super-resolution models optimized against adversarial attacks of FGSM, BIM and CW type based on a loss function comprising both a loss at the pixel level and a loss at the level of perceptual similarity.
[0100] Only the adjustment phase 220 according to the invention for the CertSR model is based on certification by median random smoothing. The adverse adjustment phases of the AD-L-PGD, AD-L-FGSM, AD-L-BIM and AD-L-CW models correspond to adverse training phases as previously described (training on adverse examples generated from the adjustment set).
[0101] For the different adjustment phases mentioned above, the DIV2K images are cropped into sub-images of size 480 x 480; an Adam optimization is used with ~ 0.9 and ~ 0.999 ( / ? respectively is a hyperparameter defining the exponential decay rate for the first moment estimates, respectively the second moment) and an initial learning rate of 10 4 for the generator and the discriminator; for the adversarial adjustments, 18k iterations are performed with 16 images per batch; for the FGSM and BIM adversarial adjustments, the parameter g = 9 / 255 is used, and T — 2 iterations for BIM; for the CW adversarial adjustment, the parameters c — 10-2^^ — 4 iterations are used. For the adjustment phase 220 according to the invention (adjustment phase with median random smoothing) 59k iterations are performed with 5 images per batch (P = 2; M = 2; — 0.03; cr2 = 0.2).
[0102] The DIV2K validation dataset is used to compare the performance of the inference phase of the different tested methods (the DIV2K validation set includes 100 low-resolution images and their associated high-resolution images). Tests are conducted respectively: - directly on the images of the DIV2K validation set (“Raw images” in the table in [Fig.8]), - on the images of the DIV2K validation set to which Gaussian noise simulating sensor noise is added, namely independent Gaussian noise at pixel level with a zero mean and a standard deviation of 0.03 (“Noisy images” in the table in [Fig.8]), - on the images of the DIV2K validation set modified by smoothing with a Gaussian kernel of size 10 and standard deviation 0.3 (“Boutée Images” in the table of [Fig.8]), - on the images of the DIV2K validation set attacked with an adversarial attack of type FGSM (“FGSM Images” in the table of [Fig.9]), - on the images of the DIV2K validation set attacked with an adversarial BIM type attack (“BIM Images” in the table of [Fig.9]), - on the images of the DIV2K validation set attacked with an adversarial attack of the PGD type (“PGD Images” in the table of [Fig.9]), - on the images of the DIV2K validation set attacked with an adversarial attack of type CW (“CW Images” in the table of [Fig.9]).
[0103] Three metrics were used to compare the performance of the different methods tested. The first metric, PSNR (Peak Signal to Noise Ratio), is described in [6]. The second metric, SSIM (Structural Similarity Index Measure), is described in [7]. The third metric, LPIPS (Leamed Perceptual Image Patch Similarity), is described in [8]. The PSNR and SSIM metrics are widely used to evaluate restoration These metrics prioritize image fidelity over visual quality. A higher value of PSNR or SSIM indicates a higher fidelity of the predicted image relative to the ground truth (this is indicated by the upward arrow next to the terms "PSNR" or "SSIM" in the tables). In contrast, the LPIPS metric places more emphasis on assessing the similarity of visual features between images. A lower value of LPIPS indicates a higher perceptual similarity between the predicted image and the ground truth (this is indicated by the downward arrow next to the term "LPIPS" in the tables).
[0104] As seen previously, the optimal value of standard deviation 17 to use for the Gaussian noise during the inference phase 300 by median random smoothing depends on the type of perturbation affecting the input images. For the results presented in the tables of Figures 8 and 9, f = 0.05 for the blurred images, cr = 0.06 for the images attacked with FGSM and PGD, <7 = 0.07 for the images attacked with BIM, and <7 = 0.03 for the images attacked with CW (these values are those providing the best results for the LPIPS metric).
[0105] For the inference phase 300 by median random smoothing with the CertSR model, the number N of noisy images whose predictions are used for the generation 330 of the median image is chosen equal to twenty-one (TV - 21). It should be noted that, for raw images and noisy images, the inference phase 300 by median random smoothing is not essential because the adjustment phase 220 has made it possible to train the model on both raw images and noisy images.
[0106] The results presented in the tables of Figures 8 and 9 show that method 100 according to the invention (CertSR) has very good performance on all the validation sets considered. In particular, method 100 according to the invention takes first place in terms of LPIPS for all the validation sets, with the exception of the “FGSM images” validation set for which it takes second place. In terms of PSNR and SSIM, it takes first place for the “noisy images”, “BIM images”, “PGD images” and “CW images” validation sets, and it takes second place for the “raw images” and “FGSM images” validation sets. It thus appears that method 100 according to the invention is the most robust method, universally, against adversarial attacks (and it is important to remember that this method does not include adversarial training or adjustment by adversarial training).These results also show that a model specifically trained against a particular adversarial attack remains vulnerable against adversarial attacks of another type (and it can also remain relatively sensitive to the adversarial attack for which it was trained).
[0107] As illustrated in [Fig.10], the method 100 according to the invention (CertSR) was also compared with two of the best performing models on the validation sets NTIRE 2020 and AIM 2019, namely the “Impressionism” model described in paper [1] and “ESRGAN-FS” described in paper [2]. These two models were pre-trained on the Flickr2K image set and then fine-tuned on the NTIRE, AIM and DPED image sets respectively. The NTIRE and AIM image sets correspond to real-world images with different types of synthetic corruptions and sensor noise in low-resolution images.
[0108] The results presented in the table of Figure 10 show that the method 100 according to the invention presents the best performances in terms of LPIPS on the NTIRE and AIM validation sets (and it is important to remember that this method does not include specific training on these images). It also presents the best performances in terms of PSNR and SSIM on the NTIRE validation set. To obtain the results presented in Figure 10, a standard deviation value <7 = 0.03 was used for the inference phase 300 of the CertSR method
[0109] Thus, the method 100 according to the invention, based on an inference phase with median random smoothing, is the most robust method, universally, both against adverse attacks and against natural noises which have not necessarily been observed during training.
[0110] The first example of implementation of the method according to the invention (CertSR) presented above comprises both an adjustment phase 220 and an inference phase based on median random smoothing (MRS).
[0111] It should be noted, however, that simply refining the training of a neural network for image super-resolution using an adjustment phase based on random median smoothing (such as the adjustment phase 220 described with reference to Figures 6 and 7) may be sufficient to improve the performance of the neural network (without necessarily resorting to random median smoothing at the time of inference).
[0112] Also, simply implementing an inference phase based on median random smoothing (such as the inference phase 300 described with reference to Figures 3 and 4) may be sufficient to improve the performance of the neural network (without necessarily resorting to an adjustment phase by median random smoothing during the training of the neural network).
[0113] The table in [Fig.l 1] illustrates this by comparing the performance of CertSR with two other exemplary implementations of the invention based on ESRGAN. In this table, “MRS-FT” corresponds to an exemplary implementation of the invention based on ESRGAN with a median random smoothing fitting phase identical to that used for CertSR, but without using median random smoothing at the time of inference. “MRS-INF” corresponds to an exemplary implementation of the invention based on ESRGAN with an identical median random smoothing inference phase. to that used for CertSR, but without having previously carried out a median random smoothing adjustment phase.
[0114] It appears in the table of [Fig.l 1] that the MRS-FT and MRS-INF methods both improve the performance of the ES RG AN model in terms of LPIPS, both on the AIM validation set and on the NTIRE validation set. However, it appears (and significantly) that CertSR remains the most efficient method.
[0115] The results presented in the table of [Fig. 12] illustrate that the method according to the invention can be applied to different types of neural networks, and not only to GAN-type neural networks. In this table, EDSR corresponds to the neural network described in document [9]; NINASR corresponds to the neural network described in
[10] ; CertEDSR (respectively CertNINASR) corresponds to an example implementation of the invention based on EDSR (respectively NINASR) and comprising both an adjustment phase 220 and an inference phase 300 based on median random smoothing.
[0116] It appears in the results presented by the table in [Fig. 12] that the CertEDSR (respectively CertNINASR) method significantly improves the performance of the EDSR (respectively NINASR) model in terms of LPIPS, both on the AIM validation set and on the NTIRE validation set.
[0117] For the 220 refinement phase of CertEDSR and CertNINASR, the same parameters are used as for CertSR (P =2; M = 2; = 0.03; cr2 = 0.2). For the phase inference 300 of CertEDSR and CertNINASR, N = 11, <7 = 0.1 for AIM and cr = 0.005 for NTIRE.
[0118] The table in [Fig. 13] compares the performances of the CertSR, ESRGAN and AD-L-PGD methods with a method based on a regularization of ESRGAN (this method is named “ESRGAN-Reg”, it is based on a regularization method similar to that described by the document
[11] ).
[0119] It appears in the results presented by the table in [Fig. 13] that CertSR provides better performances than ESRGAN-Reg, particularly in terms of LPIPS.
[0120] The table in Figure 14 shows the impact, in terms of LPIPS, of the hyperparameters and used for the CertSR 220 tuning phase on the AIM and NTIRE validation sets. It appears that the values <7} = 0.03 and <72 = 0.2 give the best results for both AIM and NTIRE.
[0121] The table in Figure 15 shows the impact, in terms of PSNR, SSIM and LPIPS, of the hyperparameter a used for the inference phase 300 of CertSR on the validation sets AIM and NTIRE. It appears that the value <7 = 0.03 gives the best results in terms of LPIPS for both AIM and NTIRE.
[0122] The above description clearly illustrates that, through its various characteristics and their advantages, the present invention achieves the set objectives. In particular, the The proposed super-resolution method is universally robust to a wide variety of disturbances (noise or adversarial attacks).
[0123] The use of median random smoothing during the inference phase is particularly well suited to the case of image super-resolution because the pixel-level variations in the predicted images remain relatively small, and it is therefore possible to use a small number N of samples. Thus, the method according to the invention is particularly inexpensive in terms of computing power, unlike methods based on adjustment by adversarial training.
[0124] It should be noted that the implementations and embodiments considered above have been described as non-limiting examples, and that other variants are therefore conceivable. In particular, the proposed method of certification by median random smoothing can be applied to different types of neural networks for super-resolution (and not only to a GAN-type neural network).
[0125] It is also possible, during the training of the neural network, to carry out both a median random smoothing adjustment phase and an adverse adjustment phase (for example an adverse adjustment of the BIM, FGSM, PGD or CW type). In particular, the inventors have observed that CertSR tends to smooth the details a little too much, whereas on the contrary AD-L-BIM tends to exaggerate them too much. Combining a median random smoothing adjustment phase and an adverse phase of the BIM type can advantageously make it possible to achieve a satisfactory compromise on this aspect. List of cited documents
[0126] [1] “Real-world super-resolution via kernel estimation and noise injection”, Xiaozhong Ji et al., Proceedings of the IEEE / CVF conference on computer vision and pattern recognition workshops, 2020;
[0127] [2] “Frequency separation for real-world super-resolution”, Manuel Fritsche et al., Proceedings of the IEEE / CVF, International Conférence on Computer Vision Workshop (ICCV), 2019 ;
[0128] [3] « ESRGAN : Enhanced Super-Resolution Generative Adversarial Network », Xintao Wang et al., Proceedings of the European conférence on computer vision (ECCV) workshops, 2018 ;
[0129] [4] « Generalized real-world super-resolution through adversarial robustness », Angela Castillo et al., Proceedings of the IEEE / CVF, International Conférence on Computer Vision (ICCV) Workshops, pages 1855-1865, 2021 ;
[0130] [5] « Détection as régression : Certified object détection with médian smoothing », Ping-Yeh Chiang et al., Advances in Neural Information Processing Systems, pages 1275-1286, 2020 ;
[0131] [6] « Single image super-resolution : A benchmark », Chih-Yuan Yang et al., European Conférence on Computer Vision (ECCV), pages 372-386, 2014 ;
[0132] [7] « Image quality assessment : from error visibility to structural similarity », Wang Zhou et al., IEEE transactions on image processing, pages 600-612, 2004 ;
[0133] [8] « The unreasonable effectiveness of deep features as perceptual metric », Richard Zhang et al., Proceedings of the IEEE conférence on computer vision and pattern récognition, pages 586-595, 2018 ;
[0134] [9] « Enhanced deep residual networks for single image super-resolution », Bee Lim et al., Proceedings of the IEE conférence on computer vision and pattern récognition workshops, pages 136-144, 2017 ;
[0135]
[10] « torchSR : A pytorch-based framework for single image super-resolution », Gabriel Gouvine, https: / / github.com / Coloquinte / torchSR / blob / main / doc / NinaSR.md, 2023 ;
[0136]
[11] « Adversarially robust deep image super-resolution using entropy regula- rization », Jun-Ho Choi et al, Proceedings of the Asian Conférence on Computer Vision, 2020.
Claims
Claims
1. Method (100) for image super-resolution, said method (100) being implemented by an image processing device (10) storing a neural network (13) previously trained to produce a high-resolution image from a low-resolution image, said method (100) being characterized in that it comprises an inference phase (300) comprising: - a generation (310) of a plurality of noisy images from a low-resolution image, each noisy image being obtained by a random application to the low-resolution image of Gaussian noise having a predefined standard deviation, - a prediction (320), by the previously trained neural network, of an image for each of the noisy images thus obtained, - a generation (330) of a high-resolution image corresponding to a median of the predicted images thus obtained.
2. Method (100) according to claim 1 wherein the pre-training of the neural network (13) comprises an adjustment phase (220) based on a training data set comprising a set of pairs of images, each pair of images comprising a high resolution version and a low resolution version of the same image, said adjustment phase (220) comprising, for each low resolution image of the training data set, and for each of a plurality of different standard deviation values: - a generation (221) of a plurality of noisy images, each noisy image being obtained by a random application to the low-resolution image considered of a Gaussian noise having a standard deviation equal to the standard deviation value considered, - a prediction (222) by the neural network of a high-resolution image for each of the noisy images thus generated, - a generation (223) of a high-resolution image corresponding to a median of the predicted images thus obtained, - a calculation (224) of a loss function from the median thus generated and the high-resolution image associated with the image low resolution considered, - an update (225) of parameters of the neural network by a gradient descent based on the loss function.
3. The method (100) of claim 2 wherein the loss function comprises a component relating to a loss at the pixel level and a component relating to a loss at the level of an image perceptual similarity.
4. A method (100) according to any one of claims 1 to 3 wherein the value of the standard deviation of the Gaussian noise to be used for the generation (310) of the plurality of noisy images during the inference phase (300) is predefined based on a perceptual image similarity metric measured for different images obtained at the output of the inference phase (300) and for different candidate standard deviation values.
5. Method (100) according to any one of claims 1 to 4 wherein the value of the standard deviation of the Gaussian noise used for the generation (310) of the plurality of noisy images is between 0.005 and 0.
1.
6. Method (100) according to any one of claims 1 to 5 in which the number of noisy images whose predictions are used for the generation (330) of the median during the inference phase (300) is between 5 and 100, or even between 10 and 50, or even between 15 and
7. ZD. Method (100) according to any one of claims 1 to 6 in which two different standard deviation values are used during the adjustment phase (220).
8. The method (100) of claim 7 wherein the two different standard deviation values used during the adjustment phase (220) comprise a first standard deviation value between 0.01 and 0.06 and a second standard deviation value between 0.1 and 0.
6.
9. Method (100) according to any one of claims 1 to 8 wherein the number of noisy images whose predictions are used for the generation (223) of a median during the adjustment phase (220) is between 2 and 8.
10. A method (100) according to any one of claims 1 to 9 wherein the neural network (13) is a generative adversarial network.
11.
12.
13.
14. Method (100) according to claim 10 wherein the neural network (13) is based on the ESRGAN model. Method (100) according to any one of claims 1 to 11 wherein the prior training of the neural network (13) comprises an adversarial training phase of the BIM, FGSM, PGD or CW type. Method for training (200) a neural network (13) for image super-resolution, said training method (200) comprising an adjustment phase (220) based on a training data set comprising a set of image pairs, each image pair comprising a high-resolution version and a low-resolution version of the same image, said adjustment phase (220) comprising, for each low-resolution image of the training data set, and for each of a plurality of different standard deviation values: - a generation (221) of a plurality of noisy images, each noisy image being obtained by a random application to the low-resolution image considered of a Gaussian noise having a standard deviation equal to the standard deviation value considered, - a prediction (222) by the neural network of a high-resolution image for each of the noisy images thus generated, - a generation (223) of a high-resolution image corresponding to a median of the predicted images thus obtained, - a calculation (224) of a loss function from the median thus generated and from the high-resolution image associated with the low-resolution image considered, - an update (225) of parameters of the neural network by a gradient descent based on the loss function. Image processing device (10) for image super-resolution, the device (10) comprises a memory (12) storing a neural network (13) previously trained to produce a high-resolution image from a low-resolution image, the device (10) comprises a processor (11) configured to implement an inference phase (300), using the neural network (13), in order to generate a high-resolution image from a low-resolution image, the image processing device (10) is characterized in that the phase inference (300) includes: - a generation (310) of a plurality of noisy images from the low-resolution image, each noisy image being obtained by a random application to the low-resolution image of a Gaussian noise having a predefined standard deviation, - a prediction (320), by the previously trained neural network (13), of an image for each of the noisy images thus obtained, - a generation (330) of the high-resolution image corresponding to a median of the predicted images thus obtained.