Method and apparatus for enhancing images
A modular enhancement system estimates enhancement parameters for images using a small neural network, enhancing images before processing by downstream models to improve accuracy and robustness, addressing performance degradation in deep learning models due to data corruptions.
Patent Information
- Application Number
- GB2024009114
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-06-25
- Publication Date
- 2025-06-11
AI Technical Summary
Deep learning models in computer vision are sensitive to mild distribution shifts caused by natural alterations or corruptions in input data, leading to performance degradation in real-world applications, and existing techniques for improving robustness are inefficient, costly, or unsuitable for resource-constrained devices.
A modular enhancement system that estimates enhancement parameters using a small, efficient neural network to enhance images before processing by downstream models, allowing for real-time, task-agnostic and noise-agnostic image improvement without requiring retraining or adaptation.
The system consistently improves downstream model accuracy by 5-10% across various tasks and datasets, with minimal computational overhead, making it suitable for resource-constrained devices and real-time applications.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Field
[001] The present application generally relates to a method and apparatus for enhancing images. In particular, the present techniques provide a method for enhancing noisy or corrupted images so that the accuracy of downstream image processing tasks, which may be implemented using computer vision machine learning, ML, models, is improved. Background
[002] Deep learning models have been widely used in several multimedia systems and applications. Nonetheless, recent studies have explored their robustness to corrupted input data, showing critical performance degradation affecting the quality of experience of many downstream applications such as extended reality and video streaming. Indeed, the data-hungry nature of deep neural network (DNN) models, as well as the ever-growing complexity of network architectures, make the generated models sensitive to even mild distribution shifts causing a severe degradation in performance. These eventualities are often naturally met in many real-world applications, where data may unavoidably encounter natural alterations or corruptions, such as sensor degradation (e.g., shot noise, defocus blur1), compression and packet loss artifacts (e.g., block artifacts, corrupted colors), or framing issues (e.g., low contrast, rotations, images too bright), to mention a few.
[003] The applicant has therefore identified the need for a way to improve the output of computer vision models. Summary
[004] In a first approach of the present techniques, there is provided a computer-implemented method for enhancing an image to improve outcomes of a downstream image processing task, the method comprising: receiving an image to be enhanced prior to implementing the downstream image processing task; estimating, using an enhancement parameter estimation module of a modular enhancement system, a set of enhancement parameters for use in enhancing the image, wherein the enhancement parameters are parameters of transformation operators that, when applied to an image, improve image content understanding, and wherein the enhancement parameter estimation module comprises a trained machine learning, ML, model; and enhancing the image using an enhancement module of the modular enhancement system to generate an enhanced image, wherein the enhancing is based on the set of enhancement parameters.
[005] Advantageously, the present techniques provide a modular enhancement system which can be used to enhance images before they are processed by a downstream machine learning model for a downstream task, where the accuracy of the downstream model is impacted by the quality of the input images. The present techniques involve processing images to enhance the images prior to being processed by the downstream model. Advantageously, the modular enhancement system works with any downstream task without requiring re-training or adaptation to suit the downstream task. In other words, the modular enhancement system is agnostic of the downstream task.
[006] The enhancement may be denoising of or noise reduction in noisy images. The enhancement may be removing, fixing, or reducing the impact of corruptions in images.
[007] The image being enhanced may be an image, a sequence of images, a frame of a video or a sequence of frames of a video.
[008] The image may be a noisy image. A noisy image comprises noise. A noisy image may be referred to as a “noised” image or a “corrupted” image.
[009] The step of estimating enhancement parameters may be referred to herein as “predicting”.
[010] The enhanced image is one which has been enhanced. The enhanced image may still comprise some noise. The enhanced image may be referred to as a “cleaned” image.
[011] A module of the modular enhancement system is a component that may be implemented in hardware and / or software.
[012] The modular enhancement system may comprise two or more than two modules.
[013] Computer vision models suffer from performance degradation when presented with noisy or corrupted input images. Regardless of the task the computer vision models perform (such as image classification, object recognition / detection, gesture recognition, and so on), such models are strongly susceptible to corrupted input images. Such models may be used in smartphones, robots, smart vacuum cleaners and other user devices. Performance of such models and user devices may be restored by cleaning the input images before further processing.
[014] Some known techniques for cleaning the input data use large, inefficient, architectures that reconstruct images from a lower-dimensional representation. Such techniques may require corrupted-clean image pairs for training. It may not be possible to apply such techniques on-device. For example, large neural networks (e.g., autoencoders) may clean an image by reconstruction from a smaller-resolution feature space. Such a ‘black-box’ solution is costly and impractical for on-device applications. A black-box solution is not useable as a forensic tool.
[015] Some known techniques involve expensive retraining on noisy or corrupted data, which may not generalise to real-world scenarios. Such techniques may suffer from low performance and may require retraining for each new corruption and / or task. For example, expensive retraining may be performed using noisy or augmented data. This does not necessarily match deployment scenarios.
[016] In some known techniques, Test-Time Adaptation (TTA) to a corrupted environment involves on-device optimization. Such techniques may suffer from low performance and may require on-device training. In other words, corruption-specific TTA strategies may require on-device training.
[017] In the present techniques, enhancement parameters are estimated instead of the enhanced image itself. In other words, parameters of filters that are applied on an image can be predicted, rather than cleaning the image directly. Thus, an input image may be cleaned via a parametrized approach. Estimating the enhancement parameters in this manner may be much faster than estimating the enhanced image. Being fast and efficient, and using explicit transformations, allows for on-device and real-time use / applications. Furthermore, estimating the enhancement parameters means the overall system is a “white box” system that is easier to interpret and analyse, i.e. to perform forensic testing that looks at how the system works rather than just its functionality. White box testing can uncover errors or unexpected outcomes. This is advantageous as many existing techniques are “black box” systems, where only the functionality can be analysed.
[018] Additionally, cleaning may be task-agnostic and noise-agnostic. By being taskagnostic, examples described herein may be applied to many, or even any, computer vision models. By being noise-agnostic, examples described herein may handle a wide variety of input noise types.
[019] Examples described herein may provide small-footprint, real-time, on-device image enhancement.
[020] The enhancement parameter estimation module may use a small, computationally efficient network to estimate parameters.
[021] The enhancement module may be a differentiable enhancement module.
[022] Without loss of generality, modular enhancement may therefore be provided via deep parametric estimation and enhancement. Denoising is a particular type of image enhancement that may be applied to an image.
[023] In an implementation experiment, a modular enhancement system in accordance with examples described herein was trained only once and was tested on multiple tasks and datasets. The modular enhancement system brought consistent relative improvement of 5-10% in all scenarios without fine-tuning.
[024] Estimating the set of enhancement parameters may comprise estimating the set of enhancement parameters by inputting a low-resolution version of the image into the ML model.
[025] For reduced computational footprint, enhancement is split from the parameter estimation. This way, resolution (for example, of an input image) can change between estimation and enhancement. For example, parameters may be estimated on a low-resolution image for faster processing, while enhancement is performed on a full-resolution image.
[026] The method may comprise downsizing the image to generate the low-resolution version of the image. Such downsizing may be performed by the modular enhancement system, for example in a downsizing layer. Alternatively, such downsizing may be performed outside the modular enhancement system.
[027] The method may comprise discarding the low-resolution version of the image after using the low-resolution version of the image to estimate the set of enhancement parameters. Storage space may be freed up once the low-resolution version of the image has been used for its intended purpose.
[028] In one example, the image being enhanced may be a single image. In this case, estimating the set of enhancement parameters may comprise estimating the set of enhancement parameters using the single image. The set of enhancement parameters may therefore be optimised to the image to which the enhancement is to be applied.
[029] In another example, the image being enhanced may be a video comprising a sequence of frames. In this case, estimating the set of enhancement parameters may comprise: selecting a frame from the sequence of frames; estimating the set of enhancement parameters using the selected frame; and applying the estimated set of enhancement parameters for the selected frame to other frames in the sequence of frames. Thus, when the images are time-correlated, as in a sequence of consecutive frames of a video, for the sake of computational efficiency, it may be advantageous to use one frame to obtain the enhancement parameters, and apply the enhancement parameters to other frames to generate the enhanced frames. For example, one in every ten frames could be used, to reduce the computation significantly. Enhancement parameters may, in effect, be reused across a sequence of image. This may reduce latency compared to enhancement parameters being estimated for each image in the sequence of images. Where the images in the sequence of images have similar characteristics, such as noise characteristics, denoising may be highly effective even where the enhancement parameters are not estimated individually for each image in the sequence of images.
[030] Estimating the set of enhancement parameters using the enhancement parameter estimation module may comprise using a convolutional neural network, CNN (such as a residual neural network, ResNet) and / or using a transformer.
[031] The enhancement parameter estimation module may be implemented as a very small residual CNN for minimal computational footprint. An experimental implementation involved 2 giga floating point operations (gigaFLOPs, GFLOPs) to process a full HD input. That is, 2GFLOPs is the number of operations required to run a single instance of the ML model. This is in the region of 10-1,000 times lower than certain other approaches.
[032] The method may comprise inputting the enhanced image into a downstream machine learning, ML, model for implementing the downstream image processing task. Thus, the efficient enhancement architecture described herein may be provided before the input of an end-task / downstream model, for example in a computer vision architecture. In some examples, the enhanced data item may be provided to any end-task model.
[033] Outputting the enhanced image for input into a downstream ML model may comprise outputting the enhanced image to a downstream ML model that is different from a pre-trained downstream ML model used to train the ML model of the enhancement parameter estimation module. (The training process is described in more detail below).
[034] Thus, training may be performed using a different downstream model from a downstream model used during inference. The actual inference-time downstream model may not need to be known when training is performed. Retraining may not be needed even if a downstream model changes.
[035] Generally speaking, estimating a set of enhancement parameters comprises estimating a set of enhancement parameters for any one or more of: a geometric transformation operator (e.g. rotation, translation); a non-geometric transformation operator (e.g. scaling, affine transformation, perspective transformation); a linear transformation operator; a non-linear transformation operator; and a Fourier transformation operator. It will be understood this is a non-limiting list of example operators.
[036] Estimating a set of enhancement parameters may comprise estimating parameters for an additive operator. In cases where the additive operator is a colour shift operator, enhancing the image using the enhancement module may comprise: enhancing the image by performing a colour shift operation on the image.
[037] Estimating a set of enhancement parameters may comprise estimating parameters for a multiplicative operator. In cases where the multiplicative operator is a colour warp operator, enhancing the image using the enhancement module may comprise: enhancing the image by performing a colour warp operation on the image.
[038] Estimating a set of enhancement parameters may comprise estimating parameters for a convolution operator. In cases where the convolution operator is a de-blurring operator, enhancing the image using the enhancement module may comprise: enhancing the image by performing a de-blurring operation on the image.
[039] Estimating a set of enhancement parameters may comprise estimating parameters for a denoising operator. In such cases, enhancing the image using the enhancement module may comprise: enhancing the image by performing a denoising operation on the image. Thus, examples described herein may be particularly effective with an input image that is heavily corrupted and / or particularly noisy. However, examples described herein may also be effective with images having any level of noise and / or corruption. Examples herein may also be compatible with images with minimal or no noise and / or corruption.
[040] Corruptions in the colour space of an image (e.g., brightness, contrast, blur, etc.) may therefore be managed. A possible parametrization supported by the modular enhancement system described herein is linear transformation of a colour space and spatial filtering via convolution. It will be understood that other types of parameterisation are also possible, including parameterisation based on predicting the coefficient of a Fourier transform and another including parameterisation based on non-linear transformation of the input image.
[041] In a second approach of the present techniques, there is provided an apparatus for enhancing an image to improve outcomes of a downstream image processing task, the apparatus comprising: an interface for receiving an image to be enhanced prior to implementing the downstream image processing task; a modular enhancement system; and at least one processor coupled to memory, for implementing the modular enhancement system by: estimating, using an enhancement parameter estimation module of the modular enhancement system, a set of enhancement parameters for use in enhancing the image, wherein the enhancement parameters are parameters of transformation operators that, when applied to an image, improve image content understanding, and wherein the enhancement parameter estimation module comprises a machine learning, ML, model; and enhancing the image using an enhancement module of the modular enhancement system to generate an enhanced image, wherein the enhancing is based on the set of enhancement parameters.
[042] The enhancement module may be a System-on-Chip, SoC, enhancement module.
[043] Such an enhancement module may use known transforms and may be implemented as an SoC. Implementing a (frozen) enhancement module in hardware may achieve faster processing, even when a parameter estimation strategy changes.
[044] The apparatus may be a home appliance. Examples of such home appliances include, but are not limited to, refrigerators, vacuum cleaners, ovens, and robotic lawnmowers.
[045] Taking the example of a refrigerator, motion blur may be reduced in an image captured when an object is inserted into the refrigerator. This makes an object recognition task easier. Storage duration may be tracked automatically with higher accuracy. Additionally, opening a refrigerator door changes illumination conditions. Examples described herein can clean images while a camera adapts.
[046] Taking the example of an oven, such as an artificial intelligence (Al) oven, steam coming off food can obscure a camera view. Examples described herein can reconstruct a scene more accurately to identify burns more easily and reliable. Additionally, a camera in an Al oven might only ordinarily work effectively when an oven light is on. In accordance with examples described herein, images captured by such a camera may be useable in darker conditions, i.e. even if the oven light is not on.
[047] Taking the example of a vacuum cleaner, such as an Al vacuum cleaner, motion blur may be reduced on fast-moving actors in a scene, such that obstacles may be avoided more effectively. Examples of such actors include, but are not limited to, people and pets. Additionally, dark conditions can ordinarily make navigation more difficult. However, examples described herein enable a scene in a captured image to be brightened.
[048] The apparatus may be an autonomous vehicle.
[049] The apparatus may be a mobile computing device. Examples of such mobile computing devices include, but are not limited to, smartphones, tablet computing devices and laptop computing devices.
[050] The features described above with respect to the first approach apply equally to the second approach and therefore, for the sake of conciseness, are not repeated.
[051] The apparatus may be a constrained-resource device, but which has the minimum hardware capabilities to use a trained neural network / ML model. The apparatus may be any one of: a smartphone, tablet, laptop, computer or computing device, virtual assistant device, a vehicle, an autonomous vehicle, a robot or robotic device, a robotic assistant, image capture system or device, an augmented reality system or device, a virtual reality system or device, a gaming system, an Internet of Things device, or a smart consumer device (such as a smart fridge or smart vacuum cleaner). It will be understood that this is a non-exhaustive and nonlimiting list of example apparatuses.
[052] In a third approach of the present techniques, there is provided a method for training a machine learning, ML, model of a modular enhancement system for enhancing an image for a downstream image processing task, the method comprising: obtaining a training dataset comprising a plurality of corrupted images, wherein each corrupted image comprises at least one image corruption, and each corrupted image comprises a ground truth label; training the ML model of an enhancement parameter estimation module of the modular enhancement system using each image of the training dataset by: estimating, using the ML model and the corrupted image, a set of enhancement parameters for use in enhancing the corrupted image; enhancing, using an enhancement module of the modular enhancement system, the corrupted image using the set of enhancement parameters; generating, using a pre-trained downstream ML model trained to perform a specific downstream task, and the enhanced image, a predicted label for the enhanced image; calculating a loss based on a difference between the predicted label for the enhanced image generated by the downstream ML model and the ground truth label for the corrupted image; and updating parameters of the ML model of the enhancement parameter estimation module based on the calculated loss.
[053] Such training may not require paired images, i.e. noisy and denoised image pairs. Instead, such training may be performed in a label-abundant task and even when large numbers of noisy-denoised image pairs are not available.
[054] The training method may comprise: enhancing the image using a smoothed version of the modular enhancement system to generate an intermediate image; enhancing the intermediate image using the modular enhancement system to generate a enhanced intermediate image; generating, using the pre-trained downstream ML model and the enhanced intermediate image, a predicted label for the enhanced intermediate image; and calculating a further loss based on: the predicted label for the enhanced intermediate image; and the ground truth label for the corrupted image, wherein updating parameters of the ML model of the enhancement parameter estimation module based on the calculated loss comprises: updating parameters of the ML model of the enhancement parameter estimation module based on the calculated loss and the calculated further loss.
[055] In this manner, an Exponential Moving Average (EMA) model may be used. This can improve stability of training. The EMA model may be used for training only and not also inference. Use of the EMA model does not impact inference speed in such situations.
[056] Using the pre-trained downstream ML model to generate a predicted label may comprise using any of the following downstream ML models: a computer vision model; a classification model; a segmentation model; an object detection model; a panoptic segmentation model; an instance segmentation model; an action recognition model; an anomaly detection model; an image captioning model; and a salient object detection image denoising model. Such models may already be label-abundant.
[057] In a related approach of the present techniques, there is provided a computer-readable storage medium comprising instructions which, when executed by a processor, causes the processor to carry out any of the methods described herein.
[058] As will be appreciated by one skilled in the art, the present techniques may be embodied as a system, method or computer program product. Accordingly, present techniques may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects.
[059] Furthermore, the present techniques may take the form of a computer program product embodied in a computer readable medium having computer readable program code embodied thereon. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable medium may be, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.
[060] Computer program code for carrying out operations of the present techniques may be written in any combination of one or more programming languages, including object oriented programming languages and conventional procedural programming languages. Code components may be embodied as procedures, methods or the like, and may comprise subcomponents which may take the form of instructions or sequences of instructions at any of the levels of abstraction, from the direct machine instructions of a native instruction set to high-level compiled or interpreted language constructs.
[061] Embodiments of the present techniques also provide a non-transitory data carrier carrying code which, when implemented on a processor, causes the processor to carry out any of the methods described herein.
[062] The techniques further provide processor control code to implement the abovedescribed methods, for example on a general purpose computer system or on a digital signal processor (DSP). The techniques also provide a carrier carrying processor control code to, when running, implement any of the above methods, in particular on a non-transitory data carrier. The code may be provided on a carrier such as a disk, a microprocessor, CD- or DVD-ROM, programmed memory such as non-volatile memory (e.g. Flash) or read-only memory (firmware), or on a data carrier such as an optical or electrical signal carrier. Code (and / or data) to implement embodiments of the techniques described herein may comprise source, object or executable code in a conventional programming language (interpreted or compiled) such as Python, C, or assembly code, code for setting up or controlling an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array), or code for a hardware description language such as Verilog (RTM) or VHDL (Very high speed integrated circuit Hardware Description Language). As the skilled person will appreciate, such code and / or data may be distributed between a plurality of coupled components in communication with one another. The techniques may comprise a controller which includes a microprocessor, working memory and program memory coupled to one or more of the components of the system.
[063] It will also be clear to one of skill in the art that all or part of a logical method according to embodiments of the present techniques may suitably be embodied in a logic apparatus comprising logic elements to perform the steps of the above-described methods, and that such logic elements may comprise components such as logic gates in, for example a programmable logic array or application-specific integrated circuit. Such a logic arrangement may further be embodied in enabling elements for temporarily or permanently establishing logic structures in such an array or circuit using, for example, a virtual hardware descriptor language, which may be stored and transmitted using fixed or transmittable carrier media.
[064] In an embodiment, the present techniques may be realised in the form of a data carrier having functional data thereon, said functional data comprising functional computer data structures to, when loaded into a computer system or network and operated upon thereby, enable said computer system to perform all the steps of the above-described method.
[065] The method described above may be wholly or partly performed on an apparatus, i.e. an electronic device, using a machine learning or artificial intelligence model. The model may be processed by an artificial intelligence-dedicated processor designed in a hardware structure specified for artificial intelligence model processing. The artificial intelligence model may be obtained by training. Here, "obtained by training" means that a predefined operation rule or artificial intelligence model configured to perform a desired feature (or purpose) is obtained by training a basic artificial intelligence model with multiple pieces of training data by a training algorithm. The artificial intelligence model may include a plurality of neural network layers. Each of the plurality of neural network layers includes a plurality of weight values and performs neural network computation by computation between a result of computation by a previous layer and the plurality of weight values.
[066] As mentioned above, the present techniques may be implemented using an Al model. A function associated with Al may be performed through the non-volatile memory, the volatile memory, and the processor. The processor may include one or a plurality of processors. At this time, one or a plurality of processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an Al-dedicated processor such as a neural processing unit (NPU). The one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (Al) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning. Here, being provided through learning means that, by applying a learning algorithm to a plurality of learning data, a predefined operating rule or Al model of a desired characteristic is made. The learning may be performed in a device itself in which Al according to an embodiment is performed, and / o may be implemented through a separate server / system.
[067] The Al model may consist of a plurality of neural network layers. Each layer has a plurality of weight values, and performs a layer operation through calculation of a previous layer and an operation of a plurality of weights. Examples of neural networks include, but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann Machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial networks (GAN), and deep Q-networks.
[068] The learning algorithm is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Brief description of the drawings
[069] Implementations of the present techniques will now be described, by way of example only, with reference to the accompanying drawings, in which:
[070] Figure 1A is a schematic diagram illustrating how corrupted input data impacts the output of a vision or vision-language model;
[071] Figure 1B is a schematic diagram showing how retraining has to be performed for each new corruption or task;
[072] Figure 1C is a schematic diagram showing how test-time adaptation techniques can be performed on-device, but this may be unsuitable for resource-constrained devices.
[073] Figure 1D is a schematic diagram shows how a model can be used to clean a noisy image or un-corrupt a corrupted image before it is processed by a vision or vision-language model, but cannot be used on-device;
[074] Figure 2 is a schematic diagram illustrating the present modular system for efficient image enhancement;
[075] Figure 3 is a schematic diagram showing how the two modules of the present modular system work together to enhance the content of an image;
[076] Figure 4 is a schematic diagram illustrating the technique to train the present model;
[077] Figure 5 shows pseudo-code for training the present modular system;
[078] Figure 6 is a table showing a breakdown of an enhancement parameter estimation mdule of the present modular system;
[079] Figure 7 shows the results for the image classification task on the ImageNetC dataset;
[080] Figure 8 is a table showing accuracy on ImageNetC-mixed via ResNet50;
[081] Figure 9 is a table of results for VizWiz with ResNet50 V2;
[082] Figure 10 is a table showing computational complexity of the present method compared to other input-level image enhancement strategies;
[083] Figure 11 is a table showing the quantitative results for semantic segmentation using a DeepLabV2 architecture with the ResNet50 backbone;
[084] Figure 12 is a table showing quantitative results on the ImageNetC-Bar dataset with ResNet50;
[085] Figure 13, which is a table of results for feature extraction from compressed video data;
[086] Figure 14 is a block diagram of the overall architecture of the present techniques;
[087] Figure 15A is a schematic diagram showing one advantage of the present techniques over existing large denoising networks;
[088] Figure 15B is a schematic diagram showing another advantage of the present techniques over existing networks that require re-training from scratch for each corruption;
[089] Figure 15C is a schematic diagram showing another advantage of the present techniques over existing networks that require per-corruption adaptation;
[090] Figure 16 is a schematic diagram showing how the present techniques can be used to handle any number and combination of corruptions and / or noise types;
[091] Figure 17 is a flowchart of example steps for image enhancing an image to improve outcomes of a downstream image processing task;
[092] Figure 18 is a flowchart of example steps for training a machine learning, ML, model of a modular enhancement system for enhancing an image for a downstream image processing task; and
[093] Figure 19 is a block diagram of an apparatus for enhancing an image to improve outcomes of a downstream image processing task. Detailed description of the drawings
[094] Broadly speaking, embodiments of the present techniques provide a method for enhancing noisy or corrupted images so that the accuracy of downstream image processing tasks, which may be implemented using computer vision machine learning, ML, models, is improved. The present techniques involve processing images to enhance the images prior to being processed by the downstream model. Advantageously, the modular enhancement system works with any downstream task without requiring re-training or adaptation to suit the downstream task. In other words, the modular enhancement system is agnostic of the downstream task.
[095] The robustness of computer vision models against distribution shifts is of vital importance in multimedia applications. Starting with the discovery of adversarial examples, many different robustness venues have been explored by the community, such as common image corruptions, images with conflicting shapes / textures, and style variations. The core idea behind robust generalization is to make multimedia models increasingly invariant to such shifts; e.g. a model trained on a training set should generalize well to samples with unseen styles, corruptions and perturbations. Arguably, the most common effects found on real data are image corruptions, which have been standardized and categorized in special benchmark datasets for the evaluation of deep learning architectures. This is the scenario on which the present techniques focus, due to the practical importance.
[096] Figure 1A is a schematic diagram illustrating how corrupted input data impacts the output of a computer vision model (such as a vision or vision-language model). Figure 1A is reproduced from Chen et al., ’’Benchmarking Robustness of Adaptation Methods on Pretrained Vision-Language Models”, published at the NeurlPS 2023 Datasets and Benchmark Track. Computer vision models suffer from performance degradation when presented with noisy or corrupted input data. Two example tasks performed by computer vision models are shown in Figure 1A: visual captioning, i.e. generating a text caption describing an input image, and Q&A, i.e. generating a question and answer pair for an input image. A computer vision model which is presented with an uncorrupted / non-noisy input image (the “original image” on the left), generally performs well at the task it has been trained to perform. This can be seen by the text caption and Q&A pair generated for the original image. However, when the same computer vision model is presented with corrupted / noisy input images, the model does not perform as well. In Figure 1A, the original image has been corrupted by applying noise or blurring techniques. The outputs of the model for the two tasks are clearly worse compared to the outputs for the original image, because the model has been unable to correctly ‘understand’ the image or extract features from the image, which leads to the errors in the outputs.
[097] Considering the increasing adoption of deep models, the issue of noisy or corrupted data has become of paramount relevance. Therefore, a new research field emerged to attempt to make models more robust under different perspectives. Existing approaches to improve model robustness can be categorized into three main branches: (i) data augmentation approaches, (ii) test-time adaptation approaches, and (iii) enhancing and denoising approaches, such as autoencoder-based, generative, and adversarial denoisers (e.g., GANs, diffusion models, etc.).
[098] Data augmentation methods are built on a simple premise; one can achieve robustness by retraining or fine-tuning a model on a training set updated with the samples that the original model failed on (e.g., unseen distributions). Adversarial training has been one of the most prominent examples of such methods, where adversarially corrupted examples are included in the training set during retraining of the model. Since then, the same core idea has been used by a plethora of methods that addressed robustness from a data augmentation perspective.
[099] Data augmentation techniques either re-train Al models from scratch or fine-tune them by applying a large set of general-purpose augmentations (often synthetic) that mimic common corruptions encountered in the real world on data. In other words, this paradigm aims to produce a model that is robust against corrupted images and that can be very effective in improving model abilities to generalize under data distribution shift. The main drawbacks of this approach are that the trained model can only perform the task that has been trained with supervision on; the trained model tends to perform well on images with training-like distortions; the types of augmentations should be decided a priori, and they cannot be easily tuned to a scenario with variable corruptions (unless via expensive retraining). Figure 1B is a schematic diagram showing how when retraining is performed using target test data A, the retrained model performs well on input images A that have the same characteristics as the test data A. However, the performance of the retrained model is poor when input images B with different characteristics are presented. Thus, the model needs retraining for each different type of corruption and / or task, which is not efficient.
[100] Although they currently hold the state-of-the-art on many benchmarks, data augmentation methods have a key disadvantage, where they require retraining or finetuning every time they encounter (and fail against) new distributions, which may or may not be unknown in practical scenarios. Assuming one can perform these expensive updates periodically, model capacity issues as well as catastrophic forgetting are likely to be new accuracy bottlenecks, which will inevitably lead to practical issues in deployment scenarios. In order to avoid such problems, another branch of methods focuses on analyzing input images to detect if they are from unseen distribution. Going further than just detecting critical samples, more recent methods aim to improve robustness in test-time. To improve model robustness to variable-type corruptions at inference time compared to training time, Test-Time Adaptation (TTA) methods have been proposed. TTA methods focus on resolving data distribution shifts directly at test-time via dynamic updates of pre-trained models based on the specific characteristics of the target test data. TTA methods enable a model to improve its performance when encountering variations, unseen examples, or changing conditions at test time. Unlike data augmentation techniques, TTA methods do not involve expensive training procedures, but they are still bounded by the same task to be used during the pre-training and testing / adaptation phases. Figure 1C is a schematic diagram showing how TTA techniques can be performed on-device, but this may be undesirable or unsuitable for resource-constrained devices.
[101] Test-time adaptation methods largely alleviate the issues inherent to data augmentation approaches; they perform partial updates at most, therefore they are fast and cheap. However, they still require a form of training, which requires periodic updates in deployment scenarios where data shift is ever-present.
[102] Another branch of methods that aims to achieve robust generalization can be roughly categorized as preprocessing methods. These approaches do not tackle the problem as a normalization issue, unlike test-time adaptation methods. Instead, they attempt to bridge the distribution gap between the query sample and the model training set, by trying to enhance and denoise the query image. In other words, enhancement and denoising are the processes of improving input samples before they are fed to the downstream network. The most promising techniques are nowadays based on deep learning models. For instance, autoencoders are used to encode compact representations of input samples and then decode them to minimize a reconstruction loss with respect to ground truth samples. Generative and / or adversarial models (e.g., diffusion models or GANs) modify input samples by iterative processes or style transfer techniques. Figure 1D is a schematic diagram showing enhancement / denoising, i.e. using a model that can clean a noisy image or un-corrupt a corrupted image. The model tries to create or construct a better version of a noisy / corrupted input image. The main drawbacks of these methods are the need for expensive inference time and for paired clean-corrupted samples at training time while being able to handle indomain corruptions only (i.e., particular corruption types seen at training time). Furthermore, these methods generally require a very large model, which is unsuitable for use on-device, i.e. on consumer electronic devices such as smartphones, smart vacuum cleaners, and so on. Instead, the models need to run on servers, because of the computing power and memory required to implement them.
[103] Overall, current techniques all come with some key limitations that are briefly discussed. Thus, the present Applicant proposes a System for Modular Parametric Image Enhancement (SyMPIE), to overcome some of these limitations. SyMPIE implements a hybrid strategy combining the best of data augmentation and enhancement approaches while being computationally efficient and fully compatible with any model and any downstream task.
[104] The present techniques tackle the above-described issue from a completely different direction, and provide a modular system that predicts parameters used by an ad-hoc module to enhance the received input. Figure 2 is a schematic diagram illustrating the modular system for efficient image enhancement targeting increased model robustness to corruptions in different multimedia tasks. SyMPIE comprises two modules - a noise estimation module (NEM) and a differential warper module (DWM). These are described in more detail below with reference to Figure 3. SyMPIE estimates parameters to clean input samples and can be integrated into any deep network architectures for multimedia understanding. SyMPIE is fully differentiable and can be trained end-to-end on an upstream task, without the need for paired clean-corrupted images to enhance images automatically. Herein, the term “upstream task” is used to mean the task used to pre-train a system (e.g., classification) and the term “downstream task” is used to mean a task performed using a trained system at inference / deployment time.
[105] The details of the present techniques are described below, but briefly, the present techniques utilize generic data augmentation pipelines and common upstream task / network (e.g., image classification with ResNet50) to pre-train the model (System for Modular Parametric Image Enhancement, SyMPIE). A key insight is that a large portion of input corruptions found in the real world can be modeled by applying either global operations on color channels (for example, a night scene can be approximated by a darkened daytime scene), or spatial noise which can be filtered out by a fast convolution operation with small kernel size (for example, a Laplacian filter or a blurring operation can be used to sharpen an image or to reduce noise on the image). Additionally, the SyMPIE model is pre-trained minimizing a standard classification task loss, alleviating the need for paired clean-corrupted samples during pre-training. In other words, the downstream task and model do not necessarily need to match the upstream task and model used for pre-training the SyMPIE model. For example, the model may be trained on a corrupted ImageNet benchmark for image classification using a Convolutional Neural Network (CNN), and may be deployed on semantic segmentation benchmarks with adverse conditions to support a transformer architecture.
[106] Therefore, SyMPIE is similar to enhancement and denoising approaches, with the key advantage of containing much smaller modules that are pre-trained without paired clean and corrupted ground truth.
[107] Some of the advantages of the present techniques include: • A lightweight modular image enhancement system, named SyMPIE, which predicts parameters of ad-hoc operators that are applied to input samples to improve the content understanding. • In the experimental analyses, SyMPIE consistently improves the accuracy of downstream models for the corresponding task with minimal impact on the number of compute operations (about 2GFLOPs on top of 343GFLOPs of a classical ResNet50 architecture for a full HD resolution input) and can be seamlessly applied on top of any competing approach improving their accuracy (up to 8% relative gain) without retraining. Unlike data augmentation methods, SyMPIE does not require long training procedures. • Unlike existing denoising approaches, SyMPIE is fast, does not require paired clean-corrupted samples during pre-training, and can be reused in multiple setups without losing its efficacy. In particular, SyMPIE proves to be fully compatible with any convolutional and transformer-based architecture, including recent foundation models (e.g., CLIP). SyMPIE enhances the accuracy of the aforementioned models on several downstream tasks (e.g., image classification, semantic segmentation) in the presence of corrupted test data improving the quality of experience of the final end users.
[108] System for Modular Parametric Image Enhancement (SyMPIE) is made of two modules, as shown in Figure 3. Figure 3 is a schematic diagram showing how the two modules work together to enhance the content of an image. The first module is the Noise Estimation Module (NEM), E, which is implemented by a small CNN to estimate a set of parameters used by the second module from training samples. The NEM receives a corrupted input and predicts a triple of parameters (K, CM,CS). The second module is the Differentiable Warping Module (DWM), D, and is used to process input samples in order to remove distortions from them. That is, the triple of parameters is used by the DWM to enhance the image using parametric operators. To obtain the final SyMPIE architecture, the two modules are combined in a single block by U = D o E where ° denotes the module composition.
[109] As shown in Figure 3, SyMPIE can be inserted between the input data and any generic downstream model. Three key features distinguish the present approach from existing image enhancement approaches, namely: (i) the removal of paired clean-corrupted images as a requirement for training; (ii) the prediction of enhancing parameters rather than of the cleaned images directly; (iii) the ability to generalize across several upstream and downstream tasks and networks.
[110] Existing works, indeed, train enhancer models as a regression task on the clean images, effectively optimizing the reconstruction PSNR (Peak Signal-to-Noise Ratio). The present techniques tackle the issue from a completely different point of view, focusing on efficiency and downstream accuracy, instead. Incidentally, this allows the present model to be trained and fine-tuned on any available dataset, especially on real-world datasets that contain naturally corrupted images with no clean counterpart ground truth.
[111] Noise Estimation Module (NEM). To estimate the parameters needed for enhancement, a small residual CNN, E, was designed which estimates explicitly the set of parameters used by the differential warping module D. Motivated by the observation that most of the distortions of the input samples affect the color space only, the present Applicant has identified a set of filters and linear transformations to be applied to the input samples for their improvement.
[112] In the present setup, the aim is to estimate the parameters of i) a filter kernel K e RKx7f (where kernel size K used in the experiments is given below; ii) a linear transformation of color channels, represented by the matrix CM 6 Rdxd; and iii) a global color shift Cs e Rd where d is the number of input channels (e.g., d = 3 for RGB images). This set of parameters allows the present architecture to model several naturally occurring corruptions, such as under / over-exposed images, sensor noise (e.g., Gaussian, impulse, shot noise), unbalanced white point, etc. without the need to employ expensive deep learning models for enhancement.
[113] To estimate the above parameters needed for enhancement and to keep the additional footprint minimal, the NEM is implemented as a small residual CNN (see the top half of Figure 3). This allows the NEM module to efficiently extract global information from an input image, and predict suitable parameters for the subsequent module. More in detail, CM, Cs) e R4, where X is a corrupted input, X denotes the space of input images, and {K, CM, Cs) are the parameters used for the cleaning process.
[114] Moreover, in the vanilla deep learning-based image enhancement formulation, the network weights used to process the images are fixed. Therefore, the models learn to partition their own parameter spaces to deal with different situations, reducing the overall efficacy and parameter efficiency. The present SyMPIE model, instead, explicitly predicts the parameters depending on the input, and therefore, the present model can seamlessly adapt to multiple types of corruptions without the need to partition its parameter space.
[115] Differentiable Warping Module (DWM). To enhance input images using predictions obtained from the NEM, the Differentiable Warping Module (DWM), D, depicted in the bottom half of Figure 3 is used. The module receives as input a corrupted image and the parameters estimated by the NEM and applies the parametric operators to the former to enhance the image. Formally, 1): (X, (K, CM, Cs)) « X, where X e X is the improved image.
[116] The operations of D are completely differentiable, allowing backpropagation of gradients through the DWM, and optimizing the predictions of the NEM in an unsupervised manner. This enables the present model to be trained without any paired data, since gradients propagated through the DWM come directly from predictions over a frozen upstream model. In other words, the NEM is trained to enhance samples in order to maximize the performance of a frozen upstream network. This is a great advantage compared to existing works which either cannot enhance the inputs to improve the performance of downstream models or cannot effectively leverage the gradient flow to produce cleaner inputs. Furthermore, when modifying input samples directly, several approaches introduce distortions that appear random and unnoticeable to the human eye but can completely change the output prediction of the model. In the present setup, this is avoided by directly estimating the parameters of transformations whose kernels are translation-invariant (that is, given a translation T, the parameters are generated by a multidimensional function n(x) = n(T(x)) for all pixel locations x). In particular, the present modules apply transformations to the whole image at the same time rather than multiple transformations on local subsets of pixels. Therefore, the present modules enforce a stronger and visible change in the input image to allow a change in the downstream prediction. In general, the present approach is optimized to enhance the content of an input and improve the performance of the upstream task, not just the perceived image quality, unlike the other approaches. The use of global operations also stabilizes the gradients received by the NEM, enabling its training with fewer samples than other approaches (e.g., when compared to augmentation strategies, the present approach requires up to 60 times fewer training samples).
[117] Training Process. A key benefit of the present module is its end-to-end training on any given upstream task (it is trained on image classification tasks in the experiments) for its deployment on any downstream task (e.g., image classification, semantic segmentation, etc.) without further need of fine-tuning. In the following, the training procedure used to optimize the model (shown in Figure 4), and the data augmentation strategy used during the tuning, are described. A detailed description is provided as pseudo-code in Figure 5.
[118] Figure 4 is a schematic diagram illustrating the technique to train the present model. A frozen model M is pre-trained on a source upstream task (e.g., classification) to provide class predictions from the cleaned input samples (P = M(X)), and these are then compared to the labels Y e y using the cross-entropy loss £ce, to obtain = £ce (W
[119] As will be explained below, a common failing point of denoisers is modal collapse upon their iterative application to the same image. To avoid this, an additional regularization term is introduced to the training loop, which employs the exponentially smoothed version of the present modules (called UEMAY The parameters of UEMA at each training iteration i >0 are computed by 9EMA i = p6EMA,i-i + (1 _ where p is the exponential smoothing rate and qema,o = 0o with St being the parameters of U at iteration i. During the optimization, UEMA is used to generate an intermediate image XEMA = UEMA(X), which will show different visual cues than X, aiding in the generalization.
[120] The intermediate sample is then fed to the modules currently being optimized ( / .e., U) to obtain XEMA = U(XEMA), which is used in the computation of the regularization loss term. The two-step sample is then processed as normal, obtaining a prediction from the classifier model PEMA = M(Xema) to compute a new loss term l2 = £ce(PEMA, Y), weighted by Aema.
[121] The present modules attempt to enhance input samples which could not otherwise be handled effectively by popular downstream networks. For this purpose, a set of augmentations is designed that serves as a proxy for the distortions experienced at deployment time, as described next.
[122] Data Augmentation Pipeline. To train the present modules, a variegate data augmentation pipeline is employed, serving two main purposes: (i) to encourage generalization of the architecture over a vast array of possible input corruptions, and (ii) to reflect conditions that are seen in the real world during deployment. In particular, the present architecture is trained on the ImageNet-lk dataset augmented using the corruptions proposed by Hendrycks and Dietterich, together with four additional ones to mimic adverse weather conditions at variable degree of severity.
[123] The four additional corruptions are the following: 1. The darken corruption mimics the effect of under-exposure of the scene by reducing the intensities of all pixels in a consistent fashion. For example, this situation may happen after encountering glare that forces cameras to reduce the exposure before re-adjusting. 2. The horizon corruption shifts the white-point of the input image to mimic haze on sunset / sunrise scenes. The new white points (RGB) are sampled uniformly in the range [255,192 + 8,192 - 5], 6 ~ 11(-32,32]. 3. The night corruption simulates nighttime acquisition. This problem is tackled in two ways: (i) changing the white point to a dark blue (uniformly sampled in the color range [32,32,64 + 5], 8 -11(-24,24]) and (ii) darkening the brightest pixels of the scene (which tend to be found in the sky region of an outdoor image). 4. The white-point corruption mimics different behaviors of cameras’ white-point balancing procedure, which could lead to images with unrealistic colors even if their relative chromaticity is consistent. In this case, a random white point is selected using a uniform distribution 11(223,287] (i.e., a variation of an eighth) for each of the 3 RGB components and then re-balance the image pixel values accordingly.
[124] Inference Process. The goal of the present system is to be as easy to use and modular as possible. Therefore, the present Applicant has aimed to keep the inference process footprint minimal. Figure 2 provides a comparison between standard inference practice on a downstream task (top half), and the modified pipeline using the present modules (bottom half). Notice how the present modules can be seamlessly embedded in any architecture to improve its final accuracy.
[125] In detail, to use the present system on a given input sample, it is first normalised and standardized. Then, for faster processing, a copy of the input image is low-resolution to have the smallest dimension equal to 232 pixels, and a central square crop of 224 pixels is extracted from it. This low-resolution cropped version is fed to the NEM, which predicts the parameters to enhance the image. Finally, the predicted parameters and the full-resolution image are fed to the DWM to clean the image and enhance its content. The normalisation, standardisation, resizing and cropping are performed to align the input image to the evaluation set-up and / or to the downstream model.
[126] The resulting output can either be used as it stands, if the final objective was attaining a cleaned image, or fed to any downstream model to attain a more accurate prediction on the considered task. Overall, the additional computational overhead is minimal, with the present modules supporting up to about 300fps at full HD resolution.
[127] Results and Discussion. To validate the generalization capability of the present modular system, its performance is evaluated on two main tasks: image classification, and semantic segmentation. In the following, the quantitative and qualitative results attained in various tasks are reported by discussing them and drawing comparisons with competing strategies.
[128] Experimental Setup. To highlight the generalization capability of the present approach, the modules were trained only once on the upstream classification task on the ImageNet dataset. The optimization lasted for 50fc steps, using batch size 384 and Adam optimizer with learning rate 10-3 scheduled according to a cosine annealing strategy. It is remarked that when optimizing SyMPIE, the upstream module can be completely frozen, reducing the computational complexity of training. This is the scenario considered here. Furthermore, the training is done once and the same pre-trained weights are used for all the downstream models and tasks.
[129] The present NEM is implemented using a 3-block architecture with strong downsampling ( / .e., with stride 4) between layers. A detailed breakdown is reported in Figure 6, which is a table showing a breakdown of one possible architecture of the enhancement parameter estimation module of the present modular system (also referred to as NEM). BatchNorm layers and ReLU activations are added after each convolution. In total, the NEM module uses seven 2D convolutions and a single fully connected layer to project the downsampled features into the space of parameters (K, CM, Cs) eR4. A kernel size K = 5 is considered, which yields A := d(d + 1) + K2 = 37 in the setup. During training, the exponential moving average rate was defined as / ? = 0.9, and the weight factor for the regularization loss was set to Aema - 0.5. Without loss of generality, the normalization of the input images at inference-time must match the normalization seen during training. In the present case, X is normalized in the range [0,1] and standardized using mean p. = [0.485,0.456,0.406] and standard deviation a = [0.229,0.224,0.225],
[130] Datasets. We employed our system on several real and synthetic datasets to verify its effectiveness in different scenarios for the image classification task. In particular, the ImageNetC and ImageNetC-Bar datasets having synthetic corruptions were used (both based on ImageNet, as well as the real dataset VizWiz having natural corruptions. Furthermore, ImageNetC-mixed is introduced, which is a new benchmark to investigate the reliability of models when presented with multiple corruptions at once. This was built by randomly applying 1-to-3 corruptions to the same image, sampled from the joint pool of augmentations of the present data augmentation pipeline and those proposed in ImageNetC.
[131] For semantic segmentation, the models were analysed using driving scenes tackling the domain adaptation problem for clear-to-adverse weather conditions. In this case, three real world datasets were employed: i) Cityscapes as the training (source) domain; ii) ACDC and iii) DarkZurich as testing downstream (target) domains.
[132] Metrics. In image classification, the per-corruption accuracies are reported (Acc, T), together with their mean (Corr. Avg., T), and the the accuracy on corruption-free data (Clean, T). In semantic segmentation, the results as per-class loll (Intersection over Union) or its mean (mloU, T) are reported. In all cases, A (%A) refers to the absolute (relative) gain with respect to the considered baseline.
[133] Results for Image Classification: In the image classification task, three main scenarios are investigated: i) single synthetic corruptions (using ImageNetC), ii) multiple synthetic corruptions (using ImageNetC-mixed), and iii) real-world corruptions (using VizWiz). Finally, the computational cost is also calculated.
[134] Results on Single Synthetic Corruptions (ImageNetC): For this discussion, referece is made to Figure 7, which is a table reporting the final accuracy attained by the present approach when mounted on multiple different backbones. Specifically, Figure 7 shows the results for the image classification task on the ImageNetC dataset (higher is better). In particular, a ResNet50 is used with different pre-training weights (TorchVision V2 and V1, HA, PRIME, and PIXMIX), as well as a VGG16, a Swin-Tiny, and finally the CLIP foundation model. Notably, the SyMPIE model is trained only once using ResNet50 with TorchVision’s V2 weights, and the trained model is used as it stands for all other experiments, highlighting the generalization capability of the present approach. It can be seen that SyMPIE consistently improves the average accuracy across the considered models with an average gain of 2.2% in absolute terms and an average relative gain of 5.0%. As anticipated, a remarkable feature of SyMPIE is the capability to improve the performance even for architectures designed to work well in the ImageNetC task, such as HA, PRIME, and PIXMIX. These are three strong state-of-the-art data augmentation approaches and, therefore, can handle corrupted samples from ImageNetC better than weaker data augmentation approaches. Keeping this in mind, the improvement brought by the present approach when used jointly with these architectures is even more striking, signifying that the effect of the present model is complementary to existing state-of-the-art approaches, and using both strategies together can improve the absolute accuracy significantly. Numerically, when using the present system together with HA, the 5 performance is improved over TorchVision’s V1 baseline by almost 20% in absolute terms (49.4% relative).
[135] Moreover, the gain is well-spread across the various corruptions and, in cases particularly suited to be tackled by the present constrained model, some higher gains can be 10 appreciated. For example, in Motion Blur and Snow, significant performance gains (up to 9% absolute points) are obtained even on already-robust backbones like those pre-trained via HA. The present Applicant believes that the motion blur kernel can be estimated accurately with the present learned filter, whereas the present affine color transforms handle the low contrast brought by snow, regardless of the localized bright spots. Slight performance drops are 15 experienced on the fog corruption, due to changes to the frequency distribution the present system may introduce via the global filter. Methods such as HA are designed with frequencyspectra changes in mind. Therefore, they expect a distribution of input images with specific frequency characteristics. Another performance drop is observed on the Brightness corruption. This corruption is modeled as a non-linear change in the HSV color space, and the 20 present affine color transforms could not approximate the underlying function accurately. Note that these limitations introduce only marginal drops in accuracy, and do not change the overall improvements brought by the present system.
[136] Lastly, the results are discussed for when the present approach is employed with 25 downstream architectures that are different than the one used for pre-training the modules. The usefulness of the present method is verified on the widely used VGG16 convolutional architecture and on a transformer-based architecture (e.g., Swin-Tiny). Then, the CLIP foundation model was considered, that was pre-trained to align image-text embeddings to the same semantic value, rather than explicitly recognizing the input image category. The present 3D modules bring improvement even in this case, highlighting how the approach is able not only to change the graphical appearance but also to highlight the semantic content, making it easier for the downstream network.
[137] Results on Mixed Synthetic Corruptions (ImageNetC-mixed): Following the tests on a 35 single corruption at a time, the more challenging setting of multiple corruptions together is considered, using the new proposed ImageNetC-mixed dataset. The results of these analyses are shown in Figure 8, which is a table showing accuracy on ImageNetC-mixed via ResNet50. Despite the existing data augmentation approaches losing about 10% accuracy compared with the single corruption case, the present approach maintains a stable gain of 2.0% in absolute terms, corresponding to 5.0% relative improvement. This proves the ability of the present approach to handle the composition of input corruptions, as it is often encountered in practice.
[138] Results on Real-World Corruptions (VizWiz): Above, the investigations on the performance of the present modules on synthetic corruptions has been described, where the standard evaluation pipeline has been followed. However, it is important to verify that the performance improvement is maintained even on real data. To this end, the corrupted real-world VizWiz dataset is used, which, to the best of the Applicant’s knowledge, is the only real-world corrupted images dataset for the considered task currently available. The numerical results are reported in Figure 9, which shows a table of results for VizWiz with ResNet50 V2. A 1.2% relative gain on corrupted images is observed, as well as an improvement in the accuracy on clean images by a relative gain of 0.4%.
[139] Computational Cost: The complexity of the present approach is reported in FLOPs (Floating Point Operations) in Figure 10, which is a table showing computational complexity of the present method compared to other input-level image enhancement strategies. Specifically, the present techniques are compared to other architectures that can be used as (or converted into) input-level denoisers. The cost of the present approach was computed using the PTFIops library at a resolution of 1920 x 1080px2.
[140] The present module is one to four orders of magnitude faster than the diffusion model competitors. Moreover, its computational complexity does not scale significantly with the input resolution - since a resizing stage is done before parameter estimation and the only operations applied on the full-resolution image are the spatial convolution and a matrix multiplication over the channels. This means that SyMPIE can handle images of arbitrary resolution, contrary to what happens for, e.g., the fixed-size diffusion models. Similar considerations also hold for auto-encoders and GANs. This confirms the efficiency of the present strategy.
[141] Finally, the throughput of the present module was computed on an NVIDIA GTX 1080Ti GPU, obtaining a speed of 289.4 images / second; corresponding to a per-image inference time of about 3ms, meaning that SyMPIE can be easily employed in real-time applications with minimal impact on the inference time of downstream architectures.
[142] Results for Semantic Segmentation: The analyses up to this point have been in the same task as training (i.e., image classification) with some changes in the downstream architecture. However, to truly show the generalization capabilities of the present approach, the downstream task is completely changed and the accuracy improvement is determined. For this investigation, the semantic segmentation task is chosen and, in particular, domain adaptation to adverse weather conditions. The baseline architecture used by Barbato et al (DeepLabV2 with ResNet50 backbone) are used on the Cityscapes dataset, and then deployed on two adverse weather datasets, namely, ACDC and DarkZurich.
[143] The quantitative results of this study are reported in Figure 11, which is a table showing the quantitative results for semantic segmentation using a DeepLabV2 architecture with the ResNet50 backbone. It can be seen that the present architecture improves the mloll score by 4.1% and 3.8% relative gain on ACDC and DarkZurich, respectively.
[144] Ablation Studies: Some ablation studies are also reported, namely: an analysis of corruptions that cannot be modeled by the present strategy (ImageNetC-Bar); a study on the iterative application of the present module, which is a common failure point for denoisers and image enhancers; and a proof of concept for applicability on video compression.
[145] Results on corruptions that cannot be modelled by SyMPIE by construction (ImageNetC-Bar): The reliability of the present method to corruptions that cannot be modeled by the present modules is also analysed. For this purpose, SyMPIE’s performance is tested on the ImageNetC-Bar dataset, which provides a corruption set complementary to ImageNetC. In general, the corruptions present in the dataset cannot be suitably modeled by the present approach, since most of them contain non-spatially-uniform distortions. Nonetheless, small accuracy gains are observed when using the present modules. This means that, even if it is not possible to model the corruptions in the dataset, SyMPIE can modify the inputs to make the classification task easier. Figure 12 is a table showing quantitative results on the ImageNetC-Bar dataset with ResNet50. Looking at the results, it is possible to identify three corruptions where the present method brings consistent and noticeable gains: Inverse Sparkle, Plasma, and Single Sine. In the first corruption, the present method improves the accuracy by an average of 2.5% in absolute terms (4.8% relative gain), in the second we improve by 0.4% (1.1%), while in the third we improve by 2.5% (7.0%). These three corruptions, in some way, are similar to those present in ImageNet-C and can be dealt with more effectively by our approach. In particular, the Plasma noise is a colored version of Fog, Inverse Sparkle is similar to Brightness in some instances, and the effect of Single Sine is similar to Contrast and Fog.
[146] Iterative Application: Furthermore, a common fallacy of denoising and enhancing models is investigated, especially those based on autoencoder architectures: when fed their own outputs iteratively, their prediction may collapse into undesirable modes, or destroy completely the input content. To evaluate qualitatively the susceptibility of the present approach to this scenario, the predictions provided by the present system are fed to itself four times. It was found that the present architecture does not suffer from the mentioned problem, and all outputs are consistent with the original input. Moreover, in many cases, a clear improvement was observed following the subsequent application of the present module. The most striking cases are the motion blur, frost, and snow corruptions highlighting the potentiality for deployment of the present architecture. In the first case, the present approach is able to completely remove the blurring artifacts, while in the other two, it removes the haze from the image and increases contrast. Also, the present method enhances significantly the overall image content and colors, while removing the blur at the same time for the Glass Blur (second column).
[147] Video Compression: Finally, a proof-of-concept investigation is provided of the present system for live-streaming applications (e.g., for scene understanding applications on broadcasted data). A freely available video was downloaded from YouTube and this was compressed multiple times using the VLC Media Player software at different compression rates, before extracting the frames from the videos for their analysis.
[148] The features extracted by the model for each frame obtained from the original, compressed, and processed videos are compared to determine how similar they are (via mean squared error, MSE, 1). This metric is representative of the overall reconstruction fidelity as well as the potential applicability to an end task, since the extracted features are strongly correlated to the final predictions, regardless of the end task employed. The results are given in Figure 13, which is a table of results for feature extraction from compressed video data.
[149] The results show that the present method brings a 4.0% (2.6%) relative improvement in the high- (low-) compression regime in the linear scale.
[150] Thus, the present techniques provide a novel, modular, and efficient system that predicts explicit parameters for content enhancement of an input image targeting improved accuracy on downstream tasks. The computational footprint of the modules is minimal, being more than 10x faster than competing approaches (2GFLOPs in total), and enjoys a throughput of about 300 images per second. A key feature of the present approach is the capability of training without any paired (clean / corrupted) input samples, but rather learning automatically the most suitable transformations of an input image using the supervision on an upstream task. Even more remarkably, once the system has been trained on a given upstream task, it can generalize to arbitrary downstream tasks without fine-tuning. To confirm the generalization claim, the present approach was validated on three classification datasets and two semantic segmentation datasets, achieving noticeable improvement in all of them.
[151] The present techniques are now described in the context of some non-limiting example use cases, and the advantages are described relative to existing systems.
[152] Figure 14 is a block diagram of the overall architecture of the present techniques. The present techniques involving inserting an efficient image denoising architecture into a pipeline for processing images using an end-task ML model, where the image denoising architecture is inserted into the pipeline before the end-task ML model. Thus, as shown in Figure 14, input images are input into the image denoising architecture first, so that their content can be enhanced (e.g. cleaned if the input image is a noisy image), before being input into the endtask ML model. The present techniques focus on corruptions in the color space of an image (e.g., brightness, constrast, blur, etc.), but it will be understood that other types of corruptions in an image could also be handled using the same principles. For minimal computational footprint, the image warping (enhancing) step is separated from the corruption estimation step. This way, the image resolution between estimation and warping / enhancing can change (e.g. estimate parameters on low-resolution image for faster processing, warp / enahnce on fullresolution). The frozen (warping / enhancing) component can be implemented in hardware for faster processing, even when the parameter estimation strategy changes.
[153] Figure 15A is a schematic diagram showing one advantage of the present techniques over existing large denoising networks. Existing denoising networks often use large networks, which may not be deployable on resource-constrained end user devices. (See also Figure 1D and corresponding description above). In contrast, as shown in Figure 15A, the present techniques provide a smaller estimation network that is computationally efficient, involving only a few convolutions, making it suitable for deployment on resource-constrained devices.
[154] Figure 15B is a schematic diagram showing another advantage of the present techniques over existing networks. Existing networks often need to be re-trained from scratch to handle corrupted or noisy data, but each retrained network may only work for a chosen end task and cannot be readily used for other end tasks. Each end task requires retraining of the network which means the process is time-consuming and computationally-expensive. (See also Figure 1B and corresponding description above). In contrast, as shown in Figure 15B, the present techniques provide a modular system that can be used to clean or denoise corrupted or noisy data for any downstream / end task.
[155] Figure 15C is a schematic diagram showing another advantage of the present techniques over existing networks. Existing networks may be adapted to work with noisy or corrupted data, but if the noise or corruption type changes, the adapted network does not work well. (See also Figure 1C and corresponding description above). In contrast, as shown in Figure 15C, the present techniques provide a corruption and noise-agnostic modular system that can be used to handle a wide variety of input data without retraining or adaptation.
[156] Figure 16 is a schematic diagram showing how the present techniques can be used to handle any number and combination of corruptions and / or noise types. That is, the present techniques support the adjustment of any parameters of the input corrupted / noisy image. Figure 16 shows some examples of linear transformations that can be be applied in the colour space of an input image, and spatial filtering via convolution. As shown, in this non-limiting example, input images may be shifted in the colour space and may be warped in the colour space to enhance the content of the input images. It will be understood that any number and type of transformation may be applied to input images to enhance their content.
[157] The present techniques may be used in a variety of end-user devices.
[158] For example, the present techniques may be deployed in a smart fridge I refrigerator. Smart fridges may comprise imaging devices (e.g. cameras) which capture images when food items are put into and taken out of the fridge. They may also be able to track when the quantity or volume of food items are running low, so that a user can order / buy more. The present techniques are suitable for deployment in a smart fridge, which is a computational resource-constrained device. The present techniques provide easier tracking and recognition of inserted food items in the fridge, which can enable storage time of those food items to be tracked more effectively (i.e. to determine when the “shelf life” of the food items is going to be expire). The present techniques enable more robust food item detection in adverse lighting conditions, or when the smart fridge’s camera has not had time to adapt to the new lights (fast door opening, night-time use, bright days, etc.) As a result, an end user is more engaged with the appliance and will explore more of its features, at the time of replacement is more likely to buy a smart appliance again. Similarly, the end user can trust the list of what is inside the refrigerator or the expiry date.
[159] In another example, the present techniques may be deployed in a smart oven. A smart oven may comprise at least one camera that is able to capture images of the inside of the oven, which may be used to monitor food items cooking inside the oven. This could be useful because a user can be alerted to the state of the food items. For example, if the food inside the oven starts to burn, the camera(s) may be able to detect this and alert the user before the food burns even more. The present techniques provide more accurate burn state identification even when steam is present, and enable the cameras to be used in darker conditions (like inside an oven). As a result, the end user will trust the appliance more, knowing that its safety features will work on adverse conditions.
[160] In another example, the present techniques may be deployed in a smart robotic, autonomous vacuum cleaner. A smart robotic vacuum cleaner may be able to move around an environment on its own and clean the floor of debris. The smart vacuum cleaner may comprise at least one camera for capturing images of the environment, so that the vacuum cleaner can navigate through the environment and avoid certain objects. The present techniques provide an onboard object recognition system that is more reliable, able to work in darker conditions and better able to avoid any obstacles or people / pets more readily. As a result, the end user can move around more freely when the robot is cleaning, and trusts the robot more around pets, knowing the device will not bother them.
[161] Figure 17 is a flowchart of example steps for image enhancing an image to improve outcomes of a downstream image processing task. The method comprises: receiving an image to be enhanced prior to implementing the downstream image processing task (step S100); estimating, using an enhancement parameter estimation module of a modular enhancement system, a set of enhancement parameters for use in enhancing the image, wherein the enhancement parameters are parameters of transformation operators that, when applied to an image, improve image content understanding, and wherein the enhancement parameter estimation module comprises a trained machine learning, ML, model (step S102); and enhancing the image using an enhancement module of the modular enhancement system to generate an enhanced image, wherein the enhancing is based on the set of enhancement parameters (step S104).
[162] Figure 18 is a flowchart of example steps for training a machine learning, ML, model of a modular enhancement system for enhancing an image for a downstream image processing task. The method comprise: obtaining a training dataset comprising a plurality of corrupted images, wherein each corrupted image comprises at least one image corruption, and each corrupted image comprises a ground truth label (step S200); training the ML model of an enhancement parameter estimation module of the modular enhancement system using each image of the training dataset by: estimating, using the ML model and the corrupted image, a set of enhancement parameters for use in enhancing the corrupted image (step S202); enhancing, using an enhancement module of the modular enhancement system, the corrupted image using the set of enhancement parameters (step S204); generating, using a pre-trained downstream ML model trained to perform a specific downstream task, and the enhanced image, a predicted label for the enhanced image (step S206); calculating a loss based on a difference between the predicted label for the enhanced image generated by the downstream ML model and the ground truth label for the corrupted image (step S208); and updating parameters of the ML model of the enhancement parameter estimation module based on the calculated loss (step S210).
[163] Figure 19 is a block diagram of provided an apparatus 100 for enhancing an image to improve outcomes of a downstream image processing task. The apparatus 100 may be a constrained-resource device, but which has the minimum hardware capabilities to use a trained neural network / ML model. The apparatus may be any one of: a smartphone, tablet, laptop, computer or computing device, virtual assistant device, a vehicle, an autonomous vehicle, a robot or robotic device, a robotic assistant, image capture system or device, an augmented reality system or device, a virtual reality system or device, a gaming system, an Internet of Things device, or a smart consumer device (such as a smart fridge or smart vacuum cleaner). It will be understood that this is a non-exhaustive and non-limiting list of example apparatuses.
[164] The apparatus comprises: an interface 108 for receiving an image to be enhanced prior to implementing the downstream image processing task. The interface may be a camera for capturing an image or video, or may be an interface for receiving images / videos captured by an external device.
[165] The apparatus comprises a modular enhancement system 106. The modular enhancement system comprises an enhancement parameter estimation module 106A. The estimation module 106A comprises a trained ML model 106C. The modular enhancement system comprises an enhancement module 106B.
[166] The apparatus comprises at least one processor 102 coupled to memory 104. The at least one processor 102 may comprise one or more of: a microprocessor, a microcontroller, and an integrated circuit. The memory 104 may comprise volatile memory, such as random access memory (RAM), for use as temporary memory, and / or non-volatile memory such as Flash, read only memory (ROM), or electrically erasable programmable ROM (EEPROM), for storing data, programs, or instructions, for example.
[167] The processor 102 may be configured for implementing the modular enhancement system by: estimating, using the enhancement parameter estimation module 106A of the modular enhancement system 106, a set of enhancement parameters for use in enhancing the image, wherein the enhancement parameters are parameters of transformation operators that, when applied to an image, improve image content understanding; and enhancing the image using the enhancement module 106B of the modular enhancement system 106 to generate an enhanced image, wherein the enhancing is based on the set of enhancement parameters.
[168] The apparatus may further comprise a downstream ML model 110, for which images are being enhanced by the modular enhancement system 106. Thus, enhanced images may be input into the downstream ML model 110, by the processor 102, to enable processing by the downstream ML model 110.
[169] References: • ImageNet / lmageNet-1k- Russakovsky, 0., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al. Imagenet large scale visual recognition challenge. International Journal of Computer Vision (IJCV) (2015). • Hendrycks D and Dietterich T (ImageNetC) - Hendrycks, D., and Dietterich, T. Benchmarking neural network robustness to common corruptions and perturbations. arXiv preprint arXiv: 1903.12261 (2019) • Adam - Kingma, D. P., and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv: 1412.6980 (2014). • ImageNetC-Bar - Mintun, E., Kirillov, A., and Xie, S. On interaction between augmentations and corruptions in natural corruption robustness. Advances in Neural Information Processing Systems 34 (2021), 3571-3583. • VizWiz - Bafghi, R. A., and Gurari, D. A new dataset based on images taken by blind people for testing the robustness of image classification models trained for imagenet categories. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (2023), pp. 16261-16270. • Cityscapes - Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. The cityscapes dataset for semantic urban scene understanding. In Proc, of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2016). • ACDC - Sakaridis, C., Dai, D., and Van Gool, L. ACDC: The adverse conditions dataset with correspondences for semantic driving scene understanding. In Proceedings of International Conference on Computer Vision (ICCV) (2021) • DarkZurich - Sakaridis, C., Dai, D., and Gool, L. V. Guided curriculum model adaptation and uncertainty-aware evaluation for semantic nighttime image segmentation. In Proceedings of the IEEE / CVF International Conference on Computer Vision (2019), pp. 7374-7383 • ResNet50 - He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition, 2015. • TorchVision v2 and v1 -TorchVision maintainers and contributors. Torch Vision: PyTorch’s Computer Vision library, https: / / github.com / pytorch / vision, 2016. • HA - Yucel, M. K., Cinbis, R. G., and Duygulu, P. Hybridaugment++: Unified frequency spectra perturbations for model robustness. In Proceedings of the IEEE / CVF International Conference on Computer Vision (2023), pp. 5718-5728. • PRIME - Modas, A., Rade, R., Ortiz-Jimenez, G., Moosavi-Dezfooli, S.-M., and Frossard, P. Prime: A few primitives can boost robustness to common corruptions. In European Conference on Computer Vision (2022), Springer, pp. 623 • PIXMIX - Hendrycks, D., Zou, A., Mazeika, M., Tang, L, Li, B., Song, D. X., and Steinhardt, J. Pixmix: Dreamlike pictures comprehensively improve safety measures. Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2021). • VGG16 - Simonyan, K., and Zisserman, A. Very deep convolutional networks for largescale image recognition. arXiv preprint arXiv: 1409.1556 (2014). • Swin-Tiny - Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE / CVF international conference on computer vision (2021), pp. 10012-10022.
[32] Lore, K. G., Akintayo, A., and Sarkar, S. Linet: A deep autoencoder approach • CLIP - Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. ArXiv:2103.00020 (2021). • PTFIops library - Sovrasov, V. ptflops: a flops counting tool for neural networks in pytorch framework, 2018-2023. • Barbato, F., Michieli, U., Toldo, M., and Zanuttigh, P. Road scenes segmentation across different domains by disentangling latent representations. arXiv preprint arXiv:2108.03021 (2021). • Barbato, F., Toldo, M., Michieli, U., and Zanuttigh, P. Latent space regularization for unsupervised domain adaptation in semantic segmentation. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops (June 2021), pp. 2835-2845. • DeepLabV2 - Chen, L.-C., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A. L. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs, 2017. • ResNet50 backbone - He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2016). • VGG19 - Simonyan, K., and Zisserman, A. Very deep convolutional networks for largescale image recognition. arXiv preprint arXiv: 1409.1556 (2014).
[170] Those skilled in the art will appreciate that while the foregoing has described what is considered to be the best mode and where appropriate other modes of performing present techniques, the present techniques should not be limited to the specific configurations and methods disclosed in this description of the preferred embodiment. Those skilled in the art will recognise that present techniques have a broad range of applications, and that the embodiments may take a wide range of modifications without departing from any inventive concept as defined in the appended claims.
Claims
1. A computer-implemented method for enhancing an image to improve outcomes of a downstream image processing task, the method comprising:receiving an image to be enhanced prior to implementing the downstream image processing task;estimating, using an enhancement parameter estimation module of a modular enhancement system, a set of enhancement parameters for use in enhancing the image, wherein the enhancement parameters are parameters of transformation operators that, when applied to an image, improve image content understanding, and wherein the enhancement parameter estimation module comprises a trained machine learning, ML, model; andenhancing the image using an enhancement module of the modular enhancement system to generate an enhanced image, wherein the enhancing is based on the set of enhancement parameters.
2. The method as claimed in claim 1, wherein estimating the set of enhancement parameters comprises:estimating the set of enhancement parameters by inputting a low-resolution version of the image into the ML model.
3. The method as claimed in claim 2, comprising:discarding the low-resolution version of the image after using the low-resolution version of the image to estimate the set of enhancement parameters.
4. The method as claimed in claim 1, 2 or 3, wherein the image being enhanced is a single image, and wherein estimating the set of enhancement parameters comprises:estimating the set of enhancement parameters using the single image.
5. The method as claimed in claim 1,2 or 3, wherein the image being enhanced is a video comprising a sequence of frames, and wherein estimating the set of enhancement parameters comprises:selecting a frame from the sequence of frames;estimating the set of enhancement parameters using the selected frame ; andapplying the estimated set of enhancement parameters for the selected frame to other frames in the sequence of frames.
6. The method as claimed in any one of claims 1 to 5, wherein estimating the set of enhancement parameters using the enhancement parameter estimation module comprises:using the trained ML model, the trained ML model comprising a convolutional neural network, CNN.
7. The method as claimed in any one of claims 1 to 5, wherein estimating the set of enhancement parameters using the enhancement parameter estimation module comprises:using the trained ML model, the trained ML model comprising a transformer.
8. The method as claimed in any one of claims 1 to 7, further comprising: inputting the enhanced image into a downstream machine learning, ML, model for implementing the downstream image processing task.
9. The method as claimed in any preceding claim wherein estimating a set of enhancement parameters comprises estimating a set of enhancement parameters for any one or more of: a geometric transformation operator; a non-geometric transformation operator; a linear transformation operator; a non-linear transformation operator; and a Fourier transformation operator.
10. The method as claimed in any preceding claim, wherein estimating a set of enhancement parameters comprises estimating parameters for an additive operator.
11. The method as claimed in claim 10, wherein the additive operator is a colour shift operator, and enhancing the image using the enhancement module comprises:enhancing the image by performing a colour shift operation on the image.
12. The method as claimed in any preceding claim, wherein estimating a set of enhancement parameters comprises estimating parameters fora multiplicative operator.
13. The method as claimed in claim 12, wherein the multiplicative operator is a colour warp operator, and enhancing the image using the enhancement module comprises:enhancing the image by performing a colour warp operation on the image.
14. The method as claimed in any preceding claim, wherein estimating a set of enhancement parameters comprises estimating parameters for a convolution operator.
15. The method as claimed in claim 14, wherein the convolution operator is a de-blurring operator, and enhancing the image using the enhancement module comprises:enhancing the image by performing a de-blurring operation on the image.
16. The method as claimed in any preceding claim, wherein estimating a set of enhancement parameters comprises estimating parameters for a denoising operator.
17. The method as claimed in claim 16 wherein enhancing the image using the enhancement module comprises:enhancing the image by performing a denoising operation on the image.
18. An apparatus for enhancing an image to improve outcomes of a downstream image processing task, the apparatus comprising:an interface for receiving an image to be enhanced prior to implementing the downstream image processing task;a modular enhancement system; andat least one processor coupled to memory, for implementing the modular enhancement system by:estimating, using an enhancement parameter estimation module of the modular enhancement system, a set of enhancement parameters for use in enhancing the image, wherein the enhancement parameters are parameters of transformation operators that, when applied to an image, improve image content understanding, and wherein the enhancement parameter estimation module comprises a machine learning, ML, model; andenhancing the image using an enhancement module of the modular enhancement system to generate an enhanced image, wherein the enhancing is based on the set of enhancement parameters.
19. The apparatus as claimed in claim 18, wherein the enhancement module is a System-on-Chip, SoC, enhancement module.
20. The apparatus as claimed in claim 18 or 19, wherein the apparatus is a home appliance.
21. The apparatus as claimed in claim 18 or 19, wherein the apparatus is an autonomous vehicle.
22. The apparatus as claimed in claim 18 or 19, wherein the apparatus is a mobile computing device.
23. A method for training a machine learning, ML, model of a modular enhancement system for enhancing an image for a downstream image processing task, the method comprising:obtaining a training dataset comprising a plurality of corrupted images, wherein each corrupted image comprises at least one image corruption, and each corrupted image comprises a ground truth label;training the ML model of an enhancement parameter estimation module of the modular enhancement system using each image of the training dataset by:estimating, using the ML model and the corrupted image, a set of enhancement parameters for use in enhancing the corrupted image;enhancing, using an enhancement module of the modular enhancement system, the corrupted image using the set of enhancement parameters;generating, using a pre-trained downstream ML model trained to perform a specific downstream task, and the enhanced image, a predicted label for the enhanced image;calculating a loss based on a difference between the predicted label for the enhanced image generated by the downstream ML model and the ground truth label for the corrupted image; andupdating parameters of the ML model of the enhancement parameter estimation module based on the calculated loss.
24. The method as claimed in claim 23, the training further comprising:enhancing the image using a smoothed version of the modular enhancement system to generate an intermediate image;enhancing the intermediate image using the modular enhancement system to generate a enhanced intermediate image;generating, using the pre-trained downstream ML model and the enhanced intermediate image, a predicted label for the enhanced intermediate image; andcalculating a further loss based on:the predicted label for the enhanced intermediate image; andthe ground truth label for the corrupted image,wherein updating parameters of the ML model of the enhancement parameter estimation module based on the calculated loss comprises:updating parameters of the ML model of the enhancement parameter estimation module based on the calculated loss and the calculated further loss.
525. The method as claimed in claim 23 or 24, wherein using the pre-trained downstream ML model to generate a predicted label comprises using any of the following downstream ML models: a computer vision model; a classification model; a segmentation model; an object detection model; a panoptic segmentation model; an instance segmentation model; an action10 recognition model; an anomaly detection model; an image captioning model; and a salient object detection image denoising model.
Citation Information
Patent Citations
Low-light image enhancement method based on noise attention map guidance
CN113643202A
Unsupervised underwater image enhancement method and related equipment
CN115660980A
Machine learning pipeline for document image quality detection and correction
US20220350996A1
A device and method for image processing
WO2021093956A1