Image quality improvement based on removing image degradation through learning from multiple machine learning models

A machine learning-based system addresses image degradation issues by training multiple models to enhance image quality, effectively removing blur, noise, and compression artifacts, resulting in sharper images.

JP2025532573APending Publication Date: 2025-10-01GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025515617
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-09-13
Publication Date
2025-10-01

AI Technical Summary

Technical Problem

Existing image capture devices struggle with removing blur, noise, and compression artifacts, which degrade image quality due to factors like incorrect focus, camera motion, sensor limitations, and image compression.

Method used

A system utilizing multiple machine learning models, including convolutional neural networks, is trained on synthetic and real-world datasets to remove image degradations such as motion blur, lens blur, noise, and compression artifacts, enhancing image quality through a process involving encoder-decoder networks with skip connections and knowledge distillation.

Benefits of technology

The system effectively improves image sharpness and quality by removing various degradations, providing clearer and more visually appealing images across diverse shooting conditions and devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025532573000001_ABST
    Figure 2025532573000001_ABST
Patent Text Reader

Abstract

The method includes receiving an input image. The method includes predicting a transformed version of the input image with an image transformation model, the image transformation model being trained to remove image degradation associated with the input image, the training including: (1) training a plurality of intermediate machine learning models to remove image degradation, each intermediate machine learning model of the plurality of intermediate machine learning models being trained with a respective training dataset of a plurality of training datasets corresponding to each of a plurality of degradation factors, each training dataset including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image; and (2) the image transformation model being trained with an additional training dataset of real images and learning from the plurality of intermediate machine learning models. The method includes providing the predicted transformed version.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Many modern computing devices, including mobile phones, personal computers, and tablets, include image capture devices such as still cameras and / or video cameras. Image capture devices can capture images, such as images containing people, animals, landscapes, and / or other subjects. Some image capture devices and / or computing devices can enhance or otherwise modify captured images. For example, some image capture devices can provide “red-eye” correction, which removes artifacts such as red-appearing eyes in people and animals that may be present in images captured using bright light, such as flash lighting. After a captured image has been enhanced, the enhanced image can be saved, displayed, transmitted, printed on paper, and / or otherwise utilized. Summary of the Invention

[0002] Removing blur, noise, and compression artifacts from images has been a long-standing problem in computational photography. Image degradation can arise from several sources: when the photographer or autofocus system incorrectly sets the focus (defocus), or when relative motion between the camera and the scene is faster than the shutter speed (motion blur). Furthermore, even under ideal acquisition conditions, there can be inherent camera blur due to sensor resolution, optical diffraction, lens aberrations, and anti-aliasing filters. Similarly, image noise is inherent in capturing a discrete number of photons (shot noise) and analog-to-digital conversion and processing (read noise). Images are typically compressed using techniques such as JPEG compression before storage or transmission. Image compression can also degrade image quality.

[0003] Equipped with a system of machine-learned components, image capture devices can be configured to allow users to create sharp images by removing blur, noise, compression artifacts, etc. In some aspects, mobile devices can be configured with such features to enable real-time image enhancement. In some examples, the mobile device can automatically enhance images. In other aspects, mobile phone users can non-destructively enhance images to suit their preferences. Also, existing images in a user's image library, for example, can be enhanced based on the techniques described herein.

[0004] In one aspect, a computer-implemented method is provided. The method includes receiving, by a computing device, a plurality of training datasets corresponding to each of a plurality of degradation factors, each training dataset including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image. The method also includes training a plurality of intermediate machine learning models to remove one or more image degradations associated with a given image, wherein each intermediate machine learning model of the plurality of intermediate machine learning models is trained with a respective training dataset of the plurality of training datasets to remove one or more image degradations associated with the given image based on each of the plurality of degradation factors. The method further includes training an image transformation model with an additional training dataset of real images to remove one or more image degradations associated with the real images, wherein the image transformation model learns from the plurality of intermediate machine learning models. The method further includes outputting, by the computing device, the trained image transformation models for removal of the image degradations.

[0005] In another aspect, a computing device is provided. The computing device includes one or more processors and data storage. The data storage has stored therein computer-executable instructions that, when executed by the one or more processors, cause the computing device to perform a function. The function includes receiving, by the computing device, a plurality of training datasets corresponding to each of a plurality of degradation factors, each training dataset including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image; training a plurality of intermediate machine learning models to remove one or more image degradations associated with a given image, wherein each intermediate machine learning model of the plurality of intermediate machine learning models is trained with a respective training dataset of the plurality of training datasets to remove one or more image degradations associated with the given image based on each of the plurality of degradation factors; training an image transformation model with an additional training dataset of real images to remove one or more image degradations associated with real images, wherein the image transformation model learns from the plurality of intermediate machine learning models; and outputting, by the computing device, the trained image transformation model for removal of the image degradations.

[0006] In another aspect, a computer program is provided, the computer program including instructions that, when executed by a computing device, cause the computing device to perform a function including receiving, by the computing device, a plurality of training datasets corresponding to each of a plurality of degradation factors, each training dataset including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image; training a plurality of intermediate machine learning models to remove one or more image degradations associated with a given image, the intermediate machine learning models being trained with a respective training dataset of the plurality of training datasets to remove one or more image degradations associated with the given image based on each of the plurality of degradation factors; training an image transformation model with an additional training dataset of real images to remove one or more image degradations associated with real images, the image transformation model learning from the plurality of intermediate machine learning models; and outputting, by the computing device, the trained image transformation model for removal of the image degradations.

[0007] In another aspect, an article of manufacture is provided, the article including one or more computer-readable media having stored thereon computer-readable instructions that, when executed by one or more processors of the computing device, cause the computing device to perform functions including receiving, by the computing device, a plurality of training datasets corresponding to each of a plurality of degradation factors, each training dataset including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image; training a plurality of intermediate machine learning models to remove one or more image degradations associated with a given image, the plurality of intermediate machine learning models being trained with a respective training dataset of the plurality of training datasets to remove one or more image degradations associated with the given image based on each of the plurality of degradation factors; training an image transformation model with an additional training dataset of real images to remove one or more image degradations associated with real images, the image transformation model learning from the plurality of intermediate machine learning models; and outputting, by the computing device, the trained image transformation model for removal of the image degradations.

[0008] In another aspect, a system is provided that includes: means for receiving, by a computing device, a plurality of training datasets corresponding to each of a plurality of degradation factors, each training dataset including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image; means for training a plurality of intermediate machine learning models to remove one or more image degradations associated with a given image, wherein each intermediate machine learning model of the plurality of intermediate machine learning models is trained with a respective training dataset of the plurality of training datasets to remove one or more image degradations associated with the given image based on each of the plurality of degradation factors; means for training an image transformation model with an additional training dataset of real images to remove one or more image degradations associated with the real images, wherein the image transformation model learns from the plurality of intermediate machine learning models; and means for outputting, by the computing device, the trained image transformation model for removal of the image degradations.

[0009] In another aspect, a computer-implemented method is provided. The method includes receiving, by a computing device, an input image including one or more image degradations. The method also includes predicting a transformed version of the input image by an image transformation model, the image transformation model being trained to remove one or more image degradations associated with the input image, the training including: (1) training a plurality of intermediate machine learning models to remove the one or more image degradations, each intermediate machine learning model of the plurality of intermediate machine learning models being trained with a respective training dataset of a plurality of training datasets corresponding to each of a plurality of degradation factors, each training dataset including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image; and (2) the image transformation model being trained with an additional training dataset of real images and learning from the plurality of intermediate machine learning models. The method further includes providing, by the computing device, the predicted transformed version of the input image.

[0010] In another aspect, a computing device is provided. The computing device includes one or more processors and data storage. The data storage has stored therein computer-executable instructions that, when executed by the one or more processors, cause the computing device to perform a function. The function includes receiving, by the computing device, an input image including one or more image degradations; predicting a transformed version of the input image with an image transformation model, the image transformation model being trained to remove one or more image degradations associated with the input image, the training including: (1) training a plurality of intermediate machine learning models to remove the one or more image degradations, each intermediate machine learning model of the plurality of intermediate machine learning models being trained with a respective training dataset of a plurality of training datasets corresponding to each of a plurality of degradation factors, each training dataset including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image; and (2) the image transformation model being trained with an additional training dataset of real images and learning from the plurality of intermediate machine learning models. Predicting includes providing, by the computing device, the predicted transformed version of the input image.

[0011] In another aspect, a computer program is provided, the computer program including instructions that, when executed by a computing device, cause the computing device to perform a function including receiving, by the computing device, an input image including one or more image degradations; predicting a transformed version of the input image with an image transformation model, the image transformation model being trained to remove one or more image degradations associated with the input image, the training including (1) training a plurality of intermediate machine learning models to remove the one or more image degradations, each intermediate machine learning model of the plurality of intermediate machine learning models being trained with a respective training dataset of a plurality of training datasets corresponding to each of a plurality of degradation factors, each training dataset including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image; and (2) the image transformation model being trained with an additional training dataset of real images and learning from the plurality of intermediate machine learning models; and providing, by the computing device, the predicted transformed version of the input image.

[0012] In another aspect, a product is provided, the product including one or more computer-readable media having stored thereon computer-readable instructions that, when executed by one or more processors of the computing device, cause the computing device to perform functions including receiving, by the computing device, an input image including one or more image degradations; predicting a transformed version of the input image with an image transformation model, the image transformation model being trained to remove one or more image degradations associated with the input image, the training including (1) training a plurality of intermediate machine learning models to remove the one or more image degradations, each intermediate machine learning model of the plurality of intermediate machine learning models being trained with a respective training dataset of a plurality of training datasets corresponding to each of a plurality of degradation factors, each training dataset including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image; and (2) providing, by the computing device, the predicted transformed version of the input image.

[0013] In another aspect, a system is provided that includes: means for receiving, by a computing device, an input image that includes one or more image degradations; means for predicting a transformed version of the input image by an image transformation model, the image transformation model being trained to remove one or more image degradations associated with the input image, the training including: (1) training a plurality of intermediate machine learning models to remove the one or more image degradations, each intermediate machine learning model of the plurality of intermediate machine learning models being trained with a respective training dataset of a plurality of training datasets corresponding to each of a plurality of degradation factors, each training dataset including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image; and means for providing, by the computing device, the predicted transformed version of the input image.

[0014] The above summary is illustrative only and is not intended to be in any way limiting. In addition to the exemplary aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the figures and the following detailed description, and accompanying drawings. [Brief explanation of the drawings]

[0015] [Figure 1A] FIG. 1 illustrates an exemplary image enhancement process, according to an exemplary embodiment. [Figure 1B] FIG. 1 illustrates an exemplary network architecture, according to an exemplary embodiment. [Figure 2A] FIG. 1 illustrates an exemplary overview of a pipeline for generating synthetic data, according to an exemplary embodiment. [Figure 2B] FIG. 1 illustrates an exemplary training of multiple intermediate models, according to an exemplary embodiment. [Figure 2C] FIG. 1 illustrates an exemplary pipeline for generating a curated dataset, according to an exemplary embodiment. [Figure 2D] FIG. 1 illustrates an exemplary training of an image transformation model, according to an exemplary embodiment. [Figure 3] FIG. 1 illustrates an exemplary pipeline for generating synthetic data, according to an exemplary embodiment. [Figure 4] 10 illustrates exemplary motion blur kernels for different scale parameters, according to an exemplary embodiment. [Figure 5] 4A-4D illustrate exemplary input and output images corresponding to simulated motion blur, according to an exemplary embodiment. [Figure 6] 10 illustrates an exemplary histogram relating to motion kernel length as a function of scale parameter, according to an exemplary embodiment. [Figure 7] 1 illustrates an exemplary generalized Gaussian blur kernel, according to an exemplary embodiment. [Figure 8] 10A-10C illustrate exemplary input and output images corresponding to simulated compression artifacts, according to an exemplary embodiment. [Figure 9] 10 illustrates exemplary images corresponding to training data generation, according to an exemplary embodiment. [Figure 10] 1 illustrates an exemplary output image of a trained image transformation model, according to an exemplary embodiment. [Figure 11] FIG. 1 illustrates the training and inference stages of a machine learning model, according to an example embodiment. [Figure 12] 1 illustrates a distributed computing architecture in accordance with an illustrative embodiment. [Figure 13] FIG. 1 is a block diagram of a computing device in accordance with an exemplary embodiment. [Figure 14] 1 illustrates a network of computing clusters arranged as a cloud-based server system, according to an example embodiment. [Figure 15] 1 is a flowchart of a method according to an example embodiment. [Figure 16] 10 is a flowchart of another method according to an example embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0016] Image blur can generally be modeled as a linear operator acting on a sharp latent image. For a shift-invariant linear operator, the blurring operation may correspond to a convolution with a blur kernel. In practice, it is generally believed that the captured image contains additive noise and compression in addition to the blurring. Thus, the following relationship may apply: v=C(S(u*k)+n) (Formula 1)

[0017] where v is the captured image, u is the underlying sharp image, k is the unknown blur kernel, * is the convolution operation, n is additive noise, S models the sensor's nonlinear response (e.g., saturation), and C represents image compression. Some existing methods deblur images by treating the problem as a "blind" deconvolution process. For example, in a first step, a blur kernel can be estimated. This estimation can be achieved by assuming a sharp image model, e.g., using a variational framework, while in a second, independent step, a "non-blind" deconvolution algorithm can be applied. However, image noise and artifacts due to compression can adversely affect both steps. Even if the blur kernel can be determined, "non-blind" deconvolution can be an ill-posed problem, and the presence of noise, compression, etc. can introduce artifacts. A significant drawback of model-based deblurring is that the degradation model generally needs to be highly accurate. This can pose significant challenges at runtime due to several unknown or partially known image transformations (e.g., unknown blur, unknown camera image signal processor (ISP), post-processing, compression, etc.).

[0018] This application relates to enhancing images using machine learning techniques, such as, but not limited to, neural network techniques. When a mobile computing device user captures an image, the resulting image may have one or more image impairments due to motion blur, lens blur, pixel saturation, image compression, etc. Therefore, an image processing-related technical problem arises that requires removing one or more image impairments to produce a sharp image.

[0019] To remove one or more image impairments, the techniques described herein apply a model based on a convolutional neural network to predict a sharp image. The techniques described herein include receiving an input image, predicting an output image that is a sharper version of the input image using a convolutional neural network, and generating an output based on the output image. The input image and the output image may be high-resolution images, such as multi-megapixel images captured by a camera of a mobile computing device. The convolutional neural network can work well with input images captured under various shooting conditions, including different cameras, different subjects and scenes, different lighting environments, etc. In some examples, the trained convolutional neural network model can function on a variety of computing devices, including, but not limited to, mobile computing devices (e.g., smartphones, tablet computers, mobile phones, laptop computers), stationary computing devices (e.g., desktop computers), and server computing devices. The convolutional neural network can solve the technical problem of removing one or more image impairments and obtaining a sharper version (e.g., higher visual quality) of an already acquired image by applying an image transformation model to the input image.

[0020] A neural network, such as a convolutional neural network, may be trained using a training dataset of images to perform one or more aspects as described herein. In some examples, the neural network may be configured as an encoder / decoder neural network.

[0021] As described herein, Deep Motion, Out-of-focus, and degradation enhancement (DeepMode) models can be applied to difficult cases where there is moderate / high blur and where the image exhibits other degradation such as noise or JPEG compression artifacts.

[0022] DeepMode can be configured to be a supervised deep learning end-to-end solution that eliminates image blur, noise, compression artifacts, and more.

[0023] In one example, (a copy of) the trained neural network can reside within a mobile computing device. The mobile computing device may include a camera capable of capturing an input image. A user of the mobile computing device may view the input image and determine that the input image should be sharpened. The user may then provide the input image to a trained neural network residing within the mobile computing device. In response, the trained neural network may generate a predicted output image that is a sharper version of the input image and then output the output image (e.g., provide the output image for display by the mobile computing device). In another example, the trained neural network does not reside within the mobile computing device. Rather, the mobile computing device provides the input image to a remotely located trained neural network (e.g., via the Internet or other data network). The remotely located convolutional neural network can process the input image and provide an output image to the mobile computing device that is a sharper version of the input image. In other examples, non-mobile computing devices may also use the trained neural network to sharpen images, including images not captured by the computing device's camera.

[0024] In some examples, the trained neural network can work in conjunction with other neural networks (or other software) and / or can be trained to recognize whether an input image has image degradation. Then, upon determining that the input image has image degradation, the trained neural network described herein can apply the trained neural network to thereby remove the image degradation in the input image.

[0025] Thus, the techniques described herein can improve images by removing image degradation, thereby increasing their actual and / or perceived quality. Improving the actual and / or perceived quality of images, including portraits, can provide emotional benefits to those who believe their photos look better. These techniques are flexible and therefore can be applied to images of human faces, as well as other subjects, scenes, etc.

[0026] Network Architecture FIG. 1A illustrates an exemplary image enhancement process 100A, according to an exemplary embodiment. In this example, the relationship between a synthetic data generation pipeline 105, a training stage 110, and an inference stage 115 is shown. Details of each of these different aspects are described in further detail below. Some embodiments include receiving, by a computing device, multiple training datasets corresponding to each of multiple degradation factors. Each training dataset may include multiple pairs of a sharp image and a corresponding synthetically degraded version of the sharp image. For example, the synthetic data generation pipeline 105 may be configured to generate synthetic data including images with image degradation. For example, the image degradation may be synthetically introduced into the sharp image based on multiple degradation factors. For example, the synthetic data generation pipeline 105 may be configured to generate multiple training datasets 120-1, 120-2, ..., 120-N. In some embodiments, each training dataset may correspond to a particular degradation factor (and different training datasets may correspond to different degradation factors). Each training dataset may include multiple pairs of a sharp image and a corresponding synthetically degraded version of the sharp image.

[0027] As used herein, the term "degradation factor" generally refers to any factor that affects the sharpness of an image, e.g., image clarity with respect to quantitative image quality parameters such as contrast, focus, etc. In some embodiments, the degradation factors may include one or more of motion blur, lens blur, image noise, image compression artifacts, or artifacts caused by saturated pixels.

[0028] As used herein, the term "motion blur" generally refers to a detrimental effect that causes one or more objects in an image to appear fuzzy and / or unclear due to the movement of a camera capturing the image, the movement of one or more objects, or a combination of the two. In some examples, motion blur may be perceived as streaking or smearing in an image. As used herein, the term "lens blur" generally refers to a detrimental effect that causes an image to appear to have a narrower depth of field than the scene being captured. For example, certain objects in an image may be in focus, while other objects may appear out of focus.

[0029] As used herein, the term "image noise" generally refers to a degradation factor that causes an image to appear to have artifacts (e.g., speckles, color dots, etc.) due to a low signal-to-noise ratio (SNR). For example, an SNR below a certain desired threshold may cause image noise. In some examples, image noise may be caused by the camera's image sensor or circuitry. As used herein, the term "image compression artifacts" generally refers to a degradation factor that results from lossy image compression. For example, image data may be lost during compression, resulting in visible artifacts in a decompressed version of the image.

[0030] As used herein, the term "saturated pixel" generally refers to a state in which a pixel is saturated with photons and leaks photons into neighboring pixels. For example, a saturated pixel may be associated with an image intensity greater than a threshold intensity (e.g., an image intensity greater than 245, or an image intensity of 255, etc.). The image intensity may correspond to a grayscale intensity or the intensity of a red, blue, or green (RGB) color component. For example, a highly saturated pixel may appear brightly colored. Thus, photon leakage from a saturated pixel into neighboring pixels may cause perceptual defects in the image (e.g., causing saturation of one or more neighboring pixels, distorting the intensity of one or more neighboring pixels, etc.).

[0031] Generating different training datasets (e.g., multiple training datasets 120-1, 120-2, ..., 120-N) enables a tailored approach for training machine learning models to remove specific image degradations. For example, during training phase 110, multiple intermediate machine learning models 125-1, 125-2, ..., 125-N may be trained to remove one or more image degradations associated with a given image. For example, during training phase 110, each intermediate machine learning model of the multiple intermediate machine learning models 125-1, 125-2, ..., 125-N may be trained with a respective training dataset of the multiple training datasets 120-1, 120-2, ..., 120-N to remove one or more image degradations associated with the given image based on each of multiple degradation factors.

[0032] Some embodiments include training an image transformation model with an additional training dataset of real images to remove one or more image degradations associated with the real images, where the image transformation model learns from multiple intermediate machine learning models. For example, the image transformation model 135 may be trained based on the additional training dataset 130 of real images by learning from multiple trained intermediate machine learning models 125-1, 125-2, ..., 125-N. The additional training dataset 130 may include real images with one or more image degradations caused by one or more degradation factors. Because each intermediate machine learning model is trained to remove image degradations caused by a particular degradation factor, the image transformation model 135 may be trained to remove any image degradations in the real images from the additional training dataset 130 by leveraging the multiple trained intermediate machine learning models 125-1, 125-2, ..., 125-N.

[0033] Some embodiments include outputting the trained image transformation model by a computing device for removal of image artifacts. For example, upon completion of the training stage 110, the trained image transformation model 135 can be provided for use in, for example, the inference stage 115.

[0034] Some embodiments include receiving, by a computing device, an input image that includes one or more image degradations. For example, during the inference stage 115, a degraded input image 135 may be received. As used herein, the term "degraded image" generally refers to an image that has one or more image degradations caused by one or more degradation factors.

[0035] Some embodiments include predicting a transformed version of an input image with an image transformation model, where the image transformation model is trained to remove one or more image degradations associated with the input image, and the training includes: (1) training a plurality of intermediate machine learning models to remove the one or more image degradations, where each intermediate machine learning model of the plurality of intermediate machine learning models is trained on a respective training dataset of a plurality of training datasets corresponding to each of a plurality of degradation factors, each training dataset including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image; and (2) the image transformation model is trained on an additional training dataset of real images and learns from the plurality of intermediate machine learning models.

[0036] For example, the trained transformation neural network 140 may receive as input the degraded input image 135. As described above, the trained transformation neural network 140 may have been trained during the training stage 110 to remove one or more image degradations associated with the degraded input image 135, which may include training a plurality of intermediate machine learning models 125-1, 125-2, ..., 125-N to remove the one or more image degradations, wherein each intermediate machine learning model of the plurality of intermediate machine learning models 125-1, 125-2, ..., 125-N is trained with a respective training dataset of a plurality of training datasets 120-1, 120-2, ..., 120-N corresponding to each of a plurality of degradation factors, each training dataset including a plurality of pairs of a sharpened image and a corresponding synthetically degraded version of the sharpened image.

[0037] Some embodiments include providing, by a computing device, a predicted transformed version of the input image. For example, the trained transformation neural network 140 can transform the degraded input image 135 to predict a transformed version of the enhanced output image 145. Generally, the enhanced output image 145 represents a version of the degraded input image 135 in which one or more image degradations associated with the degraded input image 135 have been removed.

[0038] In some embodiments, the intermediate machine learning models may be teacher networks, and the image transformation model may be a student network, and the image transformation model learns from the intermediate machine learning models based on knowledge distillation. Generally, knowledge distillation is a training process between a teacher and a student that can be used to train a smaller (or lighter), generally more accurate network based on a heavier pre-trained teacher network. For example, the heavier pre-trained teacher network may reside in a cloud server (typically associated with more computational resources), and the smaller (or lighter) network may be configured to run on a mobile device (typically with associated computational resource limitations). For example, the intermediate machine learning models 125-1, 125-2, ..., 125-N may be teacher networks, and the image transformation model 145 may be a student network trained with each of the outputs of the intermediate machine learning models 125-1, 125-2, ..., 125-N. In some embodiments, the multiple intermediate machine learning models 125-1, 125-2, ..., 125-N may be of different sizes, and the image transformation model 145 may be trained to optimize prediction quality with shorter prediction runtimes.

[0039] One or more of the intermediate machine learning models or the image transformation model may be a neural network. For example, a deep neural network model such as a convolutional neural network (CNN), a generative adversarial network (GAN), a feature space expansion and autoencoder model, or a meta-learning model may be used.

[0040] FIG. 1B illustrates an exemplary network architecture 100B according to an exemplary embodiment. In some embodiments, each intermediate machine learning model of the plurality of intermediate machine learning models may be an encoder-decoder neural network with skip connections. In some embodiments, each encoder-decoder subnetwork may include three encoder stages and three decoder stages. In some embodiments, each encoder stage may include a feature extraction module and two nonlinear transformation modules. In some embodiments, each decoder stage may include two nonlinear transformation modules and a feature reconstruction module. In some embodiments, a given image may be processed at different scales, and the output of a given scale may be upsampled and concatenated with the input of a successive scale.

[0041] For example, DeepMode may be a multi-scale encoder-decoder network with skip connections. The network can process input images 150 at different scales, with the output of a given scale being upsampled and concatenated with the input of the subsequent scale. For example, the output scale N, denoted as 170-N for the Nth scale, is upsampled and concatenated with the input scale N-1 (not shown), the output scale N-1 (not shown) for the N-1th scale is upsampled and concatenated with the input scale N-2 (not shown), and so on, until the output scale 2 (not shown) for the second scale is upsampled and concatenated with the input scale 1 (denoted as 153-1).

[0042] At each scale, input image 150 may be resized to generate resized scale 1, labeled 152-1, through resized scale N, labeled 152-N. Also, at each scale, each encoder-decoder sub-network includes three encoder stages (e.g., Enc1 at scale 1, labeled 156-1, through Enc1 at scale N, labeled 156-N; Enc2 at scale 1, labeled 158-1, through Enc2 at scale N, labeled 158-N; and Enc3 at scale 1, labeled 158-1, through Enc3 at scale N, labeled 158-N). In some embodiments, each encoder stage may include a feature extraction module and two nonlinear transformation modules, as indicated by legend 176.

[0043] Also, at each scale, each encoder-decoder subnetwork includes three decoder stages (e.g., Dec1 at scale 1 labeled 166-1 through Dec1 at scale N labeled 166-N, Dec2 at scale 1 labeled 164-1 through Dec2 at scale N labeled 164-N, and Dec3 at scale 1 labeled 162-1 through Dec3 at scale N labeled 162-N). In some embodiments, each decoder stage may include a feature reconstruction module and two nonlinear transform modules. In some embodiments, one or more of Dec3 at scale 1 labeled 162-1 through Dec3 at scale N labeled 162-N may include two nonlinear transform modules without a feature reconstruction module.

[0044] Also, for example, at each scale, the neural network may include one or more skip connections. For example, skip connection 172-1 may connect the last nonlinear transformation module of Enc2 at scale 1, labeled 158-1, with each of the first nonlinear transformation modules of Dec2 at scale 1, labeled 164-1, and so on, until skip connection 172-N connects the last nonlinear transformation module of Enc2 at scale N, labeled 158-N, with each of the first nonlinear transformation modules of Dec2 at scale N, labeled 164-N.

[0045] Similarly, skip connection 174-1 may connect the last nonlinear transformation module of Enc1 at scale 1, labeled 156-1, with each of the first nonlinear transformation modules of Dec1 at scale 1, labeled 166-1, and so on, until skip connection 174-N connects the last nonlinear transformation module of Enc1 at scale N, labeled 156-N, with each of the first nonlinear transformation modules of Dec1 at scale N, labeled 166-N.

[0046] In some embodiments, each encoder / decoder block may include a high-order residual convolution block. This architecture allows for significant computational efficiency improvements and reduced memory footprint.

[0047] In some embodiments, each intermediate machine learning model may include one or more spatial-depth (s2d) layers. For example, s2d layers 154-1, ..., 154-N may be included at each scale. Each s2d layer may enable low-resolution processing by reducing the spatial resolution of the input while increasing the number of channels for acquiring and preserving information. Also, for example, each intermediate machine learning model may include one or more depth-spatial (d2s) layers corresponding to one or more s2d layers. For example, d2s layers 168-1, ..., 168-N may be included at each scale. The s2d layers 154-1, ..., 154-N and d2s layers 168-1, ..., 168-N may be configured to accelerate processing and reduce memory usage. For example, DeepMode architecture 100B may be configured to include a spatial-depth (depth-spatial) layer that takes inputs and rearranges them as tensors with lower spatial resolution but more channels (e.g., information is preserved). Core processing may be performed at a lower resolution (e.g., 1 / 2 to 1 / 4 of the original input resolution). Some embodiments may be configured without one or more of s2d layers 154-1,..., 154-N and d2s layers 168-1,..., 168-N.

[0048] In some embodiments, the upsampled convolution layer may be implemented as a transposed convolution and / or by performing bilinear upscaling followed by regular convolution. Generally speaking, when the network is integer quantized, bilinear upscaling followed by regular convolution can reduce the total memory footprint and the amount of (gridding) artifacts.

[0049] In some embodiments, the modules of the first encoder stage (e.g., Enc1 at scale 1 labeled 156-1 to Enc1 at scale N labeled 156-N) and the first decoder stage (e.g., Dec1 at scale 1 labeled 166-1 to Dec1 at scale N labeled 166-N) may have filters of size 32.

[0050] In some embodiments, the modules of the second encoder stage (e.g., Enc2 at scale 1 labeled 158-1 to Enc2 at scale N labeled 158-N) and the second decoder stage (e.g., Dec2 at scale 1 labeled 164-1 to Dec2 at scale N labeled 164-N) may have filters of size 64.

[0051] In some embodiments, the modules of the third encoder stage (e.g., Enc3 at scale 1 labeled 158-1 to Enc3 at scale N labeled 158-N) and the third decoder stage (e.g., Dec3 at scale 1 labeled 162-1 to Dec3 at scale N labeled 162-N) may have filters of size 128.

[0052] In some embodiments, each encoder-decoder layer may perform a 3x3 convolution followed by a LeakyReLU activation. In some embodiments, the encoder may utilize a blur-pooling layer for downsampling, while the decoder may utilize bilinear resizing followed by a 3x3 convolution for upsampling.

[0053] In some embodiments, parameter sharing may be performed across different scales. For example, for the encoder stages, a scale regression structure may be used in which the same parameters are shared across scales. For example, the feature extraction modules in the first, second, and third encoder stages may share parameters across scales. Similarly, the first nonlinear transformation modules in the first, second, and third encoder stages may share parameters across scales. Similarly, the second nonlinear transformation modules in the first, second, and third encoder stages may share parameters across scales.

[0054] In some embodiments, the modified scale regression structure may be used with independent feature extraction modules. In such embodiments, similar to the scale regression structure, the first nonlinear transformation modules in the first, second, and third encoder stages may share parameters across scales. Similarly, the second nonlinear transformation modules in the first, second, and third encoder stages may share parameters across scales.

[0055] In some embodiments, the modified scale regression structure may be further modified and the feature extraction modules remain independent, however, both the first and second nonlinear transformation modules may share parameters within a stage and also between scales.

[0056] In some embodiments, each intermediate machine learning model of the plurality of intermediate machine learning models may be associated with a respective number of filters and a respective spatial-depth parameter. Various base configurations spanning different model capabilities and computational resources may be determined. For example, a light model may be configured with four spatial-depth parameters, 28 or 32 filters, and a residual convolution block of order 5, abbreviated as s4-f[28,32]-r5. In some embodiments, an image transformation model (e.g., image transformation model 140) may be configured as a light model. As another example, a base model may be configured with two spatial-depth parameters, 16 filters, and a residual convolution block of order 5, abbreviated as s2-f16-r5. Also, for example, a heavy model may be configured with two spatial-depth parameters, 32 filters, and a residual convolution block of order 5, abbreviated as s2-f32-r5. In some embodiments, the intermediate machine learning models (eg, the intermediate machine learning models 125-1, 125-2, . . . , 125-N) may be configured as a heavy model.

[0057] In some embodiments, there may be benefits from using a "swish" or "hardswish" activation function as opposed to a "ReLu" or "ReLu6" activation function.

[0058] Table 1, for example, shows basic performance information for different DeepMode configurations for the float32 model. [Table 1]

[0059] 2A-2D show the training pipeline. Training may proceed in two steps. First, different DeepMode models may be trained and then fine-tuned on a curated dataset of carefully selected real images.

[0060] 2A illustrates an exemplary overview of a pipeline 200A for generating synthetic data, according to an exemplary embodiment. For example, based on a database 202 containing a plurality of sharp images and segmentation masks, different training datasets can be generated using different configurations in a synthetic data generation pipeline 204: training dataset 1, denoted as 206-1, training dataset 2, denoted as 206-2, and so on, up to training dataset N, denoted as 206-N.

[0061] Some embodiments include training a plurality of intermediate machine learning models to remove one or more image degradations associated with a given image, where each intermediate machine learning model of the plurality of intermediate machine learning models is trained with a respective training dataset of a plurality of training datasets to remove one or more image degradations associated with the given image based on each of a plurality of degradation factors.

[0062] 2B illustrates exemplary training 200B of multiple intermediate models according to an exemplary embodiment. For example, N different DeepModels with different configurations (e.g., number of filters, spatial-depth parameters, etc.) may be trained. For example, training of Model 1, denoted as 210-1, may be performed on Training Dataset 1, denoted as 206-1, and Model Configuration 1, denoted as 208-1, to generate Trained Intermediate Model 1, denoted as 212-1. As another example, training of Model 2, denoted as 210-2, may be performed on Training Dataset 2, denoted as 206-2, and Model Configuration 2, denoted as 208-2, to generate Trained Intermediate Model 2, denoted as 212-2. Also, for example, training of Model N, denoted as 210-N, may be performed on Training Dataset N, denoted as 206-N, and Model Configuration N, denoted as 208-N, to generate Trained Intermediate Model N, denoted as 212-N. In some embodiments, trained intermediate model 1, denoted at 212-1, trained intermediate model 2, denoted at 212-2, ..., trained intermediate model N, denoted at 212-N, serve as teachers to distill and / or fine-tune student models on datasets of real images (with actual degradation not simulated in the data generation pipeline).

[0063] Some embodiments include applying the multiple trained intermediate machine learning models to a given real image of an additional training dataset of real images to generate a corresponding output image. Such embodiments include selecting an optimally transformed version of the given real image from the generated output images. Such embodiments also include generating a curated dataset including pairs of real images and corresponding optimally transformed versions of the real images. In some embodiments, applying the multiple trained intermediate machine learning models to the given real image includes applying each trained intermediate machine learning model at multiple resolutions.

[0064] FIG. 2C illustrates an exemplary pipeline 200C for generating a curated dataset, according to an example embodiment. In some embodiments, a database of actual blurred images 216 may be used. For example, the database of actual blurred images 216 may include approximately 500 images with actual impairments (e.g., motion blur, lens blur, noise, compression artifacts, etc.). These images may not have corresponding high-quality counterparts. In some embodiments, multiple trained intermediate machine learning models may be applied to a given real image of an additional training dataset of real images (e.g., from the database of actual blurred images 216) to generate a corresponding output image. For example, trained intermediate model 1, denoted 212-1, trained intermediate model 2, denoted 212-2, ..., trained intermediate model N, denoted 212-N, may be run at multiple resolutions, e.g., as teacher models. For example, model inference 1, shown as 214-1, may be performed using a database of actual blurred images 216 together with trained intermediate model 1, shown as 212-1, to produce predicted result 1, shown as 218-1. Similarly, model inference N, shown as 214-N, may be performed using a database of actual blurred images 216 together with trained intermediate model N, shown as 212-N, to produce predicted result N, shown as 218-N.

[0065] In some embodiments, the method may include selecting an optimally transformed version of a given real image from the generated output images. For example, such operations may include performing a downscaling operation, running a respective intermediate machine learning model, and performing an upscaling operation to achieve the target resolution. In some embodiments, for each image in the database of real blurred images 216, image quality and evaluation 220 may be performed on all corresponding results output by each trained intermediate model, and the best output may be selected from among the results. Thus, each input image from the database of real blurred images 216 may be associated with the best output. For example, each low-quality input image in the database of real blurred images 216 may be associated with a restored image generated by one of the trained deep intermediate models. In some embodiments, image quality and evaluation 220 may be performed using Neural Image Assessment (NIMA) scores and / or by manual selection.

[0066] This process generates a curated (or gold standard) dataset 224 of pairs of actual degraded images (e.g., from the database of actual blurred images 216) and their restored high-quality counterparts (e.g., output of image quality and assessment 220) that are stored in a database of selected images 222. Thus, the curated dataset 224 includes pairs such as "(actual image, selected image)."

[0067] 2D is a diagram illustrating an example training 200D of an image transformation model, according to an example embodiment. As described herein, a curated dataset 224 may be generated based on a database of actual blurred images 216 and a database of selected images 222. In some embodiments, training the image transformation model includes fine-tuning the image transformation model based on the curated dataset. Fine-tuning 228 of the image transformation model 226 (e.g., a model having desired latency, memory usage, etc.) may be performed based on the curated dataset 224.

[0068] In such an embodiment, fine-tuning 228 of image translation model 226 may include quantization-aware training (QAT), and, for example, a trained image translation model 230 may be generated and exported, as needed, for example, as a TensorFlow Lite integer model.

[0069] In some embodiments, each base DeepMode model may be trained using an Adaptive Estimation of Moments (ADAM) optimizer with default parameters. In some embodiments, the learning rate may be set to 1e-4 and a polynomial decay (e.g., coefficient = 1.0) may be applied. Also, for example, in some embodiments, the full optimization process may be run for 1.5M iterations. In some embodiments, a mini-batch size of 8 may be used.

[0070] In some embodiments, the model can be trained using an L1 loss function that penalizes pixel value differences at multiple output scales. Also, for example, a penalty for gradient differences between the target and the reconstruction may be applied. As another example, a penalized L1 norm of gradient differences may be applied. Additional and / or alternative advanced loss functions, such as polarization dependent loss (PDL), may be applied.

[0071] In some embodiments, training the same model multiple times in the same way but with different random initializations may yield different results on real data. This may be despite the trained models reaching similar accuracy on training and / or evaluation data (data from the synthetic data generation pipeline). Therefore, training and fine-tuning on real images may be desirable.

[0072] Training Data In general, obtaining ground truth degradation-sharpening pair data for use in supervised blind deblurring (and image restoration in general) can be challenging: for example, complex setups may be required to obtain paired photographs using different exposure times, different camera configurations, etc.

[0073] Some existing techniques that enable the synthesis of motion blur (e.g., camera or subject motion blur) may be based on two procedures. For example, emulating a (virtual) longer exposure time can make the image appear naturally blurred. This can be done by capturing high-frame-rate video and then averaging a window of consecutive frames, with a larger time window controlling the increased exposure (and therefore more blur in the capture). However, such techniques have certain limitations. For example, to accurately simulate a longer exposure, frames need to be captured for the entire inter-frame time (e.g., the camera always captures light continuously). However, this is not the case in reality. Therefore, direct averaging of consecutive frames can cause ghost-type artifacts when the motion is larger than a pixel. A common solution to avoid this problem may be to interpolate multiple intermediate frames before calculating the temporal average. However, this may involve solving the technically difficult problem of frame temporal interpolation.

[0074] Another procedure may be to simulate motion blur and other degradations of high-quality images. Such an approach may be difficult because different objects and / or different parts of a scene may undergo different motion and therefore may produce different blur. Also, for example, camera shake may cause blur that is not shift-invariant.

[0075] Therefore, as explained below, multiple impairments may be synthetically generated on a reference high quality frame.

[0076] Data Generation Pipeline FIG. 3 illustrates an exemplary pipeline 300 for generating synthetic data, according to an example embodiment. In some embodiments, as shown in block 305, a random crop of an image and an associated random segmentation (e.g., a binary mask) may be generated. Next, gamma correction values ​​may be sampled, and a corresponding inverse gamma correction 310 may be applied. A brightness shift 315 may be applied by randomly shifting pixel values ​​by a constant factor (e.g., for data augmentation). Saturated or nearly saturated pixels may then be randomly extrapolated to make them desaturated 320. Motion blur 325 and generalized Gaussian blur 330 may be simulated. For example, these operations may be performed segment by segment or on the complete crop. Also, for example, noise 335 (e.g., Poisson, Gaussian, white, colored, etc.) may be added, and gamma correction 340 may be applied. In some embodiments, JPEG compression 345 (e.g., with a random quality factor sampled from a predefined distribution) may be applied to the blurred image to generate a blurred crop 350.

[0077] Random Motion Blur Kernel In some embodiments, the multiple degradation factors may include motion blur. Such embodiments include generating one or more motion kernels that simulate hand tremor, where the amount of motion blur is associated with a scale parameter related to exposure time. For example, the motion kernel may be simulated by mimicking hand tremor. The amount of blur (e.g., expected value) may be controlled by a parameter related to exposure time, referred to herein as "scale." In some embodiments, the motion kernel may be generated from a discretized two-dimensional (2D) random walk. For example, to synthesize the motion kernel, a small Gaussian kernel (e.g., with a standard deviation = 0.35) may be rendered and accumulated at each discrete position of the random walk. In some embodiments, the kernel may be non-negative by construction and normalized to have area 1 (e.g., light-weight conservation).

[0078] In some embodiments, the strength of the blur may be controlled by a random walk characteristic. To render a kernel that is generally smooth, but may also cause certain variations, the following procedure may be applied:

[0079] Let N be the length of the random walk. Then, 2 × N independent and identically distributed Gaussian variables may be sampled from the normal distribution N(0,Id), resulting in the horizontal and vertical components of the random walk. Such random variables typically represent the acceleration at each discrete time step.

[0080] In some embodiments, to generate a sufficiently smooth random walk, a 2×N vector can be integrated ("accumulated") twice along each dimension.

[0081] In some embodiments, to represent different intensities of blur, the value of the random walk vector may be multiplied by a factor (scale / N). Generally, a larger factor corresponds to a larger movement. In some embodiments, this factor may be a random value sampled from a gamma distribution.

[0082] In some embodiments, the first half of the random walk can be removed to avoid temporal issues. Also, the remainder can be centered, for example, by subtracting the mean of each dimension. Thus, a kernel with a centered first moment can be generated (i.e., the motion kernel does not shift the image).

[0083] In some embodiments, the random kernel may be generated by rasterizing a random walk as described above.

[0084] In practice, fixing N to a large value (e.g., N=300) and controlling the "scale" of a single parameter may provide a practical way to simulate different intensities of blur while minimizing the number of adjustable parameters.

[0085] The support of the kernel may be chosen to be large enough to contain the entire kernel, in some embodiments, the kernel may then be trimmed to the minimum value that contains the support to accelerate the filtering process.

[0086] Approximate kernel length (motion blur strength) The kernel length may be calculated as the inverse of the L2 norm of the kernel. The motivation for this may be for a perfect uniform motion kernel of length L.

number

[0087] In some embodiments, the L2 norm of the kernel measures the spread of the non-negative blur kernel.

number

[0088] The limiting case, L2(k)=1, occurs when the kernel is delta (sharpest), and as the kernel gets larger, L2(k)→0.

[0089] 4 shows exemplary motion blur kernels for different scale parameters, according to an exemplary embodiment. For example, the images in column 4C1 correspond to a scale parameter s=0.5, the images in column 4C2 correspond to a scale parameter s=1.0, the images in column 4C3 correspond to a scale parameter s=1.5, the images in column 4C4 correspond to a scale parameter s=2.0, the images in column 4C5 correspond to a scale parameter s=2.5, and the images in column 4C6 correspond to a scale parameter s=3.0.

[0090] Table 2 summarizes the statistics of the kernels generated as a function of the "scale" parameter for different motion levels. In some embodiments, results may be generated with approximately 1000 sampled kernels per scale. [Table 2]

[0091] 5 shows exemplary input and output images corresponding to simulated motion blur, according to an exemplary embodiment. For example, the image in column 5C1 corresponds to the sharp image in column 5C2 modified by simulated motion blur as described herein. For example, the image of a pretzel in row 5R1 and column 5C1 corresponds to a synthetically blurred version of the corresponding image in row 5R1 and column 5C2.

[0092] FIG. 6 shows an example histogram relating to motion kernel length as a function of the scale parameter, according to an example embodiment. The illustrated histogram may, in some embodiments, be based on results generated with approximately 1000 sampled kernels per scale. The distribution of (approximate) lengths of kernels generated for different "scale" parameters may be a known quantity. In some embodiments, the distribution of generated motion blur kernels may be important in the learning process for deblurring images. The generated distribution may include images with different levels of blur (e.g., sharp blur, moderate blur, and very blurry blur) for the network to understand the deblurring problem. As shown, histogram 605 corresponds to a scale parameter s=0.5, histogram 610 corresponds to a scale parameter s=1.0, histogram 615 corresponds to a scale parameter s=1.5, histogram 620 corresponds to a scale parameter s=2.0, histogram 625 corresponds to a scale parameter s=2.5, and histogram 630 corresponds to a scale parameter s=3.0.

[0093] Generalized Gaussian Blur In some embodiments, the plurality of degradation factors may include lens blur. Such embodiments include generating one or more generalized Gaussian blur kernels that simulate lens blur. The generalized Gaussian blur may be used to simulate lens blur and / or mild defocus effects. In some embodiments, a parametric generalized Gaussian blur kernel may be used. The shape and size of the parametric generalized Gaussian blur kernel may be controlled by four parameters: σ, ρ, θ, and γ.

number

[0094] where:

number

[0095] As explained, this can be a fairly general model for mild lens blur, however, it may not capture camera-dependent bokeh effects, since the kernel for such bokeh effects corresponds to a donut-shaped kernel.

[0096] 7 illustrates exemplary generalized Gaussian blur kernels according to an exemplary embodiment. The generalized Gaussian kernels shown in different columns correspond to different gamma values. For example, the kernel in column 7C1 corresponds to a gamma value of γ=0.3, the kernel in column 7C2 corresponds to a gamma value of γ=0.5, the kernel in column 7C3 corresponds to a gamma value of γ=1.0, the kernel in column 7C4 corresponds to a gamma value of γ=1.5, and the kernel in column 7C5 corresponds to a gamma value of γ=2.0. Also shown are random orientations (theta, θ) and non-unit eccentricities (rho, ρ), and an increase in size from left to right across each row.

[0097] Image noise / JPEG compression artifacts In some embodiments, the multiple degradation factors include one or more of additive noise, signal-dependent noise, or colored noise. Image noise in RGB space can be very complex because it is signal-dependent and can be signal-correlated (e.g., colored). Also, for example, the noise fingerprint can be highly dependent on the camera sensor and / or the camera's ISP, and other post-processing (e.g., denoising, sharpening, resizing, etc.).

[0098] In general, DeepMode can be configured to transform any image, so that for training purposes, general noise models can be simulated.

[0099] In some embodiments, additive noise and signal dependent noise may be simulated that can mimic read noise and shot noise.

[0100] In some embodiments, colored (e.g., correlated) noise can be simulated by filtering the noise layer with a Gaussian blur kernel of random standard deviation and adding it back into the noise layer (e.g., after multiplying with a random coefficient).

[0101] In some embodiments, the plurality of degradation factors includes one or more image compression artifacts. Generally, the shared image may be compressed with JPEG compression. There are several ways to implement the JPEG standard. In particular, the JPEG quality factor may affect the quantization of discrete cosine transform (DCT) coefficients and / or may affect the way chroma subsampling is performed. Thus, in a synthetic data generation pipeline, if a quality factor of 90 or higher is used, greater control of image generation can be achieved without performing chroma subsampling. Alternatively, standard 420 subsampling may be employed.

[0102] 8 shows exemplary input and output images corresponding to simulated compression artifacts, according to an exemplary embodiment. JPEG compression may affect the image by introducing artifacts, but JPEG compression may also introduce noise correlation. The reference image is shown in column 8C1, the noisy image is shown in column 8C2, and the image with colored noise is shown in column 8C3. The final row 8R3 shows the difference (amplified by 5x) between the first row 8R1 and the second row 8R2.

[0103] The synthetic data generation pipeline may be configured to enable compression of blurred images using a random quality factor sampled from a uniform distribution, where the two extremes are configurable parameters. Also, JPEG compression may introduce perturbations to the employed noise distribution.

[0104] non-saturated In some embodiments, the plurality of degradation factors includes one or more artifacts caused by saturated pixels. Saturated pixels are problematic because blurring effects can spill over into neighboring pixels. However, a reference frame that includes saturated pixels does not provide an indication of the actual radiance of the saturated pixels in the reference frame. Some embodiments include determining whether a pixel value exceeds a saturation threshold. For example, a saturation threshold may be determined. Such embodiments also include multiplying the pixel value by a random coefficient (sampled from a predetermined uniform interval) based on the determination that the pixel value exceeds the saturation threshold. Pixel values ​​that exceed the threshold may be multiplied by a random coefficient (sampled from a predetermined uniform interval).

[0105] Thus, when blur is applied to a sharp, non-saturated image that may later be clipped to mimic sensor saturation, saturated pixels will contribute more than saturated values, resulting in a more realistic blur simulation.

[0106] In general, it may be desirable to use the saturation threshold conservatively (e.g., not too far from the maximum value (i.e., 1.0 or 255)). In situations with a large margin, simulated training data can produce models with thin saturated regions.

[0107] In some embodiments, images from a database such as the Open Image Dataset (OID) may be used to synthetically generate blurred-sharpened image pairs. To use high-quality images as input, images may be filtered based on blur and NIMA score. Images from the OID include object segmentation that can be used to render different blurs to different segments. To simplify the synthetic data generation pipeline, images from the OID, such as .jpg and / or .png files, may be used directly with their respective segmentation masks.

[0108] In some embodiments, approximately 7000 images from the OID dataset with a NIMA score of 7 or greater and a polysharp blur score of less than 0.65 can be used. Semantic-aware sampling may not be necessary.

[0109] A large number of (visually) different training datasets can be designed and evaluated, spanning various compromises between blur strength, noise distribution, among many other parameters. The visual quality of the results depends on the capabilities of the model. Because the training datasets are synthetically generated, different models with different capabilities may produce optimal results using different training datasets.

[0110] result 9 shows example images corresponding to training data generation according to an example embodiment. The image in column 9C1 is a synthetically degraded version of the image in column 9C2. Similarly, the image in column 9C3 is a synthetically degraded version of the image in column 9C4.

[0111] 10 shows exemplary output images of a trained image transformation model according to an exemplary embodiment. The images in column 10C2 are transformed (predicted) versions of the images in column 10C1 that are input to the trained image transformation model described herein.

[0112] As described herein, DeepMode is an end-to-end solution for removing image degradation such as blur, noise, and compression artifacts in images trained on pairs of synthetically generated degraded images and clean, sharp images. Training a machine learning model to generate inferences / predictions

[0113] FIG. 11 shows a diagram 1100 illustrating the training stage 1102 and inference stage 1104 of trained machine learning model(s) 1132, according to an example embodiment. Machine learning techniques include training one or more machine learning algorithms with an input set of training data to recognize patterns in the training data and output inferences and / or predictions about (the patterns in) the training data. The resulting trained machine learning algorithms may be referred to as trained machine learning models. For example, FIG. 11 shows the training stage 1102 in which one or more machine learning algorithms 1120 are trained with training data 1110 to become trained machine learning model(s) 1132. Then, during the inference stage 1104, the trained machine learning model(s) 1132 can receive input data 1130 and one or more inference / prediction requests 1140 (perhaps as part of the input data 1130) and, in response, provide one or more inferences and / or prediction(s) 1150 as output.

[0114] Thus, the trained machine learning model(s) 1132 may include one or more models of one or more machine learning algorithms 1120. The machine learning algorithm(s) 1120 include, but are not limited to, artificial neural networks (e.g., convolutional neural networks described herein, recurrent neural networks, Bayesian networks, hidden Markov models, Markov decision processes, logistic regression functions, support vector machines, suitable statistical machine learning algorithms, and / or heuristic machine learning systems). The machine learning algorithm(s) 1120 may be supervised or unsupervised and may perform any suitable combination of online and offline learning.

[0115] In some examples, the machine learning algorithm(s) 1120 and / or the trained machine learning model(s) 1132 may be accelerated using on-device coprocessors, such as graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), and / or application-specific integrated circuits (ASICs). Such on-device coprocessors may be used to accelerate the machine learning algorithm(s) 1120 and / or the trained machine learning model(s) 1132. In some examples, the trained machine learning model(s) 1132 may be trained to provide inference on, reside on, execute, and / or otherwise perform inference for a particular computing device.

[0116] During the training phase 1102, the machine learning algorithm(s) 1120 may be trained by providing at least the training data 1110 as training inputs using unsupervised, supervised, semi-supervised, and / or reinforcement learning techniques. Unsupervised learning involves providing some (or all) of the training data 1110 to the machine learning algorithm(s) 1120, where the machine learning algorithm(s) 1120 determine one or more output inferences based on the provided portion (or all) of the training data 1110. Supervised learning involves providing some (or all) of the training data 1110 to the machine learning algorithm(s) 1120, where the machine learning algorithm(s) 1120 determine one or more output inferences based on the provided portion (or all) of the training data 1110, where the output inference(s) are either accepted or modified based on the correct results associated with the training data 1110. In some examples, the supervised learning of the machine learning algorithm(s) 1120 may be controlled by a set of rules and / or a set of labels for the training inputs, and the set of rules and / or the set of labels may be used to modify the inference of the machine learning algorithm(s) 1120.

[0117] Semi-supervised learning involves having correct results for some, but not all, of the training data 1110. During semi-supervised learning, supervised learning is used for portions of the training data 1110 that have correct results, and unsupervised learning is used for portions of the training data 1110 that do not have correct results. Reinforcement learning includes machine learning algorithm(s) 1120 receiving a reward signal related to a prior inference, where the reward signal may be a numerical value. During reinforcement learning, the machine learning algorithm(s) 1120 can output an inference and receive a reward signal in response, where the machine learning algorithm(s) 1120 are configured to attempt to maximize the numerical value of the reward signal. In some examples, reinforcement learning also utilizes a value function that provides a numerical value representing the expected sum of the numerical values ​​provided by the reward signal over time. In some examples, the machine learning algorithm(s) 1120 and / or the trained machine learning model(s) 1132 can be trained using other machine learning techniques, including, but not limited to, incremental learning and curriculum learning.

[0118] In some examples, the machine learning algorithm(s) 1120 and / or the trained machine learning model(s) 1132 can use a transfer learning approach. For example, a transfer learning approach may include the trained machine learning model(s) 1132 being pre-trained on a set of data and additionally trained using training data 1110. More specifically, the machine learning algorithm(s) 1120 may be pre-trained on data from one or more computing devices, and the resulting trained machine learning model(s) may be provided to the computing device CD1, which is intended to perform the trained machine learning during the inference stage 1104. Then, during the training stage 1102, the pre-trained machine learning model(s) may be additionally trained using training data 1110, which may be derived from kernel data and non-kernel data of the computing device CD1. This further training of the machine learning algorithm(s) 1120 and / or the pre-trained machine learning model(s) using the training data 1110 of the data of CD1 can be performed using either supervised learning or unsupervised learning. The training phase 1102 can be completed once the machine learning algorithm(s) 1120 and / or pre-trained machine learning model(s) have been trained on at least the training data 1110. The resulting trained machine learning model can be utilized as at least one of the trained machine learning model(s) 1132.

[0119] Specifically, once the training phase 1102 is complete, the trained machine learning model(s) 1132 may be provided to the computing device if not already on the computing device. The inference phase 1104 may begin after the trained machine learning model(s) 1132 are provided to the computing device CD1.

[0120] During the inference stage 1104, the trained machine learning model(s) 1132 can receive the input data 1130 and generate and output one or more corresponding inferences and / or prediction(s) 1150 regarding the input data 1130. Thus, the input data 1130 can be used as input to the trained machine learning model(s) 1132 to provide the inference(s) and / or prediction(s) 1150 corresponding to the kernel and non-kernel components. For example, the trained machine learning model(s) 1132 can generate the inference(s) and / or prediction(s) 1150 in response to one or more inference / prediction requests 1140. In some examples, the trained machine learning model(s) 1132 can be executed by another piece of software. For example, the trained machine learning model(s) 1132 can be executed by an inference or prediction daemon and readily available to provide inferences and / or predictions upon request. The input data 1130 may include data from the computing device CD1 running the trained machine learning model(s) 1132 and / or input data from one or more computing devices other than CD1.

[0121] The input data 1130 can include training data described herein, such as real blurred images, synthetically generated images, images in a curated dataset, etc. Other types of input data are possible as well. For example, the training data may include data collected to train an image transformation model.

[0122] The inference(s) and / or prediction(s) 1150 can include task outputs, numerical values, and / or other output data generated by the trained machine learning model(s) 1132 operating on the input data 1130 (and training data 1110). In some examples, the trained machine learning model(s) 1132 can use the output inference(s) and / or prediction(s) 1150 as input feedback 1160. The trained machine learning model(s) 1132 can also rely on past inferences as input to generate new inferences.

[0123] After training, the trained version of the neural network may be an example of trained machine learning model(s) 1132. In this approach, an example of one or more inference / prediction request(s) 1140 may be a request to predict a transformed (e.g., deblurred, denoised, etc.) image, and a corresponding example of inference and / or prediction(s) 1150 may be the predicted transformed (e.g., deblurred, denoised, etc.) image.

[0124] In some examples, one computing device CD_SOLO may include, perhaps after training, a trained version of the neural network. The computing device CD_SOLO may then receive a request to transform (e.g., deblur, denoise, etc.) an image and use the trained version of the neural network to predict the transformed (e.g., deblur, denoise, etc.) image.

[0125] In some examples, two or more computing devices CD_CLI and CD_SRV may be used to provide output. For example, a first computing device CD_CLI may generate a request to a second computing device CD_SRV to predict a transformed (e.g., deblurred, denoised, etc.) image. CD_SRV may then use a trained version of the neural network to predict the transformed (e.g., deblurred, denoised, etc.) image and respond to the request from CD_CLI. Then, upon receiving the response to the request, CD_CLI may provide the requested output (e.g., using a user interface and / or display, printed material, electronic communication, etc.).

[0126] Data Network Example 12 illustrates a distributed computing architecture 1200, according to an example embodiment. The distributed computing architecture 1200 includes server devices 1208, 1210 configured to communicate with programmable devices 1204a, 1204b, 1204c, 1204d, and 1204e via a network 1206. The network 1206 may correspond to a local area network (LAN), a wide area network (WAN), a WLAN, a WWAN, a corporate intranet, the public Internet, or any other type of network configured to provide a communication path between networked computing devices. The network 1206 may also correspond to a combination of one or more LANs, WANs, corporate intranets, and / or the public Internet.

[0127] While FIG. 12 shows only five programmable devices, the distributed application architecture can serve tens, hundreds, or thousands of programmable devices. Furthermore, programmable devices 1204a, 1204b, 1204c, 1204d, and 1204e (or any additional programmable devices) may be any type of computing device, such as a mobile computing device, a desktop computer, a wearable computing device, a head-mountable device (HMD), a network terminal, a mobile computing device, etc. In some implementations, such as illustrated by programmable devices 1204a, 1204b, 1204c, and 1204e, the programmable devices may be directly connected to network 1206. In other implementations, such as illustrated by programmable device 1204d, the programmable devices may be indirectly connected to network 1206 through an associated computing device, such as programmable device 1204c. In this example, programmable device 1204c can serve as an associated computing device for passing electronic communications between programmable device 1204d and network 1206. In other examples, such as shown by programmable device 1204e, the computing device may be part of and / or within a vehicle, such as a car, truck, bus, boat or watercraft, airplane, etc. In other examples not shown in Figure 12, the programmable device may be connected both directly and indirectly to the network 1206.

[0128] Server devices 1208, 1210 can be configured to perform one or more services as requested by programmable devices 1204a-1204e. For example, server devices 1208 and / or 1210 can provide content to programmable devices 1204a-1204e. The content can include, but is not limited to, web pages, hypertext, scripts, binary data such as compiled software, images, audio, and / or video. The content can include compressed and / or uncompressed content. The content may be encrypted and / or decrypted. Other types of content are possible as well.

[0129] As another example, server devices 1208 and / or 1210 may provide programmable devices 1204a-1204e with access to software for database, search, calculation, graphics, audio, video, World Wide Web / Internet utilization, and / or other functions. Many other examples of server devices are possible as well.

[0130] Computing Device Architecture 13 is a block diagram of an exemplary computing device 1300, according to an example embodiment. Specifically, the computing device 1300 shown in FIG. 13 may be configured to perform at least one function of and / or associated with the neural networks and / or methods 1500, 1600 described herein.

[0131] Computing device 1300 may include a user interface module 1301, a network communication module 1302, one or more processors 1303, data storage 1304, one or more camera(s) 1318, one or more sensors 1320, and a power system 1322, all of which may be linked together via a system bus, network, or other connection mechanism 1305.

[0132] The user interface module 1301 may be operable to transmit data to and / or receive data from external user input / output devices. For example, the user interface module 1301 may be configured to transmit data to and / or receive data from user input devices such as a touchscreen, a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a voice recognition module, and / or other similar devices. The user interface module 1301 may also be configured to provide output to a user display device such as one or more cathode ray tubes (CRTs), liquid crystal displays, light emitting diodes (LEDs), displays using digital light processing (DLP) technology, printers, light bulbs, and / or other similar devices, either now known or later developed. The user interface module 1301 may also be configured to generate audible output using devices such as speakers, speaker jacks, audio output ports, audio output devices, earphones, and / or other similar devices. User interface module 1301 may further comprise one or more haptic devices capable of generating haptic outputs, such as vibrations and / or other outputs detectable by touch and / or physical contact with computing device 1300. In some examples, user interface module 1301 may be used to provide a graphical user interface (GUI) for utilizing computing device 1300, such as, for example, the graphical user interface of a mobile phone device.

[0133] The network communication module 1302 may include one or more devices providing one or more wireless interface(s) 1307 and / or one or more wired interface(s) 1308 configurable to communicate over a network. The wireless interface(s) 1307 may include one or more wireless transmitters, receivers, and / or transceivers, such as a Bluetooth™ transceiver, a Zigbee™ transceiver, a Wi-Fi™ transceiver, a WiMAX™ transceiver, an LTE™ transceiver, and / or other types of wireless transceivers configurable to communicate over a wireless network. The wired interface(s) 1308 may include one or more wired transmitters, receivers, and / or transceivers, such as an Ethernet transceiver, a Universal Serial Bus (USB) transceiver, or similar transceiver configurable to communicate over twisted pair wire, coaxial cable, fiber optic link, or similar physical connection to a wired network.

[0134] In some examples, the network communications module 1302 can be configured to provide reliable, secure, and / or authenticated communications. For each communication described herein, information to facilitate reliable communications (e.g., guaranteed message delivery) may be provided, perhaps as part of the message header and / or footer (e.g., packet / message ordering information, encapsulation header and / or footer, size / time information, and transmission verification information such as a cyclic redundancy check (CRC) and / or parity check value). Communications may be protected (e.g., encoded or encrypted) and / or decrypted / decoded using one or more cryptographic protocols and / or algorithms, including, but not limited to, the Data Encryption Standard (DES), the Advanced Encryption Standard (AES), the Rivest-Shamir-Adelman (RSA) algorithm, the Diffie-Hellman algorithm, a secure socket protocol such as Secure Sockets Layer (SSL) or Transport Layer Security (TLS), and / or the Digital Signature Algorithm (DSA). Other cryptographic protocols and / or algorithms may be used to protect (and decrypt / decrypt) communications similarly to or in addition to those described herein.

[0135] The one or more processors 1303 may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., digital signal processors, tensor processing units (TPUs), graphics processing units (GPUs), application-specific integrated circuits, etc.) The one or more processors 1303 may be configured to execute computer-readable instructions 1306 contained in data storage 1304 and / or other instructions described herein.

[0136] Data storage 1304 may include one or more non-transitory computer-readable storage media that can be read and / or accessed by at least one of the one or more processors 1303. The one or more computer-readable storage media can include volatile and / or non-volatile storage components, such as optical, magnetic, organic, or other memory or disk storage, which may be integrated, in whole or in part, with at least one of the one or more processors 1303. In some embodiments, data storage 1304 may be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disk storage unit), while in other embodiments, data storage 1304 may be implemented using two or more physical devices.

[0137] Data storage 1304 may include computer-readable instructions 1306 and possibly additional data. In some examples, data storage 1304 may include storage necessary to execute at least some of the methods, scenarios, and techniques described herein and / or at least some of the functionality of the devices and networks described herein. In some examples, data storage 1304 may include storage for trained neural network models 1312 (e.g., models of trained neural networks, such as the neural networks described herein). In particular, in these examples, computer-readable instructions 1306 may include instructions that, when executed by one or more processors 1303, enable computing device 1300 to provide some or all of the functionality of trained neural network models 1312.

[0138] In some embodiments, computing device 1300 may include one or more camera(s) 1318. Camera(s) 1318 may include one or more image capture devices, such as still cameras and / or video cameras, equipped to capture light and record the captured light into one or more images. That is, camera(s) 1318 may generate image(s) of the captured light. The one or more images may be one or more still images and / or one or more images utilized in video capture. Camera(s) 1318 may capture light and / or electromagnetic radiation emitted as visible light, infrared radiation, ultraviolet light, and / or as light of one or more other frequencies.

[0139] In some examples, computing device 1300 may include one or more sensors 1320. Sensors 1320 may be configured to measure conditions within computing device 1300 and / or conditions in the environment of computing device 1300 and provide data regarding these conditions.For example, the sensors 1320 may include (i) sensors for acquiring data about the computing device 1300, such as, but not limited to, a thermometer for measuring the temperature of the computing device 1300, a battery sensor for measuring the power of one or more batteries in the power system 1322, and / or other sensors for measuring the condition of the computing device 1300; (ii) identification sensors for identifying other objects and / or devices, such as RFID tags, barcodes, QR codes, and / or other devices that can be configured to read identifiers and / or objects configured to be read, and provide at least identifying information, including, but not limited to, radio frequency identification (RFID) readers, proximity sensors, one-dimensional barcode readers, two-dimensional barcode (e.g., quick response (QR) code) readers, and laser trackers; and (iii) tilt sensors, gyroscopes, accelerometers, Doppler sensors, GPS devices, and sonar. The sensors 1320 may include one or more of the following: (i) sensors that measure the position and / or movement of the computing device 1300, such as, but not limited to, a laser sensor, a radar device, a laser displacement sensor, and a compass; (ii) environmental sensors that acquire data indicative of the environment of the computing device 1300, such as, but not limited to, an infrared sensor, an optical sensor, a biosensor, a capacitive sensor, a touch sensor, a temperature sensor, a wireless sensor, a radio sensor, a movement sensor, a microphone, a sound sensor, an ultrasonic sensor, and / or a smoke sensor; and / or (iii) force sensors that measure one or more forces (e.g., inertial forces and / or gravitational acceleration) acting about the computing device 1300, such as, but not limited to, one or more sensors that measure forces in one or more dimensions, torque, ground force, friction, and / or a ZMP sensor that identifies and / or identifies the location of a zero moment point (ZMP). Many other examples of sensors 1320 are possible as well.

[0140] The power supply system 1322 may include one or more batteries 1324 and / or one or more external power interfaces 1326 for providing power to the computing device 1300. Each battery of the one or more batteries 1324 may function as a source of stored power for the computing device 1300 when electrically coupled to the computing device 1300. The one or more batteries 1324 of the power supply system 1322 may be configured to be portable. Some or all of the one or more batteries 1324 may be easily removable from the computing device 1300. In other examples, some or all of the one or more batteries 1324 may be internal to the computing device 1300 and therefore may not be easily removable from the computing device 1300. Some or all of the one or more batteries 1324 may be rechargeable. For example, rechargeable batteries may be recharged via a wired connection between the battery and another power source, such as by one or more power sources external to the computing device 1300 and connected to the computing device 1300 via one or more external power interfaces. In other examples, some or all of the one or more batteries 1324 may be non-rechargeable batteries.

[0141] The one or more external power interfaces 1326 of the power system 1322 may include one or more wired power interfaces, such as a USB cable and / or a power cord, that enable a wired power connection to one or more power sources external to the computing device 1300. The one or more external power interfaces 1326 may include one or more wireless power interfaces, such as a Qi wireless charger, that enable a wireless power connection to these one or more external power sources, such as via a Qi wireless charger. Once a power connection to an external power source is established using the one or more external power interfaces 1326, the computing device 1300 may draw power from the external power source via the established power connection. In some examples, the power system 1322 may include associated sensors, such as battery sensors or other types of power sensors associated with one or more batteries.

[0142] Cloud-based Server FIG. 14 illustrates a cloud-based server system according to an example embodiment. In FIG. 6, neural network functionality and / or computing devices can be distributed among computing clusters 1409a, 1409b, and 1409c. Computing cluster 1409a can include one or more computing devices 1400a, cluster storage array 1410a, and cluster router 1411a connected by a local cluster network 1412a. Similarly, computing cluster 1409b can include one or more computing devices 1400b, cluster storage array 1410b, and cluster router 1411b connected by a local cluster network 1412b. Similarly, computing cluster 1409c can include one or more computing devices 1400c, cluster storage array 1410c, and cluster router 1411c connected by a local cluster network 1412c.

[0143] In some embodiments, computing clusters 1409a, 1409b, 1409c may be a single computing device residing within a single computing center. In other embodiments, computing clusters 1409a, 1409b, 1409c may include multiple computing devices within a single computing center, or even multiple computing devices at multiple computing centers located in various geographic locations. For example, FIG. 14 illustrates each of computing clusters 1409a, 1409b, 1409c residing in different physical locations.

[0144] In some embodiments, the data and services in computing clusters 1409 a, 1409 b, 1409 c may be stored on a non-transitory, tangible computer-readable medium (or computer-readable storage medium) and encoded as computer-readable information accessible by other computing devices. In some embodiments, computing clusters 1409 a, 1409 b, 1409 c may be stored on a single disk drive or other tangible storage medium or may be implemented on multiple disk drives or other tangible storage media located in one or more diverse geographic locations.

[0145] In some embodiments, each of computing clusters 1409a, 1409b, and 1409c may have an equal number of computing devices, an equal number of cluster storage arrays, and an equal number of cluster routers. However, in other embodiments, each computing cluster may have a different number of computing devices, a different number of cluster storage arrays, and a different number of cluster routers. The number of computing devices, cluster storage arrays, and cluster routers in each computing cluster may depend on one or more computing tasks assigned to each computing cluster.

[0146] In computing cluster 1409a, for example, computing device 1400a may be configured to perform various computing tasks of a conditioned axial self-attention-based neural network and / or computing devices. In one embodiment, various functions of the neural network and / or computing devices may be distributed among one or more of computing devices 1400a, 1400b, and 1400c. Computing devices 1400b and 1400c in each computing cluster 1409b and 1409c may be configured similarly to computing device 1400a in computing cluster 1409a. Meanwhile, in some embodiments, computing devices 1400a, 1400b, and 1400c may be configured to perform different functions.

[0147] In some embodiments, computing tasks and stored data associated with the neural network and / or computing devices may be distributed across computing devices 1400a, 1400b, and 1400c based at least in part on the processing requirements of the neural network and / or computing devices, the processing capabilities of computing devices 1400a, 1400b, and 1400c, the latency of network links between computing devices within each computing cluster and between the computing clusters themselves, and / or other factors that may contribute to the cost, speed, fault tolerance, resilience, efficiency, and / or other design goals of the overall system architecture.

[0148] The cluster storage arrays 1410a, 1410b, 1410c of the computing clusters 1409a, 1409b, 1409c may be data storage arrays that include disk array controllers configured to manage read and write access to groups of hard disk drives. The disk array controllers, alone or in conjunction with their respective computing devices, may also be configured to manage backup or redundant copies of data stored in the cluster storage arrays to protect against disk drive or other cluster storage array failures and / or network failures that prevent one or more computing devices from accessing one or more cluster storage arrays.

[0149] Similar to how the functionality of a conditioned axial self-attention-based neural network and / or computing devices may be distributed across computing devices 1400a, 1400b, 1400c of computing clusters 1409a, 1409b, 1409c, various active and / or backup portions of these components may be distributed across cluster storage arrays 1410a, 1410b, 1410c. For example, some cluster storage arrays may be configured to store a portion of the data for a first layer of a neural network and / or computing devices, while other cluster storage arrays may store other portions of the data for a second layer of a neural network and / or computing devices. Also, for example, some cluster storage arrays may be configured to store data for an encoder of a neural network, while other cluster storage arrays may store data for a decoder of the neural network. Furthermore, some cluster storage arrays may be configured to store backup versions of data stored in other cluster storage arrays.

[0150] Cluster routers 1411a, 1411b, 1411c in computing clusters 1409a, 1409b, 1409c may include network equipment configured to provide internal and external communications for the computing clusters. For example, cluster router 1411a in computing cluster 1409a may include one or more Internet switching and routing devices configured to provide (i) local area network communications between computing device 1400a and cluster storage array 1410a via local cluster network 1412a, and (ii) wide area network communications between computing cluster 1409a and computing clusters 1409b and 1409c via wide area network link 1413a to network 1206. Cluster routers 1411b and 1411c may include network equipment similar to cluster router 1411a, and cluster routers 1411b and 1411c may perform similar networking functions for computing clusters 1409b and 1409b as cluster router 1411a performs for computing cluster 1409a.

[0151] In some embodiments, the configuration of the cluster routers 1411a, 1411b, 1411c may be based at least in part on the data communication requirements of the computing devices and cluster storage arrays, the data communication capabilities of the network equipment within the cluster routers 1411a, 1411b, 1411c, the latency and throughput of the local cluster networks 1412a, 1412b, 1412c, the latency, throughput, and cost of the wide area network links 1413a, 1413b, 1413c, and / or other factors that may contribute to the cost, speed, fault tolerance, resilience, efficiency, and / or other design criteria of the moderation system architecture.

[0152] Exemplary Methods of Operation 15 is a flowchart of a method 1500 according to an example embodiment. The method 1500 may be performed by a computing device such as the computing device 1300.

[0153] Method 1500 may begin at block 1510, which includes receiving, by a computing device, a plurality of training data sets corresponding to each of a plurality of degradation factors, each training data set including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image.

[0154] At block 1520, the method includes training a plurality of intermediate machine learning models to remove one or more image degradations associated with the given image, where each intermediate machine learning model of the plurality of intermediate machine learning models is trained with a respective training dataset of the plurality of training datasets to remove one or more image degradations associated with the given image based on each of the plurality of degradation factors.

[0155] At block 1530, the method includes training an image transformation model with an additional training dataset of real images to remove one or more image impairments associated with the real images, where the image transformation model learns from the plurality of intermediate machine learning models.

[0156] At block 1540, the method includes outputting, by the computing device, the trained image transformation model for removal of image artifacts.

[0157] In some embodiments, each intermediate machine learning model of the plurality of intermediate machine learning models may include an encoder-decoder neural network with skip connections, where a given image may be processed at different scales, and the output of a given scale may be upsampled and concatenated with the input of a successive scale.

[0158] In some embodiments, each intermediate machine learning model may include (1) one or more spatial-depth (s2d) layers, where each s2d layer enables low-resolution processing by reducing the spatial resolution of the input while increasing the number of channels for acquiring the input and preserving information, and (2) one or more depth-spatial (d2s) layers corresponding to the one or more s2d layers.

[0159] In some embodiments, each intermediate machine learning model of the plurality of intermediate machine learning models is associated with a respective filter number and a respective spatial-depth parameter.

[0160] Some embodiments include applying the multiple trained intermediate machine learning models to a given real image of an additional training dataset of real images to generate a corresponding output image. Such embodiments include selecting an optimally transformed version of the given real image from the generated output images. Such embodiments also include generating a curated dataset including pairs of real images and corresponding optimally transformed versions of the real image. In some embodiments, applying the multiple trained intermediate machine learning models to the given real image includes applying each trained intermediate machine learning model at multiple resolutions. In some embodiments, training the image transformation model includes fine-tuning the image transformation model based on the curated dataset. In such embodiments, fine-tuning the image transformation model includes quantization-aware training (QAT).

[0161] Some embodiments include generating a synthetically degraded version of a sharp image.

[0162] In some embodiments, the plurality of degradation factors includes motion blur. Such embodiments include generating one or more motion kernels that simulate hand tremor, where the amount of motion blur is related to a scale parameter related to exposure time.

[0163] In some embodiments, the plurality of degradation factors includes lens blur. Such embodiments include generating one or more generalized Gaussian blur kernels that simulate lens blur.

[0164] In some embodiments, the plurality of degradation factors includes one or more of additive noise, signal-dependent noise, or colored noise.

[0165] In some embodiments, the plurality of degradation factors includes one or more image compression artifacts.

[0166] In some embodiments, the plurality of degradation factors includes one or more artifacts caused by saturated pixels. Such embodiments include determining whether the pixel value exceeds a saturation threshold. Such embodiments also include multiplying the pixel value by a random factor based on determining that the pixel value exceeds the saturation threshold.

[0167] In some embodiments, the intermediate machine learning models may be teacher networks and the image transformation model may be a student network, and the image transformation model learns from the intermediate machine learning models based on knowledge distillation.

[0168] 16 is a flowchart of a method 1600 according to an example embodiment. The method 1600 may be performed by a computing device such as the computing device 1300.

[0169] The method 1600 may begin at block 1610, which includes receiving, by a computing device, an input image that includes one or more image impairments.

[0170] At block 1620, the method includes predicting a transformed version of the input image using an image transformation model, where the image transformation model is trained to remove one or more image degradations associated with the input image, and the training includes: (1) training a plurality of intermediate machine learning models to remove the one or more image degradations, where each intermediate machine learning model of the plurality of intermediate machine learning models is trained on a respective training dataset of a plurality of training datasets corresponding to each of a plurality of degradation factors, each training dataset including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image; and (2) the image transformation model is trained on an additional training dataset of real images and learns from the plurality of intermediate machine learning models.

[0171] At block 1630, the method includes providing, by the computing device, a predicted transformed version of the input image.

[0172] In some embodiments, each intermediate machine learning model of the plurality of intermediate machine learning models may include an encoder-decoder neural network with skip connections, where a given image may be processed at different scales, and the output of a given scale may be upsampled and concatenated with the input of a successive scale.

[0173] In some embodiments, the image transformation model includes an encoder-decoder neural network with skip connections, and the image transformation model further includes: (1) one or more spatial-depth (s2d) layers, each s2d layer enabling low-resolution processing by reducing the spatial resolution of the input while increasing the number of channels for acquiring and storing information; and (2) one or more depth-spatial (d2s) layers corresponding to the one or more s2d layers.

[0174] In some embodiments, the plurality of degradation factors include one or more of motion blur, lens blur, image noise, image compression artifacts, or artifacts caused by saturated pixels.

[0175] The present disclosure should not be limited in terms of the specific embodiments described in this application, which are intended as illustrations of various aspects. As will be apparent to those skilled in the art, many modifications and variations can be made without departing from the spirit and scope thereof. Functionally equivalent methods and apparatuses within the scope of the present disclosure, in addition to those enumerated herein, will be apparent to those skilled in the art from the above description. Such modifications and variations are intended to be included within the scope of the appended claims.

[0176] The foregoing detailed description describes various features and functions of the disclosed systems, devices, and methods with reference to the accompanying drawings. In the figures, like symbols generally identify like elements unless context dictates otherwise. The illustrative embodiments described in the detailed description, figures, and claims are not intended to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that aspects of the present disclosure, as generally described herein and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are expressly contemplated herein.

[0177] With respect to any or all of the ladder diagrams, scenarios, and flowcharts in the figures and as discussed herein, each step, block, and / or communication may represent the processing of information and / or the transmission of information according to the exemplary embodiments. Alternative embodiments are included within the scope of these exemplary embodiments. In these alternative embodiments, for example, functions described as blocks, transmissions, communications, requests, responses, and / or messages may be performed out of the order shown or discussed, including substantially simultaneously or in reverse order, depending on the functionality involved. Furthermore, more or fewer blocks and / or functions may be used in any of the ladder diagrams, scenarios, and flowcharts discussed herein, and these ladder diagrams, scenarios, and flowcharts may be combined with each other, either in part or in whole.

[0178] The blocks representing processing of information may correspond to circuitry that can be configured to perform specific logical functions of the methods or techniques described herein. Alternatively or additionally, the blocks representing processing of information may correspond to modules, segments, or portions of program code (including associated data). The program code may include one or more instructions executable by a processor to implement specific logical functions or operations in the method or technique. The program code and / or associated data may be stored on any type of computer-readable medium, such as a storage device, including a disk, or hard drive, or other storage medium.

[0179] Computer-readable media may also include non-transitory computer-readable media, such as register memory, processor cache, and non-transitory computer-readable media that store data for a short period of time, such as random access memory (RAM). Computer-readable media may also include non-transitory computer-readable media that store program code and / or data for a longer period of time, such as secondary or permanent long-term storage devices, such as read-only memory (ROM), optical or magnetic disks, and compact disk read-only memory (CD-ROM). Computer-readable media may also be any other volatile or non-volatile storage system. Computer-readable media may be considered, for example, to be a computer-readable storage medium or a tangible storage device.

[0180] Additionally, blocks representing one or more information transfers may correspond to information transfers between software and / or hardware modules within the same physical device, although other information transfers may be between software and / or hardware modules in different physical devices.

[0181] While various aspects and embodiments are disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are provided by way of example and are not intended to be limiting, the true scope of which is related to the claims that follow.

Claims

1. 1. A computer-implemented method comprising: receiving, by a computing device, a plurality of training data sets corresponding to each of a plurality of degradation factors, each training data set including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image; training a plurality of intermediate machine learning models to remove one or more image degradations associated with a given image, wherein each intermediate machine learning model of the plurality of intermediate machine learning models is trained with a respective training dataset of the plurality of training datasets to remove the one or more image degradations associated with the given image based on each of the plurality of degradation factors; training an image transformation model with an additional training dataset of real images to remove the one or more image impairments associated with the real images, wherein the image transformation model learns from the plurality of intermediate machine learning models; outputting the trained image transformation model by the computing device for removal of image artifacts; A computer-implemented method comprising:

2. 2. The computer-implemented method of claim 1, wherein each intermediate machine learning model of the plurality of intermediate machine learning models comprises an encoder-decoder neural network with skip connections, the given image is processed at different scales, and the output of the given image is upsampled and concatenated with inputs of successive scales.

3. 3. The computer-implemented method of claim 2, wherein each intermediate machine learning model further includes: (1) one or more spatial-depth (s2d) layers, each s2d layer enabling low-resolution processing by reducing the spatial resolution of the input while increasing the number of channels for acquiring the input and preserving information; and (2) one or more depth-spatial (d2s) layers corresponding to the one or more s2d layers.

4. 10. The computer-implemented method of claim 1, wherein each intermediate machine learning model of the plurality of intermediate machine learning models is associated with a respective number of filters and a respective spatial-depth parameter.

5. applying the plurality of trained intermediate machine learning models to a given real image of the additional training dataset of real images to generate a corresponding output image; selecting an optimally transformed version of the given real image from the generated output images; generating a curated dataset comprising pairs of real images and corresponding optimally transformed versions of said real images; The computer-implemented method of claim 1 further comprising:

6. 6. The computer-implemented method of claim 5, wherein applying the plurality of trained intermediate machine learning models to the given real image comprises applying each trained intermediate machine learning model at multiple resolutions.

7. The computer-implemented method of claim 5 , wherein the training of the image transformation model comprises fine-tuning the image transformation model based on the curated dataset.

8. The computer-implemented method of claim 7 , wherein the fine-tuning of the image transformation model comprises quantized awareness training (QAT).

9. The computer-implemented method of claim 1 , further comprising generating the synthetically degraded version of the sharpened image.

10. The plurality of degradation factors includes motion blur, and the generating further includes: generating one or more motion kernels simulating camera shake due to hand tremors, the amount of motion blur being related to a scale parameter related to exposure time; 10. The computer-implemented method of claim 9, comprising:

11. The plurality of degradation factors includes lens blur, and the generating further includes: generating one or more generalized Gaussian blur kernels that simulate the lens blur; 10. The computer-implemented method of claim 9, comprising:

12. The computer-implemented method of claim 9 , wherein the plurality of degradation factors comprises one or more of additive noise, signal-dependent noise, or colored noise.

13. The computer-implemented method of claim 9 , wherein the plurality of degradation factors includes one or more image compression artifacts.

14. The plurality of degradation factors include one or more artifacts caused by saturated pixels, and the generating further includes: determining whether a pixel value exceeds a saturation threshold; multiplying the pixel value by a random coefficient based on a determination that the pixel value exceeds the saturation threshold; 10. The computer-implemented method of claim 9, comprising:

15. 2. The computer-implemented method of claim 1, wherein the intermediate machine learning models are teacher networks and the image transformation model is a student network, and the image transformation model learns from the intermediate machine learning models based on knowledge distillation.

16. 1. A computer-implemented method comprising: receiving, by a computing device, an input image including one or more image impairments; predicting a transformed version of the input image with an image transformation model, the image transformation model being trained to remove the one or more image degradations associated with the input image, the training including: (1) training a plurality of intermediate machine learning models to remove the one or more image degradations, each intermediate machine learning model of the plurality of intermediate machine learning models being trained with a respective training dataset of a plurality of training datasets corresponding to each of a plurality of degradation factors, each training dataset including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image; and (2) the image transformation models being trained with an additional training dataset of real images and learning from the plurality of intermediate machine learning models. providing, by the computing device, the predicted transformed version of the input image; A computer-implemented method comprising:

17. 17. The computer-implemented method of claim 16, wherein each intermediate machine learning model of the plurality of intermediate machine learning models comprises an encoder-decoder neural network with skip connections, a given image is processed at different scales, and outputs of the given image are upsampled and concatenated with inputs of successive scales.

18. 17. The computer-implemented method of claim 16, wherein the image transformation model includes an encoder-decoder neural network with skip connections, and the image transformation model further includes: (1) one or more spatial-to-depth (s2d) layers, each s2d layer taking an input and reducing the spatial resolution of the input while increasing the number of channels for preserving information, thereby enabling low-resolution processing; and (2) one or more depth-spatial (d2s) layers corresponding to the one or more s2d layers.

19. 17. The computer-implemented method of claim 16, wherein the plurality of degradation factors comprises one or more of motion blur, lens blur, image noise, image compression artifacts, or artifacts caused by saturated pixels.

20. one or more processors; Data storage; 1. A computing device comprising: the data storage storing computer-executable instructions that, when executed by the one or more processors, cause the computing device to perform operations; The operation is receiving, by the computing device, an input image including one or more image impairments; predicting a transformed version of the input image with an image transformation model, the image transformation model being trained to remove the one or more image degradations associated with the input image, the training including: (1) training a plurality of intermediate machine learning models to remove the one or more image degradations, each intermediate machine learning model of the plurality of intermediate machine learning models being trained with a respective training dataset of a plurality of training datasets corresponding to each of a plurality of degradation factors, each training dataset including a plurality of pairs of a sharp image and a corresponding synthetically degraded version of the sharp image; and (2) the image transformation models being trained with an additional training dataset of real images and learning from the plurality of intermediate machine learning models. providing, by the computing device, the predicted transformed version of the input image; a computing device,