Machine learning model-based trigger mechanism for image enhancement
A machine learning-based quality assessment model efficiently identifies images that can benefit from enhancement, addressing image degradation challenges by predicting quality improvement likelihood, thus improving image quality and reducing computational overhead.
Patent Information
- Application Number
- JP2025519826
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-05
- Filing Date
- 2023-10-04
- Publication Date
- 2025-10-22
AI Technical Summary
Existing image capture devices struggle with removing blur, noise, and compression artifacts due to challenges in accurately identifying and addressing image degradation, which degrades image quality and requires significant computational resources.
A machine learning-based approach using a two-step semi-supervised learning method trains a quality assessment model to predict a quality improvement likelihood score, enabling efficient identification of images that can benefit from enhancement algorithms, reducing the need for large labeled datasets and computational overhead.
The method effectively identifies images that can be enhanced, improving their visual quality by selectively applying enhancement techniques, thereby enhancing user experience and reducing unnecessary computational resources.
Smart Images

Figure 2025535065000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE / INCORPORATION BY REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 378,386, filed October 5, 2022, which is incorporated herein by reference in its entirety. [Background technology]
[0002] Many modern computing devices, including mobile phones, personal computers, and tablets, include image capture devices such as still cameras and / or video cameras. Image capture devices can capture images, such as images including people, animals, landscapes, and / or objects. Some image capture devices and / or computing devices can enhance or otherwise modify the captured images. For example, some image capture devices can provide "red-eye" correction to remove artifacts, such as red-appearing eyes in people and animals, that may be present in images captured using bright lights, such as flash lighting. After the captured image is enhanced, the enhanced image can be saved, displayed, transmitted, printed on paper, and / or otherwise utilized. Summary of the Invention
[0003] Removing blur, noise, and compression artifacts from images is a long-standing challenge in computational photography. Image degradation can occur for several reasons: when the photographer or autofocus system sets the focus incorrectly (out-of-focus), or when relative motion between the camera and the scene is faster than the shutter speed (motion blur). Furthermore, even under ideal shooting conditions, there can be inherent camera blur due to sensor resolution, light diffraction, lens aberrations, and anti-aliasing filters. Similarly, image noise is inherent in capturing a discrete number of photons (shot noise) and analog-to-digital conversion and processing (read noise). Images are typically compressed using techniques such as JPEG compression before storage or transmission. Image compression can also degrade image quality.
[0004] Driven by a system of machine-learned components, an image capture device can be configured to generate a trigger based on a determination that an image should be enhanced. The trigger can alert a user, who can be provided with recommendations for removing blur, noise, compression artifacts, etc. to create a sharp image. In some aspects, a mobile device can be configured with these capabilities so that images can be enhanced in real time. In some examples, images can be automatically enhanced by the mobile device. In other aspects, a mobile phone user can non-destructively enhance an image to match their preferences. Also, for example, existing images in a user's image library can be enhanced based on the techniques described herein.
[0005] In one aspect, a computer-implemented method is provided. The method includes determining a respective delta quality score associated with each of a plurality of images, where determining the delta quality score includes predicting an enhanced image corresponding to a given image of the plurality of images using an image enhancement model, the image enhancement model being trained to remove one or more image degradations associated with the given image. Determining the delta quality score further includes determining a first quality score associated with the given image and a second quality score associated with the predicted enhanced image, the delta quality score being based on a difference between the second quality score and the first quality score, the delta quality score indicating a degree of image enhancement in the predicted enhanced image. The method includes generating, by a computing device, a training dataset including a plurality of images associated with respective delta quality scores. The method includes training, based on the generated training dataset, a quality assessment model for predicting a quality improvement likelihood score associated with an input image, the quality improvement likelihood score indicating a likelihood of improving the perceptual quality of the input image based on the removal of one or more image degradation factors. The method includes outputting, by the computing device, the trained quality assessment model.
[0006] In another aspect, a computing device is provided that includes one or more processors and data storage having computer-executable instructions stored thereon that, when executed by the one or more processors, cause the computing device to perform functions. The function includes determining a respective delta quality score associated with each of a plurality of images, where determining the delta quality score includes predicting an enhanced image corresponding to a given image of the plurality of images by an image enhancement model, the image enhancement model being trained to remove one or more image degradations associated with the given image, where determining the delta quality score further includes determining a first quality score associated with the given image and a second quality score associated with the predicted enhanced image, where the delta quality score is based on a difference between the second quality score and the first quality score, the delta quality score indicating a degree of image enhancement in the predicted enhanced image, where the function further includes generating, by a computing device, a training dataset including a plurality of images associated with respective delta quality scores, and training a quality assessment model to predict a quality improvement likelihood score associated with an input image based on the generated training dataset, where the quality improvement likelihood score indicates a likelihood of improving the perceived quality of the input image based on the removal of one or more image degradation factors, whereby the function further includes outputting, by the computing device, the trained quality assessment model.
[0007] In another aspect, a computer program is provided, the computer program including instructions that, when executed by a computing device, cause the computing device to perform functions. The functions include determining a respective delta quality score associated with each of a plurality of images, where determining the delta quality score includes predicting an enhanced image corresponding to a given image of the plurality of images using an image enhancement model, the image enhancement model being trained to remove one or more image degradations associated with the given image. Determining the delta quality score further includes determining a first quality score associated with the given image and a second quality score associated with the predicted enhanced image, the delta quality score being based on a difference between the second quality score and the first quality score, the delta quality score indicating a degree of image enhancement in the predicted enhanced image. The functions further include generating, by the computing device, a training dataset including a plurality of images associated with respective delta quality scores, and training a quality assessment model based on the generated training dataset to predict a quality improvement likelihood score associated with an input image, the quality improvement likelihood score indicating a likelihood of improving the perceptual quality of the input image based on the removal of one or more image degradation factors. The functions further include outputting, by the computing device, the trained quality assessment model.
[0008] In another aspect, an article of manufacture is provided. The article includes one or more computer-readable media having stored thereon computer-readable instructions that, when executed by one or more processors of the computing device, cause the computing device to perform functions. The functions include determining a respective delta quality score associated with each of a plurality of images, where determining the delta quality score includes predicting an enhanced image corresponding to a given image of the plurality of images with an image enhancement model, the image enhancement model being trained to remove one or more image degradations associated with the given image, determining a first quality score associated with the given image and a second quality score associated with the predicted enhanced image, the delta quality score being based on a difference between the second quality score and the first quality score, the delta quality score indicating a degree of image enhancement in the predicted enhanced image, generating, by the computing device, a training dataset including a plurality of images associated with respective delta quality scores, and training a quality assessment model based on the generated training dataset to predict a quality improvement likelihood score associated with an input image, the quality improvement likelihood score indicating a likelihood of improving the perceptual quality of the input image based on the removal of one or more image degradation factors, and outputting, by the computing device, the trained quality assessment model.
[0009] In another aspect, a system is provided, the system including: means for determining a respective delta quality score associated with each of a plurality of images, where determining the delta quality score includes predicting an enhanced image corresponding to a given image of the plurality of images with an image enhancement model, the image enhancement model being trained to remove one or more image degradations associated with the given image, determining the delta quality score further includes determining a first quality score associated with the given image and a second quality score associated with the predicted enhanced image, the delta quality score being based on a difference between the second quality score and the first quality score, the delta quality score indicating a degree of image enhancement in the predicted enhanced image, the system further includes: means for generating, by a computing device, a training data set including a plurality of images associated with respective delta quality scores; and means for training a quality assessment model to predict, based on the generated training data set, a quality improvement likelihood score associated with an input image, the quality improvement likelihood score indicating a likelihood of improving the perceptual quality of the input image based on the removal of one or more image degradation factors, and means for outputting, by the computing device, the trained quality assessment model.
[0010] In another aspect, a computer-implemented method is provided. The method includes receiving an input image by a computing device. The method includes predicting, by a quality assessment model, a quality improvement likelihood score associated with the input image, the quality improvement likelihood score indicating a likelihood of improving the perceptual quality of the input image based on removal of one or more image degradation factors, the quality assessment model being trained with a training dataset including a plurality of images associated with respective delta quality scores, the delta quality score being determined by predicting an enhanced image corresponding to a given image of the plurality of images by an image enhancement model trained to remove one or more image degradations associated with the given image, the delta quality score being further determined by determining a first quality score associated with the given image and a second quality score associated with the predicted enhanced image, the delta quality score being based on a difference between the second quality score and the first quality score, the delta quality score indicating a degree of image enhancement in the predicted enhanced image. Additionally, the method includes providing, by the computing device, an alert notification based on the predicted quality improvement likelihood score.
[0011] In another aspect, a computing device is provided. The computing device includes one or more processors and data storage. The data storage has stored therein computer-executable instructions that, when executed by the one or more processors, cause the computing device to perform functions. The functions include receiving, by the computing device, an input image; and predicting, by a quality assessment model, a quality improvement likelihood score associated with the input image, the quality improvement likelihood score indicating a likelihood of improving the perceptual quality of the input image based on removal of one or more image degradation factors, the quality assessment model being trained with a training dataset including a plurality of images associated with respective delta quality scores. The delta quality score is determined by predicting, by an image enhancement model trained to remove one or more image degradations associated with the given image, an enhanced image corresponding to a given image of the plurality of images. The delta quality score is further determined by determining a first quality score associated with the given image and a second quality score associated with the predicted enhanced image, the delta quality score being based on a difference between the second quality score and the first quality score, the delta quality score indicating a degree of image enhancement in the predicted enhanced image. The functions further include providing, by the computing device, an alert notification based on the predicted quality improvement likelihood score.
[0012] In another aspect, a computer program is provided, the computer program including instructions that, when executed by a computing device, cause the computing device to perform functions. The functions include receiving, by the computing device, an input image; and predicting, by a quality assessment model, a quality improvement likelihood score associated with the input image, the quality improvement likelihood score indicating a likelihood of improving the perceptual quality of the input image based on removal of one or more image degradation factors, the quality assessment model being trained with a training dataset including a plurality of images associated with respective delta quality scores, the delta quality score being determined by predicting, by an image enhancement model, an enhanced image corresponding to a given image of the plurality of images, the image enhancement model being trained to remove one or more image degradations associated with the given image, the delta quality score being further determined by determining a first quality score associated with the given image and a second quality score associated with the predicted enhanced image, the delta quality score being based on a difference between the second quality score and the first quality score, the delta quality score indicating a degree of image enhancement in the predicted enhanced image, and the functions further include providing, by the computing device, an alert notification based on the predicted quality improvement likelihood score.
[0013] In another aspect, an article of manufacture is provided. The article includes one or more computer-readable media having stored thereon computer-readable instructions that, when executed by one or more processors of the computing device, cause the computing device to perform functions. The functions include receiving, by the computing device, an input image; and predicting, by a quality assessment model, a quality improvement likelihood score associated with the input image, the quality improvement likelihood score indicating a likelihood of improving the perceptual quality of the input image based on removal of one or more image degradation factors, the quality assessment model being trained with a training dataset including a plurality of images associated with respective delta quality scores, the delta quality score being determined by predicting, by an image enhancement model, an enhanced image corresponding to a given image of the plurality of images, the image enhancement model being trained to remove one or more image degradations associated with the given image, the delta quality score being further determined by determining a first quality score associated with the given image and a second quality score associated with the predicted enhanced image, the delta quality score being based on a difference between the second quality score and the first quality score, the delta quality score indicating a degree of image enhancement in the predicted enhanced image, and the functions further include providing, by the computing device, an alert notification based on the predicted quality improvement likelihood score.
[0014] In another aspect, a system is provided, including means for receiving, by a computing device, an input image and means for predicting, by a quality assessment model, a quality improvement likelihood score associated with the input image, the quality improvement likelihood score indicating a likelihood of improving the perceptual quality of the input image based on removal of one or more image degradation factors, the quality assessment model being trained with a training dataset including a plurality of images associated with respective delta quality scores, the delta quality score being determined by predicting an enhanced image corresponding to a given image of the plurality of images by an image enhancement model, the image enhancement model being trained to remove one or more image degradations associated with the given image, the delta quality score being further determined by determining a first quality score associated with the given image and a second quality score associated with the predicted enhanced image, the delta quality score being based on a difference between the second quality score and the first quality score, the delta quality score indicating a degree of image enhancement in the predicted enhanced image, and means for providing, by the computing device, an alert notification based on the predicted quality improvement likelihood score.
[0015] The above summary is illustrative only and is not intended to be in any way limiting. In addition to the exemplary aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the figures and the following detailed description, and accompanying drawings. [Brief explanation of the drawings]
[0016] [Figure 1] 1 illustrates an exemplary framework for generating a delta quality score, according to an exemplary embodiment. [Figure 2] 1 illustrates an exemplary framework for training a baseline quality assessment model, according to an exemplary embodiment. [Figure 3] 1 illustrates an exemplary framework for fine-tuning a baseline quality assessment model, according to an exemplary embodiment. [Figure 4] 10 illustrates an exemplary inference by a quality assessment model, according to an exemplary embodiment. [Figure 5] 10 is a table illustrating correlation values obtained from a baseline quality assessment model during training and testing, according to an example embodiment. [Figure 6] 1 illustrates an exemplary comparison between a baseline quality assessment model and a fine-tuned quality assessment model, according to an exemplary embodiment. [Figure 7] 1 illustrates an exemplary application of a quality assessment model, according to an exemplary embodiment. [Figure 8] FIG. 1 illustrates the training and inference stages of a machine learning model, according to an example embodiment. [Figure 9] 1 illustrates a distributed computing architecture in accordance with an exemplary embodiment. [Figure 10] FIG. 1 is a block diagram of a computing device in accordance with an exemplary embodiment. [Figure 11] 1 illustrates a network of computing clusters arranged as a cloud-based server system, according to an example embodiment. [Figure 12] 1 is a flowchart of a method according to an example embodiment. [Figure 13] 10 is a flowchart of another method according to an example embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0017] Techniques are described for developing a quality assessment model to predict a quality improvement likelihood score for an input image. The quality improvement likelihood score indicates whether the image can benefit from image enhancement techniques. In some embodiments, a trigger model can be trained based on the quality improvement likelihood score, and the trigger model can be used in conjunction with an image enhancement algorithm. Also, for example, an image ranking model can be trained based on the quality improvement likelihood score to rank images that can most benefit from image enhancement.
[0018] Photo restoration operations, such as denoising and deblurring, improve the visual quality of distorted images. However, identifying such images may not be a simple task. Given an input image, a reliable trigger model needs to predict the degree of visual improvement that applying a particular restoration and / or enhancement algorithm will achieve. In addition, it may be impractical to run an enhancement model and use its output to make trigger decisions, typically due to computational overhead. Also, it may be desirable, for example, to know how much an image can be enhanced before applying an image enhancement model. This may be to avoid potentially degrading image quality, to save computational resources by not applying image enhancements when the degree of possible enhancement is minimal, and / or to perceptually improve image quality. As described herein, a framework is described for developing lightweight trigger models that can be reliably used to surface images that would most benefit from enhancement algorithms, such as, but not limited to, motion deblurring, noise removal, and compression artifact removal.
[0019] In one example, (a copy of) the trained quality assessment model may reside on a mobile computing device. The mobile computing device may include a camera capable of capturing an input image. The trained quality assessment model (e.g., residing on the mobile computing device) may predict an image quality improvement possibility score for the input image, and a user of the mobile computing device may be provided with a recommendation that the input image should be sharpened. The user may then select to enhance the image, and the input image may be provided to a trained image enhancement model (e.g., residing on the mobile computing device or on a remote server) for image enhancement. In response, the trained image enhancement model may generate a predicted output image that is a sharper version of the input image and then output the output image (e.g., provide the output image for display by the mobile computing device). In another example, the trained quality assessment model does not reside on the mobile computing device; rather, the mobile computing device provides the input image to a remotely located trained quality assessment model (e.g., via the Internet or other data network). The remotely located trained quality assessment model may process the input image and provide an output quality improvement possibility score to the mobile computing device. In other examples, the non-mobile computing device may also use the trained quality assessment model to predict a quality improvement potential score, including for images not captured by the computing device's camera.
[0020] In some examples, the trained quality assessment model can work in conjunction with other neural networks (or other software) and / or can be trained to recognize whether an input image has image degradation. Then, upon determining that the input image has image degradation, the trained quality assessment model described herein can provide the input image to a trained image enhancement model, thereby removing the image degradation in the input image.
[0021] Thus, the techniques described herein can improve images by removing image degradation (e.g., automatically or in response to user instruction), thereby enhancing their actual and / or perceived quality. Enhancing the actual and / or perceived quality of images, including portraits of people, can provide emotional benefits to people who believe their photos look better. These techniques are flexible and can be applied to images of human faces and other objects, scenes, etc.
[0022] overview Typically, for triggering purposes, image enhancement algorithms rely on either handcrafted features or deep machine learning (ML) models. Obtaining reliable handcrafted features, such as noise or blur estimation, can be challenging, especially when the camera pipeline is unknown. Meanwhile, training a deep ML trigger model requires curating large amounts of labeled data.
[0023] The approach described herein overcomes such challenges by relying on existing perceptual quality assessment models, requiring only a few hundred labeled examples (as opposed to the large number of labeled examples required by existing assessment models). The proposed approach can be a two-step semi-supervised learning approach in which a deep trigger model is first trained on image quality scores (e.g., Neural Image Assessment (NIMA) scores), and then the trigger model can be fine-tuned with a small amount of labeled data. This enables knowledge transfer from NIMA to the underlying trigger task (which is sensitive to blur, noise, and other degradations) without the need to curate thousands of ratings from human raters. Note that while NIMA can be generalized to real-world image degradations, any robust image quality assessment model can be used as part of the framework described herein.
[0024] Training an example baseline model FIG. 1 illustrates an exemplary framework 100 for generating a delta quality score, according to an exemplary embodiment. A process for training a baseline trigger model is illustrated in FIG. 1. As illustrated, an image enhancement model 110 (e.g., a DeepMode model) may be run on a plurality of images, such as a dataset of unlabeled data, input image data 105. In some embodiments, approximately 500,000 images may be used. The image enhancement model 110 predicts their respective enhanced counterparts, e.g., an enhanced image corresponding to a given image of the plurality of images, collected as enhanced image data 115. The image enhancement model 110 may be trained to remove one or more image impairments associated with the given image.
[0025] In some embodiments, the image enhancement model may be, among others, a deblurring model, a colorization model, an image artifact removal model, or a noise removal model. The image enhancement model, for example, a convolutional neural network (different from the CNN described above with respect to the quality assessment model), may be trained using a training dataset of images to perform one or more aspects described herein. In some examples, the neural network may be configured as an encoder / decoder neural network.
[0026] In some embodiments, a deep motion, defocus, and degradation enhancement (DeepMode) model may be applied in challenging cases where the amount of blur is moderate or large and the image exhibits noise or other degradation such as JPEG compression artifacts. DeepMode may be configured as a supervised deep learning end-to-end solution for removing blur, noise, compression artifacts, etc. on an image.
[0027] In some embodiments, a first quality score 120 associated with a given image of the plurality of images can be determined, and a second quality score 125 associated with the predicted enhanced image can be determined. In some embodiments, the first quality score 120 and the second quality score 125 can be Neural Image Assessment (NIMA) scores. For example, a NIMA score ranging from 1 to 10 can be used, with 10 indicating the highest quality image. In some embodiments, the NIMA model can be trained on approximately 250,000 images evaluated by human raters who evaluate images for various image degradation factors, such as blur, exposure, noise, and / or compression artifacts.
[0028] In some embodiments, a delta quality score may be determined based on the difference 140 between each second quality score (e.g., stored in the enhanced image quality score database 135) and the corresponding first quality score (e.g., stored in the input image quality score database 130), and the delta quality score may be stored in a database of delta quality scores 145. The delta quality score indicates the degree of image enhancement for the predicted enhanced image. In some embodiments, the delta quality score may also indicate the degree of degradation, for example, if an attempt to denoise a noise-free image results in excessive smoothing. For example, the delta quality score may be determined as follows: Delta quality score = Enhanced image quality score - Input image quality score (Formula 1)
[0029] In particular, when applied to NIMA scores, a delta NIMA score, denoted as Δ-NIMA, may be determined as follows: Δ-NIMA = NIMA(reinforcement) - NIMA(input) (Formula 2)
[0030] In some embodiments, the Δ-NIMA score can range from -9 to 9. The larger the Δ-NIMA, the higher the visual quality of the enhanced image. Some examples with different Δ-NIMA scores are shown in Figure 7.
[0031] Once the delta quality scores (e.g., Δ-NIMA scores) are calculated, they can be used to train a baseline quality assessment model (e.g., a deep neural network such as the MobileNet-V2 model).
[0032] In some embodiments, the trigger model may be trained based on the quality assessment model. For example, the trigger model may be a binary classifier trained to determine whether to enhance an image. In some embodiments, the quality assessment model may be used to train an image ranking model that ranks multiple images based on image quality.
[0033] FIG. 2 illustrates an exemplary framework 200 for training a baseline quality assessment model, according to an exemplary embodiment. In some embodiments, the quality assessment model may be a convolutional neural network (CNN). In some embodiments, the CNN may include a MobileNet architecture. In some embodiments, image data may be sized 224×224, and input image data 205 to the quality assessment model 210 may be resized to 448×448. This may help reduce the impact of resizing on input degradation. A delta quality score 215 may be provided to the quality assessment model 210. In some embodiments, if the quality assessment model 210 is a MobileNet model, a fully connected layer may be introduced in layer 16 of the MobileNet model to predict the delta quality score (e.g., the Δ-NIMA score). The MobileNet model may also be warm-started with weights from a JFT-trained checkpoint. For example, the JFT-300M dataset may be used to train an image classification model. Images are labeled using an algorithm that uses a combination of web signals, connections between web pages, and user feedback. Over 1 billion labels can be generated for 300 million images, and a single image can be associated with multiple labels. An example of exemplary correlation values obtained from the baseline quality assessment model during training and testing is shown in Figure 6.
[0034] Fine-tuning with labels Once the baseline quality assessment model is trained, fine-tuning of the baseline quality assessment model can be performed on data evaluated by human annotators to further improve the trigger model. The baseline model is a good approximation of the desired quality assessment model because it captures the effects of image enhancements.
[0035] FIG. 3 illustrates an exemplary framework 300 for fine-tuning a baseline quality assessment model, according to an exemplary embodiment. For example, approximately 1,000 images processed by an image enhancement algorithm (e.g., DeepMode) can be curated, and human raters can be asked to compare the enhanced images with the input image before enhancement. Each pair of image data 305 and the corresponding enhanced image can be evaluated to provide a label to generate human annotations 315. For example, the human annotator may be a photographer or other expert experienced in identifying the perceptual quality of images. In some embodiments, the label may be “great improvement,” corresponding to a score of 2; “moderate improvement,” corresponding to a score of 1; “neutral,” corresponding to a score of 0; or “degradation,” corresponding to a score of −1. The image data 305 and human annotations 315 can be provided to a baseline quality assessment model 310 and used to fine-tune the baseline quality assessment model 310. In some embodiments, the data can be split into training data and test data (e.g., an 80%-20% split). In some embodiments, fine-tuning may include fine-tuning the final layer (e.g., a layer with fewer than 120 training parameters) of a MobileNet-V2 model trained with Δ-NIMA data. Thus, instead of training hundreds of thousands of parameters, only a few parameters are trained for fine-tuning, thereby requiring a relatively minimal amount of training data. In some embodiments, the remaining weights may be loaded from the baseline quality assessment model 310 (e.g., a Δ-NIMA predictor) and kept frozen during training. Once the quality assessment model is fine-tuned, it can be used to evaluate its performance on human-rated data.
[0036] 4 illustrates an example inference 400 by a quality assessment model, according to an example embodiment. As shown, an input image 405 can be provided to a quality assessment model 410, and a quality improvement likelihood score 415 can be predicted. For example, the quality assessment model 410 predicts a quality improvement likelihood score 415 having a value of 0.49 for the input image 405. The output quality improvement likelihood score 415 can be used to identify moderately and / or significantly improved images in a test set.
[0037] 5 is a table 500 showing correlation values obtained from a baseline quality assessment model during training and testing, according to an example embodiment. These values indicate that the baseline MobileNet is effective at predicting quality improvement potential (e.g., Δ-NIMA) scores. The correlation results in table 500 verify that the quality assessment model (e.g., Δ-NIMA predictor) performs as intended. The fine-tuning step occurs after the quality assessment model is trained.
[0038] Two trained models, the baseline quality assessment model and the fine-tuned quality assessment model, can be compared.
[0039] 6 illustrates an exemplary comparison graph 600 between a baseline quality assessment model and a fine-tuned quality assessment model, according to an exemplary embodiment. The graph 600 displays precision values (along the vertical axis) against recall values (along the horizontal axis). A precision-recall analysis of the baseline and fine-tuned trigger models is shown. Ground truth data is evaluated by human raters. As expected, the fine-tuned model 605 performs better than the baseline model 610. The fine-tuned model 605 exhibits an AUC-PR of 0.755. However, the baseline model 610 also exhibits a robust AUC-PR of 0.688.
[0040] 7 illustrates an exemplary application of a quality assessment model according to an exemplary embodiment. A visual example is shown in FIG. 7 in which a blurry or noisy image (e.g., in row 7R1) exhibits a higher quality improvement potential score compared to a sharp, in-focus image (e.g., in row 7R2) at the bottom. The score threshold for triggering the enhancement model is shown as 0.4. Thus, an alert notification may be triggered for the input image in row 7R1 (which has a quality improvement potential score (QIS) of 0.99, exceeding the threshold score of 0.4), while no alert notification may be triggered for the input image in row 7R2 (which has a quality improvement potential score of −0.04, not exceeding the threshold score of 0.4).
[0041] Several examples with different delta-NIMA scores are also shown in Figure 7. For example, row 7R1 shows an enhanced output image with a delta score of 0.77, while row 7R2 shows another enhanced output image with a delta score of -0.8. These outputs are consistent with quality improvement potential scores. For example, the image in row 7R1 has a quality improvement potential score of 0.99, indicating a significant potential improvement under image enhancement, and the corresponding enhanced output image has a delta score of 0.77, indicating a significant improvement after image enhancement is performed. Similarly, the image in row 7R3 has a quality improvement potential score of -0.04, indicating a low potential improvement under image enhancement (because the image is already of high quality), and the corresponding enhanced output image has a delta score of -0.08, indicating no improvement after image enhancement is performed.
[0042] Example image degradation Image blurring can generally be modeled as a linear operator acting on a sharp latent image. For a shift-invariant linear operator, the blurring operation may correspond to a convolution with a blur kernel. In practice, there is a general assumption that the captured image contains additive noise and compression in addition to blurring. Therefore, the following relationship may apply: v=C(S(u*k)+n), (Formula 3)
[0043] where v is the captured image, u is the underlying sharp image, k is the unknown blur kernel, * is the convolution operation, n is the additive noise, S models the sensor's nonlinear response (e.g., saturation), and C represents image compression. Some existing techniques perform image deblurring by treating the problem as a "blind" deconvolution process. For example, in the first step, a blur kernel may be estimated. This can be achieved by assuming a sharp image model, e.g., by using a variational framework, while in a second, independent step, a "non-blind" deconvolution algorithm may be applied. However, image noise and artifacts due to compression can adversely affect both steps. Even when the blur kernel is determined, "non-blind" deconvolution may be an ill-posed problem, and the presence of noise, compression, etc. may lead to artifacts. A significant drawback of model-based deblurring is that the degradation model generally requires a high degree of accuracy. This can pose significant challenges in practice due to several unknown or partially known image transformations (e.g., unknown blur, unknown camera image signal processor (ISP), post-processing, compression, etc.).
[0044] To remove one or more image impairments, the techniques described herein may apply an image enhancement model (e.g., based on a convolutional neural network) to predict a sharp image. Although a particular image enhancement model is described for illustrative purposes, the quality assessment model described with reference to Figures 1-7 may be implemented in conjunction with any image enhancement model.
[0045] As used herein, the term "degradation factor" generally refers to any factor that affects the sharpness of an image, such as image clarity with respect to quantitative image quality parameters, such as contrast, focus, etc. In some embodiments, the one or more degradation factors may include one or more of motion blur, lens blur, image noise, image compression artifacts, or artifacts caused by saturated pixels.
[0046] As used herein, the term "motion blur" generally refers to a degrading factor that causes one or more objects in an image to appear fuzzy and / or unclear due to the movement of a camera capturing the image, the movement of one or more objects, or a combination of the two. In some instances, motion blur may be perceived as streaking or smearing in an image. As used herein, the term "lens blur" generally refers to a degrading factor that causes an image to appear to have a narrower depth of field than the scene being captured. For example, certain objects in an image may be in focus, while other objects may appear out of focus.
[0047] As used herein, the term "image noise" generally refers to degradation factors that cause an image to appear to have artifacts (e.g., speckles, color dots, etc.) due to a low signal-to-noise ratio (SNR). For example, an SNR below a certain desired threshold can cause image noise. In some examples, image noise can be generated by circuitry within the image sensor or camera. As used herein, the term "image compression artifacts" generally refers to degradation factors that result from lossy image compression. For example, image data can be lost during compression, which can result in visible artifacts in a decompressed version of the image.
[0048] As used herein, the term "saturated pixel" generally refers to a state in which a pixel is saturated with photons that spill over into neighboring pixels. For example, a saturated pixel may be associated with an image intensity greater than a threshold intensity (e.g., greater than 245, or at 255, etc.). The image intensity may correspond to a grayscale intensity or the intensity of a red, blue, or green (RGB) color component. For example, a more highly saturated pixel may appear as a brighter color. Thus, photon spillover from a saturated pixel into neighboring pixels may cause perceptual defects in the image (e.g., causing saturation of one or more neighboring pixels, distorting the intensity of one or more neighboring pixels, etc.).
[0049] Training a machine learning model to generate inferences / predictions FIG. 8 shows a diagram 800 illustrating the training stage 802 and inference stage 804 of trained machine learning model(s) 832, according to an example embodiment. Some machine learning techniques involve training one or more machine learning algorithms with an input set of training data to recognize patterns in the training data and provide output inferences and / or predictions regarding (the patterns in) the training data. The resulting trained machine learning algorithms may be referred to as trained machine learning models. For example, FIG. 8 shows a training stage 802 in which one or more machine learning algorithms 820 are trained with training data 810 to become trained machine learning model(s) 832. Then, during the inference stage 804, the trained machine learning model(s) 832 may receive input data 830 and one or more inference / prediction requests 840 (perhaps as part of the input data 830) and, in response, provide one or more inferences and / or predictions 850 as output.
[0050] For example, the one or more machine learning algorithms 820 may include a quality assessment model (e.g., a deep model such as a MobileNet-V2 model), a delta scoring model (e.g., a Δ-NIMA predictor), an image enhancement model (e.g., DeepMode, a deblurring model, a colorization model, an artifact removal model, etc.), a trigger model, an image ranking model, etc. The trained machine learning model(s) 832 may be trained versions of each of these one or more machine learning algorithms 820.
[0051] Thus, the trained machine learning model(s) 832 may include one or more models of one or more machine learning algorithms 820. The machine learning algorithm(s) 820 include, but are not limited to, artificial neural networks (e.g., convolutional neural networks described herein, recurrent neural networks, Bayesian networks, hidden Markov models, Markov decision processes, logistic regression functions, support vector machines, suitable statistical machine learning algorithms, and / or heuristic machine learning systems). The machine learning algorithm(s) 820 may be supervised or unsupervised and may perform any suitable combination of online and offline learning.
[0052] In some examples, the machine learning algorithm(s) 820 and / or the trained machine learning model(s) 832 may be accelerated using an on-device coprocessor, such as a graphics processing unit (GPU), a tensor processing unit (TPU), a digital signal processor (DSP), and / or an application-specific integrated circuit (ASIC). Such on-device coprocessors may be used to accelerate the machine learning algorithm(s) 820 and / or the trained machine learning model(s) 832. In some examples, the trained machine learning model(s) 832 may be trained to provide inference on, reside, execute, and / or otherwise perform inference for a particular computing device.
[0053] During the training phase 802, the machine learning algorithm(s) 820 may be trained by providing at least the training data 810 as training inputs using unsupervised, supervised, semi-supervised, and / or reinforcement learning techniques. Unsupervised learning involves providing some (or all) of the training data 810 to the machine learning algorithm(s) 820, with the machine learning algorithm(s) 820 determining one or more output inferences based on the provided portion (or all) of the training data 810. Supervised learning involves providing some (or all) of the training data 810 to the machine learning algorithm(s) 820, with the machine learning algorithm(s) 820 determining one or more output inferences based on the provided portion (or all) of the training data 810, with the output inference(s) either accepted or corrected based on the correct results associated with the training data 810. In some examples, the supervised learning of the machine learning algorithm(s) 820 may be governed by a set of rules and / or a set of labels for the training inputs, which may be used to correct the inferences of the machine learning algorithm(s) 820.
[0054] Semi-supervised learning involves having correct results for some, but not all, of the training data 810. During semi-supervised learning, supervised learning is used for portions of the training data 810 that have correct results, and unsupervised learning is used for portions of the training data 810 that do not have correct results. Reinforcement learning includes machine learning algorithm(s) 820 receiving a reward signal related to a prior inference, where the reward signal may be a numerical value. During reinforcement learning, the machine learning algorithm(s) 820 can output an inference and receive a reward signal in response, where the machine learning algorithm(s) 820 are configured to attempt to maximize the numerical value of the reward signal. In some examples, reinforcement learning also utilizes a value function that provides a numerical value representing the expected sum of the numerical values provided by the reward signal over time. In some examples, the machine learning algorithm(s) 820 and / or the trained machine learning model(s) 832 may be trained using other machine learning techniques, including, but not limited to, incremental learning and curriculum learning.
[0055] In some examples, the machine learning algorithm(s) 820 and / or the trained machine learning model(s) 832 can use transfer learning techniques. For example, transfer learning techniques may include the trained machine learning model(s) 832 being pre-trained on a set of data and additionally trained using training data 810. More specifically, the machine learning algorithm(s) 820 may be pre-trained on data from one or more computing devices and the resulting trained machine learning model provided to computing device CD1, which is intended to execute the trained machine learning model during inference stage 804. Then, during training stage 802, the pre-trained machine learning model may be additionally trained using training data 810, which may be derived from kernel data and non-kernel data of computing device CD1. This further training of the machine learning algorithm(s) 820 and / or the pre-trained machine learning model using the training data 810 of CD1's data may be performed using either supervised learning or unsupervised learning. The training phase 802 may be completed once the machine learning algorithm(s) 820 and / or pre-trained machine learning model(s) have been trained with at least the training data 810. The resulting trained machine learning model may be utilized as at least one of the trained machine learning model(s) 832.
[0056] In particular, once the training phase 802 is complete, the trained machine learning model(s) 832 may be provided to the computing device if not already on the computing device. The inference phase 804 may begin after the trained machine learning model(s) 832 are provided to the computing device CD1.
[0057] During the inference stage 804, the trained machine learning model(s) 832 may receive the input data 830 and generate and output one or more corresponding inferences and / or predictions 850 regarding the input data 830. Thus, the input data 830 may be used as input to the trained machine learning model(s) 832 to provide the corresponding inference(s) and / or prediction(s) 850 to the kernel and non-kernel components. For example, the trained machine learning model(s) 832 may generate the inference(s) and / or prediction(s) 850 in response to one or more inference / prediction requests 840. In some examples, the trained machine learning model(s) 832 may be executed by another piece of software. For example, the trained machine learning model(s) 832 may be executed by an inference or prediction daemon so that they are readily available to provide inferences and / or predictions upon request. The input data 830 may include data from the computing device CD1 running the trained machine learning model(s) 832 and / or input data from one or more computing devices other than CD1.
[0058] The input data 830 may include training data as described herein, such as images associated with delta quality scores, human-annotated data, real blurred images, synthetically generated images, images in a curated dataset, etc. Other types of input data are possible as well. For example, the training data may include data collected to train an image translation model.
[0059] The inference(s) and / or prediction(s) 850 may include task outputs, numerical values, and / or other output data generated by the trained machine learning model(s) 832 operating on the input data 830 (and training data 810). In some examples, the trained machine learning model(s) 832 may use the output inference(s) and / or prediction(s) 850 as input feedback 850. The trained machine learning model(s) 832 may also rely on past inferences as input to generate new inferences.
[0060] After training, the trained version of the neural network may be an example of trained machine learning model(s) 832. In this approach, one or more examples of inference / prediction requests 840 may be requests to predict a quality improvement possibility score and / or a transformed (e.g., deblurred, denoised, etc.) image, and corresponding examples of inference and / or prediction(s) 850 may be the predicted quality improvement possibility score and / or the transformed (e.g., deblurred, denoised, etc.) image.
[0061] In some examples, one computing device CD_SOLO may include a trained version of the neural network, perhaps after training. The computing device CD_SOLO may then receive a request to predict a quality improvement likelihood score and / or a request to transform (e.g., deblur, denoise, etc.) an image and use the trained version of the neural network to predict a quality improvement likelihood score and / or the transformed (e.g., deblurred, denoised, etc.) image.
[0062] In some examples, two or more computing devices CD_CLI and CD_SRV may be used to provide output; for example, a first computing device CD_CLI may generate a request to a second computing device CD_SRV to predict a quality improvement possibility score and / or a transformed (e.g., deblurred, denoised, etc.) image. CD_SRV may then respond to the request from CD_CLI using a trained version of the neural network to predict a quality improvement possibility score and / or a transformed (e.g., deblurred, denoised, etc.) image. Further, upon receiving a response to the request, CD_CLI may provide the requested output (e.g., using a user interface and / or display, a printed copy, electronic communication, etc.).
[0063] Data Network Example 9 illustrates a distributed computing architecture 900, according to an example embodiment. The distributed computing architecture 900 includes server devices 908, 910 configured to communicate with programmable devices 904a, 904b, 904c, 904d, and 904e via a network 906. The network 906 may correspond to a local area network (LAN), a wide area network (WAN), a WLAN, a WWAN, a corporate intranet, the public Internet, or any other type of network configured to provide a communication path between networked computing devices. The network 906 may also correspond to a combination of one or more LANs, WANs, corporate intranets, and / or the public Internet.
[0064] While FIG. 9 shows only five programmable devices, the distributed application architecture may serve tens, hundreds, or thousands of programmable devices. Furthermore, programmable devices 904a, 904b, 904c, 904d, and 904e (or any additional programmable devices) may be any type of computing device, such as a mobile computing device, a desktop computer, a wearable computing device, a head-mountable device (HMD), a network terminal, or a mobile computing device. In some examples, such as shown by programmable devices 904a, 904b, 904c, and 904e, the programmable devices may be directly connected to network 906. In other examples, such as shown by programmable device 904d, the programmable devices may be indirectly connected to network 906 through an associated computing device, such as programmable device 904c. In this example, programmable device 904c may serve as an associated computing device for passing electronic communications between programmable device 904d and network 906. In other examples, such as shown by programmable device 904e, the computing device may be part of and / or within a vehicle, such as a car, truck, bus, boat or watercraft, airplane, etc. In other examples not shown in Figure 9, the programmable device may be connected both directly and indirectly to the network 906.
[0065] Server devices 908, 910 may be configured to perform one or more services requested by programmable devices 904a-904e. For example, server devices 908 and / or 910 may provide content to programmable devices 904a-904e. The content may include, but is not limited to, web pages, hypertext, scripts, binary data such as compiled software, images, audio, and / or video. The content may include compressed and / or uncompressed content. The content may be encrypted and / or decrypted. Other types of content are possible as well.
[0066] As another example, server devices 908 and / or 910 may provide programmable devices 904a-904e with access to software for database, search, calculation, graphics, audio, video, World Wide Web / Internet usage, and / or other functions. Many other examples of server devices are possible as well.
[0067] Computing Device Architecture 10 is a block diagram of an exemplary computing device 1000, according to an example embodiment. Specifically, the computing device 1000 shown in FIG. 10 may be configured to perform at least one function of and / or associated with the neural networks and / or methods 1200, 1300 described herein.
[0068] The computing device 1000 may include a user interface module 1001, a network communication module 1002, one or more processors 1003, data storage 1004, one or more cameras 1018, one or more sensors 1020, and a power system 1022, all of which may be linked to each other via a system bus, network, or other connection mechanism 1005.
[0069] The user interface module 1001 may be operable to transmit data to and / or receive data from external user input / output devices. For example, the user interface module 1001 may be configured to transmit data to and / or receive data from user input devices such as a touchscreen, a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a voice recognition module, and / or other similar devices. The user interface module 1001 may also be configured to provide output to a user display device such as one or more cathode ray tubes (CRTs), liquid crystal displays, light emitting diodes (LEDs), displays using digital light processing (DLP) technology, printers, light bulbs, and / or other similar devices now known or later developed. The user interface module 1001 may also be configured to generate audible output using devices such as speakers, speaker jacks, audio output ports, audio output devices, earphones, and / or other similar devices. User interface module 1001 may further be configured with one or more haptic devices that may generate tactile output, such as vibration and / or other output detectable by touch and / or physical contact with computing device 1000. In some examples, user interface module 1001 may be used to provide a graphical user interface (GUI) for utilizing computing device 1000, such as, for example, the graphical user interface of a mobile phone device.
[0070] The network communication module 1002 may include one or more devices providing one or more wireless interfaces 1007 and / or one or more wired interfaces 1008 configurable to communicate over a network. The wireless interface(s) 1007 may include one or more wireless transmitters, receivers, and / or transceivers, such as a Bluetooth® transceiver, a Zigbee® transceiver, a Wi-Fi™ transceiver, a WiMAx™ transceiver, an LTE™ transceiver, and / or other type of wireless transceiver configurable to communicate over a wireless network. The wired interface(s) 1008 may include one or more wired transmitters, receivers, and / or transceivers, such as an Ethernet® transceiver, a Universal Serial Bus (USB) transceiver, or similar transceiver configurable to communicate over a twisted pair wire, coaxial cable, fiber optic link, or similar physical connection to a wired network.
[0071] In some examples, the network communications module 1002 may be configured to provide reliable, secure, and / or authenticated communications. For each communication described herein, information to facilitate reliable communications (e.g., guaranteed message delivery) may be provided, perhaps as part of the message header and / or footer (e.g., packet / message ordering information, encapsulation header and / or footer, size / time information, and transmission verification information such as a cyclic redundancy check (CRC) and / or parity check value). Communications may be protected (e.g., encoded or encrypted) and / or decrypted / decrypted using one or more cryptographic protocols and / or algorithms, such as, but not limited to, the Data Encryption Standard (DES), the Advanced Encryption Standard (AES), the Rivest-Shamir-Adelman (RSA) algorithm, the Diffie-Hellman algorithm, a secure socket protocol such as Secure Sockets Layer (SSL) or Transport Layer Security (TLS), and / or the Digital Signature Algorithm (DSA). Other encryption protocols and / or algorithms may be used similar to or in addition to those listed herein to secure (and subsequently decrypt / decrypt) communications.
[0072] The one or more processors 1003 may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., digital signal processors, tensor processing units (TPUs), graphics processing units (GPUs), application-specific integrated circuits, etc.) The one or more processors 1003 may be configured to execute computer-readable instructions 8306 contained in data storage 1004 and / or other instructions described herein.
[0073] Data storage 1004 may include one or more non-transitory computer-readable storage media that can be read and / or accessed by at least one of the one or more processors 1003. The one or more computer-readable storage media may include volatile and / or non-volatile storage components, such as optical storage, magnetic storage, organic storage, or other memory or disk storage, which may be integrated, in whole or in part, with at least one of the one or more processors 1003. In some examples, data storage 1004 may be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disk storage unit), while in other examples, data storage 1004 may be implemented using two or more physical devices.
[0074] The data storage 1004 may include computer-readable instructions 1006 and possibly additional data. In some examples, the data storage 1004 may include storage necessary to perform at least some of the methods, scenarios, and techniques described herein and / or at least some of the functionality of the devices and networks described herein. In some examples, the data storage 1004 may include storage of trained neural network models 1012 (e.g., models of trained neural networks, such as the neural network models described herein). In particular, in these examples, the computer-readable instructions 8306 may include instructions that, when executed by the one or more processors 1003, enable the computing device 1000 to provide some or all of the functionality of the trained neural network models 1012.
[0075] In some examples, computing device 1000 may include one or more cameras 1018. Camera(s) 1018 may include one or more image capture devices, such as still cameras and / or video cameras, equipped to capture light and record the captured light into one or more images. That is, camera(s) 1018 may generate image(s) of the captured light. The one or more images may be one or more still images and / or one or more images utilized in video capture. Camera(s) 1018 may capture light and / or electromagnetic radiation emitted as visible light, infrared radiation, ultraviolet light, and / or as light of one or more other frequencies.
[0076] In some examples, computing device 1000 may include one or more sensors 1020. Sensors 1020 may be configured to measure conditions within computing device 1000 and / or conditions in the environment of computing device 1000 and provide data regarding these conditions.For example, the sensors 1020 may be (i) sensors for acquiring data about the computing device 1000, such as, but not limited to, a thermometer for measuring the temperature of the computing device 1000, a battery sensor for measuring the power of one or more batteries of the power supply system 1022, and / or other sensors for measuring the state of the computing device 1000; (ii) identification sensors for identifying other objects and / or devices, such as, but not limited to, a radio frequency identification (RFID) reader, a proximity sensor, a one-dimensional barcode reader, a two-dimensional barcode (e.g., a quick response (QR) code) reader, and a laser tracker, which may be configured to read identifiers such as RFID tags, barcodes, QR codes, and / or other devices and / or objects configured to read and provide at least identification information; (iii) tilt sensors, gyroscopes, accelerometers, The sensors 1020 may include one or more of the following: (i) sensors that measure the position and / or movement of the computing device 1000, such as, but not limited to, a Doppler sensor, a GPS device, a sonar sensor, a radar device, a laser displacement sensor, and a compass; (ii) environmental sensors that acquire data indicative of the environment of the computing device 1000, such as, but not limited to, an infrared sensor, an optical sensor, a light sensor, a biosensor, a capacitive sensor, a touch sensor, a temperature sensor, a wireless sensor, a radio sensor, a movement sensor, a microphone, a sound sensor, an ultrasonic sensor, and / or a smoke sensor; and / or (iii) force sensors that measure one or more forces (e.g., inertial forces and / or G-forces) acting about the computing device 1000, such as, but not limited to, one or more sensors that measure force, torque, ground force, friction in one or more dimensions, and / or a zero moment point (ZMP) sensor that identifies the ZMP and / or the location of the ZMP. Many other examples of sensors 1020 are possible as well.
[0077] The power supply system 1022 may include one or more batteries 1024 and / or one or more external power interfaces 1026 for providing power to the computing device 1000. Each battery of the one or more batteries 1024, when electrically coupled to the computing device 1000, can serve as a source of stored power for the computing device 1000. The one or more batteries 1024 of the power supply system 1022 may be configured to be portable. Some or all of the one or more batteries 1024 may be easily removable from the computing device 1000. In other examples, some or all of the one or more batteries 1024 may be internal to the computing device 1000 and therefore may not be easily removable from the computing device 1000. Some or all of the one or more batteries 1024 may be rechargeable. For example, a rechargeable battery may be recharged via a wired connection between the battery and another power source, such as by one or more power sources external to the computing device 1000 and connected to the computing device 1000 via one or more external power interfaces. In other examples, some or all of the one or more batteries 1024 may be non-rechargeable batteries.
[0078] The one or more external power interfaces 1026 of the power system 1022 may include one or more wired power interfaces, such as a USB cable and / or a power cord, that enable a wired power connection to one or more power sources external to the computing device 1000. The one or more external power interfaces 1026 may include one or more wireless power interfaces, such as a Qi wireless charger, that enable a wireless power connection to one or more external power sources, such as via a Qi wireless charger. Once a power connection to an external power source is established using the one or more external power interfaces 1026, the computing device 1000 may draw power from the external power source via the established power connection. In some examples, the power system 1022 may include associated sensors, such as battery sensors or other types of power sensors associated with one or more batteries.
[0079] Cloud-based Server FIG. 11 illustrates a cloud-based server system according to an example embodiment. In FIG. 14, neural network functionality and / or computing devices can be distributed across computing clusters 1109a, 1109b, and 1109c. Computing cluster 1109a may include one or more computing devices 1100a, cluster storage array 1110a, and cluster router 1111a connected by a local cluster network 1113a. Similarly, computing cluster 1109b may include one or more computing devices 1100b, cluster storage array 1110b, and cluster router 1111b connected by a local cluster network 1113b. Similarly, computing cluster 1109c may include one or more computing devices 1100c, cluster storage array 1110c, and cluster router 1111c connected by a local cluster network 1113c.
[0080] In some embodiments, computing clusters 1109a, 1109b, 1109c may be a single computing device residing in a single computing center. In other embodiments, computing clusters 1109a, 1109b, 1109c may include multiple computing devices within a single computing center, or even multiple computing devices located in multiple computing centers located in various geographic locations. For example, FIG. 11 illustrates each of computing clusters 1109a, 1109b, 1109c residing in different physical locations.
[0081] In some embodiments, the data and services in computing clusters 1109a, 1109b, 1109c may be stored on a non-transitory, tangible computer-readable medium (or computer-readable storage medium) and encoded as computer-readable information accessible by other computing devices. In some embodiments, computing clusters 1109a, 1109b, 1109c may be stored on a single disk drive or other tangible storage medium, or may be implemented on multiple disk drives or other tangible storage media located in one or more diverse geographic locations.
[0082] In some embodiments, each of computing clusters 1109a, 1109b, and 1109c may have an equal number of computing devices, an equal number of cluster storage arrays, and an equal number of cluster routers. However, in other embodiments, each computing cluster may have a different number of computing devices, a different number of cluster storage arrays, and a different number of cluster routers. The number of computing devices, cluster storage arrays, and cluster routers in each computing cluster may depend on one or more computing tasks assigned to each computing cluster.
[0083] In computing cluster 1109a, for example, computing device 1100a may be configured to perform various computing tasks of a conditioned axial self-attention-based neural network and / or computing devices. In one embodiment, various functions of the neural network and / or computing devices may be distributed among one or more of computing devices 1100a, 1100b, and 1100c. Computing devices 1100b and 1100c in respective computing clusters 1109b and 1109c may be configured similarly to computing device 1100a in computing cluster 1109a. Meanwhile, in some embodiments, computing devices 1100a, 1100b, and 1100c may be configured to perform different functions.
[0084] In some embodiments, computing tasks and stored data associated with the neural network and / or computing devices may be distributed across the computing devices 1100a, 1100b, and 1100c based at least in part on the processing requirements of the neural network and / or computing devices, the processing capabilities of the computing devices 1100a, 1100b, and 1100c, the latency of network links between computing devices within each computing cluster and between the computing clusters themselves, and / or other factors that may contribute to the cost, speed, fault tolerance, resilience, efficiency, and / or other design goals of the overall system architecture.
[0085] The cluster storage arrays 1110a, 1110b, 1110c of the computing clusters 1109a, 1109b, 1109c may be data storage arrays that include a disk array controller configured to manage read and write access to a group of hard disk drives. The disk array controllers, alone or in combination with their respective computing devices, may also be configured to manage backup or redundant copies of data stored in the cluster storage arrays to protect against disk drive or other cluster storage array failures and / or network failures that prevent one or more computing devices from accessing one or more cluster storage arrays.
[0086] Similar to how the functionality of a conditioned axial self-attention-based neural network and / or computing devices may be distributed across computing devices 1100a, 1100b, 1100c of computing clusters 1109a, 1109b, 1109c, various active and / or backup portions of these components may be distributed across cluster storage arrays 1110a, 1110b, 1110c. For example, some cluster storage arrays may be configured to store some portions of data for a first layer of a neural network and / or computing devices, while other cluster storage arrays may store other portions(s) of data for a second layer of a neural network and / or computing devices. Also, for example, some cluster storage arrays may be configured to store data for an encoder of a neural network, while other cluster storage arrays may store data for a decoder of the neural network. Additionally, some cluster storage arrays may be configured to store backup versions of data stored in other cluster storage arrays.
[0087] Cluster routers 1111a, 1111b, 1111c in computing clusters 1109a, 1109b, 1109c may include network equipment configured to provide internal and external communications for the computing clusters. For example, cluster router 1111a in computing cluster 1109a may include one or more Internet switching and routing devices configured to provide (i) local area network communications between computing device 1100a and cluster storage array 1110a via local cluster network 1113A, and (ii) wide area network communications between computing cluster 1109a and computing clusters 1109b and 1109c via wide area network link 1113a to network 906. Cluster routers 1111b and 1111c may include network equipment similar to cluster router 1111a, and cluster routers 1111b and 1111c may perform networking functions for computing clusters 1109b and 1109b similar to that performed by cluster router 1111a for computing cluster 1109a.
[0088] In some embodiments, the configuration of the cluster routers 1111a, 1111b, 1111c may be based at least in part on the data communication requirements of the computing devices and cluster storage arrays, the data communication capabilities of the network equipment within the cluster routers 1111a, 1111b, 1111c, the latency and throughput of the local cluster networks 1113A, 1113B, 1113C, the latency, throughput, and cost of the wide area network links 1113a, 1113b, 1113c, and / or other factors that may contribute to the cost, speed, fault tolerance, resilience, efficiency, and / or other design criteria of the mitigation system architecture.
[0089] Exemplary Methods of Operation 12 is a flowchart of a method 1200 according to an example embodiment. The method 1200 may be performed by a computing device such as the computing device 1000.
[0090] The method 1200 may begin at block 1210, where the method includes determining a respective delta quality score associated with each of a plurality of images, where determining the delta quality score includes predicting an enhanced image corresponding to a given image of the plurality of images by an image enhancement model, where the image enhancement model is trained to remove one or more image degradations associated with the given image, where determining the delta quality score further includes determining a first quality score associated with the given image and a second quality score associated with the predicted enhanced image, where the delta quality score is based on a difference between the second quality score and the first quality score, where the delta quality score indicates a degree of image enhancement in the predicted enhanced image.
[0091] At block 1220, the method includes generating, by a computing device, a training data set including a plurality of images associated with respective delta quality scores.
[0092] At block 1230, the method includes training a quality assessment model to predict a quality improvement possibility score associated with the input image based on the generated training dataset, the quality improvement possibility score indicating a possibility of improving the perceptual quality of the input image based on removal of one or more image degradation factors.
[0093] At block 1240, the method includes outputting, by the computing device, the trained quality assessment model.
[0094] In some embodiments, the quality assessment model may be a convolutional neural network, and training the quality assessment model includes receiving labeled data indicating the degree of image enhancement in the predicted enhanced image as perceived by a human annotator. Such embodiments include fine-tuning a final layer of the convolutional neural network using the received labeled data.
[0095] In some embodiments, the convolutional neural network comprises a MobileNet architecture.
[0096] In some embodiments, the convolutional neural network includes a fully connected layer configured to determine a delta quality score.
[0097] In some embodiments, the first quality score and the second quality score may be Neural Image Assessment (NIMA) scores.
[0098] In some embodiments, the first quality score and the second quality score may be generated by an AlexNet-based convolutional neural network (CNN) trained with an aesthetic visual analysis (AVA) rank-based loss function.
[0099] In some embodiments, the one or more image degradation factors include one or more of motion blur, lens blur, image noise, image compression artifacts, or artifacts caused by saturated pixels.
[0100] 13 is a flowchart of a method 1300 according to an example embodiment. The method 1300 may be performed by a computing device such as the computing device 1000.
[0101] The method 1300 may begin at block 1310, where the method includes receiving an input image by a computing device.
[0102] At block 1320, the method includes predicting a quality improvement possibility score associated with the input image by a quality assessment model, the quality improvement possibility score indicating a possibility of improving the perceptual quality of the input image based on removal of one or more image degradation factors, the quality assessment model being trained with a training dataset including a plurality of images associated with respective delta quality scores, the delta quality score being determined by predicting an enhanced image corresponding to a given image of the plurality of images by an image enhancement model, the image enhancement model being trained to remove one or more image degradations associated with the given image, the delta quality score being further determined by determining a first quality score associated with the given image and a second quality score associated with the predicted enhanced image, the delta quality score being based on a difference between the second quality score and the first quality score, and the delta quality score indicating a degree of image enhancement in the predicted enhanced image.
[0103] At block 1330, the method includes providing, by the computing device, an alert notification based on the predicted quality improvement likelihood score.
[0104] In some embodiments, the quality assessment model may be a convolutional neural network.
[0105] In some embodiments, the convolutional neural network comprises a MobileNet architecture.
[0106] In some embodiments, the convolutional neural network includes a fully connected layer configured to determine a delta quality score.
[0107] In some embodiments, the first quality score and the second quality score may be Neural Image Assessment (NIMA) scores.
[0108] In some embodiments, the first quality score and the second quality score may be generated by an AlexNet-based convolutional neural network (CNN) trained with an aesthetic visual analysis (AVA) rank-based loss function.
[0109] Some embodiments include determining whether the predicted quality improvement possibility score exceeds a threshold score. Such embodiments include, based on a determination that the predicted quality improvement possibility score exceeds the threshold score, providing the input image to an image enhancement model to enhance the quality of the input image.
[0110] In some embodiments, the one or more image degradation factors include image blurring, and the threshold score may be a threshold deblurring score.
[0111] In some embodiments, the one or more image degradation factors include noise in the image, and the threshold score may be a threshold denoising score.
[0112] In some embodiments, the one or more image degradation factors include image compression artifacts, and the threshold score may be a threshold compression artifact removal score.
[0113] In some embodiments, the one or more image degradation factors include artifacts caused by saturated pixels, and the threshold score may be a threshold saturated pixel artifact removal score.
[0114] In some embodiments, providing the alert notification includes triggering the alert notification upon determining that the predicted quality improvement likelihood score exceeds a threshold score.
[0115] In some embodiments, providing the alert notification includes providing a recommendation to the user for enhancing the input image upon determining that the predicted quality improvement potential score exceeds the threshold score.
[0116] Some embodiments include receiving a user instruction to enhance the input image. Such embodiments include, in response to the user instruction, providing the input image to an image enhancement model to enhance the input image.
[0117] In some embodiments, the image enhancement model may be one or more of a deblurring model, a colorization model, an image artifact removal model, or a denoising model.
[0118] In some embodiments, the one or more image degradation factors include one or more of motion blur, or lens blur.
[0119] The present disclosure should not be limited by the specific embodiments described in this application, which are intended as illustrations of various aspects. As will be apparent to those skilled in the art, many modifications and variations can be made without departing from the spirit and scope thereof. Functionally equivalent methods and apparatuses within the scope of the present disclosure, in addition to those enumerated herein, will be apparent to those skilled in the art from the above description. Such modifications and variations are intended to be included within the scope of the appended claims.
[0120] The foregoing detailed description describes various features and functions of the disclosed systems, devices, and methods with reference to the accompanying drawings. In the figures, like symbols generally identify like elements unless context dictates otherwise. The illustrative embodiments described in the detailed description, figures, and claims are not intended to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that aspects of the present disclosure, as generally described herein and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are expressly contemplated herein.
[0121] With respect to any or all of the ladder diagrams, scenarios, and flowcharts in the figures and as discussed herein, each step, block, and / or communication may represent the processing of information and / or the transmission of information according to the exemplary embodiments. Alternative embodiments are included within the scope of these exemplary embodiments. In these alternative embodiments, for example, functions described as blocks, transmissions, communications, requests, responses, and / or messages may be performed out of the order shown or discussed, including substantially simultaneously or in reverse order, depending on the functionality involved. Furthermore, more or fewer blocks and / or functions may be used in any of the ladder diagrams, scenarios, and flowcharts discussed herein, and these ladder diagrams, scenarios, and flowcharts may be combined with each other, either in part or in whole.
[0122] The blocks representing processing of information may correspond to circuitry that can be configured to perform specific logical functions of the methods or techniques described herein. Alternatively or additionally, the blocks representing processing of information may correspond to modules, segments, or portions of program code (including associated data). The program code may include one or more instructions executable by a processor to implement specific logical functions or operations in the method or technique. The program code and / or associated data may be stored on any type of computer-readable medium, such as a storage device, including a disk or hard drive, or other storage medium.
[0123] Computer-readable media may also include non-transitory computer-readable media, such as register memory, processor cache, and non-transitory computer-readable media that store data for a short period of time, such as random access memory (RAM). Computer-readable media may also include non-transitory computer-readable media that store program code and / or data for a longer period of time, such as secondary or permanent long-term storage devices, such as read-only memory (ROM), optical or magnetic disks, and compact disk read-only memory (CD-ROM). Computer-readable media may also be any other volatile or non-volatile storage system. Computer-readable media may be considered, for example, to be a computer-readable storage medium or a tangible storage device.
[0124] Additionally, blocks representing one or more information transfers may correspond to information transfers between software and / or hardware modules within the same physical device, although other information transfers may be between software and / or hardware modules in different physical devices.
[0125] While various aspects and embodiments are disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are provided by way of example and are not intended to be limiting, with the true scope being set forth in the claims that follow.
Claims
1. 1. A computer-implemented method, the method comprising: determining a respective delta quality score associated with each of the plurality of images, wherein determining the delta quality scores comprises: an image enhancement model for predicting an enhanced image corresponding to a given image of the plurality of images, the image enhancement model being trained to remove one or more image impairments associated with the given image, and determining the delta quality score further comprising: determining a first quality score associated with the given image and a second quality score associated with the predicted enhanced image; The delta quality score is based on a difference between the second quality score and the first quality score, the delta quality score indicating a degree of image enhancement in the predicted enhanced image, the method further comprising: generating, by a computing device, a training data set comprising the plurality of images associated with respective delta quality scores; and training a quality assessment model to predict a quality improvement possibility score associated with an input image based on the generated training dataset, the quality improvement possibility score indicating a possibility of improving the perceptual quality of the input image based on removal of one or more image degradation factors, the method further comprising: The computer-implemented method includes the computing device outputting the trained quality assessment model.
2. The quality assessment model is a convolutional neural network, and training the quality assessment model includes: receiving labeled data indicating the degree of the image enhancement in the predicted enhanced image as perceived by a human annotator; and fine-tuning a final layer of the convolutional neural network using the received labeled data.
3. 3. The computer-implemented method of claim 2, wherein the convolutional neural network comprises a MobileNet architecture.
4. 3. The computer-implemented method of claim 2, wherein the convolutional neural network includes a fully connected layer configured to determine the delta quality score.
5. The computer-implemented method of claim 1 , wherein the first quality score and the second quality score are Neural Image Assessment (NIMA) scores.
6. 2. The computer-implemented method of claim 1, wherein the first quality score and the second quality score are generated by an AlexNet-based convolutional neural network (CNN) trained with an aesthetic visual analysis (AVA) rank-based loss function.
7. 2. The computer-implemented method of claim 1, wherein the one or more image degradation factors include one or more of motion blur, lens blur, image noise, image compression artifacts, or artifacts caused by saturated pixels.
8. 1. A computer-implemented method, the method comprising: a computing device receiving an input image; and a quality assessment model predicting a quality improvement possibility score associated with the input image, the quality improvement possibility score indicating a possibility of improving the perceptual quality of the input image based on removal of one or more image degradation factors, the quality assessment model being trained with a training dataset including a plurality of images associated with respective delta quality scores, the delta quality scores being: an image enhancement model is determined by predicting an enhanced image corresponding to a given image of the plurality of images, the image enhancement model being trained to remove one or more image impairments associated with the given image, and the delta quality score further comprising: by determining a first quality score associated with the given image and a second quality score associated with the predicted enhanced image; The delta quality score is based on a difference between the second quality score and the first quality score, the delta quality score indicating a degree of image enhancement in the predicted enhanced image, the method further comprising: The computer-implemented method includes the computing device providing an alert notification based on the predicted quality improvement likelihood score.
9. 9. The computer-implemented method of claim 8, wherein the quality assessment model is a convolutional neural network.
10. 10. The computer-implemented method of claim 8, wherein the convolutional neural network comprises a MobileNet architecture.
11. 9. The computer-implemented method of claim 8, wherein the convolutional neural network includes a fully connected layer configured to determine the delta quality score.
12. 9. The computer-implemented method of claim 8, wherein the first quality score and the second quality score are Neural Image Assessment (NIMA) scores.
13. 9. The computer-implemented method of claim 8, wherein the first quality score and the second quality score are generated by an AlexNet-based CNN trained with an aesthetic visual analysis (AVA) rank-based loss function.
14. determining whether the predicted quality improvement potential score exceeds a threshold score; 10. The computer-implemented method of claim 8, further comprising: based on a determination that the predicted quality improvement possibility score exceeds the threshold score, providing the input image to the image enhancement model to enhance the quality of the input image.
15. 15. The computer-implemented method of claim 14, wherein the one or more image degradation factors include image blurring and the threshold score is a threshold deblurring score.
16. 15. The computer-implemented method of claim 14, wherein the one or more image degradation factors include image noise and the threshold score is a threshold denoising score.
17. 15. The computer-implemented method of claim 14, wherein the one or more image degradation factors include image compression artifacts and the threshold score is a threshold compression artifact removal score.
18. 15. The computer-implemented method of claim 14, wherein the one or more image degradation factors include artifacts caused by saturated pixels, and the threshold score is a threshold saturated pixel artifact removal score.
19. providing the alert notification, 15. The computer-implemented method of claim 14, comprising triggering the alert notification upon a determination that the predicted quality improvement potential score exceeds the threshold score.
20. providing the alert notification, 15. The computer-implemented method of claim 14, comprising, upon a determination that the predicted quality improvement potential score exceeds the threshold score, providing a user with recommendations for enhancing the input image.
21. receiving user instructions to enhance the input image; 21. The computer-implemented method of claim 20, further comprising: in response to the user instruction, providing the input image to the image enhancement model for enhancing the input image.
22. 21. The computer-implemented method of claim 20, wherein the image enhancement model is one or more of a deblurring model, a colorization model, an image artifact removal model, or a denoising model.
23. The computer-implemented method of claim 8 , wherein the one or more image degradation factors include one or more of motion blur or lens blur.
24. 1. A computing device comprising: one or more processors; and a data storage storing computer-executable instructions that, when executed by the one or more processors, cause the computing device to perform functions including the computer-implemented method of any one of claims 1 to 23.
25. The computing device of claim 24, wherein the computing device is a mobile device.
26. A computer program comprising instructions which, when executed by a computer, cause the computer to carry out the steps according to the method of any one of claims 1 to 23.
27. 24. An article of manufacture comprising one or more non-transitory computer readable media having stored thereon computer readable instructions that, when executed by one or more processors of a computing device, cause the computing device to perform functions comprising the computer implemented method of any one of claims 1 to 23.
28. 1. A system comprising: A system comprising means for performing the computer-implemented method of any one of claims 1 to 23.
Citation Information
Patent Citations
Difference metric for machine learning-based processing systems
US20180240017A1