Cascaded Multiresolution Machine Learning for Image Processing with Improved Computational Efficiency
The cascaded multi-resolution approach for image processing optimizes resource usage by using two machine learning models to process low-resolution images and upscaled subsets, enhancing image quality on resource-constrained devices.
Patent Information
- Application Number
- JP2024519724
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-01
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2041-10-01
AI Technical Summary
The high computational demands of machine learning models for image processing at high resolutions make them impractical for resource-constrained devices, leading to degraded image quality when operating at lower resolutions.
A cascaded multi-resolution approach using two machine learning models, where the first model processes a low-resolution image to capture semantic information and the second model processes upscaled subsets for detailed corrections, optimizing resource usage and quality.
This method achieves high-quality image processing with reduced computational resources by leveraging semantic information across the entire image while conserving processor and memory usage.
Smart Images

Figure 0007715937000001 
Figure 0007715937000002 
Figure 0007715937000003
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to image processing such as image correction. More particularly, the present disclosure relates to systems and methods for cascade multi-resolution machine learning for image processing with improved computational efficiency.
Background Art
[0002] Image processing can include the modification of digital images to have a modified appearance. Exemplary image corrections include smoothing, blurring, deblurring, and / or many other operations. Some image corrections include generative corrections where new image data is generated and inserted into the image as a replacement for the original image data. Some exemplary generative corrections may be referred to as "inpainting".
[0003] Image processing can also include the analysis of images to identify or determine characteristics of the images. For example, image processing can include techniques such as semantic segmentation, object detection, object recognition, edge detection, human keypoint estimation, and / or various other image analysis algorithms or tasks.
[0004] One of the major issues related to the use of machine learning models for image processing is the constraints on the input and output image resolutions. Specifically, the higher the resolution, the greater the increase in memory usage and latency. Thus, operating a machine learning model to perform image processing on any reasonably sized image consumes a significant amount of computing resources, such as memory usage and processor usage. This makes it significantly difficult to use machine learning models at high resolutions and even impossible in some resource-constrained environments, such as "on-device" on a computing device (e.g., a smartphone) with few or limited computing resources. As an example, the standard resolution of a typical machine learning model can be in the range of 512×512, which is already extremely large for operating on a smartphone.
[0005] One solution to the computational challenges described above is to operate the machine learning model on images with lower resolutions. This can conserve or reduce the amount of resources consumed. However, processing images at lower resolutions degrades the quality of the processing output and thus has its own drawbacks. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0006] Aspects and advantages of embodiments of the present disclosure are described in part in the following description, or can be learned from the description, or can be learned through the practice of the embodiments.
[0007] One exemplary aspect of the present disclosure is directed to a computing system for image correction with improved computational efficiency. The computing system includes one or more processors and one or more non-transitory computer-readable media that store instructions collectively. The instructions, when executed by the one or more processors, cause the computing system to perform operations. The operations include obtaining a low-resolution version of an input image, where the low-resolution version of the input image has a first resolution and includes one or more image elements to be corrected using predicted image data. The operations include processing the low-resolution version of the input image using a first machine learning model to generate an extended image having the first resolution, where the extended image includes first predicted image data for replacing one or more image elements. The operations include extracting a portion of the extended image, where the portion of the extended image includes the first predicted image data. The operations include upscaling the extracted portion of the extended image to generate an upscaled image portion having an upscaled resolution. The operations include processing the upscaled image portion using a second machine learning model to generate an improved portion, where the improved portion includes second predicted image data for correcting at least a portion of the first predicted image data. The operations include generating an output image based on the improved portion and a high-resolution version of the input image, where both the output image and the high-resolution version of the input image have a second resolution greater than the first resolution. The operations include providing the output image as an output.
[0008] Another exemplary aspect of the present disclosure is directed to a computer-implemented method for training a machine learning model to perform image correction. The method includes receiving, by a computing system comprising one or more processors, a low-resolution version of an input image and a ground truth image, the low-resolution version of the input image having a first resolution and the ground truth image having a second resolution greater than the first resolution, the low-resolution version of the input image comprising one or more image elements not present in the ground truth image. The method includes processing, by the computing system, the low-resolution version of the input image using a first machine learning model to generate a low-resolution version of an enhanced image having the first resolution, the low-resolution version of the enhanced image comprising first prediction data replacing one or more image elements. The method includes upscaling, by the computing system, the low-resolution version of the enhanced image to generate a high-resolution version of the enhanced image having the second resolution. The method includes processing, by the computing system, at least a portion of the high-resolution version of the enhanced image using a second machine learning model to generate a predicted image having the second resolution. The method includes evaluating, by the computing system, a loss function that evaluates a difference between the predicted image and the ground truth image. The method includes adjusting one or more parameters of at least one of the first machine learning model or the second machine learning model based at least in part on the loss function.
[0009] Another exemplary aspect of the present disclosure is directed to one or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more processors, cause a computing system to perform operations. The operations include obtaining a low-resolution version of an input image, the low-resolution version of the input image having a first resolution. The operations include processing the low-resolution version of the input image using a first machine learning model to generate a first predicted image having the first resolution, the first predicted image comprising first predicted image data. The operations include extracting a portion of the first predicted image, the portion of the first predicted image comprising first predicted image data. The operations include upscaling the extracted portion of the first predicted image to generate an upscaled image portion having an upscaled resolution. The operations include processing the upscaled image portion using a second machine learning model to generate a second predicted image, the second predicted image comprising second predicted image data that modifies at least a portion of the first predicted image data.
[0010] Other aspects of the present disclosure are directed to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.
[0011] These and other features, aspects, and advantages of the various embodiments of the present disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated herein and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the relevant principles.
[0012] A detailed description of embodiments directed to those of ordinary skill in the art is set forth herein with reference to the accompanying figures. BRIEF DESCRIPTION OF THE DRAWINGS
[0013]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 3C
[0014] Reference numbers repeated across multiple figures are intended to identify the same features in various implementations.
[0015] **Overview** Generally, the present disclosure is directed to systems and methods for image processing such as image correction. More particularly, exemplary aspects of the present disclosure are directed to systems and methods for cascaded multi - resolution machine learning for performing image processing on resource - constrained devices.
[0016] In one exemplary approach, an image processing system includes two machine - learning components. Specifically, a first machine - learning model can perform image processing (e.g., image correction such as inpainting) on the entire input image at a lower resolution. A second machine - learning model can perform image processing (e.g., image correction such as inpainting) on only one or more selected subsets ("crops") of the output of the first model that have been upscaled to a higher resolution.
[0017] In such a way, the first model can utilize the context information and / or semantic information contained across the entire image to perform an initial trial in the image processing task. However, since the first model operates at a lower resolution, the computational consumption of the first model may be relatively small.
[0018] Next, the second model can perform more detailed and higher-quality image processing on a selected subset of the output of the first model. Specifically, since the second model operates at a higher resolution, the output of the second model is generally of higher quality and / or more detailed compared to the output of the first model. However, since the second model operates only on the selected subset, the computational consumption of the second model can be kept at a lower reduced level (compared to, for example, operating the second model on the entire input at the higher resolution).
[0019] In some implementations, the output of the second model can be used alone. In other implementations, the output of the second model can be combined with the original higher-resolution input to produce a complete higher-resolution output. In other implementations, the output of the second model can be combined with an upscaled version of the output of the first model to generate a complete higher-resolution output.
[0020] In some implementations, both the first model and the second model are trained together. For example, a loss can be determined based on the output of the second model. The loss can be backpropagated through the second model and then through the first model to train the second model and / or the first model.
[0021] The systems and methods of the present disclosure provide several technical effects and advantages. As one exemplary technical effect, the systems and methods of the present disclosure provide an improved trade-off between image processing quality and computing resource usage. For example, compared to a system that performs image processing only on a high-resolution crop of an input image, the proposed system can provide improved quality. This is because, in many cases, high-quality image processing requires access to semantic information not only from the information contained within a smaller crop but also from the entire image. Thus, by processing a lower-resolution version of the entire image using a first model before processing a higher-resolution version of the crop using a second model, the proposed system can have access to semantic information contained not only in the cropped portion but also across the entire image while maintaining an acceptable level of computing resource usage for all. Similarly, compared to a system that processes the entire input image at a higher resolution (which may not be possible or desirable in some computing environments), the proposed system can save computing resources such as processor usage and memory usage. Thus, high-quality image processing results can be obtained even in a computing environment with constrained computing resources.
[0022] In one example, the systems described herein can be implemented as part of or in cooperation with a camera application. For example, a camera can capture an image, and the systems and methods described herein can be used to process (e.g., modify) the image as part of a camera application or as a service for a camera application. This can enable a user to process an image that the user captures or uploads or otherwise provides as input (e.g., modify it to remove unwanted objects therefrom).
[0023] Exemplary embodiments of the present disclosure are described in further detail with reference to the figures herein.
[0024] Exemplary Image Processing Flow FIG. 1 shows an exemplary flow for performing image correction with improved computational efficiency. As a specific example, the image correction task may be a restoration where selected (e.g., user-selected) elements of the input image are “filled in” based on information from areas surrounding the input image. This may be used, for example, to “fill in” selected defects, malfunctions, etc., for example, to expand the image. FIG. 1 provides an exemplary flow in the context of an exemplary image processing task of image correction (e.g., restoration), but the disclosed techniques may be applied to other image processing tasks.
[0025] As shown in FIG. 1, the computing system can obtain a low-resolution version 16 of the input image. The low-resolution version 16 of the input image can have a first resolution. The low-resolution version 16 of the input image can include one or more image elements to be corrected using predicted image data. As an example, in FIG. 1, the low-resolution version 16 of the input image includes an undesirable image element 14, and the system attempts to replace the image element 14 via restoration.
[0026] In some implementations, the computing system can obtain a low-resolution version 16 of the input image by downscaling a high-resolution version 12 of the input image. For example, the high-resolution version 12 of the input image can be obtained from the imaging pipeline of a camera system, uploaded or selected by a user, and / or obtained via various other means by which the input image may be subjected to the illustrated process. It can be the original version of the input image.
[0027] Referring further to FIG. 1, the computing system can process the low-resolution version 16 of the input image using the first machine learning model 20 to generate an enhanced image 22 having a first resolution. The enhanced image can include first predicted image data that modifies one or more image elements 14.
[0028] The first machine learning model 20 can be various forms of machine learning models such as neural networks. In one example, the first machine learning model 20 can be a convolutional neural network. In one example, the first machine learning model 20 can be a transformer model that uses self-attention. In one example, the first machine learning model 20 can have an encoder-decoder architecture.
[0029] In some implementations, the first machine learning model 20 can perform image correction tasks such as, for example, inpainting, blur correction, color restoration, or smoothing of one or more image elements 14.
[0030] Thus, in some implementations, as shown in FIG. 1, processing the low-resolution version 16 of the input image using the first machine learning model 20 to generate the enhanced image 22 can include processing the low-resolution version 16 of the input image and a mask 18 that identifies one or more image elements 14 using the first machine learning inpainting model to generate an enhanced image 22 having first inpainting image data that modifies one or more image elements.
[0031] In some implementations, one or more image elements 14 to be replaced can include one or more user-specified image elements specified based on one or more user inputs (e.g., inputs to a graphical user interface). Alternatively or additionally, one or more image elements 14 to be replaced can include one or more computer-specified image elements. For example, one or more computer-specified image elements can be computer-specified by processing an input image using at least one of one or more classification sub-blocks of the first machine learning model 20 or the second machine learning model 28.
[0032] In other implementations, in addition to or as an alternative to the exemplary image correction task shown in FIG. 1, an image analysis task can be performed. As an example, in some implementations, the output of the first machine learning model can be a first prediction image including prediction data such as semantic segmentation data, object detection data, object recognition data, face recognition data, human keypoint detection data, edge detection data, and / or other prediction data.
[0033] Referring further to FIG. 1, the computing system can extract a portion 24 of the extended image 22. The extracted portion can comprise an image region corresponding to one or more image elements 14 and thus can be a region specified by one or more user inputs and / or the mask 18. The portion 24 of the extended image can include first prediction image data with one or more image elements 14 modified.
[0034] The computing system can upscale the extracted portion 24 of the extended image 22 to generate an upscaled image portion 26 having an upscaled resolution. Upscaling can include upsampling and / or other forms of increasing the resolution of the extracted portion 24.
[0035] The computing system can process the upscaled image portion 26 using a second machine learning model 28 to generate an improved portion 30. The improved portion 30 can include second predicted image data that modifies at least a portion of the first predicted image data.
[0036] The second machine learning model 20 can be various forms of machine learning models such as a neural network. In one example, the second machine learning model 20 can be a convolutional neural network. In one example, the second machine learning model 20 can be a transformer model that uses self-attention. In one example, the second machine learning model 20 can have an encoder-decoder architecture.
[0037] In some implementations, as shown in FIG. 1, processing the upscaled image portion 26 using a second machine learning model 28 to generate an improved portion 30 can include processing the upscaled image portion 26 using a second machine learning restoration model to generate an improved portion 30 having second restored image data that modifies at least a portion of the first restored image data.
[0038] However, in other implementations, in addition to or as an alternative to the exemplary image correction task shown in FIG. 1, an image analysis task can be performed. As an example, in some implementations, the output of the second machine learning model can be prediction data (e.g., improved prediction data) such as semantic segmentation data, object detection data, object recognition data, face recognition data, human keypoint detection data, edge detection data, and / or other prediction data.
[0039] Referring further to FIG. 1, the computing system can generate an output image 32 based on the improved portion 30 and the high-resolution version 12 of the input image. In some implementations, both the output image 32 and the high-resolution version 12 of the input image have a second resolution that is greater than the first resolution.
[0040] In some implementations, generating the output image 32 based on the improved portion 30 and the high-resolution version 12 of the input image can include inserting the improved portion 30 into the high-resolution version 12 of the input image (e.g., at corresponding locations).
[0041] In some implementations, upscaling the extracted portion 24 of the enhanced image 22 to generate an upscaled image portion 26 having an upscaled resolution can include upscaling the extracted portion 24 of the enhanced image 22 such that the upscaled resolution matches the corresponding resolution of the corresponding portion of the high-resolution version 12 of the input image, where the corresponding portion corresponds proportionally to the extracted portion 24 of the enhanced image. In such a manner, the improved portion 30 can be inserted back into the high-resolution version 12 of the input image having an appropriate size / resolution.
[0042] The computing system can provide the output image 32 as an output. For example, providing an image as an output can include storing the image in memory, sending the image to an additional device, and / or displaying the image.
[0043] In some implementations, the input image can include a plurality of image elements to be modified, replaced, etc. In some such implementations, the computing system can process a low-resolution version 16 of the input image using a first machine learning model only once to generate one output for the entire image. Thereafter, the computing system can separately perform, for each object among a plurality of different objects, extracting, upscaling, and processing the upscaled image portion using a second machine learning model 28. In such a way, multiple object crops can be improved in parallel, reducing latency.
[0044] In some implementations, the computing system can pass one or more internal feature vectors from a first machine learning model 20 to a second machine learning model 28. Thus, latent space information can be shared between models.
[0045] In some implementations, the enhanced image and / or other model outputs can further include a predicted depth channel (e.g., depth data can also be output by a first machine learning model 20 and / or a second machine learning model 28).
[0046] Exemplary training flow FIG. 2 shows a block diagram of an exemplary technique for training cascade multi-resolution machine learning for image processing (e.g., inpainting) according to an exemplary embodiment of the present disclosure.
[0047] As shown in FIG. 2, the computing system can receive a low-resolution version 216 of the input image and a ground truth image 202. The low-resolution version 216 of the input image can have a first resolution, and the ground truth image 202 can have a second resolution that is greater than the first resolution. The low-resolution version 216 of the input image can include one or more image elements 214 that are not present in the ground truth image 202 (e.g., marks of vertical and horizontal lines).
[0048] In some implementations, the low-resolution version 216 of the input image can be obtained by downscaling a high-resolution version 212 of the input image. In some implementations, the high-resolution version 212 of the input image can be obtained by adding one or more image elements 214 to the ground truth image 202.
[0049] Referring further to FIG. 2, the computing system can process the low-resolution version 216 of the input image using a first machine learning model 220 to generate a low-resolution version 222 of an extended image having the first resolution. The low-resolution version 222 of the extended image can include first prediction data that replaces one or more image elements 214.
[0050] In some implementations, a mask 218 can also be supplied as an input to the first machine learning model. The mask 218 can indicate the location of the image elements 214.
[0051] In some implementations, as an alternative or addition to modifying replacing the image elements, the model 220 can predict additional data about the input image, such as semantic segmentation data, object detection data, object recognition data, human keypoint detection data, face recognition data, etc.
[0052] Referring further to FIG. 2, the computing system can upscale the low-resolution version 222 of the enhanced image to generate a high-resolution version 226 of the enhanced image having a second resolution.
[0053] The computing system can process at least a portion of the high-resolution version 226 of the enhanced image using a second machine learning model 228 to generate a predicted image 230 having a second resolution.
[0054] The computing system can evaluate a loss function 232 that evaluates the difference between the predicted image 230 and the ground truth image 202. Exemplary loss terms that can be included in the loss function 232 can include visual loss (e.g., pixel-level loss), VGG loss, GAN loss, and / or other loss terms.
[0055] The computing system can adjust one or more parameters of at least one of the first machine learning model 220 or the second machine learning model 228 based at least in part on the loss function. For example, the loss function 232 can be backpropagated through the second model 228 and then through the first model 220 to train the second model 228 and / or the first model 220.
[0056] Exemplary Devices and Systems FIG. 3A shows a block diagram of an exemplary computing system 100 according to an exemplary embodiment of the present disclosure. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled via a network 180.
[0057] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a gaming console or game controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0058] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor or a plurality of processors operably connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.
[0059] In some implementations, the user computing device 102 can store or include one or more machine learning models 120. For example, the machine learning model 120 can be or can include various machine learning models such as a neural network (e.g., a deep neural network) that includes a non-linear model and / or a linear model, or other types of machine learning models. The neural network can include a feedforward neural network, a recurrent neural network (e.g., a long short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks. Some exemplary machine learning models can utilize an attention mechanism such as self-attention. For example, some exemplary machine learning models can include a multi-head self-attention model (e.g., a transformer model). The exemplary machine learning model 120 will be described with reference to FIGS. 1 and 2.
[0060] In some implementations, one or more machine learning models 120 can be received from the server computing system 130 via the network 180, stored in the user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, the user computing device 102 can implement multiple parallel instances of a single machine learning model 120 (e.g., to perform parallel image processing across multiple instances of an image or image elements).
[0061] Additionally or alternatively, one or more machine learning models 140 can be included within, or otherwise stored and implemented by, a server computing system 130 that communicates with user computing device 102 according to a client-server relationship. For example, a machine learning model 140 can be implemented by server computing system 130 as part of a web service (e.g., an image processing service). Thus, one or more models 120 can be stored and implemented on user computing device 102 and / or one or more models 140 can be stored and implemented on server computing system 130.
[0062] User computing device 102 can also include one or more user input components 122 that receive user input. For example, user input component 122 can be a touch-sensitive component (e.g., a touch-sensitive display screen or touch pad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). The touch-sensitive component can serve to implement a virtual keyboard. Other exemplary user input components include a microphone, a conventional keyboard, or other means by which a user can provide user input.
[0063] Server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be a single processor or multiple processors operably connected. The memory 134 can include one or more non-transitory computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, and combinations thereof. The memory 134 can store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.
[0064] In some implementations, the server computing system 130 includes or is implemented by one or more server computing devices. In cases where the server computing system 130 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0065] As described above, the server computing system 130 can store one or more machine learning models 140 or, alternatively, can include one or more machine learning extended models 140. For example, the model 140 can be or can include various machine learning models. Exemplary machine learning models include neural networks or other multi-layer non-linear models. Exemplary neural networks include feed-forward neural networks, deep neural networks, regression neural networks, and convolutional neural networks. Some exemplary machine learning models can utilize attention mechanisms such as self-attention. For example, some exemplary machine learning models can include multi-head self-attention models (e.g., Transformer models). Exemplary model 140 is described with reference to FIGS. 1 and 2.
[0066] The user computing device 102 and / or the server computing system 130 can train the model 120 and / or 140 via interaction with a training computing system 150 communicatively coupled via the network 180. The training computing system 150 can be separate from the server computing system 130 or can be a part of the server computing system 130.
[0067] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor or multiple processors operably connected. The memory 154 can include one or more non-transitory computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 154 can store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes or is implemented by one or more server computing devices.
[0068] The training computing system 150 can include a model trainer 160 that trains a machine learning model 120 and / or 140 stored in the user computing device 102 and / or the server computing system 130 using various training or learning techniques such as, for example, backpropagation of error. For example, a loss function can be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions can be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques can be used to iteratively update the parameters over a number of training iterations.
[0069] In some implementations, performing backpropagation can include performing truncated backpropagation over time. The model trainer 160 can perform some generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.
[0070] Specifically, the model trainer 160 can train the machine learning models 120 and / or 140 based on a set of training data 162. In some implementations, when the user gives consent, training examples can be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 with respect to user-specific data received from the user computing device 102. In some cases, this process may be referred to as personalizing the model.
[0071] The model trainer 160 includes computer logic utilized to provide the desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, the model trainer 160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions stored on a tangible computer-readable storage medium such as a RAM hard disk or optical or magnetic media.
[0072] Network 180 can be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication via Network 180 can be carried via any type of wired and / or wireless connection using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).
[0073] In some implementations, the input to the machine learning model of the present disclosure can be image data comprising pixel data including a plurality of pixels. The machine learning model can process the pixel data to generate an output. As an example, the machine learning model can process the image data to generate a modified and / or extended image. As another example, the machine learning model can process the image data to generate an image recognition output (e.g., recognition of image data, embedding of potentiality of image data, encoded representation of image data, hash of image data, etc.). As another example, the machine learning model can process the image data to generate an image segmentation output. As another example, the machine learning model can process the image data to generate an image classification output. As another example, the machine learning model can process the image data to generate an image data modification output (e.g., modification of image data, etc.). As another example, the machine learning model can process the image data to generate an encoded image data output (e.g., an encoded and / or compressed representation of image data, etc.). As another example, the machine learning model can process the image data to generate an upscaled image data output. As another example, the machine learning model can process the image data to generate a prediction output.
[0074] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data for one or more images and the task is an image processing task. For example, the image processing task can be image classification where the output is a set of scores, each score corresponding to a different object class and representing the likelihood that one or more images show an object belonging to the object class. The image processing task can be object detection, where the image processing output identifies one or more regions of one or more images and, for each region, the likelihood that the region depicts an object of interest. As another example, the image processing task can be image segmentation, where the image processing output defines, for each pixel in one or more images, a respective likelihood for each category in a predetermined set of categories. For example, the set of categories can be foreground and background. As another example, the set of categories can be object classes. As another example, the image processing task can be depth estimation, where the image processing output defines a respective depth value for each pixel in one or more images. As another example, the image processing task can be motion estimation, where the network input includes multiple images and the image processing output defines, for each pixel in one of the input images, the motion of the scene depicted at the pixel between the images in the network input.
[0075] FIG. 3A shows one exemplary computing system that can be used to implement the present disclosure. Other computing systems can also be used. For example, in some implementations, user computing device 102 can include model trainer 160 and training data set 162. In such implementations, model 120 can be both locally trained and used on user computing device 102. In some of such implementations, user computing device 102 can implement model trainer 160 to customize model 120 based on user-specific data.
[0076] Figure 3B shows a block diagram of an exemplary computing device 10 for execution, according to an exemplary embodiment of the present disclosure. The computing device 10 can be a user computing device or a server computing device.
[0077] The computing device 10 includes several applications (e.g., applications 1 to N). Each application includes its own machine learning library and machine learning model. For example, each application can include a machine learning model. Exemplary applications include text messaging applications, email applications, dictation note applications, virtual keyboard applications, browser applications, and the like.
[0078] As shown in Figure 3B, each application can communicate with several other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.
[0079] Figure 3C shows a block diagram of an exemplary computing device 50 for execution, according to an exemplary embodiment of the present disclosure. The computing device 50 can be a user computing device or a server computing device.
[0080] Computing device 50 includes several applications (e.g., applications 1 - N). Each application communicates with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a voice note application, a virtual keyboard application, a browser application, and the like. In some implementations, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a common API across all applications).
[0081] The central intelligence layer includes several machine - learning models. For example, as shown in FIG. 3C, each machine - learning model can be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine - learning model. For example, in some implementations, the central intelligence layer can provide a single model to all of the applications. In some implementations, the central intelligence layer is included within or otherwise implemented by the operating system of computing device 50.
[0082] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for computing device 50. As shown in FIG. 3C, the central device data layer can communicate with some other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0083] Additional Disclosure The techniques described herein refer to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent between such systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functionality among components. For example, the processes described herein may be implemented using a single device or component or multiple devices or components working in combination. Databases and applications may be implemented on a single system or distributed across multiple systems. Distributed components may operate continuously or in parallel.
[0084] The subject matter has been described in detail with respect to various specific exemplary embodiments, but each example is provided by way of illustration and not limitation of the disclosure. Those of ordinary skill in the art, upon attaining the understanding above, will readily be able to produce modifications, variations, and equivalents of such embodiments. Accordingly, the disclosure is not intended to exclude such modifications, variations, and / or additions to the subject matter that will be readily apparent to those of ordinary skill in the art. For example, features illustrated or described as part of one embodiment may be used with another embodiment to further produce additional embodiments. Accordingly, the disclosure is intended to cover such modifications, variations, and equivalents.
Description of the Reference Numerals
[0085] 10 Computing Device 12 High-Resolution Version of the Input Image 14 Image Element 16 Low-Resolution Version of the Input Image 18 Mask 20 First Machine Learning Model 22 Extended Image 24 Portion of the Extended Image 26 Upscaled Image Portion 28 Second Machine Learning Model 30 Improved part 32 Output image 50 Computing device 100 Computing system 102 User computing device 112 Processor 114 Memory 116 Data 118 Instruction 120 Model, machine learning model 122 User input component 130 Server computing system 132 Processor 134 Memory 136 Data 138 Instruction 140 Model, machine learning extended model 150 Training computing system 152 Processor 154 Memory 156 Data 158 Instruction 160 Model trainer 162 Training data, training dataset 180 Network 202 Ground truth image 212 Higher resolution version of the input image 214 Image element 216 Lower resolution version of the input image 218 Mask 220 First machine learning model 222 Lower resolution version of the enhanced image 226 Higher resolution version of the enhanced image 228 Second machine learning model 230 Predicted image 232 Loss function
Claims
1. A computing system for image correction with improved computing efficiency, the computing system comprising: one or more processors; one or more non-transitory computer-readable media for storing instructions collectively, the instructions, when executed by the one or more processors, causing the computing system to perform operations, the operations comprising: obtaining a low-resolution version of an input image, the low-resolution version of the input image having a first resolution and the low-resolution version of the input image comprising one or more image elements to be corrected using predicted image data; processing the low-resolution version of the input image using a first machine learning model to generate an extended image having the first resolution, the extended image comprising first predicted image data for replacing the one or more image elements; extracting a portion of the extended image, the portion of the extended image comprising the first predicted image data; upscaling the extracted portion of the extended image to generate an upscaled image portion having an upscaled resolution; processing the upscaled image portion using a second machine learning model to generate an improved portion, the improved portion comprising second predicted image data for correcting at least a portion of the first predicted image data; generating an output image based on the improved portion and a high-resolution version of the input image, both the output image and the high-resolution version of the input image having a second resolution greater than the first resolution; providing the output image as an output; the operations further comprising passing one or more internal feature vectors from the first machine learning model to the second machine learning model; A computing system.
2. A computing system for image correction with improved computing efficiency, the computing system comprising: one or more processors; One or more non-transitory computer-readable media that collectively store instructions, which, when executed by the one or more processors, cause the computing system to perform operations, the operations being Obtaining a low-resolution version of an input image, the low-resolution version of the input image having a first resolution and the low-resolution version of the input image comprising one or more image elements to be corrected using predicted image data; Processing the low-resolution version of the input image using a first machine learning model to generate an enhanced image having the first resolution, the enhanced image comprising first predicted image data that replaces the one or more image elements; Extracting a portion of the enhanced image, the portion of the enhanced image comprising the first predicted image data; Upscaling the extracted portion of the enhanced image to generate an upscaled image portion having an upscaled resolution; Processing the upscaled image portion using a second machine learning model to generate an improved portion, the improved portion comprising second predicted image data that modifies at least a portion of the first predicted image data; Generating an output image based on the improved portion and a high-resolution version of the input image, both the output image and the high-resolution version of the input image having a second resolution greater than the first resolution; Providing the output image as an output; The enhanced image further comprising a predicted depth channel output by the first machine learning model; A computing system. **Claim 3**: A computing system for image correction with improved computational efficiency, the computing system comprising One or more processors; One or more non-transitory computer-readable media that collectively store instructions, which, when executed by the one or more processors, cause the computing system to perform operations, the operations being Obtaining a low-resolution version of the input image, wherein the low-resolution version of the input image has a first resolution and the low-resolution version of the input image comprises one or more image elements to be corrected using prediction image data, Processing the low-resolution version of the input image using a first machine learning model to generate an extended image having the first resolution, wherein the extended image comprises first prediction image data for replacing the one or more image elements, Extracting a portion of the extended image, wherein the portion of the extended image comprises the first prediction image data, Upscaling the extracted portion of the extended image to generate an upscaled image portion having an upscaled resolution, Processing the upscaled image portion using a second machine learning model to generate an improved portion, wherein the improved portion comprises second prediction image data for correcting at least a portion of the first prediction image data, Generating an output image based on the improved portion and a high-resolution version of the input image, wherein both the output image and the high-resolution version of the input image have a second resolution greater than the first resolution, Providing the output image as an output, wherein the first prediction image data and the second prediction image data comprise human keypoint estimation image data indicating human keypoints detected in the input image, A computing system. Claim 4 The computing system according to any one of claims 1 to 3, wherein obtaining the low-resolution version of the input image comprises downscaling the high-resolution version of the input image to obtain the low-resolution version of the input image. Claim 5 Processing the low-resolution version of the input image using the first machine learning model to generate the extended image includes processing the low-resolution version of the input image and a mask identifying the one or more image elements using a first machine learning restoration model to generate the extended image having first restoration image data for modifying the one or more image elements. Processing the upscaled image portion using the second machine learning model to generate the improved portion includes processing the upscaled image portion using a second machine learning restoration model to generate the improved portion having second restoration image data for modifying at least a portion of the first restoration image data. The computing system according to any one of claims 1 to 4.
6. Upscaling the extracted portion of the extended image to generate the upscaled image portion having the upscaled resolution, wherein the upscaled resolution matches the corresponding resolution of the corresponding portion of the high-resolution version of the input image, and the corresponding portion corresponds proportionally to the extracted portion of the extended image. The computing system according to any one of claims 1 to 5.
7. Generating the output image based on the improved portion and the high-resolution version of the input image includes inserting the improved portion into the high-resolution version of the input image. The computing system according to any one of claims 1 to 6.
8. The computing system according to any one of claims 1 to 7, wherein the one or more image elements to be replaced comprise one or more user-specified image elements specified based on one or more user inputs.
9. The one or more image elements to be replaced are one or more computer-specified image elements, and the one or more computer-specified image elements are specified by processing the input image using one or more classification sub-blocks of at least one of the first machine learning model or the second machine learning model. The computing system according to any one of claims 1 to 8.
10. The first predicted image data and the second predicted image data correspond to one or more of restoration, blur correction, color restoration, or smoothing of the one or more image elements. The computing system according to any one of claims 1 to 9.
11. One or more objects comprise a plurality of objects, The processing of the low-resolution version of the input image using the first machine learning model to generate the extended image is performed once, The extracting, upscaling, and processing the upscaled image portion using the second machine learning model are performed separately for each object among the plurality of objects. The computing system according to any one of claims 1 to 10.
12. One or more non-transitory computer-readable media storing instructions collectively, wherein when the instructions are executed by one or more processors, cause a computing system to perform operations, the operations being Obtaining a low-resolution version of an input image, wherein the low-resolution version of the input image has a first resolution; Processing the low-resolution version of the input image using a first machine learning model to generate a first predicted image having the first resolution, wherein the first predicted image comprises first predicted image data; Extracting a portion of the first predicted image, wherein the portion of the first predicted image comprises the first predicted image data; Upscaling the extracted portion of the first predicted image to generate an upscaled image portion having an upscaled resolution; To generate a second predicted image, processing the upscaled image portion using a second machine learning model, the second predicted image comprising second predicted image data that modifies at least a portion of the first predicted image data. The operation further comprising passing one or more internal feature vectors from the first machine learning model to the second machine learning model. One or more non-transitory computer-readable media. [
13. ] One or more non-transitory computer-readable media storing instructions collectively, which when executed by one or more processors, cause a computing system to perform operations, the operations comprising: Obtaining a low-resolution version of an input image, the low-resolution version of the input image having a first resolution. Processing the low-resolution version of the input image using a first machine learning model to generate a first predicted image having the first resolution, the first predicted image comprising first predicted image data. Extracting a portion of the first predicted image, the portion of the first predicted image comprising the first predicted image data. Upscaling the extracted portion of the first predicted image to generate an upscaled image portion having an upscaled resolution. To generate a second predicted image, processing the upscaled image portion using a second machine learning model, the second predicted image comprising second predicted image data that modifies at least a portion of the first predicted image data. The first predicted image further comprising a predicted depth channel output by the first machine learning model. One or more non-transitory computer-readable media. [
14. ] One or more non-transitory computer-readable media storing instructions collectively, which when executed by one or more processors, cause a computing system to perform operations, the operations comprising: Obtaining a low-resolution version of an input image, the low-resolution version of the input image having a first resolution. Processing the low-resolution version of the input image using a first machine learning model to generate a first predicted image having the first resolution, wherein the first predicted image comprises first predicted image data; Extracting a portion of the first predicted image, wherein the portion of the first predicted image comprises the first predicted image data; Upscaling the extracted portion of the first predicted image to generate an upscaled image portion having an upscaled resolution; Processing the upscaled image portion using a second machine learning model to generate a second predicted image, wherein the second predicted image comprises second predicted image data that modifies at least a portion of the first predicted image data; The first predicted image and the second predicted image comprise a human keypoint estimation image indicating human keypoints detected in the input image; One or more non-transitory computer-readable media. Claim 15 The one or more non-transitory computer-readable media according to any one of claims 12 to 14, wherein the first predicted image and the second predicted image comprise an edge recognition image indicating recognized edges in the input image. Claim 16 The one or more non-transitory computer-readable media according to any one of claims 12 to 14, wherein the first predicted image and the second predicted image comprise an object detection image indicating objects detected in the input image. Claim 17 The one or more non-transitory computer-readable media according to any one of claims 12 to 14, wherein the first predicted image and the second predicted image comprise a face recognition image indicating recognized faces in the input image.
Citation Information
Patent Citations
Image processing apparatus and program
JP2020154605A
Image processing method, image processing device, program, image processing system, and learned model manufacturing method
JP2020166628A
Image processing device, image processing method, and image processing program
WO2018216207A1
Artificial intelligence systems and methods for interior design
WO2021008566A1