Method and system for creating image enhanced training data pair
By performing geometric alignment and color brightness matching during image capture, the problem of inconsistency in the image pair in the prior art is solved, high-quality training data set generation is achieved, and the training effect of machine learning models is improved.
Patent Information
- Application Number
- CN202510127823.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-05
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to effectively capture and process high-quality degraded images when generating data sets for training machine learning models, resulting in inconsistency in training data and artifact problems in model training.
By positioning the image capture device with respect to the display device, a downgraded version of the high-resolution image is captured and pixel-level geometric alignment, color and brightness matching is performed to form a consistent image pair to generate a training data set.
The pixel-level consistency of image pairs is achieved, the artifact problems in model training are reduced, the quality and consistency of training data are improved, and the image capture process is simplified.
Smart Images

Figure CN120047352A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to machine learning, and more particularly to generating a dataset for training a machine learning model. Background Art
[0002] Traditional machine learning models can be trained to generate a specific type of output from a given input. In one example, a machine learning model can be trained to generate an enhanced image output from a degraded input image. Generally, a machine learning model is trained using a set of training data. For example, training pairs can be used to represent known inputs and outputs. Once trained, the model generates an output in response to a given input. Summary of the Invention
[0003] This specification describes techniques for generating image pairs for use as training data to train a machine learning model. These techniques generally involve capturing a degraded version of a source image and performing various transformations that provide geometric alignment as well as color and luminance matching for each image pair. A set of image pairs associated with specific image features can form a dataset that serves as training data for a machine learning model configured to perform a specific machine learning task.
[0004] According to one aspect, the described techniques relate to a method for generating training data. The generation includes generating a set of image pairs. To generate the image pairs, the technique includes positioning an image capture device relative to a display device and sequentially capturing multiple images. To capture each image, the described technique includes displaying a high-resolution image on the display device; and capturing a degraded version of the high-resolution image using the image capture device. Each image pair in the generated image pairs has a high-resolution image and a corresponding degraded version of the high-resolution image.
[0005] To capture a degraded version of the high-resolution image, the technique can include positioning a structure between the image capture device and the display device that changes the center of focus of the image capture device. Alternatively or additionally, the technique can include modifying the ambient lighting around the display device to introduce noise into the captured image. In some embodiments, the system can also include modifying image capture device parameters including one or more of exposure time or ISO settings, and / or introducing motion into the image capture device to generate blur in the captured image.
[0006] To position the image capture device relative to the display device, the described technique can position the image capture device at a distance from the display device such that the resolution of the displayed image is a specified multiple of the resolution of the image capture device.
[0007] As another aspect, the described technology relates to a method that includes receiving a set of image pairs, each image pair including a high-resolution source image and a captured image corresponding to a downscaled version of the high-resolution source image. Each captured image is obtained by capturing an image of the corresponding displayed high-resolution source image. For each image pair in the set of image pairs, the described technology includes performing pixel-level geometric alignment between the high-resolution source image and the captured image, and performing color and luminance matching between the geometrically aligned high-resolution source image and the captured image. A set of geometrically aligned and color- and luminance-matched image pairs is then formed.
[0008] In some embodiments, each captured image corresponds to a particular type of image downscaling. The technology further includes using the set of geometrically aligned and color- and luminance-matched image pairs to train a machine learning model configured to generate enhanced images from input images having the particular type of downscaling.
[0009] The described technology includes one or more of denoising an input image, deblurring an image, or generating a super-resolution version of the input image.
[0010] To perform pixel-level geometric alignment, the described technology includes denoising the captured image; performing color correction on the source image; determining adjustments for pixel-level geometric alignment between the denoised captured image and the color-corrected source image; and applying the determined adjustments to the uncorrected source image. In some embodiments, the technology may include scaling the source image to match the dimensions of the captured image.
[0011] To perform color and luminance matching, the described technology may first separate the geometrically aligned source image into a low-frequency part and a high-frequency part. The described technology may determine a ratio basis between the low-frequency part of the source image and the captured image. A color and luminance matching ratio between the low-frequency source image and the low-frequency part of the captured image may be calculated. The described technology then applies the color and luminance ratio to the low-frequency source image to form a color- and luminance-matched low-frequency part; and combines the color- and luminance-matched low-frequency part of the source image with the high-frequency part of the source image to form a single color- and luminance-matched source image. The described technology may further sharpen the high-frequency part of the source image before combining the high-frequency part with the color- and luminance-matched low-frequency part.
[0012] To determine the color and luminance matching ratio between the low-frequency portion of a source image and the corresponding captured image in an image pair, the technique can perform neighborhood mean mapping. Neighborhood mean mapping can include applying a mean filter to both the low-frequency portion of the geometrically aligned source image and the low-frequency portion of the captured image. In some embodiments, the described technique can perform histogram mapping, in which a histogram mapping curve is estimated between the low-frequency portion of the geometrically aligned source image and the corresponding captured image. In some embodiments, the described technique can perform edge-preserving mapping, in which one or more edge-preserving low-pass filters are applied to the low-frequency portion of the geometrically aligned source image and the corresponding captured image.
[0013] Other embodiments of this aspect include corresponding computer systems, devices, computer program products, and computer programs recorded on one or more computer storage devices, each computer program configured to perform the actions of the method. A system of one or more computers can be configured to perform particular operations or actions by means of software, firmware, hardware, or a combination thereof installed on the running system, causing or enabling the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by means of instructions that, when executed by a data processing device, cause the device to perform the actions.
[0014] The subject matter described in this specification can be implemented in particular embodiments so as to achieve one or more of the following advantages. Degraded images can be captured using various types of image capture devices. The appropriate device can depend on the overall application for which the image is captured. The structured hardware setup of the display device and the image capture device can allow for consistent image capture of an image set.
[0015] Compared to capturing both images in the pair, a set of source images can be obtained as a selected set of high-quality images from which the second image in the pair can be captured. This can allow for the use of a greater variety of hardware to capture, including image capture devices with non-adjustable focal lengths. Additionally, capturing high-resolution images with a particular image capture device compared to a known set of high-quality images may result in one or more artifacts in the image, which could cause the model training to include such artifacts in the enhanced image.
[0016] Furthermore, compared to systems that digitally simulate degraded images, capturing a degraded version from known high-quality images can provide an improved set of captured images. In particular, it can be difficult to simulate specific defects without knowledge of all the image signal processor characteristics.
[0017] In addition, using a simple image capture device such as from a mobile phone eliminates the extra steps of capturing high-resolution images, making the process of forming image pairs more concise and efficient.
[0018] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a diagram of an example architecture for generating and aligning training data pairs;
[0020] Figure 2 is a diagram showing an example image capture system;
[0021] Figure 3 is a flowchart of an example process for performing geometric alignment;
[0022] Figure 4 is a flowchart of an example process for performing color and luminance alignment;
[0023] Like reference numerals and designations in the various drawings indicate like elements. DETAILED DESCRIPTION
[0024] A variety of image capture devices can be used to capture images including, for example, discrete images and individual video frames. Image capture devices can include digital cameras (including cameras that are part of mobile phones), image sensors, video recording devices, and the like. Images captured by these devices may have different qualities due to, for example, the environment or conditions in which the image is captured. For example, an image may have noise due to ambient lighting or device exposure. In another example, an image may become blurry due to movement of the device during image capture.
[0025] A machine learning model can be used to generate an output image from a given input image, such as generating an enhanced image output. Once trained, an input image can be provided to the model to generate an enhanced output image according to the specific configuration of the machine learning model. Image enhancement of the input image can be, for example, generating a de-blurred, noise-reduced, or super-resolution version of the input image. In some embodiments, the training dataset for training the machine learning model includes image pairs. Each image pair can correspond to a high-resolution version and a degraded version of a scene.
[0026] This specification describes techniques for generating image pairs based on specific image attributes to form a dataset. The dataset can be used as training data for a specific machine learning task based on the image attributes. In particular, this specification describes generally capturing a degraded version of a high-resolution source image and performing various transformations to align each image pair for use as a training pair. Geometric alignment as well as color and luminance matching can be performed on the source image to ensure pixel-level consistency between the source image and the captured degraded image. One or more sets of training data pairs can be formed for each type of enhancement task. The corresponding set is then used to train a machine learning model for one or more image enhancement tasks, including, for example, noise reduction or deblurring tasks.
[0027] Figure 1 FIG. is a diagram of an example architecture 100 for generating and aligning training data pairs. Architecture 100 includes an original or source image 102 and a captured image 104. The source image 102 is a high-quality image from which the corresponding degraded image is captured. For example, the source image 102 can be a high-resolution image that has been evaluated as a high-quality image meeting the resolution criteria and determined to be within the noise and blur tolerance ranges. Each captured image 104 is a degraded version of the corresponding source image 102. For example, a particular source image 102 can be displayed on a display device. The displayed source image 102 can then be captured by an image capture device such as a camera. A collection of source image 102 and captured image 104 pairs can form a specific dataset 106. Different datasets 106 can be generated. For example, each dataset 106 can include captured images 104 having one or more specific characteristics, such as introduced noise or blur. Below, for Figure 2 An example system for capturing an image of a source image is described.
[0028] The image pairs within a particular dataset 106 are processed to generate a collection of geometrically aligned and color and luminance matched image pairs 114, which can be provided as training data, for example, for training a machine learning model 116. The original image pairs 108 from the dataset 106 are geometrically aligned 110 and color and luminance matched 112 to form pixel-level aligned pairs 114. In particular, for geometric alignment 110, the source image in each pair is modified to provide pixel-level alignment with the corresponding captured image. For color and luminance matching 112, the source image in each pair is modified to provide color and luminance matching with the captured image to ensure no systematic color or luminance differences between the source image and the captured image. Geometric alignment 110 is described in more detail for Figure 3 Color and luminance matching 112 is described in more detail for Figure 4 Color and luminance matching 112 is described in more detail.
[0029] Dataset creation
[0030] A dataset that can generate image pairs can be used to train a model for different machine learning tasks. For example, image pairs can be created, where one image in the pair exhibits features for training to be enhanced by a machine learning model. In some embodiments, a specific dataset is generated for training a machine learning model for a single image enhancement task. In some alternative embodiments, a single dataset can be used to train a machine learning model configured to perform one or more different image enhancements.
[0031] Figure 2 FIG. shows an exemplary image capture system 200. The image capture system 200 includes a computing device 202, a display device 204, and an image capture device 206. The computing device 202 is communicatively coupled to both the display device 204 and the image capture device 206 via, for example, wired and / or wireless communication signals.
[0032] The display device 204 can be a stand-alone display, such as an LCD or CRT monitor or a television display. In some embodiments, the display device 204 is integrated into another device, for example, a computer. The image capture device 206 can be a camera device or a camera within another device, for example, as part of a mobile phone. Although the computing device 202, the image capture device 206, and the display device 204 are illustrated as different components, in some embodiments, one or more of the computing device 202, the display device 204, and the image capture device 206 can be integrated together. For example, the display device 204 can be integrated as part of the computing device 202, such as the display of a laptop computer that serves as the computing device. In another embodiment, the image capture device 206 can be a camera coupled to the computing device 202 as a peripheral device (such as a webcam device).
[0033] The computing device 202 is configured to control the display of images on the display device 204. For example, the computing device 202 can include (e.g., in memory or coupled to an external memory or data storage device) a collection of high-resolution images to be used as source images for the image pairs in the generated dataset. The high-resolution images can be regarded as representing the ground truth images of the image enhancement outputs required in the trained machine learning model.
[0034] The computing device 202 is also configured to instruct the image capture device 206 to capture an image. In particular, the image capture device 206 is positioned to capture an area of the display device 204 corresponding to the display area of the image provided from the computing device 202 to the display device 204. Thus, the image capture device 206 captures an image corresponding to the image currently being displayed on the display device 204. In some embodiments, the image capture device 206 is assigned to a specific location to reduce or eliminate the appearance of moiré patterns in the captured image of the display device 204.
[0035] Positioning the image capture device 206 may include mounting the image capture device at a designated location to maintain a specific position relative to the display device 204. For example, based on the default field of view of the image capture device 206, the distance from the display device 204 and the height of the image capture device 206 may be determined such that the field of view corresponds to the image displayed on the display device 204. In some embodiments, the display dimensions are determined, at least in part, based on the image capture device 206 used and the location to calibrate the captured field of view with the display size of the image. In some other embodiments, the display resolution on the display device 204 is specified to provide specific image capture characteristics, as described in more detail below.
[0036] In some alternative embodiments, control signals from the computing device 202 provide control of one or more image capture parameters (such as exposure settings or optical and digital zoom) of the image capture device 206.
[0037] In some embodiments, the image capture device 206 is mounted to a controllable structure to provide specific image capture characteristics. Additionally, other environmental parameters may be controlled, for example, manually or via the computing device 202. For example, a light box may be positioned and adjusted relative to the display device 204 to control ambient lighting.
[0038] In some embodiments, to generate a data set for training an image enhancement model configured to correct image blur, the captured image may be a blurred image of the display image. For example, blur may be introduced by vibrating the image capture device during image capture, with or without adjusting the exposure time. For example, the image capture device 206 may be mounted to a structure that can vibrate randomly to introduce motion blur into the captured image. In another embodiment, an opaque object may be positioned between the image capture device 206 and the display device 204. The image capture device may then focus on the opaque object rather than the image displayed on the display device 204. Blur corresponding to being out of focus may then be captured by the image capture device 206.
[0039] In some embodiments, motion can be introduced into the displayed image rather than into the image capture device 206 to produce motion blur in the captured image. For example, the displayed image can be replaced with a video of the image, where different frames include some small movement of the image on the display device 204, which is then captured by the image capture device 206.
[0040] In another example, to generate a data set for training an image enhancement model configured to correct noise from an input image, each of the captured images can be an image of the displayed image such that noise is introduced into the captured image. For example, the ambient lighting can be adjusted. For instance, a darker environment can increase the noise in the captured image. For example, the light box around the display device 204 can be dimmed to provide a lower ambient light environment and / or the brightness of the display device can be decreased. In some embodiments, the image capture parameters are modified, such as increasing the exposure time (shutter speed) or increasing the ISO sensitivity, resulting in increased noise in the captured image. For another example, a variable digital filter can be positioned between the image capture device 206 and the display device 204. By adjusting the range of the filter, the noise profile of the image captured by the image capture device 206 is also changed.
[0041] In another example, to generate a data set for training an image enhancement model configured to generate super-resolution images, each of the captured images can have a lower resolution than the corresponding displayed image. For example, the resolution of the displayed image can be a specific multiple, such as 5 times or 10 times, of the resolution of the captured image, where the specific multiple depends on the requirements of the image enhancement model. One example technique for obtaining a specific resolution level is to adjust the distance between the image capture device 206 and the display device 204 to produce a specified magnification between the size of the image displayed on the display device 204 and the size of the image captured by the image capture device 206. In particular, the pixel level of the display device 204 should be higher than the pixel level of the image capture device 206 to ensure that details are not lost due to a lower display pixel level. In one example, the display device 204 can present an image with a pixel size of 3000x4000. The distance to the image capture device 206 can be set such that the size of the image captured by the image capture device 206 has a pixel size of approximately 1500x2000. Thus, the model can be trained to magnify the input image by 2 times. A range of pixel size ratios can be used to provide a greater enhancement range.
[0042] The computing device 202 is also configured to provide instructions to the image capture device 206 to transmit the captured images back to the computing device 202. The computing device 202 coordinates the operations of the display device 204 and the image capture device 206 such that for each image displayed on the display device 204, one or more images are captured by the image capture device 206. For example, the command sequence may include instructions to present an image on the display device 204, instructions to capture an image of the image presented on the display device 204, and instructions to present the next image on the display device 204. The instructions to send the captured images to the computing device 202 may be sent individually after each image capture or may be sent as a batch of all captured images after the last image is captured by the image capture device 206.
[0043] In some embodiments, multiple images may be captured for a single displayed image on the display device 204. When multiple images are captured, adjustments may be made between successive captures of the displayed image to generate images for a dataset with variations in characteristics relevant to the task to be performed by the image enhancement model. For example, the adjustments made between multiple image captures of the same displayed image may include displacement, brightness changes, etc. These adjustments may be mechanical adjustments to the environment or simulated adjustments applied to the displayed image, such as changing the image brightness or displacing a specified number of pixels.
[0044] The computing device 202 then pairs each captured image with the corresponding displayed image (source image) to create image pairs for the dataset. In embodiments where multiple images are captured for a single displayed image, copies of the displayed image may be generated such that each copy is paired with a captured image. Alternatively, the image pairs may be stored as a data structure where multiple images reference the same source image.
[0045] The set of image pairs may form an initial dataset. The dataset may be further processed as described below for Figures 3 to 4 to form a final dataset. This further processing may include performing one or more transformations on the source images in each pair in the dataset, the one or more transformations including geometric alignment and color and brightness matching.
[0046] The image enhancement model may require pixel-level alignment between the images in each training image pair. To ensure that the training data pairs are aligned, the source images are operated on to ensure that they are pixel-level aligned with the corresponding captured images in the pair. In some embodiments, the captured images may be operated on instead of the source images. Figure 3 An example process for geometrically aligning each source image is described. Figure 4 An example process for color and brightness matching of the source images is described.
[0047] Geometric alignment
[0048] Figure 3 is a flowchart of an example process 300 for performing geometric alignment. For convenience, process 300 will be described as being executed by a system of one or more computers located at one or more locations and appropriately programmed in accordance with this specification. For example, Figure 3 the image processing system 301 in
[0049] can perform process 300. Figure 2 The system receives an image pair 302 of a specific raw data set. The image pair includes a high-resolution source image and a corresponding degraded image. The degraded image is a version of the source image that has been captured or modified to achieve specific characteristics based on an image enhancement task, such as introducing noise, blur, or reducing the resolution. The image pair can be generated as described above with respect to
[0050] For each image pair, the system performs a geometric alignment operation to provide pixel-level alignment between the source image and the captured image of the pair. The geometric alignment operation will be described for the first image pair.
[0051] The system scales (304) the source image. The size ratio between the source image and the captured image may not match. The system geometrically scales the source image to match the size of a specific size value. The size can be a size that matches the size of the captured image. For example, using the techniques described above with respect to Figure 2 the captured image can be captured in a specific size (e.g., the width and length of the image in pixels). A variety of different scaling algorithms can be used to change the image size by a specific scale amount, including interpolation, resampling, sampling, transformation, vectorization, and machine learning techniques. An appropriate scaling technique can be selected based on the type of image, the amount of image scaling, and the desired scaling quality (e.g., artifacts such as aliasing need to be limited).
[0052] In some other embodiments, both the source image and the captured image can be geometrically scaled to ensure that they have the same specified size. For example, in one embodiment, the captured image is enlarged by a factor of two, while the source image is reduced to match. In this way, the pair of images generally match at the midpoint of the scaling. In some embodiments, geometric scaling is optional because all source images have the same scaling that matches the captured image.
[0053] The system performs noise reduction on the captured image (306). If geometric scaling is performed, noise reduction is performed on the scaled captured image. Noise reduction is performed to improve the accuracy of geometric alignment between the two images. In some embodiments, the image capture device captures multiple image frames as part of the image capture. For example, a mobile phone camera often captures multiple frames as part of a single photo capture request. In this case, noise reduction can be performed by averaging multiple frames of the captured image to obtain a noise-reduced captured image. Although the described technique uses multiple frames of images for noise subtraction (which may be more accurate than using a single image frame), the described technique is capable of performing noise reduction by using a single image frame. More specifically, the described technique may include one or more noise reduction algorithms, e.g., NLM, BM3D, or a deep learning noise reduction model, to reduce noise as long as the selected algorithm can preserve the main texture features of the image. Since the noise-reduced image is only used for alignment purposes, other noise reduction algorithms can be implemented as long as they can preserve the main texture of the original image.
[0054] The system performs color correction on the source image (308). During color correction, the color and brightness of the source image (or the scaled source image if geometric scaling is performed) are adjusted to match the noise-reduced captured image in the pair. In some embodiments, histogram mapping is used for color correction to update the source image color and brightness values to match the color and brightness values in the captured image. For example, histograms of the RGB channels can be used for color matching of the individual red, green, and blue components of the source image. A brightness histogram can be used for brightness matching of the source image.
[0055] The system performs geometric alignment between the noise-reduced captured image and the color-corrected source image (310). Specifically, the color-corrected source image is spatially adjusted to match the noise-reduced captured image at the pixel level. In some embodiments, optical flow mapping and homography operations are performed on the color-corrected source image. Then the original (uncorrected) source image in the pair is adjusted by the determined adjustment to align the images. Geometric image alignment can also be referred to as image registration.
[0056] Optical flow mapping is used to estimate the motion of pixels between images. Techniques can be used to calculate the flow vectors between pixels in a reference frame (e.g., a denoised captured image) and a target frame (e.g., a color-corrected source image). Each flow vector indicates the direction and magnitude of the change in pixel position between the two images. The flow vectors are then used to determine the adjustments to be applied to the color-corrected source image to match the denoised captured image. The optical flow mapping can be dense or sparse, where sparse optical flow determines the flow vectors for specific regions of the image rather than the entire image. Example techniques for calculating optical flow include Lucas-Kanade, Horn-Schunck, Farneback, and an appropriate technique can be selected to determine the optical flow between the denoised captured image and the color-corrected source image.
[0057] Homography operations provide a perspective transformation for two-dimensional images to align the spatial relationship between the perceived observer and the image, e.g., to ensure that the source image and the captured image represent views from the same observer position. The homography matrix provides the transformation matrix for transferring points from the perspective of the denoised captured image to the perspective of the color-corrected source image so that the two images have aligned perspectives.
[0058] Optical flow mapping and homography operations may not need to be performed on every pair of images. In some embodiments, one or more pairs (color-corrected source image, denoised captured image) are set as calibration images. The average values of the optical flow map and / or the homography matrix are extracted from the calibration images. These values can then be applied to other image pairs for geometric alignment.
[0059] In some embodiments, a dataset of pairs (color-corrected source image, denoised captured image) is divided into multiple groups. Calibration images can be determined for each group, and the average values of the optical flow map and / or the homography matrix are applied to the image pairs of the group. In such embodiments, large displacements determined in the calibration images can cause a particular group to be filtered out. For example, the relative amount of displacement can be determined based on a comparison of adjacent groups. If the average absolute value of the optical flow map vectors exceeds a certain amount, e.g., one pixel, it indicates that vibration may have occurred in the group, and thus the entire group of data can be filtered out. Similarly, homography estimation can also be performed between the calibration images of adjacent groups. If the average offset value of the homography exceeds a certain amount, the group can be filtered out.
[0060] Using the information determined from optical flow and homography, which is applied to the color-corrected source image and the captured image, adjustments are applied to the original source images of each pair. The alignment of the captured image and the source image pairs provides an aligned dataset that can be further processed. For example, then, the system performs color and brightness matching (312) on the geometrically aligned image pairs.
[0061] Color and brightness matching
[0062] Figure 4 It is a flowchart of an example process 400 for performing color and luminance alignment. For convenience, process 400 will be described as being executed by a system of one or more computers located at one or more locations and appropriately programmed in accordance with this specification. For example, Figure 4 the image processing system 401 in
[0063] The system performs frequency separation (402) on the geometrically aligned source image. Frequency separation allows for the separation of different features of the image. For example, a Fourier transform can be used to transform the image components into discrete frequencies or groups of frequencies. The separated components can be processed individually. For example, color information can be separated from other details of the image. Generally, high-frequency components include fine details such as lines and textures within the image. Low-frequency components include image information such as color and hue. Techniques for frequency separation include, for example, applying Gaussian blur, bilateral filtering, or guided filtering to the input image to obtain the low-frequency part, and then subtracting between the original aligned source image and the generated blurred image to obtain the high-frequency part. The filter to be used can depend on the evaluation of specific types of features and output quality in the images of the dataset. In particular, the specific frequency separation used can separate the aligned source image into a high-frequency part and a mid to low-frequency part (MLF). The techniques described can use different quantifications to classify the "low", "mid", and "high" frequencies of the image. For example, the techniques described can classify the frequencies based on the texture size in a specific image. In this example, the "low" frequency can include textures or features with characteristic sizes between a single pixel and a few pixels (e.g., 2, 3, 5 pixels, etc.). The "mid" frequency can have characteristic sizes between a few pixels and ten or twenty pixels, while the "high" frequency can have characteristic sizes greater than twenty pixels.
[0064] The system optionally sharpens (404) the high-frequency part of the aligned source image. In some embodiments, the high-frequency part is magnified by a specified amount for sharpening. Different sharpening techniques can be used, such as unsharp masking which increases edge contrast. Since the high-frequency part includes fine lines and other edge features, it is suitable for sharpening the line features of the image.
[0065] The system further separates the MLF part into a low-frequency part and a mid-frequency part (406). The mid-frequency part and the low-frequency part can be separated in a manner similar to that described above for separating the high-frequency and MLF parts. Additionally, in some embodiments, this is part of the frequency separation performed at (402), rather than a separate subsequent step.
[0066] The system determines a ratio basis (408) between the source image and the captured image. Determining the ratio basis includes matching the pixel dimensions of the low-frequency source image to the captured image. Based on any scaling differences, one of the images is modified, but not both. Thus, if there is a scaling difference between the source image and the captured image, one image can be resized to match the other before proceeding with color and luminance matching. If they are the same size, no modification is required.
[0067] The system calculates a color and luminance matching ratio (410). After appending the ratio basis, a ratio map is calculated between the low-frequency portion of the low-frequency source image and the captured image. The color and luminance matching ratio for the low-frequency portion of the source image is calculated according to Formula 1, which is as follows:
[0068]
[0069] Where "Mapped_GT" refers to the color and luminance matching ratio, "IN_LF" refers to the low-frequency portion of the captured image, "GT_LF" refers to the low-frequency portion of the source image, "GT" refers to the geometrically aligned source image, and "Base" refers to the base value. A certain base value is added to both the numerator and denominator in Formula 1 to avoid an undefined result in the case where the low-frequency portion of the source image is zero. In some embodiments, the base value is 10 or less (for bitmap images, the range is 0 to 255).
[0070] The system can use different techniques to determine the color and luminance matching ratio (ratio basis) between the low-frequency portion of the geometrically aligned source image and the corresponding low-frequency portion of the captured image in the image pair. Similar to those described above for the source image, the system can implement techniques including, for example, histogram mapping, neighborhood mean mapping, or edge-preserving mapping.
[0071] In histogram mapping, a histogram mapping curve is estimated between the low-frequency source image and the captured image. The estimated histogram curve is applied to the low-frequency portion of the geometrically aligned high-resolution source image to obtain a color mapping result, for example, based on separate RGB histograms.
[0072] In neighborhood mean mapping, a mean filter is applied to both the low-frequency portion of the geometrically aligned high-resolution source image and the corresponding captured image in the pair. In edge-preserving mapping, an edge-preserving low-pass filter (including, for example, a bilateral filter or a guided filter) is applied to the low-frequency portions of both the geometrically aligned high-resolution source image and the captured image. Similar to neighborhood mean mapping, a ratio map is then calculated and applied to the low-frequency portion of the source image.
[0073] The system applies color and luminance ratios to a low-frequency source image combined with a medium-frequency source image to generate a color and luminance-matched low-frequency source image (412). In some embodiments, applying the color and luminance ratios to the low-frequency source image includes applying the color values and luminance values of the low-frequency source image pixel-by-pixel to the calculated ratios, e.g., multiplying the corresponding values by the calculated ratios.
[0074] The system combines the color and luminance-matched low-frequency / medium-frequency source image with a high-frequency (and optionally sharpened) source image to form a final matched source image (414) paired with the corresponding captured image. This generates a final data pair of geometrically aligned and color and luminance-matched datasets. This process is performed for each pair in the dataset to generate a final dataset of pixel-level geometric, color, and luminance-aligned data pairs. The final dataset can then be used as training data for one or more machine learning models being trained, e.g., for image enhancement.
[0075] Training a machine learning model
[0076] Training a machine learning model can include using the training data of the dataset to optimize the parameters of the machine learning model, which is based on an image input (represented by the captured image of the pair) to obtain an enhanced output result (represented by the source image of the pair). Different types of machine learning models (including neural network models such as convolutional neural networks or attention-based neural networks) can be used for image enhancement.
[0077] In some embodiments, training includes generating a specific machine learning model. The system generates a machine learning model. The machine learning model can be based on one or more existing machine learning models as a foundation and is configured to use specified training data and hyperparameters to train the model for a specific task. In particular, the specific task can be related to image processing, e.g., noise reduction, deblurring, or generating super-resolution images. One or more machine learning models can be trained to perform the specific task using the corresponding image pairs as described in this specification. The output generated by the trained one or more machine learning models can include, for example, noise reduction, deblurring, or a super-resolution image for the input image.
[0078] The training system uses specific training data (e.g., a specific dataset configured for training a machine learning model for a specific type of enhancement) to train the machine learning model. For example, the dataset can be obtained by the process described above for Figures 1 to 4 In some embodiments, the obtained training data is used as an input to a training engine that trains the machine learning model based on the training data and a set of model parameter values.
[0079] As part of model training, training data is iteratively processed to optimize one or more model parameter values. The training system can generate one or more output image predictions for each image pair of the training dataset. The training system analyzes the output image predictions and compares the output image predictions with the known output image (i.e., the ground truth high-resolution image) in the pair. Then, the training system updates one or more model parameter values based on the result of the comparison, for example, by using an appropriate update technique such as stochastic gradient descent with backpropagation. Once the machine learning model is trained, the final set of parameters can be used to generate a predicted output image based on a given degraded input image. For example, a blurred image can be input into a machine learning model configured to perform deblurring, and a predicted deblurred version of the image can be generated. Embodiments and functional operations of the subject matter described in this specification can be implemented in digital electronic circuitry, in tangible computer software or firmware, in computer hardware, in combinations including structures disclosed in this specification and their structural equivalents, or in one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. A computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Optionally or additionally, the program instructions may be encoded on an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiving device for execution by the data processing apparatus.
[0080] The term “data processing apparatus” refers to data processing hardware and includes all types of devices, apparatus, and machines for processing data, including examples of programmable processors, computers, or multiple processors or computers. The apparatus may also be or further include special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus may optionally include, in addition to hardware, code that creates an execution environment for computer programs, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0081] A computer program, which may also be referred to as or described as a program, software, software application, application, module, software module, script, or code, can be written in any form of programming language (including assembly language or interpreted language, or declarative or procedural language); and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may or may not correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data, for example, one or more scripts in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.
[0082] The processes and logical flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. These processes and logical flows can also be performed by special-purpose logic circuitry, such as an FPGA or ASIC, or by a combination of special-purpose logic circuitry and one or more programmed computers.
[0083] Computers suitable for executing a computer program can be based on general-purpose or special-purpose microprocessors or on both, or any other type of central processing unit. Generally, the central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a central processing unit for performing or executing instructions and one or more storage devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to, receive data from and / or transfer data to (or both) one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer need not have such devices. In addition, a computer can be embedded in another device, to name but a few, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive).
[0084] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0085] To provide for interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device, as well as a keyboard and a pointing device (e.g., a mouse or a trackball), where the display device (e.g., a CRT (cathode ray tube) or an LCD (liquid crystal display) monitor) is used to display information to the user, and the user may provide input to the computer through the keyboard and the pointing device. Other types of devices may also be used to provide for interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including auditory, speech, or tactile input. Additionally, the computer may interact with the user by sending documents to and receiving documents from the device used by the user; for example, in response to a request received from a web browser, by sending a web page to the web browser on the user device.
[0086] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes a backend component (e.g., as a data server) or includes a middleware component (e.g., an application server) or includes a frontend component (e.g., a client computer having a graphical user interface, a web browser, or an application through which a user may interact with an implementation of the subject matter described in this specification), or any combination of one or more of such backend, middleware, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN) and a wide area network (WAN), such as the Internet.
[0087] A computing system may include a client and a server. The client and the server are typically far apart from each other and typically interact via a communication network. The relationship between the client and the server arises from computer programs running on their respective computers and they have a client-server relationship with each other. In some embodiments, the server transmits data (e.g., an HTML page) to a user device, e.g., for displaying the data to a user interacting with the device acting as the client and receiving user input from the user. The server may receive data generated by the user device (e.g., the result of a user interaction).
[0088] Although this specification contains many specific implementation details, these details should not be construed as limitations on the scope of any invention or the scope that may be claimed, but rather as descriptions of specific features of particular embodiments of a particular invention. Certain features described in the context of separate embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Additionally, although the above features may be described as acting in a certain combination and even initially claimed as such, in some cases, one or more features from the claimed combination may be deleted from the combination, and the claimed combination may refer to a sub-combination or a variant of the sub-combination.
[0089] Similarly, although operations are depicted in the figures in a particular order, this should not be construed as requiring that the operations must be performed in the particular order shown or in a sequential order, or that all of the operations shown must be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of the various system modules and components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0090] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the acts recited in the claims may be performed in a different order but still achieve the desired result. As an example, the processes depicted in the figures do not necessarily need the particular order or sequential order shown to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A method for creating training data pairs for image enhancement, comprising: Generate a collection of image pairs, including: Position the image capture device relative to the display device: Capturing a plurality of images, wherein capturing each image comprises: displaying a high-resolution image on the display device; and capturing a degraded version of the high resolution image using the image capture device; and Image pairs are generated, each image pair comprising a high-resolution image and a corresponding degraded version of the high-resolution image.
2. The method for creating training data pairs for image enhancement according to claim 1, wherein: Capturing a degraded version of the high resolution image includes at least one of: positioning a structure between the image capture device and the display device, the structure changing the center of focus of the image capture device; modifying ambient lighting around the display device to introduce noise into a captured image; modifying image capture device parameters including one or more of exposure time or ISO setting; Motion is introduced into the image capture device to generate blur in the captured image.
3. The method for creating training data pairs for image enhancement according to claim 1 or 2, wherein: Positioning the image capture device relative to the display device includes positioning the image capture device at a distance from the display device such that a resolution of a displayed high resolution image is a specified multiple of a resolution of the image capture device.
4. A method for creating training data pairs for image enhancement, comprising: receiving a set of image pairs, each image pair comprising a high-resolution source image and a captured image corresponding to a degraded version of the high-resolution source image, and wherein each captured image is obtained by image capturing a corresponding displayed high-resolution source image; For each image pair in the set of image pairs: performing pixel-level geometric alignment between the high-resolution source image and the captured image to generate a geometrically aligned high-resolution source image, and performing color and brightness matching between the geometrically aligned high-resolution source image and the captured image; and A collection of geometrically aligned and color and brightness matched image pairs is formed.
5. The method for creating training data pairs for image enhancement according to claim 4, wherein: Each captured image corresponds to a specific type of image degradation, the method further comprising: Using the collection of geometrically aligned and color and brightness matched image pairs, a machine learning model configured to generate enhanced images from input images having the particular type of degradation is trained.
6. The method for creating training data pairs for image enhancement according to claim 5, wherein: Generating an enhanced image from a particular input image includes one or more of denoising the input image, deblurring the input image, or generating a super-resolution version of the input image.
7. The method for creating training data pairs for image enhancement according to any one of claims 4 to 6, wherein: Performing pixel-level geometric alignment includes: performing noise reduction on the captured image to generate a noise-reduced captured image; performing color correction on the high-resolution source image to generate a color-corrected source image; determining adjustments for pixel-level geometric alignment between the denoised captured image and the color-corrected source image; and The determined adjustments are applied to the high-resolution source image.
8. The method for creating training data pairs for image enhancement according to claim 7, further comprising: The high resolution source image is scaled to match the size of the captured image.
9. The method for creating training data pairs for image enhancement according to any one of claims 4 to 6, wherein: Performing color and brightness matching involves: separating the geometrically aligned high-resolution source image into a low-frequency portion and a high-frequency portion; determining a ratio basis between a low frequency portion of the geometrically aligned high resolution source image and the captured image; calculating, based on the ratio basis, a color and brightness matching ratio between a low frequency portion of the geometrically aligned high resolution source image and a low frequency portion of the captured image; applying the color and brightness matching ratio to the low frequency portion of the geometrically aligned high resolution source image to form a color and brightness matched low frequency portion; and The color and brightness matched low frequency portions of the geometrically aligned high resolution source images are combined with the high frequency portions of the geometrically aligned high resolution source images to form a single color and brightness matched source image.
10. The method for creating image enhanced training data pairs according to claim 9, further comprising: The high frequency portion of the geometrically aligned high resolution source image is sharpened before combining the high frequency portion with the color and brightness matched low frequency portion.
11. The method for creating training data pairs for image enhancement according to claim 9, wherein: Determining the color and brightness matching ratio includes at least one of: performing neighborhood mean mapping, the neighborhood mean mapping comprising applying a mean filter to both a low frequency portion of the geometrically aligned high resolution source image and a low frequency portion of the captured image; performing a histogram mapping in which a histogram mapping curve between a low frequency portion of the geometrically aligned high resolution source image and a low frequency portion of a corresponding captured image is estimated; Edge-preserving mapping is performed in which one or more edge-preserving low-pass filters are applied to low-frequency portions of the geometrically aligned high-resolution source images and low-frequency portions of the corresponding captured images.
12. A system for creating image-enhanced training data pairs, comprising one or more computers and one or more storage devices storing instructions, wherein when the instructions are executed by one or more computers, the one or more computers perform respective operations, and the operations include a method for creating image-enhanced training data pairs according to any one of claims 1 to 11.