Image processing method, method for generating a machine learning model, image processing apparatus, image processing system, and program
Patent Information
- Application Number
- JP2021066380
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-04-09
- Publication Date
- 2025-06-02
- Estimated Expiration
- 2041-04-09
AI Technical Summary
Existing image processing methods struggle to correct multiple correlated components, such as noise and blur, without causing unnatural changes in the image, leading to increased computational load and artifacts due to the lack of consideration for component correlations.
An image processing method using machine learning to generate residual maps for each component, allowing for independent correction of noise and blur components by modifying residual maps based on the input image and existing residual maps to achieve appropriate correction strengths.
This approach enables the production of natural-looking corrected images by adjusting correction strengths for multiple correlated components, reducing computational load and minimizing artifacts.
Smart Images

Figure 00000015_0000 
Figure 00000016_0000 
Figure 00000017_0000
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing method for correcting a plurality of components using machine learning.
Background Art
[0002] When correcting a plurality of components such as aberration components and noise components included in a captured image, it is desirable to perform correction for each component in order to achieve an appropriate correction intensity for each component. Conventionally, when the components are correlated, it has been difficult to correct one component from a captured image without changing the other component. However, in recent years, it has become possible to correct only one component with high accuracy using deep learning. Patent Document 1 discloses a method in which a neural network is used to correct blur due to aberration and diffraction from a captured image without changing the noise component, and then noise reduction processing is performed.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The method described in Patent Document 1 performs noise reduction processing on an image captured with adjusted blur correction intensity. Therefore, if the blur correction intensity is weakened to adjust sharpness or undershoot, the aberration components will change if noise reduction processing is performed without using deep learning. Generally, noise reduction processing is a blurring process, so the blur component spreads and chromatic aberration also spreads and becomes noticeable. Furthermore, since the blur component is not acquired as a residual map, the blur correction intensity cannot be adjusted after noise reduction processing, and noise reduction processing must be performed every time the blur correction intensity is adjusted. Deep learning estimation processing is computationally intensive, so performing noise reduction processing with deep learning every time the blur correction intensity is adjusted significantly increases the computational load. In addition, since the correlation between the blur component and the noise component is not considered, the noise component does not follow the change in brightness value in areas where the blur component has changed, resulting in an unnatural appearance. Furthermore, since the noise component is not acquired as a residual map, it is not possible to correct the noise component while considering the blur component.
[0005] The present invention aims to provide an image processing method that can obtain a natural-looking corrected image by performing image correction corresponding to multiple correlated components with appropriate correction intensity. [Means for solving the problem]
[0006] One aspect of the present invention is an image processing method characterized by comprising the steps of: inputting an input image to at least one machine learning model and obtaining a first residual map corresponding to a first component and a second residual map corresponding to a second component; modifying the second residual map to a third residual map with a different distribution based on the input image and the first residual map; and obtaining an output image based on the input image, the first residual map, and the third residual map. [Effects of the Invention]
[0007] According to the present invention, it is possible to provide an image processing method that can obtain a natural corrected image by performing image correction corresponding to multiple correlated components with appropriate correction intensity. [Brief explanation of the drawing]
[0008] [Figure 1] This is a block diagram of the image processing system in Example 1. [Figure 2] This is an external view of the image processing system in Example 1. [Figure 3] This is a flowchart for learning the weights in Example 1. [Figure 4] This diagram shows the learning process for the neural network weights in Example 1. [Figure 5] This is a flowchart for generating the corrected image in Example 1. [Figure 6] This is a block diagram of the image processing flow in Example 1. [Figure 7] This is a block diagram of the image processing flow of a modified example of Example 1. [Figure 8] This is a block diagram of the image processing system in Example 2. [Figure 9] This is an external view of the image processing system in Example 2. [Figure 10] This is a block diagram of the image processing system in Example 3. [Figure 11] This is a flowchart relating to the image processing in Example 3. [Modes for carrying out the invention]
[0009] Embodiments of the present invention will be described in detail below with reference to the drawings. In each drawing, the same reference numerals are used for the same components, and redundant explanations are omitted.
[0010] First, before describing the specific examples, the gist of the present invention will be explained. The present invention relates to image processing that performs image correction corresponding to multiple correlated components with appropriate correction intensity to obtain a natural corrected image.
[0011] Each component can be represented as a two-dimensional map corresponding to the correction amount at each pixel position in the image. These components include, for example, noise, blur due to the imaging system, blur due to defocus, camera shake, and subject blur, scattering due to fog, relighting, and background blur correcting. Noise is random noise that occurs randomly in the captured image with a predetermined variance, and has a different distribution each time an image is captured. Noise includes dark current noise, shot noise, and readout noise. Blur due to the imaging system includes aberrations, diffraction, low-pass filters, and aperture effects. Relighting is the virtual modification of the image's light source environment after imaging, and includes adding or changing light sources. The luminance change component due to relighting will be referred to as the relighting component below.
[0012] Image correction corresponding to each component involves modifying each component, such as reducing or adding noise, and correcting or adding blur, in order to create a desirable image or an image that reflects the user's editing intent. In this embodiment, a residual map corresponding to each component is obtained. A residual map is a two-dimensional map in which the correction amount (or the value of the component itself) of a specific component contained in the image is extracted, and the correction corresponding to each component is performed by adding or subtracting these maps.
[0013] Multiple components may be inherently correlated. For example, the noise component has different variances depending on shooting conditions such as ISO sensitivity and the difference in imaging devices. When shot noise is included, it has different variances depending on the luminance value. When the luminance value of a certain pixel changes due to correction of the blur component, the scattering component due to fog, etc., the variance of the noise that would originally occur in that pixel when there is no blur or scattering becomes different. Performing these corrections without considering the correlation results in an image with unnatural noise remaining. For example, for pixels or image regions where diffraction blur has been removed and the luminance has become low, noise with a large variance corresponding to the luminance value before the blur was removed is superimposed. Also, for pixels or image regions that have been relit and become brighter, unnaturally small noise corresponding to the dark luminance value before relighting is superimposed. Furthermore, the blur resulting from the optical characteristics of the imaging system, defocus, camera shake, and subject blur causes blur that spreads to surrounding pixels depending on the brightness of the subject. Therefore, when the brightness of the subject changes due to relighting, the blur also changes simultaneously. Thus, even when relighting is performed by image processing, correcting the blur without considering the correlation between components results in an image with unnatural blur remaining or an image with excessive blur correction.
[0014] Conventionally, it has been difficult to correct one component in an image without changing the other component when multiple components are correlated. For example, when correcting the blur component caused by the imaging system, performing sharpening processing multiplies the frequency characteristics by a gain, and the noise component is emphasized depending on the blur component or the frequency characteristics of the sharpening processing. Conversely, when increasing the background blur, blurring processing is applied, and the noise component is reduced. Even when correcting the noise component first, the noise reduction processing becomes a blurring process using the information of the surrounding image region, so the blur component increases. Also, when performing relighting processing to make the luminance value brighter or darker, the blur component, noise component, and scattering component due to fog increase or decrease. To separately obtain these correlated components, it is necessary to capture the changes in each component according to the scene and perform processing according to the subject in the image.
[0015] In recent years, machine learning, particularly deep learning, has enabled the understanding of scenes and the separation of components, allowing for highly accurate correction of only one component. This makes it possible to accurately correct each component in a sequence of correlated components. For example, consider the case where noise reduction and sharpening processing using blur correction are applied to an image containing both noise and blur components. By first using machine learning to accurately correct only the blur component, an image containing only noise components and no blur can be obtained. By applying conventional noise reduction processing to the obtained image, the noise component can be reduced. However, even if the image does not contain blur components, blurring will occur when noise reduction processing is performed, so it is preferable to perform noise reduction processing using machine learning. Furthermore, since the appropriate degree of correction varies depending on the image capture device manufacturer's intentions and the user's editing intentions, it is preferable to be able to adjust the correction strength. It is also preferable to be able to adjust the correction strength because artifacts may occur due to the correction processing depending on the scene. Artifacts include blurring due to noise reduction, undershoot and ringing due to sharpening, etc., and can occur in some scenes even with high-precision correction processing using machine learning. For example, artifacts are more likely to occur in scenes with reflections of light sources or high-contrast subjects. When adjusting the correction strength in this way, the blur correction strength is adjusted so that some of the blur component remains. Therefore, the subsequent noise reduction processing can be performed with high precision without being affected by the blur component by using machine learning.
[0016] As described above, each component can be corrected by correcting each component individually. When the correction strength in the preceding step is changed, the correction process in the subsequent step must be executed again. However, machine learning generally has a high computational load and requires re-execution, making it difficult to adjust the correction strength. Therefore, in this invention, each component is acquired as a corresponding residual map. By adjusting the correction strength based on the residual map, a corrected image can be obtained in which each component is corrected with an appropriate strength. Since it is necessary to acquire each component separately, a machine learning model must be used for each. When correcting multiple correlated components, even if the other component is not changed when correcting one component, the expected size of the other component changes. Therefore, in this invention, the residual map of one component is modified based on the residual map of the other component. This makes it possible to obtain a natural corrected image while adjusting the correction strength for multiple correlated components. The residual maps before and after correction have different distributions. Here, the two residual maps having different distributions means that they are not related by multiplying each pixel by a uniform value or adding an offset. This is because the correction is not a uniform process across the entire screen, but rather a process for each pixel or image region, and the value of one component for each pixel or image region corrects the other component. This is because the correlation between components in this invention is local, and one component is affected by the local value of the other component. Furthermore, by being able to obtain each component separately as a residual map, it is possible to make physically appropriate corrections rather than empirically. There may be a physical order to multiple correlated components, such as when a subject whose brightness is changed by a light source is blurred by the optical characteristics of the imaging system, and noise is generated when sensing the blurred light. Since components that occur physically later are affected by components that occur earlier, it is preferable to correct the residual map corresponding to the later-occurring component based on the residual map corresponding to the earlier-occurring component.
[0017] In the case of machine learning, even if learning is not performed to correct only one component in the previous-stage correction, it is also possible to perform learning assuming the previous-stage processing in the subsequent-stage correction. In this case, since learning is performed assuming the change in the component to be corrected in the subsequent stage, correction of the residual map becomes unnecessary. However, if learning is performed corresponding to the presence or absence of the previous-stage processing and the change in the correction intensity, the learning conditions increase, leading to increased learning effort, increased data volume, and deterioration of the correction performance. Also, when there is a change in the previous-stage processing or when a new processing is to be added, it is necessary to review the subsequent-stage processing.
Example
[0018] In this example, an image processing system that learns and executes correction of a blur component (first component) and a noise component (second component) derived from the optical characteristics of an imaging system using a multi-layer neural network (machine learning model) will be described. The blur component and the noise component correspond to different tasks of image correction. Note that the present invention is not limited to image processing that combines correction of the blur component and correction of the noise component, and is also applicable to other image processing.
[0019] FIG. 1 is a block diagram of an image processing system 100 of this example. FIG. 2 is an external view of the image processing system 100. The image processing system 100 includes a learning device 101, an imaging device 102, an image estimation device (image processing device) 103, a display device 104, a recording medium 105, an output device 106, and a network 107.
[0020] The learning device 101 includes a storage unit 101a, an acquisition unit 101b, a generation unit 101c, and an update unit 101d.
[0021] The imaging device 102 includes an optical system 102a and an image sensor 102b. The optical system 102a collects light incident on the imaging device 102 from the subject space. The image sensor 102b receives (photoelectrically converts) the optical image (subject image) formed through the optical system 102a to acquire an image. The image sensor 102b is, for example, a CCD (Charge Coupled Device) sensor or a CMOS (Complementary Metal-Oxide Semiconductor) sensor. The image acquired by the imaging device 102 includes blurring due to aberrations and diffraction of the optical system 102a and noise from the image sensor 102b.
[0022] The image estimation device 103 includes a storage unit 103a, an acquisition unit (first acquisition unit) 103b, and a correction unit (correction unit, second acquisition unit) 103c. The image estimation device 103 acquires an image from the imaging device 102 and, using at least one machine learning model, acquires a blur residual map (first residual map) corresponding to the blur component and a noise residual map (second residual map) corresponding to the noise component. Subsequently, it corrects the noise residual map based on the image (input image) and the blur residual map and acquires a corrected noise residual map (third residual map). Then, it generates an estimated image (corrected image, output image) in which blur and noise are corrected to an appropriate intensity based on the image (input image), the blur residual map, and the corrected noise residual map. The image estimation device 103 also has a function to perform development processing and other image processing as needed.
[0023] A multi-layered neural network is used to acquire residual maps, and weight information (machine learning model weights) is read from the memory unit 103a. The weight information is learned by the learning device 101. The image estimation device 103 has previously read the weight information from the memory unit 101a via the network 107 and stored it in the memory unit 103a. The weight information may be the numerical weights themselves or in an encoded format. Details regarding weight learning and the acquisition of each residual map using the weights will be described later.
[0024] The corrected image is output to at least one of the display device 104, the recording medium 105, and the output device 106. The display device 104 is, for example, a liquid crystal display or a projector. The user can perform editing work, etc., while checking the image in progress via the display device 104. The recording medium 105 is, for example, a semiconductor memory, a hard disk, or a server on a network. The output device 106 is, for example, a printer.
[0025] The following describes the method for learning weights (weight information) (method for manufacturing a trained model) performed by the learning device 101 of this embodiment, with reference to Figures 3 and 4. Figure 3 is a flowchart relating to weight learning. The processing of each step in Figure 3 is mainly performed by the acquisition unit 101b, generation unit 101c, or update unit 101d of the learning device 101. Figure 4 is a diagram showing the flow of learning the weights of a neural network.
[0026] In step S101, the acquisition unit 101b acquires the original image (subject image). In this embodiment, the original image is a high-resolution (high-quality) image with minimal blurring due to aberrations and diffraction of the optical system 102a. Multiple original images are acquired, representing various subjects, i.e., images with edges, textures, gradients, and flat areas of varying strengths and directions. The original image may be a real-life photograph or an image generated by computer graphics (CG). In particular, when using a real-life photograph as the original image, blurring has already occurred due to aberrations and diffraction, so reducing the image size reduces the effect of blurring and allows for a high-resolution (high-quality) image. Note that if the original image contains sufficient high-frequency components, reduction may not be necessary.
[0027] It is preferable that the original image has a signal value higher than the brightness saturation value of the image sensor 102b. This is because, even with real subjects, there are subjects that do not fall within the brightness saturation value when photographed by the imaging device 102 under specific exposure conditions. When using a real photograph as the original image, it can be obtained by HDR shooting or by shooting with an imaging device that has a higher dynamic range than the imaging device 102. When using an image taken with an imaging device that has a dynamic range equivalent to that of the imaging device 102 as the original image, it is also possible to make the signal value higher by proportionally multiplying the signal value. However, it is preferable to do so within a range where the reduction in gradation due to proportional multiplication does not affect the learning results. In addition, the original image may contain noise components. In this case, since the noise contained in the original image can be considered as part of the subject, the noise in the original image is not particularly problematic.
[0028] In step S102, the acquisition unit 101b acquires the blur used for the imaging simulation described later. First, the acquisition unit 101b acquires the shooting conditions corresponding to the lens state (zoom, aperture, and focus distance) of the optical system 102a. Then, the acquisition unit 101b acquires the blur determined by the shooting conditions and the screen position. Here, the blur is the PSF (point image intensity distribution) or OTF (optical transfer function) of the optical system 102a. The blur can be acquired by optical simulation or measurement in the optical system 102a. Note that different lens states, image height, azimuth aberration, and diffraction-induced blur are acquired for each original image. This makes it possible to perform imaging simulations corresponding to multiple shooting conditions, image height, and azimuth. In addition, components such as the optical low-pass filter included in the imaging device 102 may be added to the blur to be applied as needed.
[0029] In step S103, the generation unit 101c generates learning data, which is a combination of ground truth data consisting of ground truth patches (ground truth images) and training data consisting of training patches (training images). The ground truth patches and training patches are changed depending on the function or effect to be learned, and corresponding images should be used as the ground truth patches and training patches. Images with different first and second components are generated to be used as the ground truth patches and training patches. In this embodiment, a no-component patch generated from the original image and a blur patch with the blur component added in step S102 are generated. In addition, a noise patch and a noise-blur patch are generated by adding a noise component to the no-component patch and the blur patch. Multiple no-component patches and blur patches are generated, and one or more patches are generated corresponding to one original image. In this embodiment, the no-component patch and the blur patch are images of the same subject. In this embodiment, the no-component patch and the blur patch are used as the ground truth patch and the training patch, respectively. Then, multiple combinations of these are combined and used as training data to train a machine learning model that acquires a blur residual map to correct aberrations and diffractions of the optical system 102a. For this reason, the non-component patch must be an image with less blur compared to the blur patch. However, as will be described later, correction may not be performed depending on the conditions of the original image, so the training data may include cases where the non-component patch and the blur patch are the same image.
[0030] A patch refers to an image with a predetermined number of pixels (e.g., 64 x 64 pixels). Furthermore, the number of pixels in the ground truth patch and the training patch do not necessarily have to match. In this embodiment, mini-batch learning is used to train the weights of a multi-layer neural network. Therefore, in step S103, multiple sets of ground truth patches and training patches are generated. However, the present invention is not limited to this, and online learning or batch learning may also be used.
[0031] In this embodiment, multiple pairs of uncomponent images and blurred images, each with relatively different effects of aberration and diffraction blur, are generated by performing an imaging simulation using multiple original images stored in the memory unit 101a as subjects. At this point, the uncomponent images and blurred images have the same or greater number of pixels as the patch used as training data. Then, multiple uncomponent patches and blurred patches are obtained by extracting subregions of a specified pixel size at the same position from multiple pairs of uncomponent images and blurred images. In this embodiment, the original images are undeveloped RAW images, and the uncomponent patches and blurred patches are also RAW images. However, the present invention is not limited to this, and developed images may also be used. Furthermore, the position of the subregion refers to the center of the subregion. In this embodiment, uncomponent patches and blurred patches are obtained by the method described above, but the present invention is not limited to this.
[0032] Furthermore, since noise is generated in the image sensor 102b, noise is also added to the training data. Random numbers corresponding to the noise characteristics of the image sensor 102b should be generated and added, and shooting conditions such as ISO sensitivity may also be considered. By adding the same noise to the blurred image and the uncomponent image, a noisy blurred image (noise blur patch) and a noise image (noise patch) are obtained. The noise to be added is represented by σ(x,y)·r(x,y) and satisfies the following equation (1). σ 2 (x,y)=[k1(S1(x,y)-OB)+k0)]×ISO / 100 (1) (x,y) is a 2D spatial coordinate, S1(x,y) is the signal value of the pixel at coordinate (x,y) of the blur patch before noise is added. r(x,y) is the numerical value at coordinate (x,y) of the random number map with a standard deviation of 1, and σ(x,y) is the standard deviation of the noise (σ 2(x,y) is the variance. OB is the signal value of the optical black (black level image), ISO is the ISO sensitivity, and k1 and k0 are the proportionality constant and constant for the signal value at ISO sensitivity 100. The proportionality constant k1 represents the effect of shot noise, and the constant k0 represents the effect of dark current and readout noise. The values of k1 and k0 are determined by the noise characteristics of the image sensor 102b. This applies common noise to the corresponding pixels (corresponding pixels) of the no-component patch and the blur patch. Corresponding pixels are pixels that capture the same position in the subject space, or pixels at the same position in the no-component patch and the blur patch. To accommodate various ISO sensitivities of the image sensor 102b, noise of different ISO sensitivities is applied to multiple no-component patches 201 and blur patches 202.
[0033] Note that when reducing a real-world image to obtain the original image, the order of reduction and blurring can be reversed. If blurring is applied first, the sampling rate for the blur needs to be finer, taking reduction into account. For PSF (Point Sensorship), the sampling points in space should be finer, and for OTF (Optical Transfer Function), the maximum frequency should be increased.
[0034] It is preferable that the added bokeh does not include distortion. This is because, in sharpening processing (aberration correction processing, correction of the bokeh component), if the distortion is large, the position of the subject changes, and the subject may differ between the no-component patch and the bokeh patch. For this reason, the bokeh residual map used in this embodiment does not include distortion. Distortion is corrected individually after bokeh correction using bilinear interpolation, bicubic interpolation, etc.
[0035] In step S104, the generation unit 101c inputs the noise-blurred patch as input data 212 (training patch, training image) in Figure 4 into a multilayer neural network and generates an estimated patch (estimated image) 213. For mini-batch learning, it generates estimated patches 213 corresponding to multiple input data 212. Figure 4 shows the flow from step S104 to step S105.
[0036] In this embodiment, a noise patch is used as the ground truth data 211. The estimated patch 213 is a noise-blurred patch with blur correction applied, and ideally matches the ground truth data (ground truth patch, ground truth image) 211. The neural network outputs an estimated residual map 214 that corresponds to the difference between the input data 212 (noise-blurred patch in this embodiment) and the ground truth data 211 (noise patch in this embodiment). The estimated residual map 214 is the estimated blur residual map. In this embodiment, the neural network configuration shown in Figure 3 is used, but the present invention is not limited thereto.
[0037] In Figure 3, CN represents a convolutional layer and DC represents a deconvolutional layer. In both CN and DC, the sum of the input, the filter convolution, and the bias is calculated, and the result is nonlinearly transformed by an activation function. The initial values of each filter component and the bias are arbitrary and are determined by random numbers in this embodiment. For example, ReLU (Rectified Linear Unit) or the sigmoid function can be used as the activation function. The output of each layer except the final layer is called a feature map. Skip connections 222 and 223 synthesize feature maps output from non-contiguous layers. The feature maps may be synthesized by taking element-wise sums or by concatenation in the channel direction. In this embodiment, element-wise sums are taken. Skip connection 221 takes the sum of the estimated residual map 214 estimated from the input data 212 and the input data 212 to generate an estimated patch 213. An estimated patch 213 is generated for each of the multiple input data 212.
[0038] In step S105, the update unit 101d updates the neural network weights based on the error between the estimated patch 213 and the ground truth data 211. The weights include the filter components and biases of each layer. Backpropagation is used to update the weights, but the present invention is not limited to this. For mini-batch learning, the error between multiple noise patches input as ground truth data 211 and their corresponding estimated patches 213 is calculated, and the weights are updated. For the loss function, for example, the L2 norm or L1 norm may be used.
[0039] In step S106, the update unit 101d determines whether weight learning is complete. Completion can be determined by whether the number of iterations of learning (weight update) has reached a predetermined value, or whether the amount of change in weights during the update is less than a predetermined value. If it is determined that learning is incomplete, the process returns to step S104 and acquires multiple new noise patches and noise blur patches. On the other hand, if it is determined that learning is complete, the learning device 101 (update unit 101d) terminates learning and stores the weight information in the storage unit 101a. By learning noise patches and noise blur patches with the same random noise as the ground truth data 211 and input data 212, the neural network can learn to separate the subject, blur component, and noise component. Therefore, it is possible to obtain a blur residual map corresponding to the blur component of the subject only, while suppressing noise fluctuations.
[0040] In this embodiment, the estimation error was obtained from the noise patch and estimated patch 213 used in the ground truth data 211. However, the estimation error may also be obtained from the estimated residual map 214 and the ground truth of the blur residual map without outputting the estimated patch 213. The ground truth of the blur residual map is the difference between the noise patch and the noise-blur patch, which is equal to the difference between the no-component patch and the blur patch. In this case, a machine learning model can be trained that directly outputs the estimated blur residual map as the estimated residual map 214.
[0041] Furthermore, while this embodiment describes the training of a machine learning model for acquiring blur residual maps, a machine learning model for acquiring noise residual maps can be trained in a similar manner. Specifically, in Figure 4, the input data 212 remains a noise blur patch, and a blur patch is used instead of a noise patch as the ground truth data 211. In this case, by training with a patch in which the blur component remains unchanged and only the noise component differs, the neural network can learn to separate the subject, the blur component, and the noise component. Therefore, it is possible to acquire a noise residual map corresponding to the noise component while suppressing changes in the blur component.
[0042] In this case as well, the estimation error may be obtained from the estimated residual map 214 and the ground truth of the noise residual map without outputting the estimated patch 213. The ground truth of the noise residual map is the difference between noise-blurred patches and blurred patches, which is equal to the difference between noise patches and no-component patches. In this case, a machine learning model can be trained that directly outputs the estimated noise residual map as the estimated residual map 214.
[0043] If the correct answer for the blur residual map is obtained from the difference between the no-component patch and the blur patch, and the correct answer for the noise residual map is obtained from the difference between the noise patch and the no-component patch, then the noise blur patch is not used. Therefore, in this case, it is not necessary to obtain the noise blur patch.
[0044] Furthermore, in this embodiment, since the machine learning model that acquires the blur residual map and the machine learning model that acquires the noise residual map are trained separately, they may each have different network configurations.
[0045] Furthermore, since neural networks can learn by separating the subject, blur component, and noise component, both blur residual maps and noise residual maps can be output using a common network (a single network). In this case, training and estimation in the machine learning model only need to be done once, thus reducing the computational load.
[0046] The generation of corrected images (estimated images, output images) performed by the image estimation device 103 will be explained below with reference to Figures 5 and 6. Figure 5 is a flowchart of the generation of corrected images. Figure 6 is a block diagram of the image processing flow. The processing of each step in Figure 5 is mainly performed by the acquisition unit 103b and the correction unit 103c of the image estimation device 103.
[0047] In step S201, the acquisition unit 103b acquires the input image 401 and weight information. The input image 401 is an undeveloped RAW image, which in this embodiment is transmitted from the imaging device 102. The weight information is transmitted from the learning device 101 and stored in the storage unit 103a, and consists of the weights of the machine learning model that acquires the blur residual map and the machine learning model that acquires the noise residual map.
[0048] In step S202, the correction unit 103c inputs the input image 401 into a machine learning model that acquires the blur residual map obtained in step S201 and acquires a blur residual map (first residual map) 402. The correction unit 103c also inputs the input image 401 into a machine learning model that acquires the noise residual map and acquires a noise residual map (second residual map) 403.
[0049] In step S203, the correction unit 103c acquires the correction intensity using the blur residual map and the correction intensity using the noise residual map. The correction intensity can be obtained using the correction ratio of each component, and in this embodiment, predetermined values are used, but values specified by the user may also be acquired.
[0050] In step S204, the correction unit 103c modifies the noise residual map 403 based on the input image 401 and the blur residual map 402, and obtains a corrected noise residual map (third residual map) 404. First, the correction unit 103c adds the blur residual map 402 to the input image 401 by multiplying it by the correction ratio of the blur component obtained in step S203. For example, if the correction ratio is 0.5, an intermediate corrected image 410 with 50% of the blur component corrected is obtained. Next, the standard deviations of noise σ1(x,y) and σ2(x,y) are obtained for the case where the signal value S1(x,y) of equation (1) is the input image 401 and the intermediate corrected image 410. The noise residual map 403 corresponds to the noise component contained in the input image 401. By multiplying the noise residual map 403 by the ratio σ2(x,y) / σ1(x,y) for each pixel, a corrected noise residual map 404 is obtained. The corrected noise residual map 404 corresponds to the noise components contained in the intermediate correction image 410. Since the luminance values and blur residual map values of the input image 401 usually differ depending on the screen position, the ratio applied to each pixel differs. Therefore, the corrected noise residual map 404 is multiplied by different coefficients for each pixel of the noise residual map 403, resulting in different distributions. This makes it possible to obtain a residual map that enables natural correction reflecting the correlation of each component for each pixel.
[0051] In step S205, the correction unit 103c obtains an estimated image 405 by multiplying the corrected noise residual map 404 by the correction ratio of the noise component acquired in step S203 and adding it to the intermediate correction image 410. Since the acquired image is a RAW image, development processing is performed as needed.
[0052] In this embodiment, the correction intensity of each component may be adjusted while viewing the estimated image.
[0053] Furthermore, the configuration of this embodiment is not limited to the above. For example, the correction intensity can be obtained at any time before using the correction intensity of each component.
[0054] Furthermore, in this embodiment, the corrected noise residual map 404 was multiplied by a correction ratio and added to the intermediate correction image 410 to obtain the estimated image 405, but the present invention is not limited thereto. The estimated image 405 may also be obtained by adding the blur residual map 402 and the corrected noise residual map 404 to the input image 401 according to the correction ratio of each component.
[0055] Alternatively, both the blur residual map 402 and the noise residual map 403 may be output using a common machine learning model (a single machine learning model).
[0056] Furthermore, while the signal value S1(x,y) used to obtain the standard deviation σ1(x,y) was the input image 401, it may also be the image after correcting for noise components. That is, an image with a noise residual map added to the input image 401 at a correction ratio of 1 may be used. This method can improve accuracy, especially when the noise component is large.
[0057] Furthermore, in this embodiment, each component is corrected (removed) by adding a residual map, but the sign can also be defined in a way that allows each component to be removed by subtracting the residual map.
[0058] Alternatively, the processing may be carried out according to the image processing flow in Figure 7. Figure 7 is a block diagram of the image processing flow of a modified example. First, the input image 401 is input to a machine learning model (second learning model) that has been trained on the noise component to obtain a noise-corrected image 411 and a noise residual map 403. The noise-corrected image 411 is an image obtained by subtracting 100% of the acquired noise component from the input image 401. Next, the noise-corrected image 411 is input to a machine learning model (first learning model) that has been trained on the blur component to obtain a noise-blur-corrected image 412 and a blur residual map 402. The noise-blur-corrected image 412 is an image obtained by subtracting 100% of the acquired blur component from the noise-corrected image 411. Next, the blur residual map 402 is multiplied by (1 - blur component correction ratio) and subtracted from the noise-blur-corrected image 412 to weaken the correction strength of the blur component and obtain an intermediate correction image 413. As a result, the intermediate correction image 413 becomes a corrected image that reflects the correction ratio of the blur component. The noise residual map 403 corresponds to the noise component contained in the input image 401. Therefore, similar to this embodiment, the corrected noise residual map 404 is obtained by multiplying the noise correction image 411 and the intermediate correction image 413 by the ratio of each standard deviation of noise to the luminance value. The correction intensity of the noise component is weakened by multiplying the corrected noise residual map 404 by (1 - correction ratio of the blur component) and subtracting it from the intermediate correction image 413, and the estimated image 405 is obtained to complete the process.
[0059] According to this embodiment and its modifications, there is no need to recalculate the machine learning model even when adjusting the correction intensity of each component. When adjusting the blur correction intensity, it is sufficient to repeat the process from the step of obtaining a corrected noise residual map based on the brightness values before and after blur correction. Therefore, the computational load can be significantly reduced, and it becomes possible to adjust the correction intensity while viewing the image.
[0060] In this embodiment, we dealt with correcting the blur component originating from the optical characteristics of the imaging system, but the present invention is also effective for blur caused by other factors (such as defocus and blurring). By changing the blur applied to the blur patch during training to defocus, blurring, etc., it is possible to obtain a residual map that separates the noise component from the blur caused by these factors.
[0061] Furthermore, the present invention can also be applied to the transformation of defocus blur (background blur) instead of correcting the blur component. Transformation of defocus blur is the process of transforming the defocus blur in an captured image into a shape and distribution desired by the user. Defocus blur in captured images includes vignetting, double-line blur, annular patterns due to cutting marks of aspherical lenses, and central occlusion due to catadioptic optics. These defocus blurs are transformed by a neural network into a shape and distribution desired by the user (e.g., a flat circle or a normal distribution function). The neural network that realizes the transformation of defocus blur can be trained in the following way: For the same original image, an image equivalent to the captured image with the defocus blur occurring in the captured image applied, and an ideal equivalent image with the user's desired defocus blur applied, are generated for multiple defocus amounts. However, since it is desirable that the transformation of defocus blur does not cause any change to the subject at the focal distance, an image equivalent to the captured image and an ideal equivalent image with zero defocus amount are also generated. Multiple first and second defocus patches are extracted from the generated image-equivalent and ideal-equivalent images, respectively. Noise is then added to these to obtain noise-generated first and second defocus patches. By performing the same learning process as in this embodiment, the correction component of the defocus blur conversion can be separated from the noise component and obtained as a residual map.
[0062] Furthermore, the present invention can also be applied to lighting transformation (relighting) instead of correcting the blur component. The neural network that realizes lighting transformation can be trained in the following way: An image equivalent to the captured image is generated by rendering the original image, which has the same normal map, in the light source environment assumed for the captured image. Similarly, an ideal equivalent image is generated by rendering the normal map in the light source environment desired by the user. Multiple first lighting patches and second lighting patches are extracted from the image equivalent and the ideal equivalent image, respectively, and noise is added to these to obtain a noise first lighting patch and a noise second lighting patch. The first lighting corresponds to the lighting before relighting, and the second lighting corresponds to the lighting after relighting. By performing training in the same manner as in this embodiment, the relighting component can be separated from the noise component and obtained as a residual map.
[0063] Furthermore, the present invention can also be applied to correcting blur components using relighting instead of blur components, and to correcting blur components instead of noise components. Blur originating from the optical characteristics of the imaging system, defocus, and camera shake causes blur to spread to surrounding pixels depending on the brightness of the subject. Therefore, when the brightness of the subject changes due to lighting, the blur also changes simultaneously. In RAW images, the brightness of the subject and the blur component change in a proportional relationship. In this embodiment, the change in the standard deviation of noise is obtained based on the change in brightness before and after camera shake correction, and a corrected noise residual map is obtained by multiplying by the ratio of the change in standard deviation. Similarly, the blur residual map is corrected for each pixel based on the ratio of brightness values before and after light source correction by the relighting component. This obtains a corrected blur residual map. A neural network that separates and obtains blur components and relighting components can be trained in the following way: An image equivalent to the captured image is generated by rendering the original image, which has the same normal map, in the light source environment assumed for the captured image. Similarly, an ideal equivalent image is generated by rendering the normal map in the light source environment desired by the user. Multiple first and second lighting patches are extracted from the captured equivalent image and the ideal equivalent image, respectively, and blurred first and second lighting patches are obtained by adding blur to these. Here, the second component becomes the blur component, so the blurred first and second lighting patches correspond to the noise blur patch and noise patch in this embodiment. By performing learning in the same manner as in this embodiment, the relighting component can be separated from the blur component and obtained as a residual map.
[0064] Furthermore, when performing image correction by acquiring blur and relighting components, the residual map of the relighting component may be modified based on the ratio of luminance values before and after correction of the blur component.
[0065] Furthermore, the present invention is also applicable to three or more components. Even with three or more components, each component can be separated and obtained as a residual map, so, as with the case of two components, the residual map of each component can be modified as appropriate while considering the correlation between the components. For example, when correcting the blur component, relighting, and noise component, first, the blur residual map is modified for each pixel based on the ratio of the luminance values before and after light source correction by the relighting component. Then, the noise component map can be modified for each pixel based on the ratio of the luminance value when neither light source correction by the relighting component nor the blur component correction is applied to the luminance value when both are applied.
[0066] In this embodiment, the case in which the learning device 101 and the image estimation device 103 are separate entities has been described, but the present invention is not limited to this. The learning device 101 and the image estimation device 103 may be integrated. That is, learning (processing shown in Figure 3) and estimation (processing shown in Figure 5) may be performed within an integrated device.
[0067] With the above configuration, it is possible to obtain a natural-looking corrected image by performing image correction while adjusting the correction intensity for each of the multiple correlated components. [Examples]
[0068] The image processing system of this embodiment differs from the image processing system of Embodiment 1 in that it generates corrected images in the image estimation unit within the imaging device.
[0069] Figure 8 is a block diagram of the image processing system 300 in this embodiment. Figure 9 is an external view of the image processing system 300. The image processing system 300 has a learning device 301 and an imaging device 302. The learning device 301 and the imaging device 302 are connected via a network 303. The learning device 301 has a storage unit 311, an acquisition unit 312, a generation unit 313, and an update unit 314, and learns weights (weight information) for acquiring residual maps using a neural network. The imaging device 302 has an optical system 321, an image sensor 322, an image estimation unit (image processing device) 323, a storage unit 324, a recording medium 325, a display unit 326, and a system controller 327. The imaging device 302 captures the subject space and acquires an image, and uses the read weight information to acquire a blur residual map and a noise residual map from the image. The imaging device 302 generates a corrected image using a corrected noise residual map obtained by correcting the noise residual map. The image estimation unit 323 has an acquisition unit (first acquisition unit) 323a and a correction unit (correction unit, second acquisition unit) 323b, and acquires each residual map using the weight information stored in the memory unit 324, and performs correction based on each residual map.
[0070] Weight information is pre-learned by the learning device 301 and stored in the storage unit 311. The imaging device 302 reads the weight information from the storage unit 311 via the network 303 and stores it in the storage unit 324. The corrected image is stored in the recording medium 325. When the user issues a command to display the corrected image, the stored corrected image is read and displayed in the display unit 326. Alternatively, an image already stored in the recording medium 325 may be read and performance deviation correction performed by the image estimation unit 323. This series of controls is performed by the system controller 327.
[0071] The training of the machine learning model performed by the learning device 301 is equivalent to the training of the machine learning model described in Example 1.
[0072] The acquisition of corrected images performed by the image estimation unit 323 is carried out in the same manner as the flowchart in Figure 5. Instead of being performed by the acquisition unit 103b and correction unit 103c in Example 1, it is mainly performed by the acquisition unit 323a and correction unit 323b of the image estimation unit 323.
[0073] With the above configuration, it is possible to obtain a natural-looking corrected image by performing image correction while adjusting the correction intensity for each of the multiple correlated components. [Examples]
[0074] The image processing system of this embodiment differs from the image processing systems of Embodiments 1 and 2 in that it has a processing unit (computer) that transmits the captured image to be processed to the image estimation device and receives the processed output image (corrected image, estimated image) from the image estimation device.
[0075] Figure 10 is a block diagram of the image processing system 600 of this embodiment. The image processing system 600 includes a learning device 601, an imaging device 602, an image estimation device (image processing device) 603, and a processing device (computer) 604. The learning device 601 and the image estimation device 603 are, for example, servers. The processing device 604 is, for example, a user terminal (personal computer or smartphone) and is connected to the image estimation device 603 via a network 605. That is, the image estimation device 603 and the processing device 604 are configured to communicate with each other. The image estimation device 603 is connected to the learning device 601 via a network 606. That is, the learning device 601 and the image estimation device 603 are configured to communicate with each other.
[0076] The configuration of the learning device 601 is the same as that of the learning device 101 in Example 1, so its description is omitted. Similarly, the configuration of the imaging device 602 is the same as that of the imaging device 102 in Example 1, so its description is omitted.
[0077] The image estimation device 603 includes a storage unit 603a, an acquisition unit (first acquisition unit) 603b, a correction unit (correction unit, second acquisition unit) 603c, and a communication unit (receiving unit) 603d. The storage unit 603a, the acquisition unit 603b, and the correction unit 603c are the same as the storage unit 103a, the acquisition unit 103b, and the correction unit 103c of the image estimation device 103 of Embodiment 1. The communication unit 603d has the function of receiving requests transmitted from the processing unit 604 and the function of transmitting the output image generated by the image estimation device 603 to the processing unit 604.
[0078] The processing unit 604 includes a communication unit (transmitter) 604a, a display unit 604b, an image processing unit 604c, and a recording unit 604d. The communication unit 604a has the function of transmitting a request to the image estimation unit 603 to perform processing on the captured image, and the function of receiving the output image processed by the image estimation unit 603. The display unit 604b has the function of displaying various information. The information displayed by the display unit 604b includes, for example, the captured image transmitted to the image estimation unit 603 and the output image received from the image estimation unit 603. The image processing unit 604c has the function of performing image processing on the output image received from the image estimation unit 603. The recording unit 604d records the captured image acquired from the imaging unit 602 and the output image received from the image estimation unit 603, etc.
[0079] The image processing in this embodiment will be described below with reference to Figure 11. Figure 11 is a flowchart relating to image processing.
[0080] The image processing shown in Figure 11 is initiated when the user issues an instruction to start image processing via the processing unit 604. First, the operation of the processing unit 604 will be explained.
[0081] In step S701, the processing unit 604 sends a request to the image estimation unit 603 for processing the captured image. The method of sending the captured image to be processed to the image estimation unit 603 is not limited. For example, the captured image may be uploaded to the image estimation unit 603 at the same time as the processing in step S701, or it may be uploaded to the image estimation unit 603 before the processing in step S701. Also, the captured image may be an image stored on a server different from the image estimation unit 603. In step S701, the processing unit 604 may also send ID information for user authentication along with the request to process the captured image.
[0082] In step S702, the processing unit 604 receives the output image generated in the image estimation unit 603. The output image is an image in which each component of the captured image has been corrected using a residual map, similar to Example 1.
[0083] Next, the operation of the image estimation device 603 will be described.
[0084] In step S801, the image estimation device 603 receives a request for processing of the captured image transmitted from the processing device 604. The image estimation device 603 determines that processing of the captured image has been instructed and executes the processing from step S802 onward.
[0085] In step S802, the image estimation device 603 acquires weight information. The weight information is information (a trained model) learned in the same manner as in Example 1. The image estimation device 603 may acquire weight information from the learning device 601, or it may acquire weight information that has been previously acquired from the learning device 601 and stored in the storage unit 603a.
[0086] The processes from step S803 to step S806 are the same as the processes from step S202 to step S205 in Figure 5 described in Example 1, so their explanation will be omitted.
[0087] In step S807, the image estimation device 603 transmits the output image to the processing device 604.
[0088] As in this embodiment, when the correction process is performed within the image estimation device 603, the processing load due to the correction process can be handled within the image estimation device 603, thus reducing the processing capacity required on the processing device 604 side.
[0089] As described above, the image estimation device 603 may be configured to be controlled using a processing device 604 that is communicatively connected to the image estimation device 603. [Other examples] The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.
[0090] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of its gist. [Explanation of symbols]
[0091] 103 Image Estimation Device 103b Acquisition Department 103c Correction section
Claims
1. inputting an input image into at least one machine learning model to obtain a first residual map corresponding to the first component and a second residual map corresponding to the second component; modifying the second residual map to a third residual map having a different distribution based on the input image and the first residual map; obtaining an output image based on the input image, the first residual map, and the third residual map.
2. 2. The image processing method of claim 1, wherein the first component and the second component correspond to different tasks of image correction.
3. the first residual map corresponds to a difference between the input image corrected for the first component and the input image; 3. The image processing method of claim 1, wherein the second residual map corresponds to a difference between the input image corrected for the second component and the input image.
4. 4. The image processing method according to claim 1, wherein the at least one machine learning model is trained to output the first residual map by changing only the first component without changing the second component, and is trained to output the second residual map by changing only the second component without changing the first component.
5. 5. The image processing method of claim 4, wherein the at least one machine learning model is composed of one machine learning model, and the one machine learning model is trained to output the first residual map that changes only the first component without changing the second component, and is trained to output the second residual map that changes only the second component without changing the first component.
6. 5. The image processing method according to claim 4, wherein the at least one machine learning model includes a first learning model trained to output the first residual map by changing only the first component without changing the second component, and a second learning model trained to output the second residual map by changing only the second component without changing the first component.
7. 7. The image processing method according to claim 1, wherein the first component is at least one of a blur component, a scattering component due to fog, a relighting component, and a background blur component that corrects background blur.
8. 8. The image processing method according to claim 1, wherein the second component is a noise component.
9. 9. The image processing method according to claim 1, wherein the third residual map is obtained using luminance values of an image before and after correction using the first residual map.
10. 10. The image processing method according to claim 9, wherein the third residual map is obtained by multiplying a coefficient obtained for each pixel based on the luminance value by the value of the corresponding pixel in the second residual map.
11. 11. The image processing method according to claim 1, wherein the output image is obtained by weighting the first residual map and the third residual map and adding them to an image based on the input image.
12. acquiring an original image; generating a second image by adding a first component to a first image based on the original image; generating a third image and a fourth image by applying a second component to the first image and the second image; and training a machine learning model using a plurality of images from the first to fourth images.
13. a first acquisition unit that inputs an input image to at least one machine learning model to acquire a first residual map corresponding to a first component and a second residual map corresponding to a second component; a correction unit that corrects the second residual map into a third residual map having a different distribution based on the input image and the first residual map; and a second acquisition unit that acquires an output image based on the input image, the first residual map, and the third residual map.
14. An image processing system having a first device and a second device capable of communicating with each other, the first device has a transmission unit that transmits a request for processing an input image to the second device; The second device is a receiving unit that receives the request from the first device; a first acquisition unit that, in response to the request, inputs the input image into at least one machine learning model to acquire a first residual map corresponding to a first component and a second residual map corresponding to a second component; a correction unit that corrects the second residual map into a third residual map having a different distribution based on the input image and the first residual map; and a second acquisition unit that acquires an output image based on the input image, the first residual map, and the third residual map.
15. A program causing a computer to execute the image processing method according to any one of claims 1 to 11.