Image processing method, image processing device, image processing system, and program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2021-12-15
- Publication Date
- 2026-08-03
Smart Images

Figure 0007898849000001 
Figure 0007898849000002 
Figure 0007898849000003
Abstract
Description
Technical Field
[0004] , , , , , , , , , ,
[0005] , , , , , ,
[0001] The present invention relates to an image processing method, an image processing apparatus, an image processing system, and a program.
Background Art
[0006] Therefore, the present invention aims to provide an image processing method that allows for appropriate adjustment of the correction intensity according to user settings in a machine learning regression task for blurred images. [Means for solving the problem]
[0007] One aspect of the present invention is an image processing method comprising the steps of: inputting an image captured by imaging into a machine learning model to generate information regarding correction; generating a weight map based on information regarding saturated regions representing a first region and a second region in the image captured; and in the weight map, At the first set intensity The first territory region Correction strength against The second territory region Correction strength With respect to the ratio, the first region at a second setting intensity that is higher than the first setting intensity. Correction strength The second region Correction strength Increase the ratio of The process and the captured image and the correction information based on the weight map. Report and The process includes generating a first image by weighting and adding the following, wherein the correction is at least one of increasing the resolution of the captured image and transforming the shape of the defocus blur of the captured image, the information relating to the correction is either the corrected image obtained by performing the correction on the captured image, or the correction component which is the difference between the corrected image and the captured image, and the weight map is used in the weighting addition. keru Location The corrected image or the corrected componentThe map shows the proportion of the first region, wherein the first region is either the region in the captured image where the brightness value is saturated, or the region in the region where the subject in the saturated brightness value has expanded due to blurring that occurred during the imaging, and the second region is the region in the captured image other than the first region.
[0008] Other objects and features of the present invention are described in the following examples. [Effects of the Invention]
[0009] According to the present invention, it is possible to provide an image processing method that allows for appropriate adjustment of the correction intensity according to the user's settings in a machine learning regression task on blurred images. [Brief explanation of the drawing]
[0010] [Figure 1] This is an explanatory diagram of the machine learning model in Example 1. [Figure 2] This is a block diagram of the image processing system in Example 1. [Figure 3] This is an external view of the image processing system in Example 1. [Figure 4] This diagram illustrates the adverse effects of sharpening in Examples 1-3. [Figure 5] This diagram illustrates the adverse effects of adjusting the sharpening intensity in Examples 1-3. [Figure 6] This is a flowchart of the machine learning model training in Examples 1-3. [Figure 7] This is a flowchart showing the generation of model outputs in Examples 1 and 2. [Figure 8] This is a flowchart for adjusting the sharpening intensity in Example 1. [Figure 9] This is an explanatory diagram of the captured image and saturation effect map in Example 1. [Figure 10] This figure shows the relationship between optical performance indicators and weights in Examples 1 to 3. [Figure 11]It is an explanatory diagram of the weight map in Example 1. [Figure 12] It is an explanatory diagram of the user interface in Example 1. [Figure 13] It is an explanatory diagram of the correction intensity in Example 1. [Figure 14] It is a block diagram of the image processing system in Example 2. [Figure 15] It is an external view of the image processing system in Example 2. [Figure 16] It is a flowchart of the intensity adjustment of sharpening in Example 2. [Figure 17] It is an explanatory diagram of the user interface in Example 2. [Figure 18] It is a block diagram of the image processing system in Example 3. [Figure 19] It is an external view of the image processing system in Example 3. [Figure 20] It is a flowchart of the intensity adjustment of model output and sharpening in Example 3.
Mode for Carrying Out the Invention
[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In each figure, the same members are denoted by the same reference numerals, and redundant descriptions are omitted.
[0012] Before providing a detailed description of the embodiments, the gist of the present invention will be explained. The present invention generates an estimated image (blur-sharpened image) from an image captured using an optical system, in which blur caused by the optical system is sharpened using a machine learning model. Then, an upper limit of the weight map is determined based on the brightness saturation region where the brightness is saturated, and the captured image and the estimated image are weighted and added together. The weight map is used to determine the proportion of each image when weighting and adding the captured image and the estimated image, and has a continuous signal value. For example, if the value of the weight map determines the proportion of the blur-sharpened image, if the value is 0, the intensity-adjusted image remains the captured image. If the value of the weight map is 1, the intensity-adjusted image becomes the blur-sharpened image.
[0013] Blur caused by optical systems includes aberrations, diffraction, defocusing, the effects of optical low-pass filters, and degradation of the pixel aperture of the image sensor. Machine learning models include, for example, neural networks, genetic programming, and Bayesian networks. Neural networks include CNNs (Convolutional Neural Networks), GANs (Generative Adversarial Networks), and RNNs (Recurrent Neural Networks).
[0014] Blur sharpening refers to the process of restoring frequency components of a subject that have been reduced or lost due to blur. During blur sharpening, undershoot and ringing may not be suppressed depending on the captured image, and these problems may occur in the estimated image. Specifically, problems occur when the subject is significantly blurred due to optical system aberrations, or when there are luminance saturation regions in the image. Luminance saturation regions can occur in an image depending on the dynamic range of the image sensor and the exposure during imaging. In luminance saturation regions, it is difficult to obtain information about the structure of the subject space, which makes problems likely to occur. These problems can be suppressed by weighting and adding the captured image and the estimated image based on the luminance saturation region and optical performance.
[0015] However, to address the need to increase the correction strength even if it means tolerating some drawbacks, it is preferable to have a function to adjust the correction strength. Here, if the upper limit of the correction strength is set to be the same for the luminance saturation region and the non-saturation region (changing the correction strength of the luminance saturation region and the non-saturation region at a constant ratio), it becomes impossible to make appropriate strength adjustments according to the user's settings. In other words, if the correction strength is increased, the correction strength in the non-saturation region becomes insufficient (when the upper limit is set to match the allowable amount of drawbacks in the saturation region), or the correction strength in the luminance saturation region becomes excessive (when the upper limit is set to match the allowable amount of drawbacks in the non-saturation region).
[0016] Therefore, in each embodiment, an upper limit of the weight map is determined for both the luminance-saturated and unsaturated regions (by changing the ratio of the correction intensity of the luminance-saturated and unsaturated regions based on the set intensity), and the captured image and the estimated image are weighted and added together. This makes it possible to control the sharpness and detrimental effects of blur. In the following, the stage of learning the weights of the machine learning model will be referred to as the learning phase, and the stage of sharpening the blur using the machine learning model with the learned weights will be referred to as the estimation phase. [Examples]
[0017] First, the image processing system in Embodiment 1 of the present invention will be described. In this embodiment, the task performed by the machine learning model is to sharpen the blur of an captured image, including luminance saturation (increase the resolution of the captured image). The blur to be sharpened is limited to aberrations and diffractions generated in the optical system, as well as blur caused by optical low-pass filters. However, the effects of the invention can also be obtained when sharpening blur caused by pixel aperture, defocus, or blurring. Furthermore, this embodiment can also be implemented and the effects obtained for tasks other than blur sharpening. Specifically, these include upsampling to increase the number of pixels in the captured image, and tasks to transform (transform) the defocus blur of the captured image. Transformation of defocus blur includes, for example, transformation from double-line blur to Gaussian blur or spherical blur. Double-line blur has a PSF with separated peaks. As a result, a subject that is originally a single line appears to be double-blurred when defocused. Spherical blur has a PSF with flat intensity. Gaussian blur has a PSF with a Gaussian distribution. Other types of defocus blur that can be converted include, for example, defocus blur caused by vignetting and ring-shaped defocus blur caused by pupil occlusion, such as with catadioptric lenses.
[0018] Figure 2 is a block diagram of the image processing system 100 in this embodiment. Figure 3 is an external view of the image processing system 100. The image processing system 100 has a learning device 101 and an image processing device 103 connected by a wired or wireless network. An imaging device 102, a display device 104, a recording medium 105, and an output device 106 are connected to the image processing device 103 by wired or wireless. The captured image, which captures the subject space using the imaging device 102, is input to the image processing device 103. The captured image is blurred due to aberrations and diffraction by the optical system 102a in the imaging device 102 and the optical low-pass filter of the image sensor 102b, and the information of the subject is attenuated.
[0019] The image processing device 103 uses a machine learning model to sharpen the blur of an captured image and generates a saturation effect map and a blur-sharpened image (model output). Details of the saturation effect map will be described later. The machine learning model is trained by the learning device 101, and the image processing device 103 has previously acquired information about the machine learning model from the learning device 101 and stored it in the storage unit 103a. The image processing device 103 also has a function to adjust the intensity of blur sharpening by taking a weighted addition of the captured image and the blur-sharpened image. Details of the training and estimation of the machine learning model and the adjustment of the blur sharpening intensity will be described later. The user can adjust the blur sharpening intensity while checking the image displayed on the display device 104. The blur-sharpened image with adjusted intensity is stored in the storage unit 103a or the recording medium 105 and output to an output device 106 such as a printer as needed. The captured image may be grayscale or have multiple color components. It may also be an undeveloped RAW image or a developed image.
[0020] Next, referring to Figures 4(A) to (C), we will explain the decrease in estimation accuracy (detrimental effects of sharpening) that occurs when performing blur sharpening using a machine learning model. Figures 4(A) to (C) are explanatory diagrams of the detrimental effects of sharpening and show the spatial changes in the signal values of the image. Here, since the image is an 8-bit developed image, the saturation value is 255. In Figures 4(A) to (C), the solid line represents the captured image (blurred image), and the dotted line represents the blur-sharpened image obtained by sharpening the blur of the captured image using a machine learning model. Figure 4(A) shows the result of sharpening a non-luminance saturated subject with large blur due to optical system aberrations, Figure 4(B) shows a non-luminance saturated subject with small blur due to optical system aberrations, and Figure 4(C) shows the result of sharpening a luminance saturated subject with small blur due to optical system aberrations.
[0021] When blurring due to optical system aberrations is significant, undershoot occurs in the dark areas of the edges. Furthermore, even when blurring due to optical system aberrations is small, sharpening a luminance-saturated subject can result in undershoot (which did not occur with non-luminance-saturated subjects) and a decrease in the originally saturated pixel values. In luminance-saturated regions, information about the structure of the subject space is lost, false edges may appear at the boundaries of each region, and the correct features of the subject cannot be extracted. This reduces the estimation accuracy of machine learning models. From these results, it can be seen that the drawbacks associated with sharpening depend on the performance of the optical system and the luminance-saturated region.
[0022] The correction described above uses a machine learning model trained with two methods: one that uses the captured image and its corresponding brightness saturation map as input data, and another that generates a saturation effect map. In other words, while these methods can reduce the harmful effects, it is difficult to completely eliminate them. The method using the brightness saturation map and the method generating the saturation effect map will be explained in detail below.
[0023] First, let's explain the luminance saturation map. A luminance saturation map is a map that represents the luminance saturation regions in an captured image. In regions where luminance saturation occurs (luminance saturation regions), information about the structure of the subject space is lost, false edges may appear at the boundaries of each region, and it becomes impossible to extract the correct features of the subject. Therefore, by inputting a luminance saturation map, the neural network can identify the problematic regions mentioned above, thereby suppressing the decrease in estimation accuracy.
[0024] Next, we will explain saturation effect maps. Even when using luminance saturation maps, machine learning models may not make correct judgments. For example, if the area of interest is near a luminance-saturated area, the machine learning model can determine that the area of interest is affected by luminance saturation because there is a luminance-saturated area nearby. However, if the area of interest is far from the luminance-saturated area, it is not easy to determine whether or not this area is affected by luminance saturation, and the ambiguity increases. As a result, the machine learning model may make a misjudgment at locations far from the luminance-saturated area. In this case, if the task is blur sharpening, the sharpening process corresponding to the saturated blur image will be performed on the unsaturated blur image. In this case, artifacts occur in the image after the blur has been sharpened, and the accuracy of the task decreases. Therefore, it is preferable to generate a saturation effect map from the captured image in which blur has occurred using a machine learning model.
[0025] A saturation effect map is a map (a spatially arranged sequence of signals) that represents the magnitude and range of signal values that have spread due to blurring (degradation) during image acquisition in the luminance saturation region of an image. By having a machine learning model generate a saturation effect map, the machine learning model can accurately estimate the presence and magnitude of the luminance saturation effect in the image. Once a saturation effect map is generated, the machine learning model can execute the appropriate processing on the regions affected by luminance saturation and the other regions. Therefore, having a machine learning model generate a saturation effect map improves the accuracy of the task compared to cases where a saturation effect map is not generated (where only recognition labels and blur-sharpened images are generated directly from the image).
[0026] While the two methods described above are effective, it is difficult to completely eliminate the drawbacks, as explained with reference to Figures 4(A) to (C). Therefore, the drawbacks can be suppressed by performing weighted addition of the captured image and the blur-sharpened image. The dashed lines in Figures 4(A) and (C) represent the signal values obtained by weighted addition of the captured image and the blur-sharpened image. By performing weighted addition, the drawbacks of undershoot in dark areas and a decrease in pixel values that were originally saturated are mitigated while maintaining the blur sharpening effect. In this embodiment, the weight map used when weighted addition of the captured image and the blur-sharpened image is generated based on the performance of the optical system and the brightness saturation region. By basing it on the performance of the optical system and the brightness saturation region, it is possible to suppress the drawbacks only in Figures 4(A) and (C) where they occur, while maintaining the sharpness in Figure 4(B) where no drawbacks occur, thus making it possible to control the degree of blur sharpness and the drawbacks.
[0027] Next, we will explain the function for adjusting the correction intensity. To address the need to increase the correction intensity even if it means accepting some drawbacks, it is preferable to have a function to adjust the correction intensity. In that case, it is necessary to determine the upper limit of the correction intensity, but if the same upper limit is set for the luminance saturation region and the non-saturation region, as mentioned above, the correction intensity in the non-saturation region will be insufficient, or the correction intensity in the luminance saturation region will be excessive. The luminance saturation region includes the saturation effect map.
[0028] Figures 5(A) and (B) illustrate the drawbacks of adjusting the sharpening intensity. They show the change in signal values (spatial change in image signal values) when the sharpening intensity is adjusted by changing the weight map used when weight-adding the captured image and the blur-sharpened image. Since the image is developed in 8 bits, the saturation value is 255. Figure 5(A) shows the change in signal values in the luminance saturation region (region where luminance is saturated), and Figure 5(B) shows the change in signal values in the non-saturated region (region where luminance is not saturated). The dashed line 1001 is the signal value of the captured image. The dotted line 1002 is the signal value of the intensity-adjusted image, where the sharpness intensity of the blur is adjusted using the weight map. The dashed line 1003 is the maximum intensity image, where the correction intensity is doubled compared to the intensity-adjusted image.
[0029] In Figure 5(A), undershoot occurs in the dark areas of the edges in the maximum intensity image. On the other hand, in Figure 5(B), no undershoot occurs even in the maximum intensity image. In other words, as the correction intensity is increased, problems occur first in the luminance saturation region. The dashed line 1004 represents the signal value of an intensity-adjusted image in which the luminance saturation region is clipped with a correction intensity of 100%, and the unsaturated region is corrected with a correction intensity of 200%. A correction intensity of 100% is the blur-sharpened image output by the machine learning model, and a correction intensity of 200% means that the absolute value of the difference between the blur-sharpened image and the captured image is doubled and added to the captured image. According to the dashed line 1004, it is possible to suppress the problems in the luminance saturation region while maintaining the correction intensity in the unsaturated region.
[0030] Next, with reference to Figure 6, the training of the machine learning model performed by the learning device 101 will be described. Figure 6 is a flowchart of the machine learning model training. The learning device 101 has a storage unit 101a, an acquisition unit 101b, an arithmetic unit 101c, and an update unit 101d, and each step in Figure 6 is mainly performed by the respective parts of the learning device 101.
[0031] First, in step S101, the acquisition unit 101b acquires one or more original images from the storage unit 101a. The original image is an image with a signal value higher than the second signal value. The second signal value is a signal value corresponding to the brightness saturation value of the captured image. However, since the signal value may be normalized when inputting into the machine learning model, the second signal value and the brightness saturation value of the captured image do not necessarily have to match. Since the machine learning model is trained based on the original image, it is desirable that the original image is an image with various frequency components (edges with different directions and intensities, gradients, flat areas, etc.). The original image may be a real photograph or a computer graphics image.
[0032] Next, in step S102, the calculation unit 101c adds blur to the original image to generate a blurred image. The blurred image is the image input to the machine learning model during training and corresponds to the captured image during estimation. The blur added is the blur that will be sharpened. In this embodiment, blur generated by the aberrations and diffraction of the optical system 102a and the optical low-pass filter of the image sensor 102b is added. The shape of the blur due to the aberrations and diffraction of the optical system 102a changes depending on the image plane coordinates (image height and azimuth). It also changes depending on the magnification, aperture, and focus state of the optical system 102a. If you want to train a machine learning model that sharpens all of these blurs at once, it is advisable to generate multiple blurred images using multiple blurs generated by the optical system 102a. In addition, in the blurred image, signal values exceeding the second signal value are clipped. This is done to reproduce the brightness saturation that occurs during the image acquisition process of the captured image. If necessary, noise generated by the image sensor 102b may be added to the blurred image.
[0033] Next, in step S103, the calculation unit 101c sets a first region based on an image based on the original image and a threshold signal value. In this embodiment, a blurred image is used as the image based on the original image, but the original image itself may also be used. The first region is set by comparing the signal value of the blurred image with the threshold signal value. More specifically, the region where the signal value of the blurred image is equal to or greater than the threshold signal value is defined as the first region. In this embodiment, the threshold signal value is the second signal value. Therefore, the first region represents the luminance-saturated region (saturated region) of the blurred image. However, the threshold signal value and the second signal value do not have to be the same. The threshold signal value may be set to a value slightly smaller than the second signal value (for example, 0.9 times).
[0034] Next, in step S104, the calculation unit 101c generates a first region image having the signal values of the original image in the first region. The first region image has different signal values from the original image in regions other than the first region. More preferably, the first region image has the first signal value in regions other than the first region. In this embodiment, the first signal value is 0, but the invention is not limited thereto. In this embodiment, the first region image has the signal values of the original image only in the region where the blurred image is luminance saturated, and the signal values in the other regions are 0.
[0035] Next, in step S105, the calculation unit 101c applies blur to the first region image and generates a saturation effect ground truth map. The applied blur is the same as the blur applied to the blurred image. This generates a saturation effect ground truth map, which is a map (a spatially arranged sequence of signals) representing the magnitude and range of the signal values that have spread due to degradation during imaging, from the subject in the brightness-saturated region of the blurred image. In this embodiment, the saturation effect ground truth map is clipped with the second signal value, similar to the blurred image, but clipping is not always necessary.
[0036] Next, in step S106, the acquisition unit 101b acquires the ground truth model output. Since the task in this embodiment is blur sharpening, the ground truth model output is an image with less blur than the blurred image. In this embodiment, the ground truth model output is generated by clipping the original image with a second signal value. If the original image lacks high-frequency components, the ground truth model output may be an image obtained by reducing the original image. In this case, the same reduction is performed when generating the blurred image in step S102. Also, step S106 may be executed at any time, as long as it is after step S101 and before step S107.
[0037] Next, in step S107, the calculation unit 101c generates a saturation effect map and a model output based on the blurred image using a machine learning model. In this embodiment, the machine learning model shown in Figure 1 is used, but the invention is not limited to this. The blurred image 201 and the luminance saturation map 202 are input to the machine learning model. The luminance saturation map 202 is a map that shows the luminance-saturated region (where the signal value is greater than or equal to the second signal value) of the blurred image 201. For example, it can be generated by binarizing the blurred image 201 with the second signal value. However, the luminance saturation map 202 is not necessarily required. The blurred image 201 and the luminance saturation map 202 are concatenated in the channel direction and input to the machine learning model. However, the invention is not limited to this. For example, the blurred image 201 and the luminance saturation map 202 may each be converted into feature maps, and these feature maps may be concatenated in the channel direction. In addition, information other than the luminance saturation map 202 may be added to the input.
[0038] The machine learning model has multiple layers, and in each layer, a linear sum is taken of the layer's input and weights. The initial values of the weights can be determined by random numbers or the like. In this embodiment, the machine learning model is a CNN that uses the convolution of the input and the filter (the values of each element of the filter correspond to the weights, and may also include the sum with the bias) as the linear sum, but the invention is not limited to this. In addition, each layer may perform a nonlinear transformation using an activation function such as ReLU (Rectified Linear Unit) or a sigmoid function as needed. Furthermore, the machine learning model may have residual blocks or Skip Connections (also called Shortcut Connections) as needed. As a result of passing through multiple layers (16 convolutional layers in this embodiment), a saturation influence map 203 is generated.
[0039] In this embodiment, the saturation effect map 203 is obtained by taking the element-wise sum of the output of layer 211 and the luminance saturation map 202, but the configuration is not limited to this. The saturation effect map may be generated directly as the output of layer 211. Alternatively, the saturation effect map 203 may be obtained as the result of performing any processing on the output of layer 211. Next, the saturation effect map 203 and the blurred image 201 are concatenated in the channel direction and input to the subsequent layer, and the model output 204 is generated as a result of passing through multiple layers (16 convolutional layers in this embodiment). The model output 204 is also generated by taking the element-wise sum of the output of layer 212 and the blurred image 201, but the configuration is not limited to this. In this embodiment, each layer performs convolution with 64 types of 3x3 filters (however, layers 211 and 212 have the same number of filter types as the number of channels in the blurred image 201), but the configuration is not limited to this.
[0040] Next, in step S108, the update unit 101d updates the weights of the machine learning model based on the error function. In this embodiment, the error function is a weighted sum of the error between the saturation effect map 203 and the saturation effect ground truth map, and the error between the model output 204 and the ground truth model output. Mean Squared Error (MSE) is used to calculate the error. Both weights are set to 1. However, the error function and weights are not limited to these. Backpropagation or similar methods may be used to update the weights. The error may also be taken for the residual component. In the case of the residual component, the error of the difference component between the saturation effect map 203 and the luminance saturation map 202, and the difference component between the saturation effect ground truth map and the luminance saturation map 202 is used. Similarly, the error of the difference component between the model output 204 and the blurred image 201, and the difference component between the ground truth model output and the blurred image 201 are used.
[0041] Next, in step S109, the update unit 101d determines whether the training of the machine learning model is complete. Completion of training can be determined by whether the number of iterations of weight updates has reached a predetermined number, or whether the amount of change in weights during the update is less than a predetermined value. If it is determined in step S109 that training is not complete, the process returns to step S101, and the acquisition unit 101b acquires one or more new original images. On the other hand, if it is determined that training is complete, the update unit 101d terminates training and stores the configuration of the machine learning model and the weight information in the storage unit 101a.
[0042] Through the learning method described above, the machine learning model can estimate a saturation effect map that represents the magnitude and range of the signal values that have spread due to blurring in the brightness-saturated region of the subject in the blurred image (or captured image during estimation). By explicitly estimating the saturation effect map, the machine learning model can perform blur sharpening on the appropriate regions for both saturated and unsaturated blurred images, thereby suppressing the occurrence of artifacts.
[0043] Next, referring to Figure 7, we will describe the blur sharpening of captured images using a trained machine learning model, which is performed in the image processing device 103. Figure 7 is a flowchart of the generation of the model output. The image processing device 103 has a storage unit 103a, an acquisition unit 103b, and a sharpening unit 103c, and each step in Figure 7 is mainly performed by the respective parts of the image processing device 103.
[0044] First, in step S201, the acquisition unit 103b acquires the captured image and the machine learning model. Information on the configuration and weights of the machine learning model is acquired from the storage unit 103a.
[0045] Next, in step S202, the sharpening unit (first generation means) 103c generates correction information using a machine learning model. In this embodiment, the correction information is a blur-sharpened image (model output) in which the blur of the captured image has been sharpened from the captured image. Note that it may be a correction component of blur sharpening instead of a blur-sharpened image (image corrected from the captured image). The machine learning model has the same configuration as during training, as shown in Figure 1. As during training, a luminance saturation map representing the luminance-saturated region of the captured image is generated and input, and a saturation effect map and model output are generated. As an example, Figure 9(A) shows a blur-sharpened image (captured image), and Figure 9(B) shows a saturation effect map.
[0046] Next, with reference to Figure 8, the synthesis of the captured image and the model output, performed by the image processing device 103, will be described. Figure 8 is a flowchart for adjusting the sharpening intensity. Each step in Figure 8 is mainly performed by the various parts of the image processing device 103.
[0047] First, in step S211, the acquisition unit 103b acquires the shooting state from the captured image. The shooting state refers to the zoom position, aperture diameter, and subject distance of the optical system 102a (z, f, d), as well as the pixel pitch of the image sensor 102b.
[0048] Next, in step S212, the acquisition unit 103b acquires information regarding the optical performance of the optical system 102a (optical performance index) based on the shooting state acquired in step S211. The optical performance index is stored in the storage unit 103a. The optical performance index is information regarding the optical performance of the optical system 102a used to capture the image, which is independent of the subject space, and does not include information that is not independent of the subject space, such as a saturation effect map. In this embodiment, the magnitude (peak value) and range (spread) of the point image distribution function (PSF) are used as the optical performance index. The peak value is the maximum signal value that the PSF has, and the range means the number of pixels that have a value above a certain threshold. When sharpening blur with a machine learning model, even if the peak value is the same, blur with a smaller number of pixels that have a value above a certain threshold produces less harm, so the optical performance index is used in this embodiment. Also, since the peak value and range of the PSF depend on the pixel pitch of the imaging device 102, optical performance indexes corresponding to multiple pixel pitches are stored, and intermediate values are created by interpolation.
[0049] Furthermore, since the optical performance index only needs to reflect the optical performance, other values such as the optical transfer function (OTF) may be used as the optical performance index. For example, the optical performance index is calculated based on at least one of the zoom position, aperture diameter, or subject distance at the time of image capture, and the magnitude and range of the signal values of the point image distribution function for each image height of the optical system.
[0050] The memory unit 103a stores only optical performance indicators for discretely selected imaging states in order to reduce the number of optical performance indicators (data points). Therefore, if an optical performance indicator corresponding to the imaging state acquired in step S211, or an optical performance indicator corresponding to an imaging state close to the acquired state, is not stored in the memory unit 103a, an optical performance indicator as close as possible to that imaging state is selected. Then, the optical performance indicator to be actually used is created by correcting that optical performance indicator to optimize it for the imaging state acquired in step S211.
[0051] Next, in step S213, the sharpening unit 103c generates a first weight map and a second weight map from the optical performance index. When determining the weights, the weights of the blur-sharpened images are determined for each image height based on the optical performance index obtained in step S212 and the relational expression shown in Figure 10. Figure 10 is a diagram showing the relationship between the optical performance index and the weights. In Figure 10, the horizontal axis represents the optical performance index, and the vertical axis represents the weights of the blur-sharpened images.
[0052] The solid line 121 is for the non-saturated region, and the dotted line 122 is for the luminance-saturated region. Even with the same optical performance index, the weight of the blur-sharpened image is reduced in the luminance-saturated region where adverse effects are more likely to occur. In this embodiment, a linear equation is used, but the relationship is not limited to a linear equation. Furthermore, the relationship can be freely changed. Note that image heights for which the optical performance index is not held are generated by interpolation from image height points for which the index is held. A first weight map for the non-saturated region is generated from the solid line 121, and a second weight map for the luminance-saturated region is generated from the dotted line 122.
[0053] Next, in step S214 of Figure 8, the sharpening unit 103c generates a weight map based on the first weight map, the second weight map, and the saturation effect map. That is, the weight map is generated based on information about optical performance (first weight map, second weight map) and information about the saturation region of the captured image. In this embodiment, the information about the saturation region is the saturation effect map, and it is not necessary for all RGB to be saturated. The saturation effect map is normalized by the second signal value and used to combine the first weight map and the second weight map. That is, if the normalized value is between 0 and 1, both the first weight map and the second weight map contribute. By using the saturation effect map, it is possible to adjust the correction intensity up to the saturation effect region. Note that using the saturation effect map is not mandatory; a luminance saturation map may be used, or a luminance saturation map blurred for each image height may be used.
[0054] Figures 11(A) to (C) are explanatory diagrams of an example weight map. In this embodiment, the weight map shows that the correction strength is stronger as the pixel value increases and weaker as the pixel value decreases. In a typical optical system, optical performance decreases as you move off-axis, so the weight map often has a gradient like Figure 11(B). Also, even with the same optical performance, the saturation region is prone to problems, so the correction strength is reduced even at the same image height. Note that the weight map may also be set so that the correction strength is stronger as the pixel value decreases and weaker as the pixel value increases.
[0055] Next, in step S215 of Figure 8, the acquisition unit 103b acquires the sharpening intensity specified (set) by the user. An example of user-controlled intensity adjustment will be explained with reference to Figure 12. Figure 12 is an explanatory diagram of the user interface. The user can adjust the correction intensity by moving the slider to change the set intensity. Alternatively, the user may input a numerical value for the set intensity. For example, if the set intensity is 1.0, the default intensity adjustment image is created by combining the captured image and the model output based on the weight map generated in step S214. If the set intensity is set to 1.5, an intensity adjustment image is generated in which the correction intensity is 1.5 times that of the default intensity.
[0056] Next, in step S216, the sharpening unit 103c changes the weight map according to the acquired sharpening intensity. Figure 11 shows examples of weight map changes. Figure 11(B) shows the weight map when the set intensity is 1.0, which is the default weight map generated in step S211. Figure 11(C) shows the weight map when the set intensity is 2.0, which is the default weight map uniformly doubled across the screen. Figure 11(A) shows the weight map when the set intensity is 0, where all weights are 0.
[0057] Here, we will explain the change in correction intensity with reference to Figure 13. Figure 13 is an explanatory diagram of the correction intensity. 1101 shows the change in correction intensity in the non-saturated region, and 1102 shows the change in correction intensity in the luminance saturation region. Both have the same image height and are regions with high optical performance. In this embodiment, the correction intensity in the non-saturated region is defined as the first correction intensity, and the correction intensity in the luminance saturation region is defined as the second correction intensity. The correction intensity in the saturation-affected region is clipped at 100%. If the setting intensity is 1.0 as the first setting intensity and the setting intensity is 2.0 as the second setting intensity, the ratio of the first correction intensity to the second correction intensity is different for the first and second setting intensity. That is, the ratio of the second correction intensity to the first correction intensity changes based on the setting intensity.
[0058] In Figure 13, 1103 shows the change in correction intensity in the non-saturated region, and 1104 shows the change in correction intensity in the luminance saturation region. Both are at the same image height and represent a region with low optical performance. Due to the low optical performance, the default correction intensity is low, and even with a set intensity of 2.0, the correction intensity in the luminance saturation region does not reach 100%. Thus, the upper limit of the correction intensity differs depending on the optical performance. Note that the relationship between correction intensity and set intensity shown in Figure 13 is not limited to this. For example, instead of clipping, one could use the formula 1102, which asymptotically approaches 100% correction intensity at a set intensity of 2.0. Also, there is no fixed upper limit for the set intensity, and it may be possible to set an intensity of 2.0 or higher.
[0059] Next, in step S217 of Figure 8, the sharpening unit 103c weights and adds the captured image and the blur-sharpened image (model output) based on the weight map to generate an intensity adjustment image 205. Step S216 is optional; all intensity adjustment images that can be set with a slider may be generated in advance, and the image may be displayed according to the intensity specified by the user.
[0060] In this embodiment, the sharpening section 103c functions as a first generation means, a second generation means, and an adjustment means. The first generation means inputs the captured image into a machine learning model and generates information regarding correction. The second generation means generates an intensity adjustment image based on the captured image, the correction information, and a weight map. The adjustment means adjusts the correction intensity of the intensity adjustment image by changing the weight map based on the set intensity. In this embodiment, the correction intensity includes a first correction intensity in a first region of the intensity adjustment image and a second correction intensity in a second region of the intensity adjustment image, based on information regarding the optical performance of the optical system used to capture the captured image and information regarding the saturation region. The ratio of the second correction intensity to the first correction intensity changes based on the set intensity. Preferably, when the set intensity is a first set intensity, the ratio is a first ratio, and when the set intensity is a second set intensity which is stronger than the first set intensity, the ratio is a second ratio which is lower than the first ratio.
[0061] With the above configuration, this embodiment provides an image processing system that can control the sharpness and detrimental effects of blur in a machine learning-based regression task on blurred images. [Examples]
[0062] Next, an image processing system in Embodiment 2 of the present invention will be described. In this embodiment, the intensity of the non-saturated region and the brightness-saturated region is adjusted using separate sliders. Figure 14 is a block diagram of the image processing system 300 in this embodiment. Figure 15 is an external view of the image processing system 300. The image processing system 300 includes a learning device 301, an imaging device 302, and an image processing device 303. The learning device 301 and the image processing device 303, and the image processing device 303 and the imaging device 302 are connected by a wired or wireless network, respectively.
[0063] The imaging device 302 includes an optical system 321, an image sensor 322, a storage unit 323, a communication unit 324, and a display unit 325. The captured image is transmitted to the image processing device 303 via the communication unit 324. The image processing device 303 receives the captured image via the communication unit 332 and performs blur sharpening using the configuration and weight information of the machine learning model stored in the storage unit 331. The configuration and weight information of the machine learning model is learned by the learning device 301, is acquired in advance from the learning device 301, and is stored in the storage unit 331. Furthermore, the image processing device 303 has a function to adjust the intensity of blur sharpening. The blur-sharpened image (model output) with the blur of the captured image sharpened and the intensity-adjusted image with the intensity adjusted are transmitted to the imaging device 302, stored in the storage unit 323, and displayed in the display unit 325.
[0064] The generation of training data and training of weights performed by the learning device 301 (training phase) and the blur sharpening of captured images using the trained machine learning model performed by the image processing device 303 (estimation phase) are the same as in Example 1 and will therefore be omitted.
[0065] Next, with reference to Figure 16, the synthesis of the captured image and model output performed by the image processing device 303 will be described. Figure 16 is a flowchart of the sharpening intensity adjustment. Each step in Figure 16 is mainly performed by the various parts of the image processing device 303.
[0066] First, in step S311, the sharpening unit 334 generates a weight map. The method for generating the weight map is the same as in Example 1, so it will be omitted here. Next, in step S312, the acquisition unit 333 acquires the sharpening intensity specified by the user. An example of user-controlled intensity adjustment will be explained with reference to Figure 17. Figure 17 is an explanatory diagram of the user interface. The user can adjust the correction intensity by moving the slider. In Example 2, the user can adjust the correction intensity of the luminance saturation region and the non-saturation region individually.
[0067] Next, in step S313, the sharpening unit 334 changes the weight map according to the acquired sharpening intensity. Then, in step S314, the sharpening unit 334 weights and adds the captured image and the blurred sharpened image (model output) based on the weight map to generate an intensity-adjusted image.
[0068] With the above configuration, this embodiment provides an image processing system that can control the sharpness and detrimental effects of blur in a machine learning-based regression task on blurred images. [Examples]
[0069] Next, an image processing system in Embodiment 3 of the present invention will be described. Figure 18 is a block diagram of the image processing system 400 in this embodiment. Figure 19 is an external view of the image processing system 400. The image processing system 400 includes a learning device 401, a lens device 402, an imaging device 403, a control device (first device) 404, an image estimation device (second device) 405, and networks 406 and 407.
[0070] The learning device 401 and the image estimation device 405 are, for example, servers. The control device 404 is a user-operated device such as a personal computer or mobile terminal. The learning device 401 has a storage unit 401a, an acquisition unit 401b, a calculation unit 401c, and an update unit 401d, and learns the weights of a machine learning model that sharpens blur from captured images captured using the lens device 402 and the imaging device 403. The learning method, i.e., the generation of learning data and learning of weights (learning phase) performed by the learning device 401, is the same as in Example 1 and will therefore be omitted.
[0071] The imaging device 403 has an image sensor 403a, which converts the optical image formed by the lens device 402 into an optical image via photoelectric conversion to acquire an image. The lens device 402 and the imaging device 403 are detachable and can be combined with multiple types of lenses. The control device 404 has a communication unit 404a, a display unit 404b, a storage unit 404c, and an acquisition unit 404d, and controls the processing to be performed on the image acquired from the imaging device 403, which is connected by wire or wireless, according to the user's operation. Alternatively, the image captured by the imaging device 403 may be stored in the storage unit 404c in advance, and the image may be read out.
[0072] The image estimation device 405 includes a communication unit 405a, an acquisition unit 405b, a storage unit 405c, and a sharpening unit 405d, and is configured to communicate with the control device 404. The image estimation device 405 performs a blur sharpening process on the captured image in response to a request from the control device 404, which is connected via the network 406. The image estimation device 405 acquires learned weight information from the learning device 401, which is connected via the network 406, either at the time of blur sharpening estimation or in advance, and uses it to estimate the blur sharpening of the captured image. The estimated image after blur sharpening estimation is transmitted back to the control device 404 after the sharpening intensity has been adjusted, stored in the storage unit 404c, and displayed on the display unit 404b.
[0073] Next, with reference to Figure 20, we will describe the blur sharpening of the captured image performed by the control device 404 and the image estimation device 405. Figure 20 is a flowchart of the model output and sharpening intensity adjustment. Each step in Figure 20 is mainly performed by the respective parts of the control device 404 or the image estimation device 405.
[0074] First, in step S401, the acquisition unit 404d of the control device 404 acquires the captured image and the sharpening intensity specified by the user. Next, in step S402, the communication unit (transmission means) 404a transmits the captured image and a request to the image estimation device 405 for the execution of the blur sharpening estimation process.
[0075] Next, in step S403, the communication unit (receiving means) 405a of the image estimation device 405 receives and acquires the captured image and processing request transmitted from the control device 404. Next, in step S404, the acquisition unit 405b acquires the learned weight information corresponding to the captured image from the storage unit 405c. The weight information is read in advance from the storage unit 401a and stored in the storage unit 405c. Next, in step S405, the sharpening unit 405d uses a machine learning model to generate a blur-sharpened image (model output) from the captured image, in which the blur of the captured image has been sharpened. The machine learning model has the same configuration as during training, as shown in Figure 1. As during training, a luminance saturation map representing the luminance-saturated region of the captured image is generated and input, and a saturation effect map and model output are generated.
[0076] Next, in step S406, the sharpening unit 405d generates a weight map. The method for generating the weight map is the same as in Embodiment 1. The default weight map is adjusted according to the sharpening intensity specified by the user. Note that a pre-adjusted weight map may be kept within a range where the intensity can be adjusted. Next, in step S407, the sharpening unit 405d weights and adds the captured image and the blurred sharpened image (model output) based on the weight map. Next, in step S408, the communication unit 405a transmits the composite image to the control device 404. Next, in step S409, the communication unit 404a of the control device 404 acquires the estimated image transmitted from the image estimation device 405.
[0077] With the above configuration, this embodiment provides an image processing system that can control the sharpness and detrimental effects of blur in a machine learning-based regression task on blurred images.
[0078] (Other examples) The present invention can also be realized by supplying a program that implements one or more of the functions of the above embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.
[0079] According to each embodiment, it is possible to provide an image processing method, an image processing device, an image processing program, and a storage medium that can appropriately adjust the correction intensity according to the user's settings in a regression task using machine learning on blurred images.
[0080] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of its essence. [Explanation of Symbols]
[0081] 103 Image Processing Device 103c Sharpening section (first generating means, second generating means, adjustment means)
Claims
1. The process involves inputting the captured images obtained through imaging into a machine learning model to generate information related to correction, A step of generating a weight map based on information regarding saturated regions representing a first region and a second region in the captured image, In the weight map, the process involves increasing the ratio of the correction intensity of the second region to the correction intensity of the first region at a second setting intensity that is higher than the first setting intensity, compared to the ratio of the correction intensity of the second region to the correction intensity of the first region at a first setting intensity. The process includes generating a first image by weighting and adding the captured image and the correction information based on the weight map, The correction is at least one of increasing the resolution of the captured image and transforming the shape of the defocus blur in the captured image. The information relating to the correction is either the corrected image obtained by performing the correction on the captured image, or the correction component which is the difference between the corrected image and the captured image. The weight map is a map showing the proportion of the corrected image or the corrected component for each position in the weighted addition, The first region is either a region in the captured image where the brightness value is saturated, or a region in the region where the subject in the saturated brightness value is expanded due to blurring that occurred during the image capture. The image processing method is characterized in that the second region is a region in the captured image other than the first region.
2. The image processing method according to claim 1, characterized in that, in the step of generating the weight map, the weight map is further generated based on information regarding the optical performance of the optical system used for imaging.
3. The image processing method according to claim 2, characterized in that the information relating to the optical performance is calculated based on at least one of the zoom position, aperture diameter, and subject distance in the imaging, and the point image distribution function for each image height of the optical system.
4. The image processing method according to any one of claims 1 to 3, characterized in that the upper limit of the correction intensity in the weight map is determined based on information regarding the optical performance of the optical system used for imaging.
5. The image processing method according to any one of claims 1 to 4, characterized in that the second set intensity is an intensity set by the user.
6. The image processing method according to any one of claims 1 to 5, characterized in that the information regarding the saturated region is generated by the machine learning model.
7. A means for generating correction information by inputting the captured images obtained through imaging into a machine learning model, Means for generating a weight map based on information regarding saturated regions representing a first region and a second region in the captured image, In the weight map, adjustment means for increasing the ratio of the correction intensity of the second region to the correction intensity of the first region at a second setting intensity that is higher than the first setting intensity, with respect to the ratio of the correction intensity of the second region to the correction intensity of the first region at a first setting intensity. The system includes means for generating a first image by weighting and adding the captured image and the correction information based on the weight map, The correction is at least one of increasing the resolution of the captured image and transforming the shape of the defocus blur in the captured image. The information relating to the correction is either the corrected image obtained by performing the correction on the captured image, or the correction component which is the difference between the corrected image and the captured image. The weight map is a map showing the proportion of the corrected image or the corrected component for each position in the weighted addition, The first region is either a region in the captured image where the brightness value is saturated, or a region in the region where the subject in the saturated brightness value is expanded due to blurring that occurred during the image capture. The image processing apparatus is characterized in that the second region is a region in the captured image other than the first region.
8. The image processing apparatus according to claim 7 and a control device capable of communicating with the image processing apparatus, The control device has a transmission means for transmitting a request to the image processing device to perform processing on the captured image, The image processing system is characterized in that the image processing device has receiving means for receiving the request and performs processing on the captured image in response to the request.
9. A program characterized by causing a computer to execute the image processing method described in any one of claims 1 to 6.