Learning method, image processing method, and program
A multi-layer neural network-based learning method corrects blur in imaging devices by training on intermediate data, addressing overcorrection and undercorrection issues, resulting in a more accurate corrected image.
Patent Information
- Application Number
- JP2020093827
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-05-29
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2040-05-29
AI Technical Summary
Existing image processing methods for correcting blur caused by optical characteristics in imaging devices often result in overcorrection or undercorrection due to deviations between expected and actual blur characteristics, leading to artifacts like overshoot, undershoot, and ringing.
A learning method using a multi-layer neural network is employed to train a first machine learning model with intermediate correction data, optimizing weights to minimize errors between output and correct images, thereby correcting for deviations in optical characteristics.
The method effectively reduces overcorrection and undercorrection, producing a corrected image with improved blur correction by aligning the intermediate-corrected image closer to the original, minimizing adverse effects such as ringing.
Smart Images

Figure 0007738984000001 
Figure 0007738984000002 
Figure 0007738984000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a learning method and an image processing method for acquiring a corrected image from intermediate corrected data in which blur caused by the optical characteristics of a captured image has been corrected. [Background technology]
[0002] Patent Document 1 discloses a method for obtaining a high-resolution image by correcting blurring caused by aberration and diffraction from a captured image using a neural network. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2018 / 037521 Summary of the Invention [Problem to be solved by the invention]
[0004] In image processing for sharpening images with blurring due to aberrations, diffraction, or other factors caused by the optical characteristics of an imaging device, high-precision aberration correction can be achieved by correcting based on known optical characteristics. However, when the optical characteristics used for aberration correction differ from the blur characteristics in the actual captured image, overcorrection or undercorrection can occur. Overcorrection, in particular, can cause problems. Here, the term "problems" refers to structures that appear in the corrected image but are not present in the actual subject, such as overshoot, undershoot, and ringing. While aberration correction processing using deep learning can achieve higher-precision correction than conventional image processing using linear filters, overcorrection and undercorrection are more likely to be noticeable. The method disclosed in Patent Document 1 performs aberration correction using shooting conditions, which can result in overcorrection or undercorrection if the blur in the actual captured image differs from the expected performance.
[0005] Therefore, the present invention aims to provide a learning method, an image processing method, etc. that can obtain a corrected image with better blur correction based on intermediate correction data that corrects blur caused by optical characteristics. [Means for solving the problem]
[0006] A learning method according to one aspect of the present invention includes: First, based on optical properties Blur To the original image a step of obtaining a blurred image by adding a blurred image to the original image; Second The method includes a step of acquiring intermediate correction data obtained by executing a sharpening process based on optical characteristics, a step of acquiring a correct image based on the original image, and a step of training a first machine learning model using the intermediate correction data and the correct image corresponding to the blurred image, The first optical characteristic and the second optical characteristic are expressed by a point spread function or an optical transfer function, and are different from each other. .
[0007] Other objects and features of the present invention are illustrated in the following examples. [Effects of the Invention]
[0008] According to the present invention, it is possible to provide a learning method, an image processing method, etc., which can obtain a corrected image with better blur correction based on intermediate correction data that corrects blur caused by optical characteristics. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 2 is a diagram showing a flow of learning of a neural network in the first embodiment. [Figure 2] 1 is a block diagram of an image processing system according to a first embodiment. [Figure 3] 1 is an external view of an image processing system according to a first embodiment. [Figure 4] 10 is a flowchart relating to weight learning in the first embodiment. [Figure 5] 4 is a flowchart relating to generation of a corrected image in the first embodiment. [Figure 6] FIG. 10 is a block diagram of an image processing system according to a second embodiment. [Figure 7] FIG. 10 is an external view of an image processing system according to a second embodiment. [Figure 8] 10 is a flowchart relating to generation of a corrected image in the second embodiment. [Figure 9] FIG. 10 is a block diagram of an image processing system according to a third embodiment. [Figure 10] 11 is a flowchart relating to generation of a corrected image in the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the drawings, the same components are designated by the same reference numerals, and redundant explanations will be omitted.
[0011] First, before describing the specific embodiments, the gist of the present invention will be explained. The present invention relates to image processing that improves overcorrection or undercorrection that occurs in sharpening processing that corrects blur caused by the optical characteristics of an imaging device. Here, optical characteristics include aberration, diffraction, low-pass filter, pixel aperture effect, etc. In particular, the present invention relates to improving overcorrection or undercorrection that occurs when the characteristics of blur in an actual captured image differ from the assumed optical characteristics in sharpening processing based on known optical characteristics. Here, blur refers to a point spread function (PSF) or optical transfer function (OTF).
[0012] In sharpening processing, optical characteristics can be determined based on the shooting conditions (zoom, aperture, and focal length) of the optical system during image capture and the screen position of the captured image. When design values are used for optical characteristics, the optical characteristics may be better or worse than expected due to manufacturing errors in the imaging device. The optical characteristics also change depending on whether the subject distance differs from the expected value. If the subject is assumed to exist on the in-focus plane when determining the optical characteristics used in sharpening processing, subjects on out-of-focus planes (defocused) in scenes with depth will be blurred with different optical characteristics. On the other hand, even when the optical characteristics on out-of-focus planes are taken into account, the optical characteristics may differ from the expected value due to differences in subject distance. Even for subjects on out-of-focus planes, the optical characteristics may be better or worse than expected depending on the field curvature of the imaging optical system. Thus, when the optical characteristics of a captured image differ from those expected based on the shooting conditions and screen position, sharpening processing based on incorrect optical characteristics can result in overcorrection or undercorrection.
[0013] In the present invention, a multi-layer neural network (first machine learning model) is used to correct overcorrection or undercorrection when optical characteristics differ from those expected. By training the first machine learning model using training data having the characteristics described below, it is possible to obtain a corrected image in which blur is more effectively corrected based on intermediate corrected data obtained by sharpening processing.
[0014] When training the weights (filters, bias, etc.) used in a multilayer neural network, the output (intermediate correction data) obtained during sharpening processing when optical performance deviates from expectations is used as training data. An image with appropriately corrected aberrations is used as the correct image. The training data is input into the neural network, and the weights are optimized to minimize the error between the output and the correct image. At this time, the neural network is trained to improve (correct) over-correction and under-correction. Optical characteristics vary in various ways due to manufacturing errors and deviations in subject distance. Therefore, by generating training data assuming multiple different optical characteristics, it is possible to train a neural network that can improve over-correction and under-correction in all cases. Hereinafter, deviations in gloss characteristics from expectations are referred to as performance deviations, and the correction process that improves over-correction and under-correction caused by performance deviations is referred to as performance deviation correction. [Example]
[0015] First, an image processing system according to a first embodiment of the present invention will be described with reference to Figs. 2 and 3. In this embodiment, a multi-layer neural network is made to learn and execute overcorrection and undercorrection. Fig. 2 is a block diagram of an image processing system 100 according to this embodiment. Fig. 3 is an external view of the image processing system 100.
[0016] The image processing system 100 includes a learning device (image processing device) 101, an imaging device 102, an image estimation device (image processing device) 103, a display device 104, a recording medium 105, an output device 106, and a network 107. The learning device 101 includes a memory unit (storage means) 101a, an acquisition unit (acquisition means) 101b, a generation unit (generation means) 101c, and an update unit (learning means) 101d.
[0017] The imaging device 102 has an optical system (imaging optical system) 102a and an imaging element 102b. The optical system 102a collects light incident on the imaging device 102 from the subject space. The imaging element 102b receives (photoelectrically converts) an optical image (subject image) formed via the optical system 102a to obtain a captured image. The imaging element 102b is, for example, a CCD (Charge Coupled Device) sensor or a CMOS (Complementary Metal-Oxide Semiconductor) sensor. The captured image obtained by the imaging device 102 contains blur due to aberration and diffraction of the optical system 102a and noise due to the imaging element 102b.
[0018] The image estimation device 103 includes a storage unit 103a, an acquisition unit 103b, and a correction unit 103c. The image estimation device 103 acquires a captured image, performs sharpening processing, and then performs performance deviation correction to generate an estimated image (corrected image) using a first machine learning model, which will be described later. The sharpening processing uses a multi-layer neural network (second machine learning model), and weight information is read from the storage unit 103a. The shooting conditions and the screen position of the captured image (image height, azimuth) are input to the second machine learning model, and an intermediately corrected image (intermediately corrected data) is obtained as an output. Here, the sharpening processing is performed by the image estimation device 103, but an intermediately corrected image sharpened by a different image processing device may also be obtained. The performance deviation correction uses a multi-layer neural network (first machine learning model), and weight information is read from the storage unit 103a.
[0019] The weights (weight information) of the first machine learning model and the second machine learning model are learned by the learning device 101. The image estimation device 103 reads the weight information from the storage unit 101a via the network 107 in advance and stores it in the storage unit 103a. The stored weight information may be the weight's numerical value itself or in an encoded format. Weight learning and sharpening processing using the weight are well known, and details of the performance deviation correction processing will be described later. The image estimation device 103 performs performance deviation correction on the intermediate corrected image to generate a corrected image. Note that in this embodiment, an image is used as intermediate corrected data, so it is called an intermediate corrected image, but a feature map, described later, may also be used as the intermediate corrected data.
[0020] The corrected image is output to at least one of a display device 104, a recording medium 105, and an output device 106. The display device 104 is, for example, a liquid crystal display or a projector. A user can perform editing work while checking the image in the middle of processing via the display device 104. The recording medium 105 is, for example, a semiconductor memory, a hard disk, a server on a network, etc. The output device 106 is, for example, a printer. The image estimation device 103 has a function of performing development processing and other image processing as necessary.
[0021] Next, a weight (weight information) learning method (method for manufacturing a trained model) executed by the learning device 101 in this embodiment will be described with reference to Figures 1 and 4. Figure 1 is a diagram showing the flow of learning weights of a neural network (first machine learning model). Figure 4 is a flowchart related to weight learning. Each step in Figure 4 is mainly executed by the acquisition unit 101b, generation unit 101c, or update unit 101d of the learning device 101.
[0022] The learning method of the first machine learning model will be described. First, in step S101 of FIG. 4, the acquisition unit 101b acquires an original image (object image). In this embodiment, the original image is a high-resolution (high-quality) image with little blurring due to aberrations and diffraction of the optical system 102a. Multiple original images are acquired, and are images of various objects, i.e., images having edges of various strengths and directions, textures, gradations, flat areas, etc. The original image may be a real-life image or an image generated by CG (Computer Graphics). In particular, when a real-life image is used as the original image, blurring has already occurred due to aberrations and diffraction, so by reducing the image, the effect of blurring can be reduced and a high-resolution (high-quality) image can be obtained. Note that if the original image contains sufficient high-frequency components, reduction may not be necessary.
[0023] Preferably, the original image has a signal value higher than the brightness saturation value of the image sensor 102b. This is because, even in actual subjects, there are subjects that do not fall within the brightness saturation value when photographed by the image pickup device 102 under specific exposure conditions. When using a real-life image as the original image, it can be obtained by HDR photography or by photographing with an image pickup device having a higher dynamic range than the image pickup device 102. When using an image photographed with an image pickup device having the same dynamic range as the image pickup device 102 as the original image, it is possible to obtain a higher signal value by proportionally multiplying the signal value. However, it is preferable to do this within a range where the reduction in gradation due to proportional multiplication does not affect the learning results. Furthermore, the original image may contain noise components. In this case, the noise contained in the original image is considered to be part of the subject, so the noise in the original image is not a particular problem.
[0024] Next, in step S102, the acquisition unit 101b acquires blur to be used for performing an imaging simulation, which will be described later. First, the acquisition unit 101b acquires imaging conditions corresponding to the lens state (zoom, aperture, and focal length states) of the optical system 102a. Then, the acquisition unit 101b acquires blur determined by the imaging conditions and the screen position of the captured image. Here, the blur refers to the PSF (point spread function) or OTF (optical transfer function) of the optical system 102a. The blur can be acquired by optical simulation or measurement of the optical system 102a. Note that blur due to lens state, image height, azimuth aberration, and diffraction, which differ for each original image, is acquired. This makes it possible to perform imaging simulations corresponding to multiple imaging conditions, image heights, and azimuth. If necessary, the blur to be applied may include a component of a filter, such as an optical low-pass filter, included in the imaging device 102 (a filter may be applied to the optical characteristics).
[0025] Next, in step S103, the generation unit 101c generates a correct patch (correct image) and a training patch (training data). A plurality of correct patches and a plurality of training patches are generated, and one or more patches are generated corresponding to one original image. In this embodiment, the correct patch and the training patch are images of the same subject. In this embodiment, a collection of a plurality of combinations of correct patches and training patches is used as training data. The training data is used to train a machine learning model that corrects the effects of deviations in optical characteristics (performance deviations) that occur when aberrations and diffractions of the optical system 102a are sharpened based on the optical characteristics. Therefore, the correct patch is an image that suffers less from overcorrection or has improved undercorrection compared to the training patch. However, as will be described later, the machine learning model may not perform correction depending on the conditions, so the training data may include cases where the correct patch and the training patch are the same image.
[0026] Note that a patch refers to an image having a predetermined number of pixels (for example, 64 × 64 pixels). The number of pixels of the ground truth patch and the training patch do not necessarily have to be the same. In this embodiment, mini-batch learning is used to learn the weights of the multi-layer neural network. Therefore, in step S103, multiple pairs of ground truth patches and training patches are generated. However, this embodiment is not limited to this, and online learning or batch learning may also be used.
[0027] In this embodiment, the correct patch and the training patch are obtained by the following method, but the present invention is not limited to this. In this embodiment, an imaging simulation is performed using multiple original images stored in the storage unit 101a as subjects, thereby generating multiple blurred images that incorporate the effects of blurring due to aberrations and diffraction. Then, each blurred image is sharpened using a second machine learning model based on the optical characteristics to generate intermediate corrected images. In this embodiment, the intermediate corrected images are used as training data. Furthermore, a correct image is obtained based on the original image corresponding to each training data. At this time, the correct image and the training data have the same or larger number of pixels as the number of pixels in the patch used as learning data. Then, multiple correct patches and training patches are obtained by extracting partial regions of a specified pixel size at the same position from multiple pairs of the correct image and the training data. In this embodiment, the original image is an undeveloped RAW image, and the correct patch and the training patch are also RAW images. However, this embodiment is not limited to this. A developed image or a feature map converted from an image, as described below, may also be used. Furthermore, the position of a partial region refers to the center of the partial region.
[0028] The learning data used for training the first machine learning model is generated by the following method. The generation unit 101c performs the processing described below for predetermined conditions (such as shooting conditions and screen position). First, the generation unit 101c generates a blur with a smaller amount of blur based on the blur caused by aberration and diffraction of the optical system 102a, i.e., the blur acquired in step S102. The generation unit 101c also imparts (convolves) the generated blur to the original image. Thereafter, the generation unit 101c acquires a blurred image by extracting a partial region from the image clipped at the saturation brightness value of the image sensor 102b.
[0029] Next, the generation unit 101c corrects the blurred image using a second machine learning model based on the blur (optical characteristics) acquired in step S102 under predetermined conditions to acquire an intermediately corrected image. By extracting the same partial area from the intermediately corrected image and the original image, pairs of training patches and correct patches can be generated. The generation unit 101c uses multiple blur amounts to generate blurred images to which multiple different blur sizes have been added, and generates multiple pairs of training patches and correct patches for each blur size. That is, the generation unit 101c generates pairs of training patches and correct patches using blurred images to which multiple different blur sizes have been added, based on the blur acquired in step S102. In this embodiment, all of the multiple blur sizes are smaller than the blur acquired in step S102, i.e., the optical characteristics under predetermined conditions, but this is not limited to this. By performing the above-mentioned process under different conditions, blurred images corresponding to multiple performance deviations can be generated for multiple conditions of the imaging device (such as shooting conditions and screen position), and learning data having training patches and correct patches based on the blurred images can be generated.
[0030] Here, the transformation of the blur will be described. In this embodiment, the blur to be added to the blurred image is obtained by reducing the spatial extent of the PSF obtained in step S102. The reduction ratio of the PSF is randomly set between 0.1 and 0.9, and the PSF is reduced at the set reduction ratio. The reduction ratio may be different in the meridian direction and the sagittal direction. The reduction process is performed using downsampling and interpolation processing such as bicubic interpolation. As a result, the blur added to the blurred image is small, based on the blur (optical characteristics) obtained in step S102. By performing a sharpening process using the second machine learning model on the blurred image generated in this way, the intermediately corrected image will be overcorrected, and adverse effects will occur.
[0031] Since neural networks are trained to bring training patches closer to the correct patches, they are also trained to bring intermediate-corrected images closer to the original image. In other words, the first machine learning model of this embodiment suppresses overcorrection to an appropriate correction effect and also suppresses adverse effects (reducing the effects of overcorrection). Here, the appropriate correction effect is a correction effect that brings the intermediate-corrected image closer to the original image or the intermediate-corrected image without performance deviation. Even if there are no adverse effects such as ringing, excessive sharpening may result in an unnatural image. Therefore, it is advisable to suppress the correction effect to an appropriate level in order to align the correction effect with other image areas without performance deviation or when there is no performance deviation.
[0032] The blurring applied to the blurred image is not limited to the blurring described above. A blurring filter such as Gaussian blurring may also be used. Furthermore, instead of downsampling the PSF acquired in step S102, the image may be reduced by convolution with a sharpening filter or by existing reduction processing. When the blurring is acquired as the OTF in step S102, the blurring can be reduced by performing an enlargement process such as upsampling, a process for improving the OTF by applying a frequency gain, or an expansion process. The magnitude of the blurring amount in this embodiment can be expressed by a predetermined index. For example, the amount of blurring can be said to be small when the maximum value of the PSF is large, the half-width of the PSF in the mid-section is small, the MTF for a specified spatial frequency is large, or the integral value of the MTF with respect to the spatial frequency is large. Furthermore, the amount of blurring can be increased by convolving a blurring filter (low-pass filter) with the PSF or applying the frequency characteristics of a low-pass filter to the OTF. The magnitude of blur based on a predetermined blur (the magnitude of blur relative to the predetermined blur) may be expressed as a difference or ratio of the above-mentioned index, or the above-mentioned index may be used for a difference or ratio of OTF.
[0033] The sharpening process based on optical characteristics (second machine learning model) can be learned by using the blurred image obtained in step S102 as the training patch and the correct patch as the original image. When reducing a real image to create the original image, the order of reduction and blurring can be reversed. When blurring is performed first, the blur sampling rate must be finer to take the reduction into account. In the case of a PSF (point spread function), the spatial sampling points should be finer, and in the case of an OTF (optical transfer function), the maximum frequency should be increased.
[0034] It is preferable that the blur to be applied does not include distortion. This is because if distortion is large in the sharpening process (aberration correction process), the position of the subject will change, and the subject may differ between the target patch and the training patch. For this reason, the sharpening process (second machine learning model) used in this embodiment does not correct distortion. Accordingly, the performance deviation correction by the first machine learning model also does not correct distortion. Distortion is corrected separately after blur correction using bilinear interpolation, bicubic interpolation, or the like.
[0035] Note that, since noise occurs in the image sensor 102b during actual image capture, it is preferable to also add noise to the training data. Random numbers corresponding to the noise characteristics of the image sensor 102b can be generated and added, and shooting conditions such as ISO sensitivity can also be taken into consideration. If noise is added only to the training images, or if noise that is uncorrelated between the training images and the reference image is added, denoising is also learned at the same time as blur correction during learning. If the same noise is added to the training images and the reference image, blur correction that suppresses changes in noise is learned. Note that, if a feature map, described later, is used as training data here, noise is added to the blurred images instead of the training images.
[0036] Next, in step S104, the generation unit 101c inputs the training patches (training data) 212 in Fig. 1 into a multi-layer neural network to generate estimated patches (estimated images) 213. For mini-batch learning, estimated patches 213 corresponding to the multiple training patches 212 are generated. Fig. 1 shows the flow from step S104 to step S105.
[0037] The estimated patch 213 is the training patch 212 with the performance deviation corrected, and ideally, it matches the ground truth patch (ground truth image) 211. Note that in this embodiment, the neural network configuration shown in Figure 1 is used, but the present invention is not limited to this. In Figure 1, CN represents a convolutional layer, and DC represents a deconvolutional layer. In both CN and DC, the convolution of the input with the filter and the sum of the bias are calculated, and the result is nonlinearly transformed using an activation function. The initial values of each filter component and bias are arbitrary and are determined by random numbers in this embodiment. The activation function can be, for example, a rectified linear unit (ReLU) or a sigmoid function. The output of each layer except the final layer is called a feature map. Skip connections 222 and 223 combine feature maps output from discontinuous layers. Feature maps can be combined by element-by-element summation or by channel-wise concatenation. In this embodiment, element-by-element summation is used. The skip connection 221 takes the sum of the residual estimated from the training patch 212 and the correct patch 211 and the training patch 212 to generate an estimated patch 213. An estimated patch 213 is generated for each of the multiple training patches 212.
[0038] Next, in step S105, the update unit 101d updates the weights (weight information) of the neural network based on the error between the estimated patch 213 and the correct patch (correct image) 211. Here, the weights include the filter components and biases of each layer. Backpropagation is used to update the weights, but this embodiment is not limited to this. For mini-batch learning, the errors between multiple correct patches 211 and their corresponding estimated patches 213 are calculated, and the weights are updated. For example, the L2 norm or the L1 norm may be used as the loss function.
[0039] Next, in step S106, the update unit 101d determines whether weight learning is complete. Completion can be determined by, for example, whether the number of iterations of learning (weight update) has reached a specified value, or whether the amount of change in weight during update is smaller than a specified value. If it is determined that learning is not complete, the process returns to step S104, and multiple new supervised patches and training patches are acquired. On the other hand, if it is determined that learning is complete, the learning device 101 (update unit 101d) ends learning and saves weight information in the storage unit 101a.
[0040] Next, the generation of a corrected image (correction processing, estimation method) executed by the image estimation device 103 in this embodiment will be described with reference to Fig. 5. Fig. 5 is a flowchart related to the generation of a corrected image. Each step in Fig. 5 is mainly executed by the acquisition unit 103b or the correction unit 103c of the image estimation device 103.
[0041] First, in step S201, the acquisition unit 103b acquires information about a captured image and weights. The captured image is an undeveloped RAW image, similar to that used in learning, and in this embodiment, it is transmitted from the imaging device 102. The weight information is the weights of the first machine learning model and the second machine learning model transmitted from the learning device 101 and stored in the storage unit 103a.
[0042] Subsequently, in step S202, the acquisition unit 103b executes sharpening processing based on the weights of the second machine learning model acquired in step S201, and acquires an intermediate corrected image.
[0043] Next, in step S203, the correction unit 103c inputs the intermediate corrected image acquired in step S202 as an input image for the acquired first machine learning model, and generates a corrected image. The corrected image is an image obtained by correcting the performance deviation of the intermediate corrected image. In other words, the amount of correction is reduced to an appropriate level for portions that have been corrected more than expected due to overcorrection, and adverse effects are also suppressed (the effects of overcorrection are reduced). Note that when inputting a captured image to a neural network, it is not necessary to crop it to the same size as the training patch used during learning.
[0044] In this embodiment, the sharpening process is performed using a machine learning model. However, the sharpening process may also be performed using other existing methods based on optical characteristics, such as a method of convolving a sharpening filter such as a Wiener filter. Similarly, the first machine learning model can be trained by using intermediate corrected images obtained by applying other sharpening methods to blurred images as training data. On the other hand, performing the sharpening process using a machine learning model using a neural network such as CNN is preferable because it allows for more accurate correction of image quality degradation due to optical characteristics. When using a neural network, the high correction effect can cause significant adverse effects if deviations from the expected optical characteristics occur. However, by performing processing using the first machine learning model of this embodiment, a corrected image with high correction effect and reduced adverse effects can be obtained.
[0045] Furthermore, it is preferable to use a machine learning model (a second machine learning model common to multiple optical characteristics) that has been trained for the sharpening process. When using other sharpening methods or when a machine learning model for the sharpening process is trained for each optical characteristic, differences in correction effect are likely to occur for each optical characteristic, and differences in overcorrection and the resulting adverse effects are likely to occur. On the other hand, by using a machine learning model trained in common, the correction effect and the adverse effects are common, and the first machine learning model can be trained with high accuracy. As a result, performance discrepancies can be corrected more effectively.
[0046] Furthermore, it is preferable that the machine learning model for sharpening processing includes not only the captured image but also at least one of information on the shooting conditions and screen position as input data (information on the shooting conditions and screen position is input and used). In this case, the information must be input both during learning and during correction processing. This allows for more accurate correction of blurring due to aberration, diffraction, etc., and by performing processing using the first machine learning model of the present invention in combination, it is possible to obtain a corrected image with high correction effectiveness and reduced adverse effects.
[0047] Furthermore, it is preferable to add information on the shooting conditions and screen position to the input data for the first machine learning model, in addition to the intermediate correction data. In this case, information on the shooting conditions must be input both during learning and correction processing. For example, in a low-performance area where field curvature occurs at a high image height, overcorrection due to a three-dimensional subject or partial blur is likely to occur. The likelihood and extent of overcorrection and other adverse effects vary depending on the shooting conditions and optical characteristics. By inputting the shooting conditions into the first machine learning model, performance deviations can be corrected based on the shooting conditions and optical characteristics, and a good corrected image can be obtained. In particular, if information on the shooting conditions is also input into the machine learning model for sharpening processing, it can be learned using shooting condition information common to the sharpening processing and performance deviation correction, and a better corrected image can be obtained.
[0048] It is also preferable to add manufacturing error information about the imaging device as input data to the first machine learning model. By adding manufacturing error information obtained by measuring after manufacturing, it is possible to correct performance deviations based on the impact of manufacturing errors on optical characteristics. For example, if it is known that the imaging optical system has partial blur, it is possible to determine the image positions that are prone to overcorrection and undercorrection.
[0049] In this embodiment, the intermediate-corrected data and training data are intermediate-corrected images. However, when sharpening is performed using a machine learning model, a feature map of the intermediate layer may be used as the intermediate-corrected data and training data. Using an image as the intermediate-corrected data has the advantage that it is easy to check for defects on the image during learning, and that a machine learning model can be created using the difference from the captured image. On the other hand, if the feature map has a large number of channels, converting it to an image reduces the amount of information.
[0050] Therefore, by using a feature map as the intermediate correction data and training data input to the first machine learning model, it is possible to correct the performance deviation without reducing the amount of information. Even in this case, the intermediate correction image can be estimated when training the second machine learning model. For example, it is possible to train using the error between the intermediate correction image estimated from the captured image and the original image. During correction processing using the captured image, the intermediate correction data is calculated using weights up to the intermediate layer of the second machine learning model, and input to the first machine learning model to obtain a sharpened corrected image with the performance deviation corrected. Therefore, the weight information of the second machine learning model acquired in step S201 only needs to include weights for the layers before the intermediate layer used as the intermediate correction data. Note that the number of pixels and the number of channels of the training data may differ from those of the target image. Furthermore, blurred images may be used as training data, and the machine learning model for sharpening processing and the machine learning model for performance deviation correction may be connected during training. In this case, the machine learning model for sharpening processing may be trained by updating only the weights for the layer corresponding to the model that performs performance deviation correction, without updating the weights previously trained.
[0051] Furthermore, when sharpening processing is performed using a machine learning model, there is no need to output intermediate corrected data or an intermediate corrected image once during correction processing from a captured image. In other words, the sharpening processing and performance deviation correction may be performed together as a single model.
[0052] In addition, the correct image when training the first machine learning model is the original image, and the correct patch is extracted from the original image, but this embodiment is not limited to this. For example, an image obtained by performing sharpening processing based on the optical characteristics on a blurred image to which the same blur as the optical characteristics has been imparted may be used as the correct image. In addition, any image that is not affected by performance deviation and has little blur due to the optical characteristics can be used as the correct image.
[0053] Furthermore, the correct image may be used as the original image only if the intermediate-corrected data and training data used in training the first machine learning model exhibit adverse effects such as undershooting or ringing. In other words, if overcorrection does not result in unnatural-looking adverse effects such as undershooting or ringing, the correct image may be the intermediate-corrected data and training data themselves. In this case, the first machine learning model is trained to suppress only the adverse effects that appear unnatural, thereby obtaining an output image that is overcorrected but still has a sharp appearance. Furthermore, if an attempt is made to suppress the correction effect of intermediate-corrected data or training data that is overcorrected but does not exhibit unnatural-looking adverse effects, a suppression effect that is difficult to directly determine from the image will be learned. Avoiding this makes learning easier. It is not necessary to distinguish whether adverse effects actually occur. For example, the correct image may be used as the original image only for high-contrast images or images containing edges, which are prone to adverse effects.
[0054] In this embodiment, the performance deviation caused by overcorrection is corrected by making the blurred image blurred less than the optical characteristics, but this embodiment is not limited to this. By generating training data by making the blur applied to the blurred image blurred more than the optical characteristics, it is possible to learn a first machine learning model that corrects the performance deviation caused by insufficient correction. In this case, as in the case of correcting the performance deviation by overcorrection, the original image may be used as the ground truth image, or an image obtained by performing a sharpening process based on the optical characteristics on a blurred image that has been given the same blur as the optical characteristics may be used.
[0055] Furthermore, training data may be generated by adding blur to a blurred image that includes both blur smaller than the optical characteristics and blur larger than the optical characteristics. The target image may be an original image, or an image obtained by performing a sharpening process based on the optical characteristics on a blurred image that has been added with the same blur as the optical characteristics. This allows training of a first machine learning model that corrects both performance deviations due to overcorrection and performance deviations due to undercorrection. If only overcorrection correction were learned, the correction effect in undercorrected image areas may be further suppressed. However, by learning both, the first machine learning model can learn both overcorrection and undercorrection, and therefore can correct both overcorrection and undercorrection.
[0056] Furthermore, when training data is generated by adding blur to blurred images that include both blur smaller and blur larger than the optical characteristics, only one of the performance discrepancies, either overcorrection or undercorrection, may be corrected. A machine learning model that performs such correction can be trained by changing the corrected image without correction to an image obtained by performing sharpening on the corresponding blurred image. Overcorrection is a process that reduces frequency characteristics, while undercorrection is a process that improves frequency characteristics. Because the two functions are contradictory, limiting the correction of performance discrepancies to only one can result in a lighter model and improved correction effectiveness of performance discrepancies. For example, as a correct image corresponding to training data in which the blur added to the blurred image is smaller than the optical characteristics, an image obtained by adding blur to the original image based on the optical characteristics and then performing sharpening processing based on the optical characteristics, or the original image, is used. Furthermore, as a correct image corresponding to training data in which the blur added to the blurred image is larger than the optical characteristics, a blurred image corresponding to the training data or intermediate corrected data (intermediate corrected image) corresponding to the training data is used. This eliminates the adverse effects of overcorrection (reduces the impact of overcorrection), but does not correct undercorrection due to performance discrepancies. As in the case described above, compared to training with training data from only one case, by training including both, the first machine learning model can distinguish between cases where no correction is made and cases where correction is made.
[0057] Furthermore, training data may be generated including cases where the blur applied to the blurred image is the same (or substantially the same) as the optical characteristics. The correct image may be the original image, or an image obtained by performing a sharpening process based on the optical characteristics on a blurred image to which the blur is applied identically to the optical characteristics. In other words, the original image or intermediate corrected data (intermediate corrected image) corresponding to the training data may be used as the correct image corresponding to the training data in which the blur applied to the blurred image is the same as the optical characteristics. This allows the first machine learning model to learn both correction when there is a performance discrepancy and no correction when there is no performance discrepancy.
[0058] Furthermore, when performing correction using the first machine learning model, not only the intermediate correction data but also the captured image may be added to the input data. A blurred image may be added during learning. In this case, by knowing the image before sharpening processing, performance deviations can be corrected with higher accuracy. For example, if overcorrection causes adverse effects such as undershoot or ringing, the captured image can be used as information to determine whether the structure is actually present or whether it is a structure caused by the sharpening processing. Furthermore, by using the difference from the captured image, it is possible to correct performance deviations by explicitly using the changes caused by the sharpening processing.
[0059] Furthermore, the blur applied to the original image when acquiring the training data may be based on at least one of the imaging conditions of the optical system and the screen position. For example, when the optical characteristics are highly sensitive to manufacturing tolerances or subject distance, the performance deviation from the optical characteristics may be large. Therefore, by significantly changing the blur applied to the blurred image in such a case, a first machine learning model that can cope with the larger deviation in performance can be trained. To significantly change the blur, for example, the upper limit of the magnification rate may be set to a large value and the lower limit of the reduction rate may be set to a small value when multiple magnification rates and reduction rates are used when generating the blur to be applied from the PSF.
[0060] In this embodiment, the learning device 101 and the image estimation device 103 are separate devices, but the present embodiment is not limited to this. The learning device 101 and the image estimation device 103 may be integrated. That is, learning (the process shown in FIG. 4) and estimation (the process shown in FIG. 5) may be performed within an integrated device.
[0061] Furthermore, the image estimation device that performs the sharpening process and the performance deviation correction process may be separate devices, and instead of acquiring a captured image, an image that has undergone the sharpening process may be acquired. In this case, the execution of the sharpening process is skipped.
[0062] As described above, the image processing device (learning device 101) of this embodiment performs learning of a machine learning model for outputting a corrected image from an intermediate corrected image in which blur caused by the optical characteristics of the optical system has been corrected. The image processing device includes a generation unit 101c having functions as a first acquisition unit, a second acquisition unit, and a third acquisition unit, and an update unit 101d serving as learning unit. The first acquisition unit acquires a blurred image by adding blur to an original image. The second acquisition unit acquires, as a plurality of training data, intermediate corrected data obtained by performing a sharpening process based on the optical characteristics on the blurred image. The third acquisition unit acquires a plurality of ground truth images based on the original image, corresponding to each training data. The learning unit trains the first machine learning model using learning data consisting of the training data and the ground truth images. Furthermore, the blur added to the blurred image differs for each training data set based on the optical characteristics.
[0063] According to this embodiment, it is possible to provide an image processing method or the like that can obtain a corrected image with better blur correction based on intermediate correction data that corrects blur caused by optical characteristics. [Example]
[0064] Next, an image processing system according to a second embodiment of the present invention will be described with reference to Fig. 6 and Fig. 7. In this embodiment, a corrected image is generated by an image estimation unit 323 in an imaging device. In this embodiment, the flow of generating the corrected image is different from that of the first embodiment.
[0065] FIG. 6 is a block diagram of an image processing system 300 according to this embodiment. FIG. 7 is an external view of the image processing system 300. The image processing system 300 includes a learning device (image processing device) 301 and an imaging device 302 connected via a network 303. The learning device 301 includes a memory unit (storage means) 311, an acquisition unit (acquisition means) 312, a generation unit (generation means) 313, and an update unit (learning means) 314, and learns weights (weight information) for blur correction using a neural network. The imaging device 302 captures an image of the subject space, sharpens the captured image using the read weight information, and corrects performance deviations to generate a corrected image. The imaging device 302 includes an optical system 321 and an image sensor 322. The image estimation unit 323 includes an acquisition unit 323a and a correction unit 323b, and corrects performance deviations in the captured image using weight information stored in the memory unit 324.
[0066] The weight information is learned in advance by the learning device 301 and stored in the memory unit 311. The imaging device 302 reads the weight information from the memory unit 311 via the network 303 and stores it in the memory unit 324. The captured image (corrected image) in which the performance deviation has been corrected is stored in the recording medium 325. When a user issues an instruction to display the corrected image, the saved corrected image is read out and displayed on the display unit 326. Note that a captured image already saved in the recording medium 325 may be read out, and the performance deviation may be corrected by the image estimation unit 323. The above series of controls are performed by the system controller 327. Note that the learning of the first model executed by the learning device 301 of this embodiment is the same as in the first embodiment.
[0067] Next, generation of a corrected image executed by the image estimation unit 323 in this embodiment will be described with reference to Fig. 8. Fig. 8 is a flowchart related to generation of a corrected image. Each step in Fig. 8 is mainly executed by the acquisition unit 323a or the correction unit 323b of the image estimation unit 323. Steps S401 to S403 are similar to steps S201 to S203 executed by the acquisition unit 103b or the correction unit 103c in the first embodiment, respectively.
[0068] Next, in step S404, the correction unit 323b performs a weighted average of the corrected image that has undergone performance error correction and the intermediate corrected image (intermediate corrected data) that has undergone sharpening processing to generate a composite image. This allows the strength of the effect of performance error correction to be adjusted. The weighted average ratio may vary depending on the region of the image. In this embodiment, the ratio of the corrected image is reduced when the amount of change in luminance due to sharpening processing, i.e., the difference between the captured image and the intermediate corrected image, is smaller than a predetermined value. In other words, the effect of performance error correction is reduced. This is because when performing performance error correction using the first machine learning model, a small amount of change in luminance can reduce the accuracy of determining whether or not there is a performance error. In textured areas where the subject has a fine structure, whether or not performance error correction is performed varies depending on slight changes in luminance or structure, making it prone to uneven correction, which is undesirable. In areas where the amount of change in luminance due to sharpening processing is small, the impact of performance error is originally small. Therefore, by reducing the change due to performance error correction in these areas, uneven correction can be avoided.
[0069] In this embodiment, a weighted average is calculated using the corrected image and the intermediate corrected image. However, a captured image may be used instead of the intermediate corrected image. For example, the strength of sharpening can be adjusted by increasing the ratio of the captured image in the entire image or in a portion of the image. The weighted average ratio may be changed by a user, thereby achieving a sharpening effect desired by the user. Furthermore, the composite image is not limited to being generated using two images. The strength of sharpening and performance deviation correction may be adjusted by performing a weighted average using the captured image, the intermediate corrected image, and the corrected image. As for the composition method, an existing method other than the weighted average may also be used.
[0070] As described above, in this embodiment, at least one of the captured image and the intermediate correction data is synthesized with a corrected image based on a sharpening component (for example, the amount of change in luminance due to the sharpening process) acquired based on the captured image and the intermediate correction data. As described above, in this embodiment, unlike in the first embodiment, the strength of sharpening and performance deviation correction can be adjusted by performing synthesis processing between multiple images. According to this embodiment, it is possible to provide an image processing method or the like that can acquire a corrected image in which blur has been more effectively corrected based on intermediate correction data in which blur caused by optical characteristics has been corrected. [Example]
[0071] Next, an image processing system according to a third embodiment of the present invention will be described with reference to Fig. 9. The image processing system of this embodiment differs from the first and second embodiments in that it includes a processing device (computer) that transmits a captured image to be processed to an image estimation device and receives a processed output image (corrected image) from the image estimation device.
[0072] FIG. 9 is a block diagram of an image processing system 600 in this embodiment. The image processing system 600 includes a learning device 601, an imaging device 602, an image estimation device 603, and a processing device (computer) 604. The learning device 601 and the image estimation device 603 are, for example, servers. The computer 604 is, for example, a user terminal (a personal computer or a smartphone). The processing device 604 is connected to the image estimation device 603 via a network 605. The image estimation device 603 is connected to the learning device 601 via a network 606. That is, the computer 604 and the image estimation device 603 are configured to be able to communicate with each other, and the image estimation device 603 and the learning device 601 are configured to be able to communicate with each other. Note that the configuration of the learning device 601 is similar to that of the learning device 101 in the first embodiment, and therefore a description thereof will be omitted. Note that the configuration of the imaging device 602 is similar to that of the imaging device 102 in the first embodiment, and therefore a description thereof will be omitted.
[0073] The image estimation device 603 has a storage unit 603a, an acquisition unit 603b, a correction unit 603c, and a communication unit (receiving means) 603d. The storage unit 603a, the acquisition unit 603b, and the correction unit 603c are similar to the storage unit 103a, the acquisition unit 103b, and the correction unit 103c of the image estimation device 103 in Example 1. The communication unit 603d has a function of receiving a request transmitted from the computer 604 and a function of transmitting an output image generated by the image estimation device 603 to the computer 604.
[0074] The computer 604 has a communication unit (transmission means) 604a, a display unit 604b, an image processing unit 604c, and a recording unit 604d. The communication unit 604a has a function of transmitting to the image estimation device 603 a request to cause the image estimation device 603 to execute processing on a captured image, and a function of receiving an output image processed by the image estimation device 603. The display unit 604b has a function of displaying various information. The information displayed by the display unit 604b includes, for example, a captured image to be transmitted to the image estimation device 603 and an output image received from the image estimation device 603. The image processing unit 604c has a function of further performing image processing on the output image received from the image estimation device 603. The recording unit 604d records the captured image acquired from the imaging device 602, the output image received from the image estimation device 603, etc.
[0075] Next, image processing (generation of a corrected image) by the image processing system 600 will be described with reference to Fig. 10. Fig. 10 is a flowchart relating to generation of a corrected image in this embodiment. The content of the image processing (correction processing) in this embodiment is the same as the correction processing described in embodiment 1 with reference to Fig. 5. The image processing shown in Fig. 10 is started when a command to start image processing is given by a user via the computer 604.
[0076] First, the operation of the computer 604 will be described. In step S701, the computer 604 transmits a request for processing a captured image to the image estimation device 603. Note that the method for transmitting the captured image to be processed to the image estimation device 603 is not important. For example, the captured image may be uploaded to the image estimation device 603 simultaneously with step S701, or may be uploaded to the image estimation device 603 before step S701. Furthermore, the captured image may be an image stored on a server different from the image estimation device 603. Note that in step S701, the computer 604 may transmit ID information for authenticating a user together with the request for processing the captured image.
[0077] Next, in step S702, the computer 604 receives the output image generated in the image estimation device 603. As in the first embodiment, the output image is an image in which the captured image has been sharpened and then the influence of performance deviation has been corrected.
[0078] Next, a description will be given of the operation of the image estimation device 603. In step S801, the image estimation device 603 receives a request for processing a captured image transmitted from the computer 604. At this time, the image estimation device 603 determines that processing (sharpening processing and performance deviation correction processing) for the captured image has been instructed, and executes the processing from step S802 onwards.
[0079] In step S802, the image estimation device 603 acquires weight information. The weight information is information (trained model) learned by the same method as described in the first embodiment with reference to FIG. 4. The image estimation device 603 may acquire the weight information from the learning device 601, or may acquire weight information previously acquired from the learning device 601 and stored in the storage unit 603a. The subsequent steps S803 and S804 are the same as steps S202 and S203, respectively, in the first embodiment. Subsequently, in step S805, the image estimation device 603 transmits an output image to the computer 604. Note that although this embodiment has been described as performing the correction processing of the first embodiment, the correction processing of the second embodiment may also be performed.
[0080] When the performance deviation correction process is performed within the image estimation device 603 as in this embodiment, the processing load due to the correction process can be borne within the image estimation device 603, thereby making it possible to reduce the processing capacity required of the computer 604. Furthermore, as in this embodiment, the image estimation device 603 may be configured to be controlled by the computer 604 connected to the image estimation device 603 so as to be able to communicate with it.
[0081] (Other Examples) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0082] According to each embodiment, a learning method, an image processing method, a learning device, and a program can be provided that can obtain a corrected image with better blur correction based on intermediate correction data that corrects blur caused by optical characteristics.
[0083] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. [Explanation of symbols]
[0084] 101 Learning device (image processing device) 101c generation unit (first acquisition means, second acquisition means, third acquisition means) 101d Update unit (learning means)
Claims
1. obtaining a blurred image by applying blur based on the first optical characteristic to the original image; acquiring intermediate correction data obtained by performing a sharpening process on the blurred image based on a second optical characteristic of an optical system corresponding to the original image; obtaining a ground truth image based on the original image; a step of learning a first machine learning model using the intermediate corrected data and the correct image corresponding to the blurred image, The learning method, wherein the first optical characteristic and the second optical characteristic are expressed by a point spread function or an optical transfer function and are different from each other.
2. The method of claim 1 , wherein the intermediate corrected data includes at least one of an intermediate corrected image or a feature map.
3. The learning method according to claim 1 or 2, wherein the sharpening process is performed by a second machine learning model different from the first machine learning model.
4. 4. The learning method according to claim 1, wherein the sharpening process is performed based on the conditions of the optical system.
5. In the step of acquiring the blurred image, a first blur is applied to the original image to acquire a first blurred image, and a second blur is applied to the original image to acquire a second blurred image; 5. The learning method according to claim 1, wherein the first blur and the second blur have different amounts of change from the blur caused by the second optical characteristic.
6. The learning method according to claim 5 , wherein at least one of the first blur and the second blur is smaller than the blur caused by the second optical characteristic.
7. The learning method according to claim 5 , wherein at least one of the first blur and the second blur is larger than the blur caused by the second optical characteristic.
8. 8. The learning method according to claim 5, wherein the correct image is an image obtained by performing the sharpening process on the blurred image or the original image.
9. The learning method described in any one of claims 5 to 7, characterized in that the correct image corresponding to the blurred image obtained by adding a blur to the original image that is greater than the blur caused by the second optical characteristic is an image obtained by performing the sharpening process on the blurred image.
10. The learning method according to any one of claims 1 to 9, characterized in that in the step of learning the first machine learning model, the intermediate corrected data and the blurred image are input to the first machine learning model.
11. A learning method according to any one of claims 1 to 10, characterized in that in the step of learning the first machine learning model, at least one of the conditions of the optical system or the screen position of the original image is input into the first machine learning model.
12. A program causing a computer to execute the learning method according to any one of claims 1 to 11.
13. acquiring the first machine learning model obtained by the learning method according to any one of claims 1 to 11 and input data; and generating an estimated image by inputting the input data into the first machine learning model.
14. A program causing a computer to execute the image processing method according to claim 13.
15. a first acquisition means for acquiring a blurred image by adding blur based on a first optical characteristic to an original image; a second acquisition means for acquiring intermediate correction data obtained by performing a sharpening process on the blurred image based on a second optical characteristic of an optical system corresponding to the original image; a third acquisition means for acquiring a correct image based on the original image; a learning means for learning a first machine learning model using the intermediate corrected data and the correct image corresponding to the blurred image, A learning device, wherein the first optical characteristic and the second optical characteristic are expressed by a point spread function or an optical transfer function and are different from each other.
16. obtaining a blurred image by applying blur based on the first optical characteristic to the original image; acquiring intermediate correction data obtained by performing a sharpening process on the blurred image based on a second optical characteristic of an optical system corresponding to the original image; obtaining a ground truth image based on the original image; a step of learning a first machine learning model using the intermediate corrected data and the correct image corresponding to the blurred image, A method for manufacturing a trained model, characterized in that the first optical characteristic and the second optical characteristic are expressed by a point spread function or an optical transfer function and are different from each other.
Citation Information
Patent Citations
Control system, control unit, image processing apparatus, and program
JP2019215635A
Method for processing image, image processor, imaging device, image processing system, program, and storage medium
JP2020061129A
Image processing method, image processing apparatus, image capture apparatus, image processing program, and storage medium
WO2018037521A1