Learning methods, image processing methods, and programs

JP7899400B2Active Publication Date: 2026-08-03CANON KK
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
CANON KK
Filing Date
2025-05-23
Publication Date
2026-08-03

AI Technical Summary

Benefits of technology

【0008】 本発明によれば、光学特性に起因するぼけを補正した中間補正データに基づいて、より良好にぼけが補正された補正画像を取得することが可能な学習方法および画像処理方法等を提供することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007899400000001
    Figure 0007899400000001
  • Figure 0007899400000002
    Figure 0007899400000002
  • Figure 0007899400000003
    Figure 0007899400000003
Patent Text Reader

Abstract

To provide a learning method which can acquire a correction image in which blur is excellently corrected on the basis of intermediate correction data in which blur due to optical characteristics is corrected.SOLUTION: A learning method of a machine learning model for outputting a correction image from intermediate correction data in which blur due to optical characteristics of an optical system is corrected comprises the steps of: acquiring a blur image by applying blur to an original image; acquiring the intermediate correction data obtained by executing sharpening processing based on the optical characteristics to the blur image as plural pieces of training data; acquiring a plurality of correct answer images on the basis of the original image in correspondence with each training data; and learning a first machine learning model by using the learning data consisting of the training data and correct answer image. The blur applied to the blur image is different for each training data with the optical characteristics as a reference.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learning method and an image processing method for obtaining a corrected image from intermediate correction data obtained by correcting blur due to optical characteristics of a captured image.

Background Art

[0002] Patent Document 1 discloses a method of using a neural network to correct blur due to aberration and diffraction from a captured image to obtain a high-resolution image.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In image processing for sharpening an image in which blur such as aberration and diffraction occurs due to the optical characteristics of an imaging device, high-precision aberration correction can be achieved by correcting based on known optical characteristics. On the other hand, when the optical characteristics used for aberration correction are different from the blur characteristics in the actual captured image, overcorrection or undercorrection may occur. In particular, in the case of overcorrection, adverse effects occur. Here, the adverse effects refer to structures that do not exist in the original subject that occur in the corrected image, such as overshoot, undershoot, ringing, etc. In the aberration correction process using deep learning, correction can be achieved with higher precision than image processing using a conventional linear filter, but overcorrection and undercorrection are also more likely to become prominent. Since the method disclosed in Patent Document 1 performs aberration correction using shooting conditions, overcorrection or undercorrection may occur when the blur in the actual captured image is different from the assumed performance.

[0005] Therefore, the present invention aims to provide a learning method and an image processing method, etc., that can acquire a corrected image with better blur correction based on intermediate correction data that corrects blur caused by optical characteristics. [Means for solving the problem]

[0006] One aspect of the present invention Image processing The method is, The image was captured using Optical properties of the optical system Based on The sharpening process described above imaging Obtained by applying to the image While Steps to obtain inter-correction data and ,before Note Intermediate correction data from The aforementioned Correction by a first machine learning model that reduces the effects of overcorrection during sharpening. image of The steps to acquire, and the aforementioned Captured image and the above The corrected image is combined with at least one of the intermediate correction data. The steps to do vinegar ru.

[0007] Other objects and features of the present invention are described in the following examples. [Effects of the Invention]

[0008] According to the present invention, it is possible to provide a learning method and an image processing method, etc., that can acquire a corrected image with better blur correction based on intermediate correction data that corrects blur caused by optical characteristics. [Brief explanation of the drawing]

[0009] [Figure 1] This diagram shows the learning process of the neural network in Example 1. [Figure 2] This is a block diagram of the image processing system in Example 1. [Figure 3] This is an external view of the image processing system in Example 1. [Figure 4] This is a flowchart illustrating the weight learning process in Example 1. [Figure 5]It is a flowchart regarding generation of a corrected image in Example 1. [Figure 6] It is a block diagram of an image processing system in Example 2. [Figure 7] It is an external view of an image processing system in Example 2. [Figure 8] It is a flowchart regarding generation of a corrected image in Example 2. [Figure 9] It is a block diagram of an image processing system in Example 3. [Figure 10] It is a flowchart regarding generation of a corrected image in Example 3.

Modes for Carrying Out the Invention

[0010] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In each figure, the same members are denoted by the same reference numerals, and overlapping descriptions are omitted.

[0011] First, before giving a specific description of the embodiments, the gist of the present invention will be described. The present invention relates to image processing for improving overcorrection and undercorrection that occur in a sharpening process for correcting blurring caused by the optical characteristics of an imaging device. Here, the optical characteristics include aberration, diffraction, a low-pass filter, pixel aperture effects, and the like. In particular, it relates to improvement of overcorrection and undercorrection that occur when the blurring characteristics in an actual captured image are different from the assumed optical characteristics in a sharpening process based on known optical characteristics. Here, blurring is a point spread function (PSF) or an optical transfer function (OTF).

[0012] In sharpening processing, the optical characteristics can be determined from the shooting conditions of the optical system at the time of imaging (zoom, aperture, focus distance) and the position of the captured image on the screen. When design values ​​are used for the optical characteristics, the optical characteristics may be better or worse than expected due to manufacturing errors in the imaging device. Also, the optical characteristics change if the subject distance deviates from what is expected. If it is assumed that the subject is on the in-focus plane when determining the optical characteristics used in sharpening processing, then in scenes with depth, subjects on the out-of-focus plane (defocused) will blur with different optical characteristics. On the other hand, even when considering the optical characteristics on the out-of-focus plane, the optical characteristics will differ from what is expected due to deviations in subject distance. Even for subjects on the out-of-focus plane, the optical characteristics may be better or worse than expected due to the relationship with the field curvature of the imaging optical system. Thus, when the optical characteristics of the captured image differ from the optical characteristics expected from the shooting conditions and screen position, sharpening processing based on incorrect optical characteristics will result in overcorrection or undercorrection.

[0013] In this invention, a multi-layer neural network (first machine learning model) is used to correct overcorrection or undercorrection when optical characteristics differ from those assumed. By training the first machine learning model using training data having the characteristics described later, it is possible to obtain a corrected image with better blur correction based on intermediate correction data obtained by sharpening processing.

[0014] In the learning of weights (filters, biases, etc.) used in a multi-layer neural network, when the optical performance deviates from the assumption, the output (intermediate correction data) obtained by the sharpening process is used as training data. As the correct image, an image with appropriately corrected aberration is used. The training data is input into the neural network, and the weights are optimized so that the error between the output and the correct image is reduced. At this time, the neural network is trained to improve (correct) overcorrection and undercorrection. Due to manufacturing errors and subject distance deviations, the optical characteristics vary in various ways. Therefore, by generating training data assuming a plurality of different optical characteristics, a neural network that can improve overcorrection and undercorrection for all cases can be learned. Hereinafter, the case where the gloss characteristics deviate from the assumption is called performance deviation, and the correction process for improving overcorrection and undercorrection caused by the influence of performance deviation is called performance deviation correction.

Example

[0015] First, referring to FIGS. 2 and 3, the image processing system in Example 1 of the present invention will be described. In this example, a multi-layer neural network is made to learn and execute correction of overcorrection and undercorrection. FIG. 2 is a block diagram of the image processing system 100 in this example. FIG. 3 is an external view of the image processing system 100.

[0016] The image processing system 100 includes a learning device (image processing device) 101, an imaging device 102, an image estimation device (image processing device) 103, a display device 104, a recording medium 105, an output device 106, and a network 107. The learning device 101 includes a storage unit (storage means) 101a, an acquisition unit (acquisition means) 101b, a generation unit (generation means) 101c, and an update unit (learning means) 101d.

[0017] The imaging device 102 includes an optical system (imaging optical system) 102a and an image sensor 102b. The optical system 102a collects light incident on the imaging device 102 from the subject space. The image sensor 102b receives (photoelectrically converts) the optical image (subject image) formed through the optical system 102a to acquire an image. The image sensor 102b is, for example, a CCD (Charge Coupled Device) sensor or a CMOS (Complementary Metal-Oxide Semiconductor) sensor. The image acquired by the imaging device 102 includes blurring due to aberrations and diffraction of the optical system 102a and noise from the image sensor 102b.

[0018] The image estimation device 103 includes a storage unit 103a, an acquisition unit 103b, and a correction unit 103c. The image estimation device 103 acquires an captured image, performs sharpening processing, and then performs performance shift correction to generate an estimated image (corrected image) by the first machine learning model described later. The sharpening processing uses a multilayer neural network (second machine learning model), and the weight information is read from the storage unit 103a. The shooting conditions and the screen position of the captured image (image height, azimuth) are input to the second machine learning model, and an intermediate correction image (intermediate correction data) is obtained as output. Here, the sharpening processing is performed by the image estimation device 103, but the sharpened intermediate correction image may be obtained by a different image processing device. A multilayer neural network (first machine learning model) is used for performance shift correction, and the weight information is read from the storage unit 103a.

[0019] The weights (weight information) of the first and second machine learning models are learned by the learning device 101. The image estimation device 103 has previously read the weight information from the storage unit 101a via the network 107 and stored it in the storage unit 103a. The weight information stored may be the numerical weights themselves or in an encoded format. Weight learning and sharpening processing using the weights are well known, and details regarding the performance deviation correction processing will be described later. The image estimation device 103 performs performance deviation correction on the intermediate correction image to generate a corrected image. In this embodiment, an image is used as intermediate correction data and is therefore called an intermediate correction image, but a feature map, described later, may also be used as intermediate correction data.

[0020] The corrected image is output to at least one of the display device 104, the recording medium 105, and the output device 106. The display device 104 is, for example, a liquid crystal display or a projector. The user can perform editing work, etc., while checking the image in progress via the display device 104. The recording medium 105 is, for example, a semiconductor memory, a hard disk, or a server on a network. The output device 106 is, for example, a printer. The image estimation device 103 has the function of performing development processing and other image processing as needed.

[0021] Next, with reference to Figures 1 and 4, the method for learning weights (weight information) (method for manufacturing a trained model) performed by the learning device 101 in this embodiment will be described. Figure 1 is a diagram showing the flow of learning weights for a neural network (first machine learning model). Figure 4 is a flowchart related to weight learning. Each step in Figure 4 is mainly performed by the acquisition unit 101b, generation unit 101c, or update unit 101d of the learning device 101.

[0022] The training method for the first machine learning model will now be explained. First, in step S101 of Figure 4, the acquisition unit 101b acquires the original image (subject image). In this embodiment, the original image is a high-resolution (high-quality) image with minimal blurring due to aberrations and diffraction of the optical system 102a. Multiple original images are acquired, each containing various subjects, i.e., edges of varying strengths and directions, textures, gradients, flat areas, etc. The original image may be a real-life photograph or an image generated by computer graphics (CG). In particular, when using a real-life photograph as the original image, blurring has already occurred due to aberrations and diffraction, so reducing the image size can reduce the effect of blurring and produce a high-resolution (high-quality) image. Note that if the original image contains sufficient high-frequency components, reduction may not be necessary.

[0023] Preferably, the original image has a signal value higher than the luminance saturation value of the image sensor 102b. This is because, even with real subjects, when photographed by the imaging device 102 under specific exposure conditions, there are subjects that do not fall within the luminance saturation value. When using a real-world image as the original image, it can be obtained by HDR shooting or by shooting with an imaging device that has a higher dynamic range than the imaging device 102. When using an image taken with an imaging device that has a dynamic range equivalent to that of the imaging device 102 as the original image, it is possible to obtain a higher signal value by proportionally multiplying the signal value. However, it is preferable to do so within a range where the reduction in gradation due to proportional multiplication does not affect the learning results. In addition, the original image may contain noise components. In this case, since the noise contained in the original image can be considered as part of the subject, the noise in the original image is not particularly problematic.

[0024] Next, in step S102, the acquisition unit 101b acquires the blur used for the imaging simulation described later. First, the acquisition unit 101b acquires the shooting conditions corresponding to the lens state (zoom, aperture, and focus distance) of the optical system 102a. Then, the acquisition unit 101b acquires the blur determined by the shooting conditions and the screen position of the captured image. Here, the blur is the PSF (point image intensity distribution) or OTF (optical transfer function) of the optical system 102a. The blur can be acquired by optical simulation or measurement in the optical system 102a. In addition, the blur due to aberrations and diffraction of the lens state, image height, and azimuth, which differ for each original image, is acquired. This makes it possible to perform imaging simulations corresponding to multiple shooting conditions, image heights, and azimuths. Furthermore, if necessary, the components of filters such as the optical low-pass filter included in the imaging device 102 may be added to the blur to be applied (filters may be applied to the optical characteristics).

[0025] Next, in step S103, the generation unit 101c generates ground truth patches (ground truth images) and training patches (training data). Multiple ground truth patches and training patches are generated, with one or more patches generated for each original image. In this embodiment, the ground truth patches and training patches are images of the same subject. In this embodiment, multiple combinations of ground truth patches and training patches are used as training data. The training data is used to train a machine learning model that corrects the effects of optical characteristic deviations (performance deviations) that occur when aberrations and diffractions of the optical system 102a are sharpened based on the optical characteristics. For this reason, the ground truth patches are images that have fewer problems due to overcorrection or have improved correction compared to the training patches. However, as will be described later, the machine learning model may not correct depending on the conditions, so the training data may include cases where the ground truth patches and training patches are the same image.

[0026] A patch refers to an image with a predetermined number of pixels (e.g., 64 x 64 pixels). The number of pixels in the ground truth patch and the training patch do not necessarily have to match. In this embodiment, mini-batch learning is used to train the weights of a multi-layer neural network. Therefore, in step S103, multiple sets of ground truth patches and training patches are generated. However, this embodiment is not limited to this, and online learning or batch learning may also be used.

[0027] This embodiment obtains ground truth patches and training patches by the following method, but is not limited thereto. This embodiment generates multiple blurred images with the effects of aberration and diffraction by performing an imaging simulation using multiple original images stored in the memory unit 101a as subjects. Then, intermediate correction images are generated by sharpening each blurred image using a second machine learning model based on its optical properties. In this embodiment, intermediate correction images are used as training data. In addition, a ground truth image is obtained based on the original image corresponding to each training data. At this time, the ground truth image and training data have the same number of pixels as or greater than the number of pixels in the patch used as training data. Then, multiple ground truth patches and training patches are obtained by extracting subregions of a specified pixel size at the same position from multiple pairs of ground truth images and training data. In this embodiment, the original image is an undeveloped RAW image, and the ground truth patch and training patch are also RAW images. However, this embodiment is not limited to this, and developed images may also be used, or feature maps obtained by transforming images as described later. Also, the position of the subregion refers to the center of the subregion.

[0028] The training data used to train the first machine learning model is generated in the following way. The generation unit 101c performs the following processing for predetermined conditions (such as shooting conditions and screen position). First, the generation unit 101c generates a blur with a smaller amount of blur based on the blur obtained in step S102, which is the blur caused by aberrations and diffraction of the optical system 102a. The generation unit 101c then applies (convolves) the generated blur to the original image. Subsequently, the generation unit 101c obtains a blurred image by extracting a partial region from the image clipped at the brightness saturation value of the image sensor 102b.

[0029] Next, the generation unit 101c corrects the blurred image using a second machine learning model based on the blur (optical characteristics) acquired in step S102 under predetermined conditions, and acquires an intermediate correction image. By extracting the same subregion from the intermediate correction image and the original image, a set of training patches and ground truth patches can be generated. The generation unit 101c generates multiple sets of training patches and ground truth patches for each blur size, based on blurred images with multiple different sizes of blur applied to them, using multiple amounts of blur to be generated. That is, the generation unit 101c uses the blur acquired in step S102 as a reference and generates sets of training patches and ground truth patches using blurred images with multiple different sizes of blur applied to them. In this embodiment, the sizes of the multiple blurs are all smaller than the blur acquired in step S102, i.e., the optical characteristics under predetermined conditions, but it is not limited to this. By performing the aforementioned process under predetermined conditions, multiple blurred images corresponding to multiple performance deviations for multiple conditions of the imaging device (such as shooting conditions and screen position) can be generated, and training data with training patches and ground truth patches based on these blurred images can be generated.

[0030] Here, we will explain the deformation of the blur. In this embodiment, the blur applied to the blurred image is obtained by reducing the spatial extent of the PSF acquired in step S102. The reduction ratio of the PSF is randomly set between 0.1 and 0.9 times, and the PSF is reduced by the set reduction ratio. The reduction ratio may be different for the depth of field and the depth of field. The reduction process is performed using downsampling and interpolation processes such as bicubic interpolation. As a result, the blur applied to the blurred image is small, based on the blur (optical characteristics) acquired in step S102. When sharpening is performed on the blurred image generated in this way using the second machine learning model, the intermediate correction image becomes overcorrected, and other problems occur.

[0031] Since neural networks learn to make training patches closer to the correct patch, they learn to make intermediate-corrected images closer to the original image. In other words, the first machine learning model in this embodiment suppresses overcorrection to an appropriate correction effect and also suppresses harmful effects (reduces the impact of overcorrection). Here, an appropriate correction effect is a correction effect that brings the intermediate-corrected image closer to the original image or an intermediate-corrected image without performance deviation. Even without harmful effects such as ringing, over-sharpening can result in an unnatural image. Therefore, it is desirable to suppress the correction effect to an appropriate level in order to match the correction effect with other image regions that do not have performance deviation or performance deviation.

[0032] Note that the blur applied to the blurred image is not limited to the blur described above. Other blur filters such as Gaussian blur may also be used. Alternatively, instead of downsampling the PSF obtained in step S102, it may be reduced by convolution of a sharpening filter or by existing reduction processing. If the blur is obtained as OTF in step S102, the blur can be reduced by performing expansion processing such as upsampling, processing to improve the OTF by applying frequency gain, or dilation processing. The magnitude of the blur in this embodiment can be expressed by a predetermined index. For example, the amount of blur can be said to be small when the maximum value of the PSF is large, the half-width at the PSF's cross-section is small, the MTF of a specified spatial frequency is large, or the integral value of the MTF with respect to spatial frequency is large. Furthermore, the amount of blur can be increased by convolving a blur filter (low-pass filter) onto the PSF or by applying the frequency characteristics of a low-pass filter to the OTF. The magnitude of blur relative to a predetermined blur (the magnitude of blur with respect to a predetermined blur) may be expressed as the difference or ratio of the aforementioned indicators, or the aforementioned indicators may be used for the difference or ratio of OTF.

[0033] The sharpening process based on optical properties (the second machine learning model) can be trained by using the blurred image acquired in step S102 as the training patch and the ground truth patch as the original image. Note that if a real-world image is reduced in size to become the original image, the order of reduction and blurring can be reversed. If blurring is applied first, the sampling rate of the blur needs to be finer to account for the reduction. For PSF (Point Image Intensity Distribution), the sampling points in space should be finer, and for OTF (Optical Transfer Function), the maximum frequency should be increased.

[0034] It is preferable that the added blur does not include distortion. This is because, in sharpening processing (aberration correction processing), if distortion is large, the position of the subject changes, and the subject may differ between the ground truth patch and the training patch. For this reason, the sharpening processing (second machine learning model) used in this embodiment does not correct distortion. Accordingly, the performance deviation correction by the first machine learning model also does not correct distortion. Distortion is corrected individually after blur correction using methods such as bilinear interpolation or bicubic interpolation.

[0035] Furthermore, since noise is generated in the image sensor 102b during actual imaging, it is preferable to add noise to the training data as well. Random numbers corresponding to the noise characteristics of the image sensor 102b can be generated and added, and shooting conditions such as ISO sensitivity may also be considered. If noise is added only to the training images, or if noise that is not correlated between the training images and the ground truth images is added, denoising will be learned simultaneously with blur correction during training. If the same noise is added to the training images and the ground truth images, blur correction that suppresses changes in noise will be learned. In this case, if a feature map, which will be described later, is used as training data, noise should be added to the blurred images instead of the training images.

[0036] Next, in step S104, the generation unit 101c inputs the training patch (training data) 212 shown in Figure 1 into a multilayer neural network and generates estimation patch (estimated image) 213. For mini-batch learning, estimation patches 213 corresponding to multiple training patches 212 are generated. Figure 1 shows the flow from step S104 to step S105.

[0037] The estimated patch 213 is the training patch 212 with performance deviations corrected, and ideally matches the ground truth patch (ground truth image) 211. In this embodiment, the neural network configuration shown in Figure 1 is used, but the present invention is not limited to this. In Figure 1, CN represents a convolutional layer, and DC represents a deconvolutional layer. In both CN and DC, the sum of the input, the convolution of the filter, and the bias is calculated, and the result is nonlinearly transformed by an activation function. The initial values ​​of each filter component and the bias are arbitrary and are determined by random numbers in this embodiment. The activation function can be, for example, ReLU (Rectified Linear Unit) or a sigmoid function. The output of each layer except the final layer is called a feature map. Skip connections 222 and 223 synthesize feature maps output from non-contiguous layers. The feature maps may be synthesized by element-wise summation or by concatenation in the channel direction. In this embodiment, element-wise summation is adopted. The skip connection 221 generates an estimated patch 213 by summing the residuals estimated from the training patch 212 and the ground truth patch 211 with the training patch 212. An estimated patch 213 is generated for each of the multiple training patches 212.

[0038] Next, in step S105, the update unit 101d updates the neural network weights (weight information) based on the error between the estimated patch 213 and the ground truth patch (ground truth image) 211. Here, the weights include the filter components and biases of each layer. Backpropagation is used to update the weights, but this embodiment is not limited to this method. For mini-batch learning, the error between multiple ground truth patches 211 and their corresponding estimated patches 213 is calculated, and the weights are updated. For the loss function, for example, the L2 norm or L1 norm may be used.

[0039] Next, in step S106, the update unit 101d determines whether or not weight learning is complete. Completion can be determined by whether the number of iterations of learning (weight update) has reached a predetermined value, or whether the amount of change in weight at the time of update is less than a predetermined value. If it is determined that learning is not complete, the process returns to step S104 and multiple new correct answer patches and training patches are acquired. On the other hand, if it is determined that learning is complete, the learning device 101 (update unit 101d) terminates learning and stores the weight information in the storage unit 101a.

[0040] Next, with reference to Figure 5, the generation of corrected images (correction processing, estimation method) performed by the image estimation device 103 in this embodiment will be described. Figure 5 is a flowchart relating to the generation of corrected images. Each step in Figure 5 is mainly performed by the acquisition unit 103b or the correction unit 103c of the image estimation device 103.

[0041] First, in step S201, the acquisition unit 103b acquires the captured image and weight information. The captured image is an undeveloped RAW image, similar to that used for training, and in this embodiment, it is transmitted from the imaging device 102. The weight information consists of the weights of the first machine learning model and the second machine learning model, which are transmitted from the learning device 101 and stored in the storage unit 103a.

[0042] Next, in step S202, the acquisition unit 103b performs sharpening processing based on the weights of the second machine learning model acquired in step S201 and acquires an intermediate corrected image.

[0043] Next, in step S203, the correction unit 103c receives the intermediate correction image acquired in step S202 as the input image for the first machine learning model and generates a corrected image. The corrected image is an image in which performance deviation correction has been applied to the intermediate correction image. In other words, it suppresses the portion that has been corrected more than expected due to overcorrection to an appropriate amount of correction, and also suppresses the harmful effects (reduces the impact of overcorrection). Note that when inputting the captured image to the neural network, it is not necessary to crop it to the same size as the training patch used during learning.

[0044] In this embodiment, sharpening is performed using a machine learning model. However, the sharpening process may also be performed using other existing methods based on optical properties, such as a convolutional method using a sharpening filter like a Wiener filter. The first machine learning model can be similarly trained by using intermediate corrected images obtained by applying other sharpening methods to blurred images as training data. On the other hand, performing sharpening with a machine learning model using a neural network such as a CNN is preferable because it can correct image quality degradation due to optical properties with higher accuracy. When using a neural network, the high correction effect can lead to significant problems if there is a deviation from the expected optical properties. However, by performing processing with the first machine learning model in this embodiment, it is possible to obtain corrected images with a high correction effect and suppressed problems.

[0045] Furthermore, it is preferable to use a machine learning model (a second machine learning model common to multiple optical characteristics) that has been commonly trained on multiple optical characteristics for sharpening. When other sharpening methods are used, or when a machine learning model for sharpening is trained for each optical characteristic, differences in correction effects tend to occur for each optical characteristic, and differences in the occurrence of overcorrection and the resulting drawbacks also tend to occur. On the other hand, by using a commonly trained machine learning model, the correction effects and the occurrence of drawbacks are common, and the first machine learning model can be trained with greater accuracy. As a result, performance deviations can be corrected more effectively.

[0046] Furthermore, it is preferable to include information on at least one of the following in the input data for the machine learning model used for sharpening: not only the captured image, but also the shooting conditions and the screen position (information on the shooting conditions and screen position is input and used). In this case, it is necessary to input this information both during training and during the correction processing. This allows for more accurate correction of blurring due to aberrations, diffraction, etc., and by performing processing by the first machine learning model of the present invention in conjunction with this, it is possible to obtain a corrected image with a high correction effect and suppressed harmful effects.

[0047] Furthermore, it is preferable to include information on shooting conditions and screen position in the input data of the first machine learning model, in addition to intermediate correction data. In this case, it is necessary to input information on shooting conditions both during training and during the correction processing. For example, in low-performance areas where field curvature occurs at high image height, overcorrection due to three-dimensional subjects or uneven blur is likely to occur. The likelihood and manifestation of overcorrection and other problems differ depending on the shooting conditions and optical characteristics. By inputting the shooting conditions into the first machine learning model, performance deviations can be corrected based on the shooting conditions and optical characteristics, and a good corrected image can be obtained. In particular, if information on shooting conditions is also input into the machine learning model for sharpening processing, it is possible to learn using shooting condition information common to both sharpening processing and performance deviation correction, and an even better corrected image can be obtained.

[0048] Furthermore, it is preferable to include manufacturing error information of the imaging device as input data for the first machine learning model. By adding manufacturing error information obtained by measuring after manufacturing, performance deviations can be corrected based on the impact of manufacturing errors on optical characteristics. For example, if it is known that the imaging optical system is partially defocused, it is possible to identify screen positions that are prone to overcorrection and screen positions that are prone to undercorrection.

[0049] In this embodiment, the intermediate correction data and training data are intermediate correction images. However, when sharpening is performed using a machine learning model, the feature maps of the intermediate layers may be used as the intermediate correction data and training data. Using images for intermediate correction data has the advantage of making it easier to identify problems on the image during training, and also allows for the creation of a machine learning model that uses the difference between the image and the captured image. On the other hand, if the number of channels in the feature map is large, the amount of information will be reduced when converted to an image.

[0050] Therefore, by using feature maps as intermediate correction data and training data to input to the first machine learning model, performance discrepancies can be corrected without reducing the amount of information. In this case as well, when training the second machine learning model, it is sufficient to estimate the intermediate correction image. For example, it is possible to train using the error between the intermediate correction image estimated from the captured image and the original image. When performing correction processing using the captured image, intermediate correction data can be calculated using the weights up to the intermediate layer of the second machine learning model and input to the first machine learning model to obtain a sharpened and corrected image with performance discrepancies corrected. Therefore, the weight information of the second machine learning model obtained in step S201 only needs to include the weights of the layers before the intermediate layer used as intermediate correction data. Note that the number of pixels and channels of the training data may differ from those of the ground truth image. Alternatively, blurred images can be used as training data, and the machine learning model for sharpening and the machine learning model for performance discrepancy correction can be linked and trained together during training. In this case, the machine learning model for sharpening only needs to update the weights of the layer corresponding to the model that performs performance discrepancy correction, without updating the weights that were previously trained.

[0051] Furthermore, when performing sharpening using a machine learning model, it is not necessary to output intermediate correction data or intermediate correction images during the correction process from the captured image. In other words, sharpening and performance deviation correction can be performed together as a single model.

[0052] Furthermore, while the ground truth images used to train the first machine learning model are used as source images and ground truth patches are extracted from the source images, this embodiment is not limited to this. For example, a blurred image with the same blur as the optical characteristics may be used as the ground truth image after performing sharpening processing based on the optical characteristics. In addition, any image that is not affected by performance discrepancies and has little blur due to the optical characteristics can be used as the ground truth image.

[0053] Furthermore, the ground truth image may be used as the original image only if the intermediate correction data and training data used to train the first machine learning model exhibit adverse effects such as undershoot or ringing. In other words, even if the image is overcorrected, if it does not exhibit unnatural-looking adverse effects such as undershoot or ringing, the ground truth image may be the intermediate correction data and training data themselves. In this case, the first machine learning model will be trained to suppress only the unnatural-looking adverse effects, resulting in an output image that is overcorrected but still sharp. Also, attempting to suppress the correction effect of overcorrected intermediate correction data or training data, even if it does not exhibit unnatural-looking adverse effects, will result in learning a suppression effect that is difficult to judge directly from the image. Avoiding this makes training easier. It is not always necessary to distinguish whether adverse effects are actually occurring or not. For example, the ground truth image may be used as the original image only for high-contrast images or images containing edges, which are prone to adverse effects.

[0054] In this embodiment, performance deviations resulting from overcorrection are corrected by making the blurred image blurred to a degree smaller than the optical characteristics, but this embodiment is not limited to this. By generating training data with a blurred image that is greater than the optical characteristics, a first machine learning model that corrects performance deviations due to insufficient correction can be trained. In this case, as the ground truth image, the original image may be used, as in the case of correcting performance deviations due to overcorrection, or an image obtained by performing a sharpening process based on the optical characteristics on a blurred image that has been given the same blur as the optical characteristics may be used.

[0055] Furthermore, training data may be generated by including both smaller and larger blurs than the optical characteristics in the blurred images. The ground truth image may be the original image, or it may be an image obtained by applying a sharpening process based on the optical characteristics to a blurred image that has been given the same blur as the optical characteristics. This allows for the training of a first machine learning model that corrects both performance deviations due to overcorrection and performance deviations due to undercorrection. If only the correction of overcorrection is trained, the correction effect on undercorrected image areas may be further suppressed. However, by training with both included, the first machine learning model can learn both cases of overcorrection and undercorrection, and therefore can correct both overcorrection and undercorrection.

[0056] Furthermore, if training data is generated by including both smaller and larger blurs than the optical characteristics in the blurred images, only one of the performance deviations—overcorrection or undercorrection—may be corrected. By changing the ground truth image (when no correction is applied) to an image obtained by sharpening the corresponding blurred image, a machine learning model that performs such corrections can be trained. Correcting overcorrection is a process that degrades the frequency characteristics, while correcting undercorrection is a process that improves the frequency characteristics. Since these two functions are contradictory, limiting the correction to only one of the performance deviations can result in a lighter model or improve the effectiveness of the performance deviation correction. For example, as the ground truth image corresponding to training data where the blur applied to the blurred image is smaller than the optical characteristics, an image obtained by applying blur based on the optical characteristics to the original image and then sharpening it based on the optical characteristics, or the original image, can be used. Also, as the ground truth image corresponding to training data where the blur applied to the blurred image is larger than the optical characteristics, a blurred image corresponding to the training data or intermediate correction data (intermediate correction image) corresponding to the training data can be used. This eliminates the harmful effects of overcorrection (or reduces its impact), but does not correct for undercorrection due to performance discrepancies. As in the case mentioned above, compared to training with training data for only one case, training with both cases allows the first machine learning model to distinguish between cases where correction is applied and cases where it is not.

[0057] Furthermore, training data may be generated including cases where the blur applied to the blurred image is identical (or nearly identical) to the optical characteristics. The ground truth image may be the original image, or it may be an image obtained by performing a sharpening process based on the optical characteristics on a blurred image to which the same blur as the optical characteristics has been applied. In other words, the original image or intermediate correction data (intermediate correction image) corresponding to the training data can be used as the ground truth image corresponding to training data in which the blur applied to the blurred image is identical to the optical characteristics. This allows the first machine learning model to learn both to correct when there is a performance discrepancy and not to correct when there is no performance discrepancy.

[0058] Furthermore, when performing correction using the first machine learning model, the input data may include not only intermediate correction data but also captured images. During training, blurred images can be added. In this case, knowing the image before sharpening allows for more accurate performance deviation correction. For example, if overcorrection causes problems such as undershoot or ringing, using the captured image can help determine whether the structure is real or a structure created by the sharpening process. Also, by using the difference with the captured image, it is possible to explicitly correct performance deviations using the changes caused by the sharpening process.

[0059] Furthermore, the blur applied to the original image when acquiring training data may be based on at least one of the optical system's shooting conditions and the screen position. For example, if the optical properties are highly sensitive to manufacturing tolerances or subject distance, the performance deviation from the optical properties may be large. In such cases, by significantly changing the blur applied to the blurred image, a first machine learning model that can handle larger performance deviations can be trained. To significantly change the blur, for example, one can use multiple magnification and reduction ratios when generating the blur from the PSF, setting a large upper limit for the magnification ratio or a small lower limit for the reduction ratio.

[0060] In this embodiment, the learning device 101 and the image estimation device 103 are described as separate components, but this embodiment is not limited to this. The learning device 101 and the image estimation device 103 may be configured as an integrated unit. That is, learning (processing shown in Figure 4) and estimation (processing shown in Figure 5) may be performed within a single unit.

[0061] Furthermore, the image estimation device that performs sharpening and performance deviation correction may be a separate unit, and instead of acquiring the captured image, the image after sharpening may be acquired. In this case, the execution of the sharpening process is skipped.

[0062] As described above, the image processing device (learning device 101) of this embodiment performs training of a machine learning model for outputting a corrected image from an intermediate corrected image that corrects for blur caused by the optical characteristics of the optical system. The image processing device has a generation unit 101c that functions as a first acquisition means, a second acquisition means, and a third acquisition means, and an update unit 101d that functions as a learning means. The first acquisition means acquires a blurred image by adding blur to the original image. The second acquisition means acquires intermediate corrected data obtained by performing a sharpening process based on optical characteristics on the blurred image as multiple training data. The third acquisition means acquires multiple ground truth images based on the original image, corresponding to each training data. The learning means learns the first machine learning model using learning data consisting of training data and ground truth images. Furthermore, the blur added to the blurred image differs for each training data based on the optical characteristics.

[0063] According to this embodiment, it is possible to provide an image processing method that can acquire a corrected image with better blur correction based on intermediate correction data that corrects blur caused by optical characteristics. [Examples]

[0064] Next, with reference to Figures 6 and 7, the image processing system in Embodiment 2 of the present invention will be described. In this embodiment, the generation of the corrected image is performed by the image estimation unit 323 in the imaging device. Furthermore, the flow of generating the corrected image in this embodiment differs from that of Embodiment 1.

[0065] Figure 6 is a block diagram of the image processing system 300 in this embodiment. Figure 7 is an external view of the image processing system 300. The image processing system 300 includes a learning device (image processing device) 301 and an imaging device 302 connected via a network 303. The learning device 301 has a storage unit (storage means) 311, an acquisition unit (acquisition means) 312, a generation unit (generation means) 313, and an update unit (learning means) 314, and learns weights (weight information) for performing blur correction using a neural network. The imaging device 302 captures the subject space and acquires an image, sharpens the image using the read weight information, corrects performance deviations, and generates a corrected image. The imaging device 302 has an optical system 321 and an image sensor 322. The image estimation unit 323 has an acquisition unit 323a and a correction unit 323b, and performs performance deviation correction of the image using the weight information stored in the storage unit 324.

[0066] Weight information is pre-learned by the learning device 301 and stored in the storage unit 311. The imaging device 302 reads the weight information from the storage unit 311 via the network 303 and stores it in the storage unit 324. The captured image with performance deviation corrected (corrected image) is stored in the recording medium 325. When the user issues a command to display the corrected image, the stored corrected image is read and displayed in the display unit 326. Alternatively, an captured image already stored in the recording medium 325 may be read and the image estimation unit 323 may perform performance deviation correction. This series of controls is performed by the system controller 327. The learning of the first model executed by the learning device 301 in this embodiment is the same as in Embodiment 1.

[0067] Next, with reference to Figure 8, the generation of corrected images performed by the image estimation unit 323 in this embodiment will be described. Figure 8 is a flowchart of the generation of corrected images. Each step in Figure 8 is mainly performed by the acquisition unit 323a or the correction unit 323b of the image estimation unit 323. Steps S401 to S403 are the same as steps S201 to S203 performed by the acquisition unit 103b or the correction unit 103c in Embodiment 1, respectively.

[0068] Next, in step S404, the correction unit 323b performs a weighted average of the corrected image with performance deviation correction applied and the sharpened intermediate correction image (intermediate correction data) to generate a composite image. This allows for adjustment of the strength of the performance deviation correction effect. The weighted average ratio may vary depending on the region of the image. In this embodiment, the ratio of the corrected image is reduced when the amount of brightness change due to the sharpening process, i.e., the difference between the captured image and the intermediate correction image, is smaller than a predetermined value. In other words, the effect of performance deviation correction is reduced. This is because, when the first machine learning model performs performance deviation correction, if the amount of brightness change is small, the accuracy of determining whether or not there is performance deviation may be poor. In textured areas of the subject with fine structures, whether or not to perform performance deviation correction depends on slight changes in brightness or structure, and patchy correction is likely to occur, which is undesirable. In regions where the amount of brightness change due to the sharpening process is small, the influence of performance deviation is originally small, so by reducing the change due to performance deviation correction in these regions, patchy correction can be avoided.

[0069] In this embodiment, a weighted average was performed using the corrected image and the intermediate corrected image, but the captured image may be used instead of the intermediate corrected image. For example, the degree of sharpening can be adjusted by increasing the proportion of the captured image in the entire or partial area of ​​the image. The weighted average ratio may be changed according to the user's specifications, thereby achieving a degree of sharpening desired by the user. Furthermore, the composite image is not limited to using two images; the degree of sharpening and performance deviation correction can also be adjusted by performing a weighted average using the captured image, the intermediate corrected image, and the corrected image. Existing methods other than weighted average may also be used for the composite method.

[0070] As described above, in this embodiment, a corrected image is synthesized from at least one of the captured image and the intermediate correction data based on a sharpening component (for example, the amount of brightness change due to sharpening processing) obtained based on the captured image and the intermediate correction data. Thus, unlike Embodiment 1, this embodiment allows for adjustment of the strength of sharpening and performance deviation correction by performing synthesis processing between multiple images. According to this embodiment, it is possible to provide an image processing method that makes it possible to obtain a corrected image with better blur correction based on intermediate correction data that corrects blur caused by optical characteristics. [Examples]

[0071] Next, with reference to Figure 9, the image processing system in Embodiment 3 of the present invention will be described. The image processing system in this embodiment differs from Embodiments 1 and 2 in that it has a processing unit (computer) that transmits the captured image to be processed to an image estimation device and receives the processed output image (corrected image) from the image estimation device.

[0072] Figure 9 is a block diagram of the image processing system 600 in this embodiment. The image processing system 600 includes a learning device 601, an imaging device 602, an image estimation device 603, and a processing device (computer) 604. The learning device 601 and the image estimation device 603 are, for example, servers. The computer 604 is, for example, a user terminal (personal computer or smartphone). The processing device 604 is connected to the image estimation device 603 via a network 605. The image estimation device 603 is connected to the learning device 601 via a network 606. That is, the computer 604 and the image estimation device 603 are configured to communicate with each other, and the image estimation device 603 and the learning device 601 are configured to communicate with each other. Note that the configuration of the learning device 601 is the same as that of the learning device 101 in Embodiment 1, so its description is omitted. Similarly, the configuration of the imaging device 602 is the same as that of the imaging device 102 in Embodiment 1, so its description is omitted.

[0073] The image estimation device 603 includes a storage unit 603a, an acquisition unit 603b, a correction unit 603c, and a communication unit (receiving means) 603d. The storage unit 603a, the acquisition unit 603b, and the correction unit 603c are the same as those of the image estimation device 103 of Embodiment 1. The communication unit 603d has the function of receiving requests transmitted from the computer 604 and transmitting the output image generated by the image estimation device 603 to the computer 604.

[0074] The computer 604 includes a communication unit (transmission means) 604a, a display unit 604b, an image processing unit 604c, and a recording unit 604d. The communication unit 604a has the function of transmitting a request to the image estimation unit 603 to perform processing on the captured image, and the function of receiving the output image processed by the image estimation unit 603. The display unit 604b has the function of displaying various information. The information displayed by the display unit 604b includes, for example, the captured image transmitted to the image estimation unit 603 and the output image received from the image estimation unit 603. The image processing unit 604c has the function of performing further image processing on the output image received from the image estimation unit 603. The recording unit 604d records the captured image acquired from the imaging unit 602, the output image received from the image estimation unit 603, etc.

[0075] Next, with reference to Figure 10, the image processing (generation of corrected images) by the image processing system 600 will be described. Figure 10 is a flowchart relating to the generation of corrected images in this embodiment. The content of the image processing (correction processing) in this embodiment is the same as the correction processing described in Embodiment 1 with reference to Figure 5. The image processing shown in Figure 10 is initiated when the user issues an instruction to start image processing via the computer 604.

[0076] First, let's explain the operation of computer 604. In step S701, computer 604 sends a request to image estimation device 603 for processing the captured image. The method of sending the captured image to be processed to image estimation device 603 is not limited. For example, the captured image may be uploaded to image estimation device 603 at the same time as step S701, or it may have been uploaded to image estimation device 603 before step S701. Also, the captured image may be an image stored on a server different from image estimation device 603. In addition, in step S701, computer 604 may send ID information for user authentication along with the request to process the captured image.

[0077] Next, in step S702, the computer 604 receives the output image generated in the image estimation device 603. The output image is, as in Embodiment 1, an image that has been sharpened relative to the captured image and then corrected for performance deviation effects.

[0078] Next, the operation of the image estimation device 603 will be described. In step S801, the image estimation device 603 receives a request from the computer 604 for processing the captured image. At this time, the image estimation device 603 determines that processing (sharpening processing and performance shift correction processing) on ​​the captured image has been instructed, and executes the processing from step S802 onwards.

[0079] In step S802, the image estimation device 603 acquires weight information. The weight information is information (a trained model) learned in the same manner as described in Example 1 with reference to Figure 4. The image estimation device 603 may acquire weight information from the learning device 601, or it may acquire weight information that has been previously acquired from the learning device 601 and stored in the storage unit 603a. The following steps S803 and S804 are the same as S202 and S203 of Example 1, respectively. Next, in step S805, the image estimation device 603 transmits the output image to the computer 604. Although this embodiment has been described as performing the correction processing of Example 1, the correction processing of Example 2 may also be performed.

[0080] As in this embodiment, when performance deviation correction processing is performed within the image estimation device 603, the processing load due to the correction processing can be handled within the image estimation device 603, thus reducing the processing power required on the computer 604 side. Also, as in this embodiment, the image estimation device 603 may be configured to be controlled using a computer 604 that is connected to the image estimation device 603 in a manner that enables communication with the image estimation device 603.

[0081] (Other examples) The present invention can also be realized by supplying a program that implements one or more of the functions of the above embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0082] According to each embodiment, it is possible to provide a learning method, an image processing method, a learning device, and a program that can acquire a corrected image with better blur correction based on intermediate correction data that corrects blur caused by optical characteristics.

[0083] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of its essence. [Explanation of Symbols]

[0084] 101 Learning device (image processing device) 101c Generation unit (first acquisition means, second acquisition means, third acquisition means) 101d Update section (learning method)

Claims

1. The steps include: acquiring intermediate correction data obtained by performing a sharpening process on the captured image based on the optical characteristics of the optical system used to capture the captured image; The steps include: obtaining a corrected image from the aforementioned intermediate correction data using a first machine learning model that reduces the effects of overcorrection in the sharpening process; An image processing method characterized by comprising the step of combining at least one of the captured image and the intermediate correction data with the corrected image.

2. The image processing method according to Claim 1, characterized in that, in the synthesis step, the image image and the corrected image are synthesized based on a sharpening component obtained based on the image image and the intermediate correction data.

3. The image processing method according to claim 1 or 2, characterized in that, in the step of acquiring the intermediate correction data, the sharpening process is performed using a second machine learning model.

4. The image processing method according to claim 1 or 2, characterized in that, in the step of acquiring the intermediate correction data, the sharpening process is performed using a Wiener filter.

5. The image processing method according to any one of claims 1 to 4, characterized in that the intermediate correction data is at least one of an intermediate correction image or a feature map.

6. The image processing method according to claim 3, characterized in that the second machine learning model is common to a plurality of optical properties.

7. The image processing method according to claim 3, characterized in that, in the step of acquiring the intermediate correction data, the shooting conditions of the optical system are input to the second machine learning model.

8. The image processing method according to claim 3, characterized in that, in the step of acquiring the intermediate correction data, information on the screen position of the captured image is input to the second machine learning model.

9. The image processing method according to claim 8, characterized in that the information on the screen position of the captured image includes information on the image height of the captured image.

10. A program characterized by causing a computer to execute the image processing method described in any one of Claims 1 to 9.

11. A generation means for acquiring intermediate correction data obtained by performing a sharpening process on the captured image based on the optical characteristics of the optical system used to capture the captured image, A learning means for acquiring a corrected image from the aforementioned intermediate correction data using a first machine learning model that reduces the effects of overcorrection in the sharpening process, An image processing apparatus characterized by having correction means for combining at least one of the captured image and the intermediate correction data with the corrected image.