Training method, image processing method, method for manufacturing a trained model, program, and training device.

JP7899291B2Active Publication Date: 2026-08-03CANON KK
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
CANON KK
Filing Date
2024-12-27
Publication Date
2026-08-03

Smart Images

  • Figure 0007899291000003
    Figure 0007899291000003
  • Figure 0007899291000004
    Figure 0007899291000004
  • Figure 0007899291000005
    Figure 0007899291000005
Patent Text Reader

Abstract

To provide an image processing method that can suppress generation of images with low shaping accuracy in shaping of defocus blur that uses a machine learning model.SOLUTION: An image processing method includes: a first step of generating a first image by adding a first blur to an original image, and generating a second image by adding a second blur to the original image; and a second step of generating an output image by inputting the first image to a machine learning model, and performing learning of the machine learning model on the basis of the output image and the second image. The first blur and the second blur are blurs of different shapes. Optical performance corresponding to the first blur is higher than that corresponding to the second blur at a specific frequency.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing method for shaping the blur due to defocus on a captured image to obtain an image with a good blurring effect.

Background Art

[0002] In Patent Document 1, the pupil of an optical system is divided into a plurality of divided pupils, a plurality of parallax images observing a subject space from each divided pupil are captured, and by adjusting the weights when synthesizing the plurality of parallax images, a method for shaping the shape of the blur due to defocus (defocus blur) is disclosed.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the method disclosed in Patent Document 1, since the weights of each divided pupil are adjusted to synthesize a plurality of parallax images, it is impossible to reproduce the defocus blur corresponding to a pupil larger than the pupil of the optical system. That is, in the method disclosed in Patent Document 1, it is impossible to fill in the missing parts of the defocus blur due to vignetting. Further, when the weights of the synthesis of a plurality of parallax images become non-uniform, noise increases. In addition, since the structure of double-line blur and the like is fine, in order to reduce their effects, it is necessary to finely divide the pupil of the optical system. In this case, a decrease in the spatial resolution and an increase in noise occur in each parallax image.

[0005] On the other hand, to address the aforementioned issues, methods for correcting defocus blur using machine learning models such as CNNs (Convolutional Neural Networks) are known. However, methods using machine learning models to correct defocus blur do not always allow for correction to any desired shape, and the correction accuracy may decrease.

[0006] Therefore, the present invention aims to provide an image processing method that can suppress the generation of images with low processing accuracy when processing defocus blur using a machine learning model. [Means for solving the problem]

[0007] A training method as one aspect of the present invention comprises a first step of generating a first image by adding a first blur to an original image, and a second step of generating a second image by adding a second blur to the original image, and a second step of generating an output image by inputting the first image into a machine learning model, and updating the weights of the machine learning model based on the output image and the second image, wherein the first blur and the second blur have different shapes. Furthermore, the optical performance corresponding to the first blur is higher than the optical performance corresponding to the second blur at a specific frequency. .

[0008] Other objects and features of the present invention are described in the following examples. [Effects of the Invention]

[0009] According to the present invention, it is possible to provide an image processing method that can suppress the generation of images with low processing accuracy when processing defocus blur using a machine learning model. [Brief explanation of the drawing]

[0010] [Figure 1] This is a flowchart showing the method for determining the second defocus blur in Example 1. [Figure 2] This is a block diagram of the image processing system in Example 1. [Figure 3]It is an external view of the image processing system in Example 1. [Figure 4] It is a flowchart regarding the method for generating learning data in Example 1. [Figure 5] It is a flowchart regarding the learning of weights in Example 1. [Figure 6] It is a flowchart regarding the generation of an estimated image in Example 1. [Figure 7] It is a configuration diagram of the machine learning model in Example 1. [Figure 8] It is an explanatory diagram of the optical performance of the first defocus blur in Example 1. [Figure 9] It is an explanatory diagram of the optical performance of the second defocus blur in Example 1. [Figure 10] It is a configuration diagram of the machine learning model in Example 2. [Figure 11] It is a block diagram of the image processing system in Example 2. [Figure 12] It is an external view of the image processing system in Example 2. [Figure 13] It is a flowchart regarding the method for generating learning data in Example 2. [Figure 14] It is a flowchart regarding the learning of weights in Example 2. [Figure 15] It is a flowchart regarding the generation of an estimated image in Example 2. [Figure 16] It is an explanatory diagram of the optical performance of the first defocus blur in Example 2. [Figure 17] It is a block diagram of the image processing system in Example 3. [Figure 18] It is an external view of the image processing system in Example 3. <000​​​​​​​It is a diagram showing the result of defocus blur shaping in the comparative example.

Mode for Carrying Out the Invention

[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In each figure, the same members are denoted by the same reference numerals, and redundant descriptions are omitted.

[0012] Before giving a specific description of each embodiment, the gist of the present invention will be described. The present invention shapes the defocus blur of a captured image using a machine learning model. The machine learning model includes, for example, a neural network, genetic programming, or a Bayesian network. The neural network includes CNN, GAN (Generative Adversarial Network), RNN (Recurrent Neural Network), and the like. The shaping of defocus blur includes shaping from two-line blur to Gaussian blur or spherical blur.

[0013] Refer to Figure 20 to explain double-line bokeh, ball bokeh, and Gaussian bokeh. Figure 20 shows the point spread function (PSF) at the defocus distance. Figure 20(A) shows the PSF for double-line bokeh, Figure 20(B) shows the PSF for ball bokeh, and Figure 20(C) shows the PSF for Gaussian bokeh. In Figures 20(A) to (C), the horizontal axis represents spatial coordinates (position), and the vertical axis represents intensity. As shown in Figure 20(A), double-line bokeh has a PSF with separated peaks. When the PSF at the defocus distance has the shape shown in Figure 20(A), the subject, which is originally a single line, appears to be blurred twice when defocused. As shown in Figure 20(B), ball bokeh has a PSF with flat intensity. As shown in Figure 20(C), Gaussian bokeh has a PSF with a Gaussian distribution. Other types of defocus blur that can be corrected include, for example, defocus blur caused by vignetting and ring-shaped defocus blur caused by pupil occlusion, such as with catadioptric lenses. However, the shape of the defocus blur that can be corrected is not limited to these examples.

[0014] The defocus blur shaping in this invention does not involve adding defocus blur to a pan-focus image with a deep depth of field to reproduce an image with a shallow depth of field. Rather, it requires shaping an already defocused subject to a desired level of defocus blur, that is, applying a defocus blur that satisfies the difference between the already existing defocus blur and the desired defocus blur, which necessitates more advanced processing.

[0015] In image shaping using a machine learning model, if the defocus blur (first blur) in the training image (first image) does not have higher optical performance frequencies than the defocus blur (second blur) in the ground truth image (second image), the shaping accuracy may decrease. In other words, if the first blur does not have higher optical performance at a specific frequency than the second blur, the shaping accuracy may decrease. The ground truth image is the image that we aim to obtain as the output image of the machine learning model, and the training image is the input image to the machine learning model.

[0016] Figure 21 shows the results of correcting defocus blur in the comparative example. It shows the results of correcting defocus blur using weights learned from training images that do not have higher optical performance frequencies than the ground truth image. In Figure 21, the horizontal axis represents spatial coordinates (position), and the vertical axis represents intensity. The dotted line 601 is the defocus blur before correction, the solid line 602 is the ground truth defocus blur, and the dashed line 603 is the defocus blur after correction. There is a large difference between the ground truth defocus blur (solid line 602) and the defocus blur after correction (dashed line 603), indicating low correction accuracy.

[0017] Furthermore, defocus blur changes depending on the state of the optical system (lens state), image height and azimuth, and the amount of defocus used to capture the image. Therefore, if these changes are not taken into consideration and an appropriate defocus blur is not learned as the ground truth image, the learning difficulty may increase and the accuracy of defocus blur correction may decrease. Accordingly, each embodiment provides an image processing method that corrects the defocus blur of the captured image by setting an appropriate defocus blur as the ground truth image so that the defocus blur of the training image has higher optical performance than the defocus blur of the ground truth image at a specific frequency.

[0018] In the following, the stage of learning the weights of a machine learning model will be referred to as the learning phase, and the stage of using the machine learning model with the learned weights to correct the defocus blur will be referred to as the estimation phase. [Examples]

[0019] First, the image processing system 100 in Embodiment 1 of the present invention will be described. Figure 2 is a block diagram of the image processing system 100. Figure 3 is an external view of the image processing system 100.

[0020] The image processing system 100 includes a learning device (image processing device) 101, an imaging device 102, an image estimation device 103, a display device 104, a recording medium 105, an output device 106, and a network 107. The learning device 101 includes a storage unit 101a, an acquisition unit 101b, a generation unit (generation means) 101c, and an update unit (learning means) 101d, and learns the weights of a machine learning model used for defocus blur correction. Details regarding weight learning and defocus blur correction processing using the weights will be described later.

[0021] The imaging device 102 has an optical system 102a and an image sensor 102b, and captures an image of the subject space to acquire an image. The optical system 102a collects light incident from the subject space and forms an optical image (subject image). The image sensor 102b acquires an image by photoelectric conversion of the optical image. The image sensor 102b is, for example, a CCD (Charge Coupled Device) sensor or a CMOS (Complementary Metal-Oxide Semiconductor) sensor.

[0022] The image estimation device 103 includes a storage unit 103a, an acquisition unit 103b, a blur correction unit 103c, and a generation unit 103d. The image estimation device 103 generates an estimated image with defocus blur corrected from an captured image (or at least a part thereof) captured by the imaging device 102. For defocus blur correction, a machine learning model using weights learned by the learning device 101 is used. The learning device 101 and the image estimation device 103 are connected by a network 107, and the image estimation device 103 reads the learned weight information from the learning device 101 during or before defocus blur correction. The estimated image is output to at least one of the display device 104, the recording medium 105, or the output device 106.

[0023] The display device 104 is, for example, a liquid crystal display or a projector. The user can perform editing work while checking the image in progress via the display device 104. Details of the user interface during editing work will be described later. The recording medium 105 is, for example, a semiconductor memory, a hard disk, or a server on a network, and stores the estimated image. The output device 106 is, for example, a printer.

[0024] Next, with reference to Figure 4, the generation of training data performed by the learning device 101 will be described. Figure 4 is a flowchart of the method for generating training data in this embodiment.

[0025] First, in step S101, the acquisition unit 101b acquires the source image. The source image may be one or multiple images. The source image may be a real-life image or a CG (Computer Graphics) image. In the following steps, the source image is subjected to a first defocus (first blur) and a second defocus blur (second blur), respectively, to generate a training image (first image) and a ground truth image (second image). For this reason, it is desirable that the source image has edges with various intensities and directions, as well as textures, gradients, flat areas, etc., so that the shape of the defocus blur can be correctly shaped for various subjects.

[0026] Preferably, the original image has a signal value higher than the luminance saturation value of the image sensor 102b. This is because, even with real subjects, when image is captured by the imaging device 102 under specific exposure conditions, there are subjects that do not fall within the luminance saturation value. By applying defocus blur to the ground truth image and the training image, and then clipping them by the luminance saturation value of the image sensor 102b, it is possible to reproduce subjects that do not fall within the actual luminance saturation value.

[0027] Next, in step S102, the generation unit 101c generates a first defocus blur (first blur) and stores it in the memory unit 101a. The first defocus blur is the defocus blur to be shaped. The first defocus blur changes according to the lens state and image height / azimuth of the optical system 102a, as well as the amount of defocus. The lens state refers to the state of the optical system 102a, and includes, but is not limited to, the state of zoom, aperture, and focus distance. For example, the defocus blur at 5m when focused at 2.5m will differ depending on whether the aperture is set to F1.4 or F2.8. Also, even with the same lens state, the defocus blur at 5m will differ depending on whether the image height is 80% / azimuth 0 degrees or 50% / azimuth 45 degrees.

[0028] Next, in step S103, the generation unit 101c generates a second defocus blur (second blur) and stores it in the storage unit 101a. The second defocus blur is the defocus blur after shaping. The method for determining the second defocus blur in relation to the first defocus blur in a specific lens state will be described in detail with reference to Figure 1. Figure 1 is a flowchart of the method for determining the second defocus blur in this embodiment.

[0029] In this embodiment, the second defocus blur is determined to have the same shape as the first defocus blur, which changes depending on the image height and azimuth. This makes it possible to reshape the defocus blur to have the same shape across the entire field of view, even if various defocus blur shapes are included in the captured image.

[0030] Furthermore, since the first defocus blur changes with each amount of defocus, the second defocus blur is determined for each amount of defocus. This allows us to determine the appropriate defocus blur for the ground truth image according to the amount of defocus.

[0031] Furthermore, since the first defocus blur changes with each lens state of the optical system 102a, the second defocus blur is determined for each lens state. This makes it possible to determine an appropriate defocus blur as the ground truth image according to each lens state. In this embodiment, the second defocus blur is determined in one way for each amount of defocus in a specific lens state, but it may be determined in two or more ways. The case in which the second defocus blur is determined in two or more ways will be described later in Example 2. When the second defocus blur is determined in one way, the number of training images and training time can be reduced.

[0032] First, in step S1031, the acquisition unit 101b acquires the optical performance of the first defocused blur generated in step S102. In this embodiment, optical performance includes, but is not limited to, the modulation transfer function (MTF) or the optical transfer function (OTF).

[0033] Next, in step S1032, the acquisition unit 101b acquires the average value of the optical performance of the first defocus blur acquired in step S1031. In this embodiment, the acquisition unit 101b acquires the average value from half the Nyquist frequency of the MTF to the Nyquist frequency.

[0034] In this embodiment, the first defocus blur has a higher average optical performance across multiple frequencies than the second defocus blur. Here, the multiple frequencies are in the range from half the Nyquist frequency of the training image to the Nyquist frequency. Note that any interval can be selected as the calculation interval for the average value used to compare optical performance, and an appropriate calculation interval can be set according to the optical performance of the defocus blur. For example, the multiple frequencies may be in the range from frequency 0 to half the Nyquist frequency of the training image.

[0035] In this embodiment, half of the Nyquist frequency is used as f. N / 2, the Nyquist frequency is f N When the number of data points is n, the mean value M is calculated by the following formula (1).

[0036]

number

[0037] Figure 8 is an explanatory diagram of the optical performance of the first defocus blur in this embodiment, showing the MTF of the first defocus blur at 50% image height / 0 degrees azimuth and 100% image height / 0 degrees azimuth for specific lens states and defocus amounts. In Figure 8, the horizontal axis represents spatial frequency and the vertical axis represents MTF. On-axis, 50% image height / 0 degrees azimuth and 100% image height / 0 degrees azimuth correspond to dotted line 1001, dashed line 1002, and dashed line 1003, respectively. The Nyquist frequency of the image sensor 102b used for imaging is 90 [lp / mm]. For each MTF, the average value from half the Nyquist frequency to the Nyquist frequency is obtained. Similarly, the average value is obtained when the defocus amount and image height / azimuth are changed.

[0038] Next, in step S1033, the generation unit 101c generates multiple patterns of defocus blur as candidates for the second defocus blur. The generated second defocus blurs are variations in shape. Shape refers to the type and size of the defocus blur. Type refers to differences such as Gaussian blur, spherical blur, and double-line blur, which are due to differences in the intensity distribution of the PSF. Size refers to the range in which the PSF has intensity. There are no restrictions on the second defocus blurs to be generated as candidates.

[0039] Figure 9 is an explanatory diagram of the optical performance of the second defocus blur in this embodiment. In Figure 9, the horizontal axis represents spatial frequency and the vertical axis represents MTF. The solid line 1004, dotted line 1005, dashed line 1006, and dashed line 1007 in Figure 9 each represent the MTF of Gaussian blur of different magnitudes. Similar to step S1032, the acquisition unit 101b acquires the average value from half the Nyquist frequency to the Nyquist frequency for each MTF. Alternatively, multiple candidates for the second defocus blur may be generated in advance, and the average value from half the Nyquist frequency to the Nyquist frequency for each MTF may be acquired and stored in the storage unit 101a.

[0040] Next, in step S1034, the acquisition unit 101b compares the average value of the first defocus blur acquired in step S1032 with the average value of the second defocus blur acquired in step S1033, and determines and acquires the second defocus blur. The average values ​​of the dotted line 1001, dashed line 1002, and dashed line 1003 in Figure 8 are approximately 0.11, 0.1, and 0.07, respectively. The average values ​​of the solid line 1004, dotted line 1005, dashed line 1006, and dashed line 1007 in Figure 9 are approximately 0.25, 0.03, 0, and 0, respectively. Therefore, one of the dotted line 1005, dashed line 1006, or dashed line 1007 is determined to be the second defocus blur.

[0041] In this embodiment, multiple training images and ground truth images are generated. There are at least two types of first defocus blurs applied to generate multiple training images. The same second defocus blur is applied to generate multiple ground truth images. However, this embodiment is not limited to this, and there may be at least two types of second defocus blurs.

[0042] The above describes the method for determining the second defocus blur in this embodiment. In this embodiment, as described above, the second defocus blur is determined relative to the first defocus blur in a specific lens state. Similarly, for other lens states, the average value is obtained from the optical performance of the first defocus blur, and the second defocus blur is determined.

[0043] In this embodiment, it is preferable that the second defocus blur is determined according to the size of the image sensor 102b. The maximum image height of the captured image is determined according to the size of the image sensor 102b. For example, even when using the same optical system 102a, the maximum image height that can be captured is smaller for APS-C size than for full-frame. That is, the maximum image height of the first defocus blur used for comparison becomes smaller.

[0044] Next, in step S104 of Figure 4, the generation unit 101c generates a training image and stores it in the storage unit 101a. The training image is an image (first image) obtained by performing an imaging simulation by applying the first defocus blur, which is the target of shaping, to the original image. In order to accommodate all types of captured images, it is preferable to apply defocus blur corresponding to various amounts of defocus. The defocus blur can be applied by convolving the PSF onto the original image or by taking the product of the frequency characteristics of the original image and the OTF. Furthermore, since it is desirable that the image does not change before and after defocus blur shaping at the focus plane, training images and ground truth images without applying defocus blur are also generated.

[0045] Next, in step S105, the generation unit 101c generates a ground truth image and stores it in the storage unit 101a. The ground truth image is the image (second image) obtained by applying a second defocus blur, after reshaping, to the original image and performing an imaging simulation. The ground truth image and the training image may be either undeveloped RAW images or developed images. The order in which the training image and the ground truth image are generated may also be reversed.

[0046] The above explains how to generate training data. Alternatively, you can extract a subregion with a predetermined number of pixels from the training image and the ground truth image and use that for training.

[0047] Next, we will explain the weight learning (learning phase) with reference to Figure 5. Figure 5 is a flowchart of the weight learning process in this embodiment. In this embodiment, a CNN is used as the machine learning model, but other models can be similarly applied.

[0048] First, in step S111, the acquisition unit 101b acquires one or more sets of ground truth images (second images) and training images (first images) from the storage unit 101a. The training images are the input data for the learning phase of the CNN. Next, in step S112, the generation unit 101c inputs the training images into the machine learning model (CNN) and generates output images.

[0049] Here, with reference to Figure 7, the generation of the output image in this embodiment will be explained. Figure 7 is a diagram of the machine learning model in this embodiment. The training image 201 may be grayscale or may have multiple channel components. In this embodiment, CNN202 has one or more convolutional layers or fully connected layers. In the first training, the weights of CNN202 (the values ​​of each element of the filter and the bias) are generated by random numbers.

[0050] Next, in step S113 of Figure 5, the update unit 101d compares the output image with the ground truth image (second image) and updates the CNN weights based on the difference (error) between the output image and the ground truth image. In this embodiment, the Euclidean norm of the difference in signal values ​​between the output image and the ground truth image is used as the loss function. However, the loss function is not limited to this. If multiple sets of training images and ground truth images are obtained in step S111, the value of the loss function is calculated for each set. The weights are updated using the calculated loss function value by methods such as backpropagation.

[0051] Next, in step S114, the update unit 101d determines whether or not weight learning is complete. Learning completion can be determined by whether the number of iterations of learning (weight update) has reached a predetermined number, or whether the amount of change in weight at the time of update is less than a predetermined value. If it is determined that learning is not complete, the process returns to step S111, and the acquisition unit 101b acquires one or more sets of new training images and correct images. On the other hand, if it is determined that learning is complete, the update unit 101d terminates learning and stores the weight information in the storage unit 101a.

[0052] Next, with reference to Figure 6, we will explain the defocus blur correction (estimation phase) of the captured image performed by the image estimation device 103. Figure 6 is a flowchart of the generation of the estimated image in this embodiment.

[0053] First, in step S121, the acquisition unit 103b acquires the captured image and weight information. The captured image to be acquired may be only a part of the entire captured image. The weight information is read in advance from the storage unit 101a and stored in the storage unit 103a.

[0054] Next, in step S122, the blur correction unit 103c inputs the captured image into the CNN and generates an estimated image. The estimated image is an image in which the defocus blur of the captured image has been corrected. Similar to the training process, the estimated image is generated using the machine learning model (CNN) shown in Figure 7. The acquired trained weights are used in the CNN. The generated multiple estimated images are stored in the storage unit 103a.

[0055] According to this embodiment, in the process of correcting defocus blur using a machine learning model, it is possible to learn the model in a way that suppresses the generation of images with low correction accuracy, and thus correct defocus blur.

[0056] Next, preferred conditions for enhancing the effectiveness of this embodiment will be described. It is desirable that the input data further include a luminance saturation map. The luminance saturation map indicates the luminance saturation pixel region of the image and is the same size as the image. In the learning phase, a luminance saturation map is generated from the training image. In the estimation phase, a luminance saturation map is generated from the captured image. Because the luminance saturation region contains false edges that differ from the structure of the subject due to luminance saturation, it is difficult for machine learning models to distinguish these from those with edges, such as defocus blur with high-frequency components or the focal position. The luminance saturation map allows the machine learning model to distinguish between the luminance saturation region, defocus blur with high-frequency components, and the focal position, enabling highly accurate image shaping. Note that defocus blur with high-frequency components is likely to occur when PSF with sharp peaks, such as double-line blur, is applied.

[0057] Preferably, the input data also includes an image (Image A) obtained by imaging the subject space based on the light beam passing through a partial pupil (second pupil), which is part of the pupil of the optical system 102a. The image obtained by imaging the subject space based on the light beam passing through the pupil (first pupil) of the optical system 102a (Image A+B) and Image A differ in the degree of defocus blur. Since the second pupil is smaller than the first pupil, the defocus blur of Image A is smaller than that of Image A+B. By using both Image A+B and Image A, it is possible to distinguish between defocus blur in the image and the structure of the subject. That is, if there is a blurred area in the image that lacks high-frequency information, it is possible to distinguish whether this area is blurred because it is defocused, or whether it is in focus but appears blurred because the subject lacks high-frequency information. Alternatively, instead of Image A, an image (Image B) obtained by imaging the subject space based on the light beam passing through a partial pupil (third pupil), which is part of the pupil of the optical system 102a, may be used.

[0058] The input data should preferably also include a state map. The state map represents the state of the optical system 102a during imaging using (Z,F,D). In (Z,F,D), Z corresponds to zoom, F to aperture, and D to focus distance.

[0059] The input data should preferably also include a location map. The location map is a map that shows the image plane coordinates for each pixel of the image. The location map may be in a Cartesian coordinate system or a polar coordinate system (corresponding to image height and azimuth).

[0060] Defocus blur varies depending on the lens state and image height / azimuth. Since CNNs are trained to average out all defocus blur in the training data, the accuracy of shaping each defocus blur of different shapes decreases. Therefore, by inputting statemaps and position maps into the machine learning model, the machine learning model can identify the PSF acting on the captured image. As a result, in the training phase, even if the training images contain defocus blur of various shapes, the machine learning model learns weights that perform different shaping for each shape of defocus blur, rather than weights that average out all of them. This enables highly accurate shaping for each defocus blur in the estimation phase. Therefore, the decrease in shaping accuracy can be suppressed, and training data that can handle defocus blur of various shapes can be trained all at once. [Examples]

[0061] Next, the image processing system in Embodiment 2 of the present invention will be described. Figure 11 is a block diagram of the image processing system 300. Figure 12 is an external view of the image processing system 300.

[0062] The image processing system 300 includes a learning device 301, an imaging device 302, an image estimation device 303, and networks 304 and 305. The learning device 301 includes a storage unit 301a, an acquisition unit 301b, a generation unit 301c, and an update unit 301d, and learns the weights of a machine learning model used for defocus blur correction. Details regarding weight learning and defocus blur correction using the weights will be described later.

[0063] The imaging device 302 includes an optical system 302a, an image sensor 302b, an acquisition unit 302c, a recording medium 302d, and a system controller 302e. The optical system 302a collects light incident from the subject space and forms an optical image (subject image). The image sensor 302b converts the optical image into an electrical signal by photoelectric conversion and generates an image.

[0064] The image estimation device 303 includes a storage unit 303a, a blur correction unit 303b, an acquisition unit 303c, and a generation unit 303d. The image estimation device 303 generates an estimated image with defocus blur corrected from an image captured by the imaging device 302 (or at least a part thereof). The learned weight information learned by the learning device 301 is used to generate the estimated image. The weight information is stored in the storage unit 303a. The acquisition unit 302c acquires the estimated image, and the recording medium 302d saves the estimated image. The system controller 302e controls a series of operations of the imaging device 302.

[0065] Next, with reference to Figure 13, the generation of training data performed by the learning device 301 will be described. Figure 13 is a flowchart of the method for generating training data in this embodiment.

[0066] First, in step S201, the acquisition unit 301b acquires the original image. Next, in step S202, the generation unit 301c generates a first defocused blur (first blur) and stores it in the storage unit 301a. Next, in step S203, the generation unit 301c generates a second defocused blur (second blur) and stores it in the storage unit 301a.

[0067] Figure 16 is an explanatory diagram of the optical performance of the first defocus blur in this embodiment, showing the MTF of the first defocus blur at 50% image height and 0 degrees azimuth, and at 100% image height and 0 degrees azimuth, for specific lens states and defocus amounts. In Figure 16, the horizontal axis represents spatial frequency and the vertical axis represents MTF. On-axis, 50% image height and 0 degrees azimuth, and 100% image height and 0 degrees azimuth correspond to dotted line 1008, dashed line 1009, and dashed line 1010, respectively. The Nyquist frequency of the image sensor 302b used for imaging is 90 [lp / mm].

[0068] The first defocus blur in this embodiment is a defocus blur with lower optical performance than that in Embodiment 1. For example, when the focus distance is 2.5m and the defocus amount is 5m, the on-axis defocus blur will be a defocus blur with lower optical performance when using an optical system with a long focal length or when shooting with a small F-number (aperture value). Therefore, it is difficult to compare the average value from half the Nyquist frequency to the Nyquist frequency with the second defocus blur shown in Figure 9. In this embodiment, therefore, the average value from frequency 0 to half the Nyquist frequency is compared. Frequency 0 is f0, and half the Nyquist frequency is f N / 2 If the number of data points is n, the mean value M can be calculated using the following formula (2).

[0069]

number

[0070] Thus, there is no fixed calculation interval for the average value used to compare optical performance, and an appropriate calculation interval may be selected according to the optical performance of the defocus blur. Also, the second defocus blur is determined in the same way as the flowchart in Figure 1, but in this embodiment, the optical performance of the determined defocus blur is used as the upper limit, and two or more types of the second defocus blur are determined. This makes it possible to reshape the shape of the defocus blur in the captured image into various shapes. In this embodiment, among the second defocus blurs, the average value of the dashed line 1007 is lower than the average value of the first defocus blur. Therefore, the dashed line 1007 is used as the upper limit of the optical performance in the second defocus blur, and two or more types of the second defocus blur are determined.

[0071] Next, in step S204 of Figure 13, the generation unit 301c generates shape specification information and stores it in the storage unit 301a. The shape specification information is information regarding the size or type of defocus blur after shaping. In this embodiment, the input data input to the machine learning model includes the captured image and information specifying the shape of the defocus blur after shaping (shape specification information). During the training of the machine learning model, the shape specification information is input along with the training image, and a ground truth image corresponding to the shape specification information is set, allowing the machine learning model to distinguish and learn from multiple ground truth images having different defocus blur shapes for a single training image. That is, even if the ground truth image contains defocus blur of various shapes, the model learns weights that perform different shaping for each defocus blur shape, rather than weights that shape the image to an average shape of those defocus blurs. Therefore, it is possible to learn training data containing defocus blur of various shapes with high accuracy all at once.

[0072] Specifying the size is equivalent to virtually changing the F-number (aperture value) of the optical system 302a. Changing the F-number changes the size of the pupil of the optical system 302a, and therefore changes the size of the defocus blur. Through image processing to correct the defocus blur, it is also possible to change the F-number of the optical system 302a from the captured image to a value that is not physically possible.

[0073] Specifying a type of bokeh is equivalent to virtually changing the optical system 302a to a different lens configuration. The types of defocus bokeh, such as double-line bokeh, spherical bokeh, and Gaussian bokeh, depend on the pupil function determined by the lens configuration of the optical system 302a.

[0074] In other words, specifying the size or type of defocused blur after shaping is equivalent to specifying virtual lens parameters. More specifically, specifying the f-number is equivalent to changing the spread of the pupil function. Furthermore, specifying types such as double-line blur or spherical blur is equivalent to changing the amplitude or phase of the pupil function.

[0075] Next, in step S205, the generation unit 301c generates a training image and stores it in the storage unit 301a. Then, in step S206, the generation unit 301c generates multiple correct images corresponding to each of the multiple shape specification information for a single training image and stores them in the storage unit 301a. The correct images and training images may be undeveloped RAW images or developed images. Also, the order in which the training image, correct images, and shape specification information are generated may be changed.

[0076] Next, with reference to Figure 14, we will explain the weight learning (learning phase) performed by the learning device 301. Figure 4 is a flowchart of the weight learning in this embodiment. In this embodiment, a GAN is used as the machine learning model, but other models can be similarly applied. Note that in this embodiment, explanations of parts that are the same as in Embodiment 1 will be omitted. A GAN is a generative adversarial network consisting of a generator that generates images and a discriminator that identifies the generated images.

[0077] First, in step S211, the acquisition unit 301b acquires one or more sets of correct images and training input data from the storage unit 301a. The generation of correct images and training images is the same as in Example 1.

[0078] Figure 10 is a diagram of the configuration of the machine learning model (GAN) in this embodiment. The concatenation layer 406 concatenates the training image 401 and the shape specification information 402 in a predetermined order in the channel direction to generate the training input data 403.

[0079] Next, in step S212 of Figure 14, the generation unit 301c inputs the training input data 403 to the generator 407 to generate the output image 404. The generator 407 is, for example, a CNN. Next, in step S213, the update unit 301d updates the weights of the generator 407 based on the error between the output image 404 and the ground truth image 405. The Euclidean norm of the difference at each pixel is used as the loss function. Next, in step S214, the update unit 301d determines whether the first learning is complete or not. If the first learning is not complete, it returns to step S211. On the other hand, if the first learning is complete, it proceeds to step S221 and performs the second learning.

[0080] In step S221, the acquisition unit 301b acquires one or more sets of correct images 405 and training input data 403 from the storage unit 301a, similar to step S211. Next, in step S222, the generation unit 301c inputs the training input data 403 to the generator 407 to generate the output image 404, similar to step S212. Next, in step S223, the update unit 301d updates the weights of the classifier 408 based on the output image 404 and the correct images 405. The classifier 408 identifies whether the input image is a fake image generated by the generator 407 or a real image which is the correct image 405. The output image 404 or the correct image 405 is input to the classifier 408 to generate an identification label (fake or real). The weights of the classifier 408 are updated based on the error between the classified label and the ground truth label (output image 404 is fake, ground truth image 405 is real). A sigmoid cross entropy function is used as the loss function, but other loss functions may also be used.

[0081] Next, in step S224, the update unit 301d updates the weights of the generator 407 based on the output image 404 and the ground truth image 405. The loss function is the Euclidean norm and a weighted sum of the following two terms. The first term is called Content Loss, which is obtained by converting the output image 404 and the ground truth image 405 into feature maps and taking the Euclidean norm of the differences for each element. By adding the differences in the feature maps to the loss function, the more abstract properties of the output image 404 can be brought closer to those of the ground truth image 405. The second term is called Adversarial Loss, which is the sigmoid cross entropy of the classification labels obtained by inputting the output image 404 into the classifier 408. By training the classifier 408 to distinguish between real and false, it is possible to obtain an output image 404 that appears more subjectively like the ground truth image 405.

[0082] Next, in step S225, the update unit 301d determines whether the second learning is complete. If the second learning is not complete, the process returns to step S221. On the other hand, if the second learning is complete, the update unit 301d stores the learned weight information of the generator 407 in the storage unit 301a.

[0083] Next, with reference to Figure 15, the defocus blur correction (estimation phase) performed by the image estimation device 303 will be described. Figure 15 is a flowchart of the generation of the estimated image in this embodiment.

[0084] First, in step S231, the acquisition unit 303c acquires the captured image (or at least a part thereof). Next, in step S232, the generation unit 303d generates shape specification information. Next, in step S233, the acquisition unit 303c acquires the input data and the learned weight information. The input data includes the captured image and the shape specification information. The weight information is read in advance from the storage unit 301a and stored in the storage unit 303a. Next, in step S234, the blur correction unit 303b inputs the input data to the generator 407 and generates an estimated image.

[0085] According to this embodiment, in the process of correcting defocus blur using a machine learning model, it is possible to learn the model in a way that suppresses the generation of images with low correction accuracy, and thus correct defocus blur. [Examples]

[0086] Next, the image processing system in Embodiment 3 of the present invention will be described. The image processing system in this embodiment differs from Embodiments 1 and 2 in that it has a processing unit (computer) that transmits the captured image to be processed to an image estimation device and receives the processed output image from the image estimation device.

[0087] Figure 17 is a block diagram of the image processing system 500. Figure 18 is an external view of the image processing system 500. The image processing system 500 includes a learning device 501, an imaging device 502, a lens device 503, a control device (first device) 504, an image estimation device (second device) 505, and networks 506 and 507. The learning device 501 and the image estimation device 505 are, for example, servers. The control device 504 is a personal computer. The learning device 501 has a storage unit 501a, an acquisition unit 501b, a generation unit 501c, and an update unit 501d, and learns the weights of a machine learning model that corrects the defocus blur of images captured using the imaging device 502. Further details regarding the learning process will be described later.

[0088] The imaging device 502 has an image sensor 502a, which converts the optical image formed by the lens device 503 into an optical image to acquire an image. The lens device 503 and the imaging device 502 are detachable and can be combined with multiple types of each other. The control device 504 has a communication unit 504a, a display unit 504b, and a storage unit 504c, and controls the processing to be performed on the image acquired from the imaging device 502, which is connected by wire or wireless, according to the user's operation. Alternatively, the image captured by the imaging device 502 may be stored in the storage unit 504c in advance, and the image may be read out.

[0089] The image estimation device 505 includes a communication unit 505a, an acquisition unit 505b, a storage unit 505c, and a shaping unit 505d. The image estimation device 505 performs defocus blur shaping processing on captured images in response to a request from a control device 504 connected via a network 506. The image estimation device 505 acquires learned weight information from a learning device 501 connected via a network 507, either during or before defocus blur shaping, and uses it for defocus blur shaping of captured images. The estimated image after defocus blur shaping is transmitted back to the control device 504, stored in the storage unit 504c, and displayed on the display unit 504b. Note that the generation of learning data and learning of weights (learning phase) performed by the learning device 501 are the same as in Example 1, so their explanation is omitted.

[0090] Next, with reference to Figure 19, the defocus blur correction (estimation phase) of the captured image performed by the control device 504 and the image estimation device 505 will be described. Figure 19 is a flowchart relating to the generation of the estimated image in this embodiment.

[0091] First, in step S301, the communication unit (transmitting means) 504a of the control device (first device) 504 transmits the captured image and processing request to the image estimation device (second device) 505. Next, in step S302, the communication unit (receiving means) 505a receives and acquires the transmitted captured image and processing request. Next, in step S303, the acquisition unit 505b acquires the learned weight information from the storage unit 505c. The weight information is read in advance from the storage unit 501a and stored in the storage unit 505c. Next, in step S304, the shaping unit 505d inputs the input data into the CNN and generates an estimated image with the defocus blur of the captured image shaped. Next, in step S305, the communication unit 505a transmits the estimated image to the control device 504. Next, in step S306, the communication unit 504a acquires the transmitted estimated image and saves it to the storage unit 504c.

[0092] According to this embodiment, in the process of correcting defocus blur using a machine learning model, it is possible to learn the model in a way that suppresses the generation of images with low correction accuracy, and thus correct defocus blur.

[0093] (Other examples) The present invention can also be realized by supplying a program that implements one or more of the functions of the above embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0094] Each embodiment of the image processing method includes the steps of generating a first image (training image) and a second image (ground truth image) from the original image, and inputting the first image into a machine learning model and training the machine learning model by comparing the output image with the second image. The first image is the original image with a first blur (first defocus blur) added to it. The second image is the original image with a second blur (second defocus blur) added to it. Furthermore, the first blur has higher optical performance than the second blur at certain frequencies.

[0095] According to each embodiment, it is possible to set an appropriate defocus blur in the ground truth image so that the defocus blur in the ground truth image does not have higher optical performance frequencies than the defocus blur in the training image, thereby correcting the defocus blur in the captured image. Therefore, according to each embodiment, it is possible to provide an image processing method, a method for manufacturing a trained model, a program, and an image processing device that can suppress the generation of images with low correction accuracy when correcting defocus blur using a machine learning model.

[0096] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of its essence. [Explanation of Symbols]

[0097] 101 Learning device (image processing device) 101c Generation unit (generation means) 101d Update section (learning method)

Claims

1. A first step of generating a first image by adding a first blur to the original image, and a second step of generating a second image by adding a second blur to the original image, The process includes a second step of inputting the first image into a machine learning model to generate an output image, and updating the weights of the machine learning model based on the output image and the second image. The first blur and the second blur have different shapes from each other. A training method characterized in that the optical performance corresponding to the first blur is higher than the optical performance corresponding to the second blur at a specific frequency.

2. The training method according to claim 1, characterized in that the average value of the optical performance corresponding to the first blur at multiple frequencies is higher than the average value of the optical performance corresponding to the second blur at the multiple frequencies.

3. A first step of generating a first image by adding a first blur to the original image, and a second step of generating a second image by adding a second blur to the original image, The process includes a second step of inputting the first image into a machine learning model to generate an output image, and updating the weights of the machine learning model based on the output image and the second image. The first blur and the second blur have different shapes from each other. A training method characterized in that the average value of the optical performance corresponding to the first blur at multiple frequencies is higher than the average value of the optical performance corresponding to the second blur at the same multiple frequencies.

4. The training method according to claim 2 or 3, characterized in that the plurality of frequencies are in the range from half the Nyquist frequency of the first image to the Nyquist frequency.

5. The training method according to claim 2 or 3, characterized in that the plurality of frequencies are in the range from 0 to half the Nyquist frequency of the first image.

6. In the first step, a plurality of first images are generated by adding a plurality of first blurs to the original image, and a plurality of second images are generated by adding a plurality of second blurs to the original image. The multiple first blur shapes are different from each other, The training method according to any one of claims 1 to 5, characterized in that the shapes of the multiple second blurs are the same as each other.

7. In the first step, a plurality of first images are generated by adding a plurality of first blurs to the original image, and a plurality of second images are generated by adding a plurality of second blurs to the original image. The multiple first blur shapes are different from each other, The training method according to any one of claims 1 to 5, characterized in that the shapes of the multiple second blurs are different from each other.

8. The training method according to any one of claims 1 to 7, characterized in that the first blur and the second blur are each defocus blur.

9. The training method according to any one of claims 1 to 5, characterized in that the optical performance includes a modulation transfer function or an optical transfer function.

10. A program characterized by causing a computer to execute the training method described in any one of claims 1 to 9.

11. A step of obtaining a machine learning model trained by the training method described in any one of claims 1 to 9, The process includes the step of generating an estimated image by inputting an input image into the machine learning model, The image processing method is characterized in that the estimated image is an image having a blur shape different from that of the input image.

12. A first step of generating a first image by adding a first blur to the original image, and a second step of generating a second image by adding a second blur to the original image, The process includes a second step of inputting the first image into a machine learning model to generate an output image, and updating the weights of the machine learning model based on the output image and the second image. The first blur and the second blur have different shapes from each other. A method for generating a trained model, characterized in that the optical performance corresponding to the first blur is higher than the optical performance corresponding to the second blur at a specific frequency.

13. A generation means that generates a first image by adding a first blur to the original image, and generates a second image by adding a second blur to the original image, The system includes a learning means that generates an output image by inputting the first image into a machine learning model, and updates the weights of the machine learning model based on the output image and the second image. The first blur and the second blur have different shapes from each other. A training device characterized in that the optical performance corresponding to the first blur is higher than the optical performance corresponding to the second blur at a specific frequency.