Image processing method, image processing apparatus, image processing system, method for generating learned weights, and program
The image processing method employs a machine learning model and a state map to achieve high-precision sharpening or shaping of blur in captured images, addressing the challenges of existing methods by reducing learning load and data retention while handling various optical aberrations.
Patent Information
- Application Number
- JP2024015702
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-02-05
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2039-06-06
AI Technical Summary
Existing image processing methods, such as those using Wiener filters and convolutional neural networks (CNNs), face challenges in achieving high-precision sharpening or shaping of blur caused by optical systems, particularly in systems with various aberrations, leading to increased learning load and data retention requirements.
An image processing method that utilizes a machine learning model to generate a state map based on the optical system's state and combines this with the captured image to produce a sharpened or shaped image, where the state map includes numerical values for zoom, aperture, and focal length, allowing for precise sharpening or shaping of blur.
This method effectively suppresses the learning load and data retention requirements of the machine learning model while achieving high-precision sharpening or shaping of blur in captured images, even in optical systems with various aberrations.
Smart Images

Figure 0007700292000002 
Figure 0007700292000003 
Figure 0007700292000004
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing method for sharpening or shaping the blur caused by an optical system from a captured image captured using the optical system.
Background Art
[0002] Patent Document 1 discloses a method of obtaining a sharpened image by correcting the blur due to aberration from a captured image by processing based on a Wiener filter. Patent Document 2 discloses a method of correcting the blur due to defocus of a captured image using a convolutional neural network (CNN).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, since the method disclosed in Patent Document 1 uses processing (linear processing) based on a Wiener filter, high-precision sharpening of blur cannot be performed. For example, it is impossible to restore the information of a subject whose spatial frequency spectrum has decreased to zero or the same intensity as noise due to blur. In addition, for different aberrations, it is necessary to use different Wiener filters, so the amount of holding data (the capacity of data indicating a plurality of Wiener filters) for sharpening increases in an optical system in which various aberrations occur.
[0005] On the one hand, since the CNN disclosed in Patent Document 2 performs non-linear processing, it can estimate the spatial frequency spectrum of a subject that has decreased to near zero. However, when sharpening an imaging image captured by an optical system in which various aberrations occur, it causes a decrease in the accuracy of sharpening, or an increase in the learning load and the amount of data to be retained. When performing defocus sharpening with a CNN, unlearned defocus is not correctly sharpened. Since the defocus generated by the optical system changes depending on zoom, aperture, focal length, etc., the following two methods can be considered to enable all of these defocuses to be sharpened.
[0006] The first method is a method of training a CNN with training data including all possible defocuses generated by the optical system. However, in this case, since the CNN is trained to sharpen all the defocuses included in the training data on average, the accuracy of sharpening for each defocus of different shapes decreases. The second method is a method of dividing each possible defocus generated by the optical system into a plurality of similar groups and individually training the CNN with the training data of each group. However, in this case, in an optical system in which various aberrations occur, such as a high-magnification zoom lens, the number of groups becomes enormous, and the learning load and the amount of data to be retained (the capacity of the data indicating the weights of the trained CNN) increase. Therefore, it is difficult to balance the accuracy of defocus sharpening, the learning load, and the amount of data to be retained.
[0007] Therefore, an object of the present invention is to provide an image processing method, an image processing apparatus, an image processing system, a method for manufacturing learned weights, and a program that suppress the learning load and the amount of data to be retained of a machine learning model and sharpen or shape the defocus of an imaging image with high accuracy.
Means for Solving the Problems
[0008] An image processing method as one aspect of the present invention is An image processing method executed using a processor, a first step of acquiring an imaging image and first information regarding the state at the time of imaging of the optical system used for imaging the imaging image, a second step of generating a state map indicating the state of the optical system based on the first information, and the imaging image and the state map based on a machine learning modelusing a step of generating a presumed image 3 and has to perform sharpening of the captured image or shaping of blur included in the captured image, and the state map includes numerical values indicating at least two of zoom, aperture, and focal length of the optical system as elements of different channels .
[0009] An image processing apparatus as another aspect of the present invention includes an acquisition unit that acquires a captured image and first information regarding a state at the time of imaging of an optical system used for capturing the captured image, a first generation unit that generates a state map indicating the state of the optical system based on the first information, and the captured image and the state map based on machine learning model using and a generation unit that generates a presumed image to perform sharpening of the captured image or shaping of blur included in the captured image, and the state map includes numerical values indicating at least two of zoom, aperture, and focal length of the optical system as elements of different channels .
[0010] An image processing system as another aspect of the present invention includes an acquisition unit that acquires input data including a captured image and first information regarding a state at the time of imaging of an optical system used for capturing the captured image, and a generation unit that inputs the input data into a machine learning model and generates a presumed image by performing sharpening of the captured image or shaping of blur included in the captured image. The image processing system includes the image processing apparatus and a control apparatus communicable with the image processing apparatus. The control apparatus includes a transmission unit that transmits a request regarding execution of processing on the captured image to the image processing apparatus, and the second apparatus includes a reception unit that receives the request and generates the presumed image in response to the request.
[0011] An image processing method as another aspect of the present invention An image processing method executed using a processor, for imaging an image and the imaging image's imaging first information regarding the state of the optical system used for at the time of imaging and a first step, the acquire second step of generate a state map indicating the state of the optical system based on first information performing, and the and a third step of generating an estimated image by performing sharpening of the captured image or shaping of blur included in the captured image using a machine learning model based on the captured image, the state map, and second information regarding the position of each pixel of the captured image, wherein the second information includes a numerical value normalized by a length based on the image circle of the optical system.
[0014] A program as another aspect of the present invention causes a computer to execute the image processing method.
[0015] Other objects and features of the present invention will be described in the following examples. [Effect of the Invention]
[0016] According to the present invention, it is possible to provide an image processing method, an image processing apparatus, an image processing system, a method for manufacturing learned weights, and a program that suppress the learning load and the amount of data to be held of a machine learning model and sharpen or shape the blur of a captured image with high precision.
Brief Description of Drawings
[0017]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Modes for Carrying Out the Invention
[0018] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In each figure, the same members are denoted by the same reference numerals, and duplicate explanations are omitted.
[0019] Before giving a specific description of each embodiment, the gist of the present invention will be described. The present invention sharpens or shapes the blur caused by the optical system (occurring in the optical system) from the captured image captured using the optical system using a machine learning model. The "optical system" referred to here refers to what exerts an optical action related to imaging. That is, the "optical system" includes not only the imaging optical system, but also, for example, an optical low-pass filter and a microlens array of the image sensor. Therefore, the blur caused by the optical system includes blur due to aberration, diffraction, defocus, the action of the optical low-pass filter, pixel aperture degradation of the image sensor, and the like.
[0020] The machine learning model includes, for example, a neural network, genetic programming, a Bayesian network, and the like. The neural network includes a CNN (Convolutional Neural Network), a GAN (Generative Adversarial Network), an RNN (Recurrent Neural Network), and the like.
[0021] Sharpening of blur refers to a process of restoring the frequency components of a subject that have been reduced or disappeared due to blur. Blur shaping refers to a conversion of the shape of blur without restoring the frequency components. For example, conversions from two-line blur to Gaussian or disk (flat circular distribution), conversions from defocus blur lacking due to vignetting to circular defocus blur, etc. are included.
[0022] The input data input to the machine learning model includes the captured image and information regarding the state of the optical system at the time of capturing the captured image. The state of the optical system refers to the state of the device that can affect the optical action related to imaging. The state of the optical system includes, for example, the zoom, aperture, focal length, etc. of the optical system. In addition, as information regarding the state of the optical system, information regarding the presence or absence (or type) of an optical low-pass filter and the presence or absence (or type) of an accessory (e.g., a converter lens) attached to the optical system may be included.
[0023] By inputting information regarding the state of the optical system in the learning of the machine learning model and in the estimation after learning, the machine learning model can identify in which state of the optical system the blur acting on the captured image has occurred. Thereby, the machine learning model learns weights for performing different sharpening (or shaping) for each state of the optical system, rather than weights for sharpening (or shaping) those blurs on average even if various shapes of blur are included in the learning.
[0024] For this reason, high-precision sharpening (or shaping) can be performed for each blur. Therefore, a decrease in the accuracy of sharpening (or shaping) can be suppressed, and learning data including various shapes of blur can be learned collectively. As a result, the learning load and the amount of data to be retained can be suppressed, and the blur caused by the optical system of the captured image can be sharpened or shaped with high precision. The effects of the present invention will be quantitatively shown in Example 2. In the following description, the stage of learning the weights of the machine learning model is referred to as the learning phase, and the stage of performing blur sharpening or shaping using the learned weights of the machine learning model is referred to as the estimation phase.
Example
[0025] First, the image processing system in Embodiment 1 of the present invention will be described. This embodiment performs defocus sharpening, but is similarly applicable to defocus shaping. Also, in this embodiment, the object to be sharpened is defocus due to aberration and diffraction, but it is also applicable to defocus due to defocusing.
[0026] FIG. 2 is a block diagram of the image processing system 100 in this embodiment. FIG. 3 is an external view of the image processing system 100. The image processing system 100 includes a learning device (image processing device) 101, an imaging device 102, and a network 103. The learning device 101 and the imaging device 102 are connected via a network 103 that can be wired or wireless. The learning device 101 has a storage unit 111, an acquisition unit (acquisition means) 112, a calculation unit (generation means) 113, and an update unit (update means) 114, and learns weights for performing defocus sharpening with a machine learning model. The imaging device 102 captures the subject space to obtain a captured image, and sharpens the defocus of the captured image using the information of the weights read out after imaging or in advance. Details regarding the learning of the weights executed by the learning device 101 and the defocus sharpening executed by the imaging device 102 will be described later.
[0027] The imaging device 102 has an optical system (imaging optical system) 121 and an imaging element 122. The optical system 121 condenses the light incident from the subject space to the imaging device 102. The imaging element 122 receives (photoelectrically converts) the optical image (subject image) formed via the optical system 121 to generate a captured image. The imaging element 122 is, for example, a CCD (Charge Coupled Device) sensor or a CMOS (Complementary Metal-Oxide Semiconductor) sensor.
[0028] The image processing unit (image processing apparatus) 123 includes an acquisition unit (acquisition means) 123a and a sharpening unit (generation means) 123b, and generates a sharpened estimated image (sharpened image) from the captured image. For generating the estimated image, the information of the learned weights learned by the learning apparatus 101 is used. The information of the weights is read in advance from the learning apparatus 101 via a wired or wireless network 103 and stored in the storage unit 124. The stored weight information may be the weight values themselves or in an encoded format. The recording medium 125 stores the estimated image. Alternatively, the recording medium 125 may store the captured image, and the image processing unit 123 may read the captured image and generate the estimated image. The display unit 126 displays the estimated image stored in the recording medium 125 according to a user's instruction. The system controller 127 controls the above-described series of operations.
[0029] Next, with reference to FIG. 4, the weight learning (learning phase, method for manufacturing a learned model) executed by the learning apparatus 101 in the present embodiment will be described. FIG. 4 is a flowchart regarding weight learning. Each step in FIG. 4 is mainly executed by the acquisition unit 112, the calculation unit 113, or the update unit 114 of the learning apparatus 101. In the present embodiment, a CNN is used as the machine learning model, but the present invention is similarly applicable to other models.
[0030] First, in step S101, the acquisition unit 112 acquires one or more sets of correct images and training input data from the storage unit 111. The training input data is the input data in the learning phase of the CNN. The training input data includes a training image and information regarding the state of the optical system corresponding to the training image. The training image and the correct image are a pair of images in which the same subject exists and the presence or absence of blur is different. The correct image is an image without blur, and the training image is an image with blur. The blur is the combined blur of aberration and diffraction generated by the optical system 121 and the pixel aperture degradation of the image sensor 122. In one training image, the combined blur of aberration and diffraction generated by the optical system 121 and pixel aperture degradation at a specific zoom, aperture, and focal length acts. The information regarding the state of the optical system corresponding to the training image is information indicating at least any one of a specific zoom, aperture, and focal length. In other words, the information regarding the state of the optical system is information for specifying the blur acting on the training image. In the present embodiment, the information regarding the state of the optical system includes all of the zoom, aperture, and focal length. The training image is not limited to a captured image and may be an image generated by CG or the like.
[0031] Regarding the method for generating the correct image and the training input data stored in the storage unit 111, an example is shown below. The first generation method is a method of performing imaging simulation using the original image as the subject. The original image is a real photographed image, a CG (Computer Graphics) image, or the like. It is desirable that the original image is an image having edges with various intensities and directions, textures, gradations, flat portions, etc., so that correct sharpening can be performed for various subjects. The original image may be one or a plurality of images. The correct image is an image obtained by performing imaging simulation without applying blur to the original image. The training image is an image obtained by performing imaging simulation with the blur to be sharpened applied to the original image.
[0032] In this embodiment, aberration, diffraction, and blurring due to pixel aperture degradation occurring in the optical system 121 are made to act. Here, Z represents the zoom, F represents the aperture, and D represents the state of the focal length. When the imaging device 122 acquires a plurality of color components, the blurring of each color component is made to act on the original image. The action of blurring can be executed by convolving the original image with a PSF (Point Spread Function) or by taking the product of the frequency characteristics of the original image and the OTF (Optical Transfer Function). Information regarding the state of the optical system corresponding to the training image on which the blurring specified by (Z, F, D) is made to act is information specifying (Z, F, D). The correct image and the training image may be either an undeveloped RAW image or a developed image. For one or more original images, blurring of a plurality of different (Z, F, D) is made to act to generate a plurality of sets of correct images and training images.
[0033] In this embodiment, correction for all blurring occurring in the optical system 121 is learned in a batch. Therefore, (Z, F, D) is changed within the range that the optical system 121 can take to generate a plurality of sets of correct images and training images. Also, even at the same (Z, F, D), since there are a plurality of blurrings depending on the image height and azimuth, sets of correct images and training images are generated for each different image height and azimuth.
[0034] Preferably, the original image may have a signal value higher than the luminance saturation value of the imaging device 122. This is because, in an actual subject, there are subjects that do not fit within the luminance saturation value when imaging is performed by the imaging device 102 under specific exposure conditions. The correct image is generated by clipping the signal of the original image at the luminance saturation value of the imaging device 122. The training image is generated by clipping with the luminance saturation value after the blurring is made to act.
[0035] Also, when generating the correct image and the training image, reduction processing of the original image may be performed. When using a real image as the original image, since blurring already occurs due to aberration and diffraction, reducing the image can reduce the influence of blurring and generate a high-resolution correct image. In this case, in order to match the scale of the correct image, the training image is also reduced in the same way. The order of reduction and the action of blurring can be either first. When performing the blurring action first, it is necessary to make the sampling rate of blurring finer in consideration of reduction. If it is a PSF, the spatial sampling points should be made finer, and if it is an OTF, the maximum frequency should be increased. Note that when the original image sufficiently contains high-frequency components, high-precision sharpening is possible, so reduction may not be necessary.
[0036] Also, the blurring applied in the generation of the training image does not include distortion aberration. This is because if the distortion aberration is large, the position of the subject changes, and the subject contained within each of the correct image and the training image may be different. Therefore, the CNN learned in this embodiment does not correct the distortion aberration. In the estimation phase, the distortion aberration is corrected after blurring and sharpening using bilinear interpolation, bicubic interpolation, or the like. Similarly, the blurring applied in the generation of the training image does not include magnification chromatic aberration. In the estimation phase, the magnification chromatic aberration is corrected before blurring and sharpening using the shift of each color component or the like.
[0037] The second method for generating the correct image and the training input data is a method that uses a real image captured by the optical system 121 and the image sensor 122. The optical system 121 captures an image in the state of (Z, F, D) to obtain a training image. The information regarding the state of the optical system corresponding to the training image is information that specifies (Z, F, D). The correct image can be obtained, for example, by capturing the same subject as the training image using an optical system with higher performance than the optical system 121. Note that a partial region with a predetermined number of pixels may be extracted from the training image and the correct image generated by the above two methods and used for learning.
[0038] Next, in step S102 of FIG. 4, the arithmetic unit 113 inputs the training input data into the CNN and generates an output image. With reference to FIG. 1, the generation of the output image in this embodiment will be described. FIG. 1 is a configuration diagram of the machine learning model in this embodiment.
[0039] The training input data includes the training image 201 and the information (z, f, d) 202 regarding the state of the optical system. The training image 201 may be grayscale or may have a plurality of channel components. The same applies to the correct answer image. (z, f, d) is the normalized (Z, F, D). The normalization is performed based on the range that the optical system 121 can take for each of the zoom, aperture, and focal length. For example, let Z be the focal length, F be the aperture value, and D be the reciprocal of the absolute value of the distance from the imaging device 102 to the focused subject. Let the minimum value and the maximum value of the focal length of the optical system 121 be Z min and Z max , the minimum value and the maximum value of the aperture value be F min and F max , and the minimum value and the maximum value of the reciprocal of the absolute value of the focusable distance be D min and D max respectively. Here, when the focusable distance is infinite, D min = 1 / |∞| = 0. The normalized (z, f, d) is obtained by the following formula (1).
[0040]
Equation
[0041] In formula (1), x is a dummy variable indicating any one of (z, f, d), and X is a dummy variable indicating any one of (Z, F, D). Note that when X min = X max , x is regarded as a constant. Or, since there is no degree of freedom in that x, it is excluded from the information regarding the state of the optical system. Here, generally, as the focal length gets closer, the performance change of the optical system 121 becomes larger, so D is set as the reciprocal of the distance.
[0042] In this embodiment, CNN211 has a first sub-network 221 and a second sub-network 223. The first sub-network 221 has one or more convolutional layers or fully-connected layers. The second sub-network 223 has one or more convolutional layers. At the first time of learning, the weights of CNN211 (each element of the filter and the value of the bias) are generated by random numbers. The first sub-network 221 takes the information (z, f, d) 202 regarding the state of the optical system as input and generates a state map 203 converted into a feature map. The state map 203 is a map indicating the state of the optical system and has the same number of elements (pixels) as one channel component of the training image 201. The concatenation layer 222 concatenates the training image 201 and the state map 203 in a specified order in the channel direction. Note that other data may be concatenated between the training image 201 and the state map 203. The second sub-network 223 takes the concatenated training image 201 and state map 203 as input and generates an output image 204. When a plurality of sets of training input data are acquired in step S101 of FIG. 4, an output image 204 is generated for each of them. Further, the training image 201 may be converted into a feature map by a third sub-network, and the feature map and the state map 203 may be concatenated by the concatenation layer 222.
[0043] Subsequently, in step S103 of FIG. 4, the update unit 114 updates the weights of the CNN from the error between the output image and the correct image. In this embodiment, the Euclidean norm of the difference in signal values between the output image and the correct image is used as the loss function. However, the loss function is not limited to this. When a plurality of sets of training input data and correct images are acquired in step S101, the value of the loss function is calculated for each set. The weights are updated by the error backpropagation method or the like from the calculated value of the loss function.
[0044] Subsequently, in step S104, the update unit 114 determines whether the learning of the weights is completed. Completion can be determined by whether the number of iterations of learning (weight update) has reached a specified number, or whether the amount of change in the weights during update is smaller than a specified value. If it is determined in step S104 that the learning of the weights is not completed, the process returns to step S101, and the acquisition unit 112 acquires one or more sets of new training input data and correct images. On the other hand, if it is determined that the learning of the weights is completed, the update unit 114 ends the learning and stores the weight information in the storage unit 111.
[0045] Next, with reference to FIG. 5, the blurring and sharpening (estimation phase) of the captured image executed by the image processing unit 123 in the present embodiment will be described. FIG. 5 is a flowchart regarding blurring and sharpening (generation of an estimated image) in the present embodiment. Each step in FIG. 5 is mainly executed by the acquisition unit 123a or the sharpening unit 123b of the image processing unit 123.
[0046] First, in step S201, the acquisition unit 123a acquires input data and weight information. The input data includes the captured image and information regarding the state of the optical system 121 when the captured image is captured. The captured image to be acquired may be a part of the entire captured image. The information regarding the information of the optical system is (z, f, d) indicating the state of the zoom, aperture, and focal length of the optical system 121. The weight information is acquired by being read from the storage unit 124.
[0047] Subsequently, in step S202, the sharpening unit 123b inputs the input data into the CNN and generates an estimated image. The estimated image is an image in which the blurring caused by the aberration and diffraction of the optical system 121 and the pixel aperture degradation of the imaging device 122 is sharpened with respect to the captured image. Similar to the learning time, the estimated image is generated using the CNN shown in FIG. 1. The learned weights acquired are used for the CNN. In the present embodiment, the weights for blurring and sharpening are learned all at once for all possible (z, f, d) of the optical system 121. For this reason, blurring and sharpening is executed using the same CNN with the same weights for all captured images of (z, f, d).
[0048] In this embodiment, instead of the sharpening unit 123b that sharpens the captured image, a shaping unit that shapes the blur included in the captured image may be provided. This also applies to Example 2 described later. In this embodiment, the image processing apparatus (image processing unit 123) includes an acquisition unit (acquisition unit 123a) and a generation unit (sharpening unit 123b or shaping unit). The acquisition unit acquires input data including the captured image and information on the state of the optical system used for capturing the captured image. The generation unit inputs the input data into a machine learning model and generates an estimated image obtained by sharpening the captured image, or an estimated image obtained by shaping the blur included in the captured image. Also, in this embodiment, the image processing apparatus (learning apparatus 101) includes an acquisition unit (acquisition unit 112), a generation unit (calculation unit 113), and an update unit (update unit 114). The acquisition unit acquires input data including the training image and information on the state of the optical system corresponding to the training image. The generation unit inputs the input data into a machine learning model and generates an output image obtained by sharpening the training image, or an output image obtained by shaping the blur included in the training image. The update unit updates the weights of the machine learning model based on the output image and the correct image.
[0049] According to this embodiment, it is possible to realize an image processing apparatus and an image processing system that can suppress the learning load and the amount of data to be held of the machine learning model and sharpen the blur caused by the optical system of the captured image with high accuracy.
Example
[0050] Next, the image processing system in Example 2 of the present invention will be described. This example executes blur sharpening processing, but is similarly applicable to blur shaping processing.
[0051] FIG. 6 is a block diagram of the image processing system 300 in this embodiment. FIG. 7 is an external view of the image processing system 300. The image processing system 300 includes a learning device (image processing device) 301, a lens device 302, an imaging device 303, an image estimation device 304, a display device 305, a recording medium 306, an output device 307, and a network 308. The learning device 301 includes a storage unit 301a, an acquisition unit (acquisition means) 301b, a calculation unit (generation means) 301c, and an update unit (update means) 301d, and learns the weights of the machine learning model used for defocus sharpening. Details regarding the learning of the weights and the defocus sharpening process using the weights will be described later.
[0052] The lens device 302 and the imaging device 303 are detachable and can be connected to different types of lens devices 302 or imaging devices 303. The lens device 302 has different focal lengths, apertures, and focus distances depending on the type. Also, since the lens configuration varies depending on the type, the shape of the blur due to aberration and diffraction also differs. The imaging device 303 has an image sensor 303a, and the presence or absence and type (separation method, cut-off frequency, etc.), pixel pitch (including pixel aperture), color filter array, etc. of the optical low-pass filter differ depending on the type.
[0053] The image estimation device 304 includes a storage unit 304a, an acquisition unit (acquisition means) 304b, and an edge enhancement unit (generation means) 304c. The image estimation device 304 generates an estimated image in which the blur caused by the optical system is sharpened for the captured image (or at least a part thereof) captured by the imaging device 303. A plurality of types of lens devices 302 and the imaging device 303 can be connected to the image estimation device 304. For the blur sharpening, a machine learning model using the weights learned by the learning device 301 is used. The learning device 301 and the image estimation device 304 are connected by a network 308, and the image estimation device 304 reads out the information of the learned weights from the learning device 301 during or before the blur sharpening. The estimated image is output to at least one of a display device 305, a recording medium 306, or an output device 307. The display device 305 is, for example, a liquid crystal display or a projector. The user can perform an editing operation or the like while checking the image during processing via the display device 305. The recording medium 306 is, for example, a semiconductor memory, a hard disk, or a server on a network. The output device 307 is, for example, a printer.
[0054] Next, with reference to FIG. 4, the learning of the weights (learning phase) performed by the learning device 301 will be described. In this embodiment, a CNN is used as the machine learning model, but the same can be similarly applied to other models. Note that the same description as in the first embodiment is omitted.
[0055] First, in step S101, the acquisition unit 301b acquires one or more sets of correct images and training input data from the storage unit 301a. In the storage unit 301a, training images are stored for a plurality of types of combinations of the lens device 302 and the imaging device 303. In this embodiment, the learning of the weights for defocus sharpness is performed collectively for each type of the lens device 302. For this reason, first, the type of the lens device 302 for which the weights are to be learned is determined, and training images are acquired from the set of training images corresponding thereto. The sets of training images corresponding to a certain type of lens device 302 are each a set of images with different defocus effects, such as zoom, aperture, focal length, image height and azimuth, optical low-pass filter, pixel pitch, color filter array, etc.
[0056] In this embodiment, learning is performed with the configuration of the CNN shown in FIG. 8. FIG. 8 is a configuration diagram of the machine learning model in this embodiment. The training input data 404 includes a training image 401, a state map 402, and a position map 403. The generation of the state map 402 and the position map 403 is performed in this step. The state map 402 and the position map 403 are maps showing (Z, F, D) and (X, Y) corresponding to the defocus acting on the acquired training image, respectively. (X, Y) are the coordinates (horizontal direction and vertical direction) of the image plane shown in FIG. 9, and correspond to the image height and azimuth in polar coordinate representation.
[0057] In this embodiment, the coordinates (X, Y) are based on the optical axis of the lens device 302 as the origin. FIG. 9 shows the relationship between the image circle 501 of the lens device (optical system) 302, the first effective pixel region 502 and the second effective pixel region 503 of the imaging element 303a, and the coordinates (X, Y). Since the imaging device 303 has imaging elements 303a of different sizes depending on the type of the imaging device 303, the imaging device 303 includes a type having the first effective pixel region 502 and a type having the second effective pixel region 503. Among the imaging devices 303 connectable to the lens device 302, the imaging device 303 having the largest-sized imaging element 303a has the first effective pixel region 502.
[0058] The position map 403 is generated based on (x, y) obtained by normalizing the coordinates (X, Y). The normalization is performed by dividing (X, Y) by the length (radius of the image circle) 511 based on the image circle 501 of the lens device 302. Alternatively, X may be divided by the horizontal length 512 of the first effective pixel region from the origin, and Y may be divided by the vertical length 513 of the first effective pixel region from the origin for normalization. If (X, Y) is normalized such that the edge of the captured image is always 1, the positions (X, Y) indicated by the same value of (x, y) are different depending on the images captured by the imaging elements 303a of different sizes, and the correspondence between (x, y) and blur is not uniquely determined. This causes a decrease in the sharpness accuracy of the blur. The position map 403 is a two-channel map having the values of (x, y) as channel components respectively. Note that polar coordinates may be used for the position map 403, and the way of taking the origin is not limited to that in FIG. 9.
[0059] The state map 402 is a three-channel map having the values of the normalized (z, f, d) as channel components respectively. The number of elements (pixels) per channel of each of the training image 401, the state map 402, and the position map 403 is equal. Note that the configurations of the position map 403 and the state map 402 are not limited to this. For example, as shown in FIG. 10 which shows an example of the position map, the first effective pixel region 502 may be divided into a plurality of sub-regions, and the position map may be represented by one channel by assigning numerical values to each sub-region. Note that the number of divided sub-regions and the method of distributing the numerical values are not limited to those in FIG. 10. Similarly, (Z, F, D) may also be divided into a plurality of sub-regions in a three-dimensional space with each as an axis and numerical values assigned, and the state map may be represented by one channel. The training image 401, the state map 402, and the position map 403 are concatenated in a specified order in the channel direction by the concatenation layer 411 in FIG. 8 to generate the training input data 404.
[0060] Subsequently, in step S102 of FIG. 4, the arithmetic unit 301c inputs the training input data 404 to the CNN 412 and generates the output image 405. Subsequently, in step S103, the update unit 301d updates the weights of the CNN from the error between the output image and the correct image. Subsequently, in step S104, the update unit 301d determines whether the learning has been completed. The information on the learned weights is stored in the storage unit 301a.
[0061] Next, with reference to FIG. 11, the blurring and sharpening (estimation phase) of the captured image executed by the image estimation device 304 will be described. FIG. 11 is a flowchart regarding blurring and sharpening (generation of the estimated image). Each step in FIG. 11 is mainly executed by the acquisition unit 304b or the sharpening unit 304c of the image estimation device 304.
[0062] First, in step S301, the acquisition unit 304b acquires the captured image (or at least a part thereof). Subsequently, in step S302, the acquisition unit 304b acquires the weight information corresponding to the captured image. In the second embodiment, the weight information for each type of the lens device 302 is read out from the storage unit 301a in advance and stored in the storage unit 304a. Therefore, the acquisition unit 304b acquires the weight information corresponding to the type of the lens device 302 used for capturing the captured image from the storage unit 304a. The type of the lens device 302 used for capturing is specified from, for example, the metadata in the file of the captured image.
[0063] Subsequently, in step S303, the acquisition unit 304b generates a state map and a position map corresponding to the captured image, and generates input data. The state map is generated based on the number of pixels of the captured image and information on the state (Z, F, D) of the lens device 302 when the captured image is captured. The number of elements (pixels) per channel of the captured image and the state map is equal. (Z, F, D) is specified from, for example, the metadata of the captured image. The position map is generated based on the number of pixels of the captured image and information on the position of each pixel of the captured image. The number of elements (pixels) per channel of the captured image and the position map is equal. The size of the effective pixel region of the image sensor 303a used for capturing the captured image is specified from the metadata of the captured image or the like, and a normalized position map is generated using, for example, the length of the image circle of the lens device 302 specified in the same manner. The input data is generated by concatenating the captured image, the state map, and the position map in a specified order in the channel direction, as in FIG. 8. In this embodiment, the order of step S302 and step S303 does not matter. Also, the state map and the position map may be generated at the time of capturing the captured image and saved together with the captured image.
[0064] Subsequently, in step S304, the sharpness enhancement unit 304c inputs the input data to the CNN in the same manner as in FIG. 8 and generates an estimated image. FIG. 12 is a diagram showing the effect of sharpness enhancement, and shows the sharpness enhancement effect at 90% of the image height in a specific (Z, F, D) of a certain zoom lens. In FIG. 12, the horizontal axis represents the spatial frequency, and the vertical axis represents the measured value of the SFR (Spatial Frequency Response). The SFR corresponds to the MTF (Modulation Transfer Function) in a certain cross section. The Nyquist frequency of the image sensor used for imaging is 76 [lp / mm]. The solid line 601 is the captured image, and the broken line 602, the one-dot chain line 603, and the two-dot chain line 604 are the results of blurring and sharpening the captured image by the CNN. The broken line 602, the one-dot chain line 603, and the two-dot chain line 604 use a CNN that has learned blurring and sharpening using learning data that mixes all the aberrations and diffraction blurs generated by the zoom lens.
[0065] In the estimation phase (or learning phase), the input data of the dashed line 602 is only the captured image (or training image). The input data of the dashed-dotted line 603 is the captured image (or training image) and the position map. The input data of the double-dashed line 604 is the captured image (or training image), the position map, and the state map, which corresponds to the configuration of this embodiment. The CNNs used for the dashed line 602, the dashed-dotted line 603, and the double-dashed line 604 only differ in the number of channels of the filters in the first layer (because the number of channels of the input data is different), and other filter sizes, the number of filters, the number of layers, etc. are common. Therefore, the learning load and the amount of data to be held (the data capacity of the CNN weight information) for each of the dashed line 602, the dashed-dotted line 603, and the double-dashed line 604 are substantially the same. On the other hand, the double-dashed line 604 adopting the configuration of this embodiment has a high sharpening effect as shown in FIG. 12.
[0066] According to this embodiment, it is possible to realize an image processing apparatus and an image processing system that can suppress the learning load and the amount of data to be held of the machine learning model and highly accurately sharpen the blur caused by the optical system of the captured image.
[0067] Next, the preferable conditions for enhancing the effects of this embodiment will be described.
[0068] Preferably, the input data further includes information indicating the presence or absence and type of the optical low-pass filter of the imaging device 303 used for capturing the captured image. Thereby, the sharpening effect of the blur is improved. The type refers to the separation method (vertical two-point separation, horizontal two-point separation, four-point separation, etc.) and the cut-off frequency. A map having numerical values that can specify the presence or absence and the type as elements may be generated based on the number of pixels of the captured image and included in the input data.
[0069] The input data preferably further includes information regarding manufacturing variations of the lens device 302 used for imaging the captured image. This enables high-precision blurring and sharpening including manufacturing variations. In the learning phase, blurring including manufacturing variations is applied to the original image to generate training images, and the machine learning model is learned with the input data for training including information indicating manufacturing variations. As the information indicating manufacturing variations, for example, a numerical value indicating the degree of actual performance including manufacturing variations with respect to the design performance is used. For example, when the actual performance is equal to the design performance, the numerical value is set to 0, and it moves in the negative direction as the actual performance is inferior to the design performance and in the positive direction as the actual performance is superior to the design performance. In the estimation phase, as shown in FIG. 13, a map having a numerical value indicating the degree of actual performance with respect to the design performance is included in the input data for a plurality of partial regions (or for each pixel) of the captured image. FIG. 13 is a diagram showing an example of a manufacturing variation map. This map is generated based on the number of pixels of the captured image. The actual performance including manufacturing errors of the lens device 302 can be obtained by measuring during manufacturing or the like to obtain the map. Further, manufacturing variations such as overall image performance degradation (degradation of spherical aberration) and performance variations due to astigmatism (asymmetric blurring) may be classified into several categories, and the manufacturing variations may be indicated by numerical values indicating the categories.
[0070] The input data preferably further includes information on the distribution regarding the distance of the subject space at the time of imaging the captured image. This enables high-precision defocus sharpening including performance changes due to defocus. Due to axial chromatic aberration and field curvature, the optical performance of a defocused subject plane may be better than that of the focus plane. If defocus sharpening is performed with a machine learning model learned only with the blur at the focus plane without considering this, the resolution feeling becomes excessive and the image becomes unnatural. To solve this, first, in the learning phase, learning is performed using training images with defocus blur applied to the original image. At this time, a numerical value indicating the defocus amount (corresponding to the distance of the subject space) is also included in the training input data. For example, the focus plane may be set to 0, the direction away from the imaging device may be negative, and the approaching direction may be positive. In the estimation phase, a defocus map (information on the distribution regarding the distance of the subject space) of the captured image is obtained using imaging of a disparity image, DFD (Depth from Defocus), etc., and included in the input data. The defocus map is generated based on the number of pixels of the captured image.
[0071] The input data preferably further includes information regarding the pixel pitch of the imaging device 303a used for imaging the captured image or the color filter array. This enables high-precision blurring and sharpening regardless of the type of the imaging device 303a. Depending on the pixel pitch, the strength of pixel aperture degradation and the magnitude of blurring with respect to the pixel change. Also, depending on the color components constituting the color filter array, the shape of the blur changes. The color components are, for example, RGB (Red, Green, Blue) or complementary colors CMY (Cyan, Magenta, Yellow). Also, when the training image or the captured image is an undeveloped Bayer image or the like, depending on the arrangement order of the color filter array, the shape of the blur is different even for pixels at the same position. In the learning phase, information specifying the pixel pitch and the color filter array corresponding to the training image is included in the training input data. For example, it includes a map having numerical values of the normalized pixel pitch as elements. For normalization, it is preferable to use the maximum pixel pitch among a plurality of types of imaging devices 303 as the divisor. Also, a map having numerical values indicating the color components of the color filter array as elements may be included. By including a similar map in the input data also in the estimation phase, the accuracy of sharpening can be improved. The map is generated based on the number of pixels of the captured image.
[0072] The input data preferably further includes information indicating the presence or absence and type of accessories of the lens device 302. Accessories include a wide converter, a teleconverter, a close-up lens, a wavelength cut filter, and the like. Since the shape of the blur changes depending on the type of accessory, by inputting information regarding the presence or absence and type of accessory, it is possible to perform sharpening including their effects. In the learning phase, information identifying the accessory is included in the training input data, including the effect of the accessory on the blur applied to the training image. For example, a map having numerical values indicating the presence or absence and type of accessory as elements is used. In the estimation phase, similar information (map) may be included in the input data. This map is generated based on the number of pixels of the captured image.
Example
[0073] Next, the image processing system according to Embodiment 3 of the present invention will be described. This embodiment performs blurring shaping processing, but is similarly applicable to sharpness enhancement processing of blurring.
[0074] FIG. 14 is a block diagram of the image processing system 700. FIG. 15 is an external view of the image processing system 700. The image processing system 700 includes a learning device 701, a lens device (optical system) 702, an imaging device 703, a control device (first device) 704, an image estimation device (second device) 705, and networks 706 and 707. The learning device 701 and the image estimation device 705 are, for example, servers. The control device 704 is a device operated by a user such as a personal computer or a mobile terminal. The learning device 701 includes a storage unit 701a, an acquisition unit (acquisition means) 701b, a calculation unit (generation means) 701c, and an update unit (update means) 701d, and learns the weights of a machine learning model that shapes the blurring of the captured image captured using the lens device 702 and the imaging device 703. Details regarding learning will be described later. The blurring to be shaped in this embodiment is blurring due to defocus, but is similarly applicable to aberration, diffraction, etc.
[0075] The imaging device 703 includes an imaging element 703a. The imaging element 703a photoelectrically converts the optical image formed by the lens device 702 to acquire a captured image. The lens device 702 and the imaging device 703 are detachable and can be combined with each other in a plurality of types. The control device 704 includes a communication unit 704a, a storage unit 704b, and a display unit 704c, and controls the processing to be executed on the captured image acquired from the imaging device 703 connected by wire or wirelessly according to the user's operation. Alternatively, the control device 704 may store the captured image captured by the imaging device 703 in the storage unit 704b in advance and read out the captured image.
[0076] The image estimation device 705 includes a communication unit 705a, a storage unit 705b, an acquisition unit (acquisition means) 705c, and a shaping unit (generation means) 705d. The image estimation device 705 executes a blurring and shaping process on a captured image in response to a request from the control device 704 connected via the network 707. The image estimation device 705 acquires information on learned weights from the learning device 701 connected via the network 706 either during or prior to the blurring and shaping process, and uses it for the blurring and shaping of the captured image. The estimated image after blurring and shaping is transmitted back to the control device 704, stored in the storage unit 704b, and displayed on the display unit 704c.
[0077] Next, with reference to FIG. 16, the learning of weights (learning phase, method for manufacturing a learned model) executed by the learning device 701 in this embodiment will be described. FIG. 16 is a flowchart regarding the learning of weights. Each step in FIG. 16 is mainly executed by the acquisition unit 701b, the calculation unit 701c, or the update unit 701d of the learning device 701. In this embodiment, a GAN is used as the machine learning model, but the same applies to other models. In a GAN, there is a generator that generates an output image with defocus blur shaping, and a discriminator that discriminates between a correct image and the output image generated by the generator. In the learning, first, a first learning using only the generator is performed as in the first embodiment. When the weights of the generator have converged to a certain extent, a second learning using both the generator and the discriminator is performed. Hereinafter, descriptions of the same parts as in the first embodiment will be omitted.
[0078] First, in step S401, the acquisition unit 701b acquires one or more sets of correct images and training input data from the storage unit 701a. In this embodiment, the correct image and the training image are a pair of images with different defocus blur shapes. The training image is an image on which the defocus blur of the object to be shaped acts. Examples of the shaping target include secondary blur, lack of blur due to vignetting, ring-shaped blur due to pupil occlusion such as a catadioptric lens, and zonular patterns of blur caused by uneven cutting of the mold of an aspherical lens. The correct image is an image on which the defocus blur after shaping acts. The shape of the defocus blur after shaping may be determined according to the user's preference, such as Gaussian or disk (circular distribution with flat intensity). In Embodiment 3, a plurality of training images and correct images generated by applying blur to the original image are stored in the storage unit 701a. When generating a plurality of training images and correct images, blur corresponding to various defocus amounts is applied so that shaping accuracy can be ensured for various defocus amounts. Also, on the focus plane, since it is desirable that the image does not change before and after blur shaping, training images and correct images with a defocus amount of zero are also generated.
[0079] In this embodiment, weights for converting blur shaping are collectively learned for a plurality of types of lens devices 702. Therefore, the information regarding the state of the optical system includes information for specifying the type of the lens device 702. Blurred training images corresponding to the types of lens devices 702 to be collectively learned are acquired from the storage unit 701a. The information regarding the state of the optical system further includes information for specifying the zoom, aperture, and focal length of the lens device 702 corresponding to the blur acting on the training image.
[0080] FIG. 17 is a configuration diagram of a machine learning model (GAN). A lens state map 802 having numerical values (L, z, f, d) for specifying the type, zoom, aperture, and focal length of the lens device 702 as channel components is generated based on the number of pixels of the training image 801. The concatenation layer 811 concatenates the training image 801 and the lens state map 802 in a predetermined order in the channel direction to generate training input data 803.
[0081] Subsequently, in step S402 of FIG. 16, the arithmetic unit 701c inputs the training input data 803 to the generator 812 to generate the output image 804. The generator 812 is, for example, a CNN. Subsequently, in step S403, the update unit 701d updates the weights of the generator 812 from the error between the output image 804 and the correct image 805. The Euclidean norm of the difference at each pixel is used as the loss function.
[0082] Subsequently, in step S404, the update unit 701d determines whether the first learning has been completed. If the first learning has not been completed, the process returns to step S401. On the other hand, if the first learning has been completed, the process proceeds to step S405, and the update unit 701d executes the second learning.
[0083] In step S405, the acquisition unit 701b acquires one or more sets of correct images 805 and training input data 803 from the storage unit 701a in the same manner as in step S401. Subsequently, in step S406, the arithmetic unit 701c inputs the training input data 803 to the generator 812 to generate the output image 804 in the same manner as in step S402. Subsequently, in step S407, the update unit 701d updates the weights of the discriminator 813 from the output image 804 and the correct image 805. The discriminator 813 discriminates whether the input image is a fake image generated by the generator 812 or a real image that is the correct image 805. The output image 804 or the correct image 805 is input to the discriminator 813 to generate a discrimination label (fake or real). Based on the error between the discrimination label and the correct label (the output image 804 is fake and the correct image 805 is real), the update unit 701d updates the weights of the discriminator 813. Sigmoid cross entropy is used as the loss function, but other loss functions may also be used.
[0084] Subsequently, in step S408, the update unit 701d updates the weights of the generator 812 from the output image 804 and the correct image 805. The loss function is the Euclidean norm in step S403 and the weighted sum of the following two terms. The first term is called Content Loss, which is obtained by converting the output image 804 and the correct image 805 into feature maps and taking the Euclidean norm of the difference for each element. By adding the difference in the feature maps to the loss function, the more abstract properties of the output image 804 can be made closer to the correct image 805. The second term is called Adversarial Loss, which is the sigmoid cross entropy of the discrimination label obtained by inputting the output image 804 to the discriminator 813. By training the discriminator 813 to discriminate between real and fake, an output image 804 that looks more subjectively like the correct image 805 can be obtained.
[0085] Subsequently, in step S409, the update unit 701d determines whether the second learning has been completed. Similar to step S404, if the second learning has not been completed, the process returns to step S405. On the other hand, if the second learning has been completed, the update unit 701d stores the information on the weights of the trained generator 812 in the storage unit 701a.
[0086] Next, with reference to FIG. 18, the defocus shaping (estimation phase) executed by the control device 704 and the image estimation device 705 will be described. FIG. 18 is a flowchart regarding defocus shaping (generation of an estimated image). Each step in FIG. 18 is mainly executed by each part of the control device 704 or the image estimation device 705.
[0087] First, in step S501, the communication unit 704a of the control device 704 transmits a request regarding the captured image and the execution of the blurring process to the image estimation device 705. Subsequently, in step S601, the communication unit 705a of the image estimation device 705 receives and acquires the captured image and the processing request transmitted from the control device 704. Subsequently, in step S602, the acquisition unit 705c of the image estimation device 705 acquires information on the learned weights corresponding to the captured image from the storage unit 705b. The weight information has been read out from the storage unit 701a in advance and stored in the storage unit 705b.
[0088] Subsequently, in step S603, the acquisition unit 705c acquires information regarding the state of the optical system corresponding to the captured image and generates input data. From the metadata of the captured image, information for specifying the type, zoom, aperture, and focal length of the lens device (optical system) 702 when the captured image was captured is acquired, and a lens state map is generated in the same manner as in FIG. 17. The input data is generated by concatenating the captured image and the lens state map in a predetermined order in the channel direction. In this embodiment, the captured image or the feature map based on the captured image and the state map can be concatenated in the channel direction before or during the input to the machine learning model.
[0089] Subsequently, in step S604, the shaping unit 705d inputs the input data into a generator and generates a blurred and shaped estimated image. The weight information is used for the generator. Subsequently, in step S605, the communication unit 705a transmits the estimated image to the control device 704. Subsequently, in step S502, the communication unit 704a of the control device 704 acquires the estimated image transmitted from the image estimation device 705.
[0090] In this embodiment, instead of the shaping unit 705d that shapes the blur included in the captured image, a sharpening unit that sharpens the captured image may be provided. In this embodiment, the image processing system 700 includes a first device (control device 704) and a second device (image estimation device 705) that can communicate with each other. The first device has a transmission means (communication unit 704a) that transmits a request regarding the execution of processing on the captured image to the second device. The second device includes a reception means (communication unit 705a), an acquisition means (acquisition unit 705c), and a generation means (shaping unit 705d or sharpening unit). The reception means receives the request. The acquisition means acquires input data including the captured image and information regarding the state of the optical system used for capturing the captured image. The generation means inputs the input data into a machine learning model based on the request, and generates an estimated image obtained by sharpening the captured image or an estimated image obtained by shaping the blur included in the captured image.
[0091] According to this embodiment, it is possible to realize an image processing apparatus and an image processing system that can suppress the learning load and the amount of data to be held of the machine learning model and accurately shape the blur caused by the optical system of the captured image.
[0092] Note that in this embodiment, an example has been described in which in step S501 control, the communication unit 704a of the control device 704 transmits the captured image together with a request for processing the captured image to the image estimation device 705. However, the transmission of the captured image by the control device 704 is not essential. For example, the control device may transmit only a request for processing the captured image to the image estimation device 705, and the image estimation device 705 that has received the request may be configured to acquire the captured image corresponding to the request from another image storage server or the like.
[0093] (Other Embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiment to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. Further, it can also be realized by a circuit (for example, ASIC) that realizes one or more functions.
[0094] According to each embodiment, it is possible to provide an image processing method, an image processing apparatus, an image processing system, a method for manufacturing learned weights, and a program that suppress the learning load and the amount of data to be held of a machine learning model and sharpen or shape the blur of a captured image with high accuracy.
[0095] As described above, the preferred embodiments of the present invention have been described. However, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist thereof.
Description of Reference Numerals
[0096] 123 Image processing unit (image processing apparatus) 123a Acquisition unit (acquisition means) 123b Sharpening unit (generation means)
Claims
1. An image processing method carried out using a processor, comprising: A first step of acquiring a captured image and first information on a state of an optical system used to capture the captured image at the time of capturing the image; a second step of generating a state map indicating a state of the optical system based on the first information; and a third step of generating an estimated image by sharpening the captured image or shaping blur included in the captured image using a machine learning model based on the captured image and the state map, The image processing method according to the present invention, wherein the state map includes numerical values indicating at least two of the zoom, aperture, and focus distance of the optical system as elements of different channels.
2. The image processing method according to claim 1, characterized in that in the third step, the machine learning model with the same weighting is used for a first captured image obtained by capturing an image in a first state of the optical system and a second captured image obtained by capturing an image in a second state of the optical system different from the first state.
3. the first information includes a numerical value indicating at least one of a zoom, an aperture, and a focus distance of the optical system, 3. The image processing method according to claim 1, wherein the numerical values are normalized based on a possible range of the optical system with respect to at least one of zoom, aperture, and focus distance.
4. 4. The image processing method according to claim 1, wherein the state map is generated based on the number of pixels of the captured image.
5. 5. The image processing method according to claim 1, wherein elements included in the same channel in the state map have the same numerical value.
6. 6. The image processing method according to claim 1, wherein the captured image or a feature map based on the captured image is linked to the state map in a channel direction.
7. the estimated image is generated using second information related to a position of each pixel of the captured image; 7. The image processing method according to claim 1, wherein the second information includes a numerical value normalized by a length based on an image circle of the optical system.
8. A method of image processing carried out using a processor, comprising: A first step of acquiring a captured image and first information on a state of an optical system used to capture the captured image at the time of capturing the image; a second step of generating a state map indicating a state of the optical system based on the first information; and a third step of generating an estimated image by sharpening the captured image or shaping blur included in the captured image using a machine learning model based on the captured image, the state map, and second information related to the position of each pixel of the captured image, The image processing method according to claim 1, wherein the second information includes a numerical value normalized by a length based on an image circle of the optical system.
9. 9. The image processing method according to claim 1, wherein the first information includes information relating to a type of the optical system.
10. 10. The image processing method according to claim 1, wherein the first information includes information on the presence or absence of an optical low-pass filter, or information on the type of the optical low-pass filter.
11. 11. The image processing method according to claim 1, wherein the state includes information regarding the presence or absence of an accessory of the optical system, or information regarding the type of the accessory.
12. 12. The image processing method according to claim 1, wherein the first information includes information regarding manufacturing variations of the optical system.
13. 13. The image processing method according to claim 1, wherein the estimated image is generated using information about a distribution of distances in a subject space at the time of capturing the image.
14. 14. The image processing method according to claim 1, wherein the estimated image is generated using information about a pixel pitch or a color filter array of an image sensor used for capturing the image.
15. The image processing method according to any one of claims 1 to 14, characterized in that the third step is performed by generating the estimated image by processing that is executed by inputting input data including the captured image and the state map into the machine learning model.
16. The image processing method according to any one of claims 1 to 15, characterized in that the machine learning model is trained in advance to output the estimated image based on input data including the captured image and the state map.
17. In the first step, weight information of the machine learning model is acquired; The image processing method according to any one of claims 1 to 16, characterized in that in the third step, input data including the captured image and the state map is input to the machine learning model based on the weight information.
18. The image processing method according to claim 17, characterized in that the weight information is pre-trained to perform different processing depending on the first information using training input data including a training image and information regarding the state of the optical system corresponding to the training image.
19. A program causing a computer to execute the image processing method according to any one of claims 1 to 18.
20. an acquisition means for acquiring a captured image and first information relating to a state of an optical system used to capture the captured image at the time of capturing the image; a first generating means for generating a state map indicating a state of the optical system based on the first information; A generation means for generating an estimated image by sharpening the captured image or shaping blur included in the captured image using a machine learning model based on the captured image and the state map, The image processing device, wherein the state map includes numerical values indicating at least two of the zoom, aperture, and focus distance of the optical system as elements of different channels.
21. 21. An image processing system comprising the image processing device according to claim 20 and a control device capable of communicating with the image processing device, the control device has a transmission means for transmitting a request for execution of processing on the captured image to the image processing device, The image processing device includes: an image processing system comprising: a receiving means for receiving the request, and generating the estimated image in response to the request;
Citation Information
Patent Citations
Image processing method and apparatus, image capturing apparatus, and image processing program
JP2011123589A
Focus correction processing method by learning type algorithm
JP2017199235A
Distance measuring device, distance measuring system, imaging apparatus, mobile body, method for controlling distance measuring device, and program
JP2019078716A
Machine learning data generation method, machine learning data generation program, machine learning data generation system, server device, image processing method, image processing program, and image processing system
WO2019009007A1