Image processing method, image processing apparatus, and program
Patent Information
- Application Number
- JP2022164616
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-10-13
- Publication Date
- 2025-10-17
AI Technical Summary
【0008】 本発明によれば、高感度画像のアップスケールによる弊害の発生を抑制することができる。
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an image processing method for upscaling an image using a machine learning model. [Background technology]
[0002] A method of upscaling an image using a machine learning model is known. Image processing that increases the resolution by estimating high-frequency components that cannot be expressed in a low-resolution image is called upscaling. In addition, an image (high-sensitivity image) that contains a lot of noise (high-sensitivity noise) is obtained by capturing an image with a high sensor sensitivity of an imaging device. When a high-sensitivity image is upscaled using a machine learning model, there is a risk that excessive high-frequency components will appear as artifacts in the image after image processing.
[0003] Patent Document 1 discloses an image processing method that uses a machine learning model to simultaneously increase the resolution and remove high-sensitivity noise. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] US2019 / 0114742 Summary of the Invention [Problem to be solved by the invention]
[0005] In Patent Document 1, the parameters of the machine learning model are set based on the noise reduction level selectable by the user, and the strength of the resolution enhancement and the strength of the high-sensitivity noise removal are adjusted. In the image processing method of Patent Document 1, in order to increase the strength of the resolution enhancement, it is necessary to lower the strength of the high-sensitivity noise removal. Therefore, in processing a high-sensitivity image acquired when the sensor sensitivity of the imaging device is high, if the resolution is enhanced to a desired resolution, many adverse effects may occur in the image after image processing.
[0006] Therefore, an object of the present invention is to suppress the occurrence of adverse effects caused by upscaling of high-sensitivity images. [Means for solving the problem]
[0007] An image processing method according to one aspect of the present invention is characterized by comprising a first step of generating a second image by removing noise from a first image, and a second step of generating a third image based on the second image using an upscaling machine learning model. Effect of the Invention
[0008] According to the present invention, it is possible to suppress the occurrence of adverse effects caused by upscaling of high-sensitivity images. [Brief description of the drawings]
[0009] [Figure 1] 1 is a block diagram of an image processing system according to a first embodiment. [Diagram 2] 1 is an external view of an image processing system according to a first embodiment. [Diagram 3] FIG. 1 is a diagram showing a learning flow of a machine learning model in a first embodiment. [Figure 4] 1 is a flowchart showing a learning process of a machine learning model in the first embodiment. [Diagram 5] 1 is a flowchart showing an estimation process of a machine learning model in the first embodiment. [Figure 6] FIG. 11 is a block diagram of an image processing system according to a second embodiment. [Figure 7] FIG. 11 is an external view of an image processing system according to a second embodiment. [Figure 8] FIG. 11 is a diagram showing the flow of learning of a machine learning model in a second embodiment. [Figure 9] 11 is a flowchart showing a learning process of a machine learning model in the second embodiment. [Figure 10] 11 is a flowchart showing an estimation process of a machine learning model in the second embodiment. [Figure 11] FIG. 11 is a block diagram of an image processing system according to a third embodiment. [Figure 12] 13 is a flowchart showing an estimation process of a machine learning model in the third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. In each drawing, the same reference numerals are used to refer to the same components, and duplicated explanations will be omitted.
[0011] First, before concretely describing the examples, the gist of the present invention will be described. The high sensitivity noise in this embodiment is noise that occurs in an image captured when the sensor sensitivity of the imaging device is high, and includes color noise (color noise, false color) and luminance noise. The color noise is noise that randomly generates coloring of pixels such as red, blue, and green that do not exist in the subject. On the other hand, the luminance noise is noise that randomly generates achromatic roughness.
[0012] In this embodiment, for example, a neural network, a genetic programming, a Bayesian network, etc. can be adopted for the machine learning. In addition, the neural network is, for example, a convolutional neural network (CNN), a generative adversarial network (GAN), a recurrent neural network (RNN), etc.
[0013] In this embodiment, one of the features is that after removing noise from the captured image, an upscaled image is generated using a machine learning model. With this configuration, noise is not overemphasized in the upscaling of the captured image. Therefore, it is possible to suppress the occurrence of adverse effects caused by upscaling a high-sensitivity image.
[0014] The image processing method described above is merely an example, and the present invention is not limited to this. Details of other image processing methods will be described in the following examples.
[0015] [Example 1] First, an image processing system 100 according to a first embodiment will be described with reference to Fig. 1 and Fig. 2. In this embodiment, a machine learning model is used to learn and execute image processing for upscaling an image containing a lot of noise after removing the noise.
[0016] Fig. 1 is a block diagram of an image processing system 100 in this embodiment. Fig. 2 is an external view of the image processing system 100. The image processing system 100 includes a learning device 101, an imaging device 102, an image estimation device 103, a display device 104, a recording medium 105, an input device 106, an output device 107, and a network 108.
[0017] The learning device 101 includes a storage unit 101a, an acquisition unit 101b, a noise removal unit 101c, and a learning unit (learning means) 101d, and updates the weights of a machine learning model.
[0018] The imaging device 102 has an optical system 102a and an imaging element 102b. The optical system 102a collects incident light from the subject space to the imaging device 102. The imaging element 102b receives an optical image of the subject formed via the optical system 102a to obtain a captured image. The imaging element 102b is a charge coupled device (CCD) sensor or a complementary metal oxide semiconductor (CMOS) sensor. The imaging element 102b is capable of setting the sensitivity at the time of shooting. The imaging device 102 can transmit information related to the shooting or information related to development together with the acquired image to an acquisition unit 103b of the image estimation device 103 described later. The information related to the shooting includes, for example, the pixel pitch of the imaging element 102b in the shooting using the imaging device 102, the ISO sensitivity of the imaging element 102b, and the type of the optical low pass filter of the optical system 102a. The information related to development includes the luminance noise removal strength, color noise removal strength, sharpness strength, and image compression ratio during development processing of each image acquired by the imaging device 102. Note that a storage unit that stores the acquired image, a display unit that displays the image, a transmission unit that transmits the image to the outside, an output unit that stores the image in an external storage medium, and the like are not shown. Also not shown is a control unit that controls each unit of the imaging device 102.
[0019] The image estimation device 103 includes a storage unit 103a, an acquisition unit 103b, a noise removal unit 103c, a processing unit 103d, and a noise addition unit 103e. In the image estimation device 103, the noise removal unit 103c removes color noise from a low-resolution image (first image) acquired by the acquisition unit 103b, and then the processing unit 103d performs upscaling processing using a machine learning model (upscale machine learning model). In addition, the noise addition unit 103e may perform image processing for adding noise to an image (third image) obtained by upscaling as necessary. With such a configuration, it is possible to reproduce the sense of noise in the captured image after upscaling. Furthermore, information related to imaging or information related to development acquired by the acquisition unit 103b may be used when generating an estimated image. Note that the first image is not limited to an image acquired by the imaging device 102, and may be an image stored in the recording medium 105. The functions of the image estimation device 103 can be realized by one or more processors (processing means).
[0020] Weight information of the machine learning model used in the image estimation device 103 is read from the storage unit 103a. The weight information is learned by the learning device 101, and the image estimation device 103 reads the weight information from the storage unit 101a via the network 108 in advance and stores it in the storage unit 103a. The stored weight information may be the weight numerical value itself or may be in an encoded format. Details regarding weight learning and image processing using the weights will be described later.
[0021] The output image is output to at least one of a display device 104, a recording medium 105, and an output device 107. The display device 104 is, for example, a liquid crystal display or a projector. A user can check an image being processed via the display device 104 and perform image editing work or the like via the input device 106. The recording medium 105 is, for example, a semiconductor memory, a hard disk, a server on a network, etc. The input device 106 is, for example, a keyboard or a mouse, etc. The output device 107 is, for example, a printer, etc.
[0022] Next, the learning phase (manufacturing method of a trained model) executed by the learning device 101 in this embodiment will be described with reference to Fig. 3 and Fig. 4. Fig. 3 is a diagram showing the flow of updating (learning) the weights of a machine learning model. Fig. 4 is a flowchart related to the weight update. Each step in Fig. 4 is mainly executed by the acquisition unit 101b, the noise removal unit 101c, and the learning unit 101d.
[0023] In the convolution layer CN in FIG. 3, the sum of the input, the convolution of the filter, and the bias is calculated, and the result is nonlinearly transformed by the activation function. The initial values of each component of the filter and the bias are arbitrary, and in this embodiment, they are determined by random numbers. The activation function can be, for example, ReLU (Rectified Linear Unit) or a sigmoid function. The multidimensional array output in each layer except the final layer is the feature map. The feature map generally has four dimensions: batch, vertical, horizontal, and channel. Skip concatenation SC combines feature maps output from discontinuous layers. The combination of feature maps may be done by taking the sum of each element, or by concatenation in the channel direction. In this embodiment, the sum of each element is adopted.
[0024] A residual block (RB) is an element (block or module) that combines multiple convolution layers (CN). To perform more accurate learning, learning may be performed using a network in which residual blocks are multi-layered, called a residual network. In this embodiment, a residual network is used as the multi-layered network, but this is not limited to this. For example, a network may be configured by multi-layering using elements such as an inception module and a dense block.
[0025] In addition, in layers close to the output, pixel shuffle PS is used to expand the low-resolution feature map to a high-resolution feature map. In this embodiment, pixel shuffle is used to expand the low-resolution feature map, but this is not limited to this. For example, deconvolution (deconvolution or transposed convolution), interpolation, etc. may be used to expand the feature map.
[0026] In updating the weights of the machine learning model, first, in step S101, the acquisition unit 101b acquires the first correct answer image 10 and the first training image 11. In this embodiment, the first correct answer image 10 is a high-resolution patch, and the first training image 11 is a low-resolution patch corresponding to the high-resolution patch. In addition, a patch is an image having a predetermined number of pixels. For example, the low-resolution patch is 128×128 pixels, and the corresponding high-resolution patch is 256×256 pixels. In this case, the upscaling magnification is 2 times (the number of pixels is 4 times). Note that the upscaling magnification is not limited to this, and may be any magnification as long as the corresponding first correct answer image 10 and first training image 11 can be acquired. Note that the high sensitivity noise in this embodiment is color noise.
[0027] In this embodiment, the first training image 11 is an image obtained by imaging an object with an imaging device 102. The first correct answer image 10 is generated by numerical calculation so as to correspond to an image obtained by imaging the same object as the first training image 11 with an imaging device having an imaging element with a smaller pixel pitch and less influence (aberration and diffraction) of the optical system than the imaging device 102. However, the method of obtaining the first correct answer image 10 and the first training image 11 is not limited to this. For example, the first correct answer image 10 and the first training image 11 may be obtained by cutting out corresponding parts of two images obtained by imaging the same object using an optical system with different focal lengths. The first correct answer image 10 obtained by imaging an object with an arbitrary imaging device may be downsampled to generate the corresponding first training image 11.
[0028] In step S102, the noise removal unit 101c generates a high-resolution patch (second answer image) 12 from which color noise has been removed, and a low-resolution patch (second training image) 13 from which color noise has been removed. The second answer image 12 and the second training image 13 are generated by removing color noise from the first answer image 10 and the first training image 11. A known method can be used to remove color noise. For example, a method can be used in which an image is decomposed into luminance and color difference, and a discrete cosine transform is performed on the color difference to perform threshold processing on the coefficients of the frequency components obtained. Even when the high-sensitivity noise is luminance noise, noise removal can be performed by a known method.
[0029] Also, when removing high sensitivity noise, information related to imaging or information related to development may be used. For example, when the ISO sensitivity is used among the information related to imaging, color noise is removed at a high intensity in step S102 when the ISO sensitivity is high. When the color noise removal intensity is used among the information related to development, color noise is removed at a high intensity in step S102 when the intensity of color noise removal during development is weak. When the information related to imaging or information related to development is used, in step S101, the acquisition unit 101b acquires the first answer image 10 and the first training image 11, and also acquires the corresponding information related to imaging or information related to development.
[0030] In step S103, the learning unit 101d generates an estimated image 14. The estimated image 14 is generated by upscaling the second training image 13 using the machine learning model. The estimated image 14 ideally coincides with the second ground truth image 12. If necessary, information about imaging or information about development may be used during upscaling. For example, when upscaling using the ISO sensitivity among the information about imaging, the strength of upscaling can be reduced so as not to overemphasize luminance noise when the ISO sensitivity is high. At this time, by linking a feature map having information about imaging or information about development for each pixel position in the channel direction of the first training image 11, the information about imaging or information about development can be input to the machine learning model. Note that both the information about imaging and the information about development may be input to the machine learning model.
[0031] In step S104, the learning unit 101d updates (learns) the weights of the machine learning model based on the error between the second ground truth image 12 and the estimated image 14. In this embodiment, the weights are updated using mini-batch learning. The weights include the filter components and biases of each layer. In addition, backpropagation is used to update the weights. In mini-batch learning, the errors between the multiple second ground truth images 12 and the estimated images 14 corresponding to the multiple second ground truth images 12 are calculated, and the weights are updated. For example, the L2 norm or the L1 norm can be used as the loss function. Note that the weight update method (learning method) is not limited to mini-batch learning, and may be batch learning or online learning.
[0032] In step S105, the learning unit 101d determines whether the weight update is complete. Completion of the update can be determined by, for example, whether the number of iterations of learning (weight update) reaches a specified value, or whether the amount of change in the weight at the time of update is smaller than a specified value. If it is determined that the weight update is not complete, the process returns to step S101, and a plurality of new first answer images 10 and first training images 11 are obtained. On the other hand, if it is determined that the weight update is complete, the learning device 101 ends the learning, and stores the weight information in the storage unit 101a.
[0033] Next, the estimation phase in this embodiment will be described with reference to Fig. 5. Fig. 5 is a flowchart of the estimation phase in this embodiment. Each step in Fig. 5 is mainly executed by the acquisition unit 103b, the noise removal unit 103c, the processing unit 103d, and the noise addition unit 103e.
[0034] First, in step S201, the acquisition unit 103b acquires a captured image (first image). In this embodiment, the first image is transmitted from the imaging device 102, but the present invention is not limited to this. It should be noted that information related to imaging or information related to development may be acquired together with the first image. The information related to imaging is information based on the optical system or image sensor of the imaging device used when acquiring the first image. The information related to development is information based on a development process in which the first image is developed from a RAW image. When the information related to imaging or information related to development is acquired, image processing may be performed using the information related to imaging or information related to development as necessary in subsequent processing.
[0035] In step S202, the noise removal unit 103c generates an image (second image) by removing high-sensitivity noise from the first image. When high-sensitivity noise removal is performed using information related to image capture or information related to development, it can be performed in the same manner as in step S102.
[0036] In step S203, the processing unit 103d generates an estimated image (third image) using the machine learning model. The third image is generated by upscaling the second image using the machine learning model. Note that the weight information of the machine learning model is transmitted from the learning device 101 and stored in advance in the storage unit 103a. When inputting the second image to the machine learning model, it is not necessary to cut out the image to the same size as the low-resolution patch used during learning, but in order to speed up processing, the second image may be decomposed into multiple overlapping patches and then processed. In this case, the patches obtained after processing can be synthesized to form the third image.
[0037] In step S204, the noise adding unit 103e generates an image (fourth image) based on the third image and the noise map. The noise adding unit 103e generates the noise map based on the first image and the second image. The noise map is an image having a value related to the noise removed for each pixel position. For example, the noise map is generated based on the difference between each pixel value between the first image and the second image. Furthermore, the noise adding unit 103e generates the fourth image by adding the noise map to the third image. The noise map added to the third image may be subjected to an interpolation process according to the resolution of the third image. By performing processing after making the noise map the same resolution as the third image by interpolation, the noise feeling of the image before upscaling can be reproduced. Note that the resolution of the noise map is not limited to the same resolution as the third image.
[0038] The fourth image may be generated by, for example, taking a weighted average of the third image and a noise map. For example, the ISO sensitivity of the image sensor used to obtain the first image may be used as the weight of the weighted average. Since high-sensitivity noise is more likely to increase as the ISO sensitivity is increased, using the ISO sensitivity as the weight allows for a process of synthesizing the third image and noise with high accuracy.
[0039] In this embodiment, the fourth image generated in step S204 is the output image. If there is no need to add noise to the third image, step S204 may not be performed. In that case, the third image may be the output image.
[0040] In this embodiment, the learning device 101 and the image estimation device 103 are separate devices, but the present invention is not limited to this. The learning device 101 and the image estimation device 103 may be integrated. In other words, the learning phase and the estimation phase may be performed within a single device.
[0041] With the above configuration, it is possible to upscale a captured image without emphasizing noise through upscaling, thereby reducing the adverse effects of noise that occurs in a generated high-resolution image.
[0042] [Example 2] Next, an image processing system 200 according to a second embodiment will be described with reference to Fig. 6 and Fig. 7. In this embodiment, a machine learning model is used to learn and execute image processing for upscaling an image containing a large amount of high-sensitivity noise after removing the high-sensitivity noise. The image processing system 200 of this embodiment differs from the first embodiment in that an image capturing device 202 acquires a captured image and processes the image, and in that a machine learning model is used to remove noise.
[0043] Fig. 6 is a block diagram of an image processing system 200 in this embodiment. Fig. 7 is an external view of the image processing system 200. The image processing system 200 has a learning device 201 and an imaging device 202, and the learning device 201 and the imaging device 202 are connected via a network 203. Note that the learning device 201 and the imaging device 202 do not need to be always connected via the network 203.
[0044] The learning device 201 includes a storage unit 201a, an acquisition unit 201b, a noise removal unit 201c, and a learning unit (learning means) 201d, and updates the weights of the machine learning model.
[0045] The imaging device 202 has an optical system 221, an imaging element 222, an image estimation unit 223, a storage unit 224, a recording medium 225, a display unit 226, an input unit 227, and a system controller 228. The imaging device 202 acquires a captured image (first image) by capturing an image of a subject space, and executes image processing to upscale the captured image.
[0046] The optical system 221 and the image sensor 222 are similar to the optical system 102a and the image sensor 102b in the embodiment 1. The image estimation unit 223 has an acquisition unit 223a, a noise removal unit 223b, a processing unit 223c, and a noise addition unit 223d, and performs image processing (estimation) using weight information of a trained machine learning model. The weight information of the machine learning model is trained in advance by the learning device 201, and the weight information is read out via the network 203 and stored in the storage unit 224.
[0047] The generated output image is stored in the recording medium 225. When an instruction to display an upscaled image is given from the user via the input unit 227, the stored image is read out and displayed on the display unit 226. Note that the upscaled image may be generated by the image estimation unit 223 by reading out the captured image stored in the recording medium 225, information related to the capture thereof, and information related to development. The above series of controls are performed by the system controller 228.
[0048] Next, the learning phase (manufacturing method of a trained model) executed by the learning device 201 in this embodiment will be described with reference to Figs. 8 and 9. Fig. 8 is a diagram showing a flow of updating (learning) the weights of a machine learning model. Fig. 9 is a flowchart related to updating the weights. Each step in Fig. 9 is mainly executed by the acquisition unit 201b, the noise removal unit 201c, and the learning unit 201d. This embodiment differs from the first embodiment in that an image before noise removal is used as an input image for the machine learning model.
[0049] In the learning phase of the machine learning model in this embodiment, first, in step S301, the acquisition unit 201b acquires a first supervised image 20 and a first training image 21. The first supervised image 20 and the first training image 21 correspond to the first supervised image 10 and the first training image 11 in the first embodiment, respectively.
[0050] In step S302, the noise removal unit 201c generates the second supervised image 22 and the third supervised image 23. The noise removal unit 201c generates the second supervised image 22 and the third supervised image 23 by removing high-sensitivity noise from the first supervised image 20 and the first training image 21. The second supervised image 22 and the third supervised image 23 correspond to the second supervised image 12 and the second training image 13 in the first embodiment, respectively.
[0051] In step S303, the learning unit 201d generates a noise-removed estimated image (first estimated image) 25. The first estimated image 25 is generated by removing high-sensitivity noise from the first training image 21 using a machine learning model (noise-removing machine learning model). The first estimated image 25 ideally matches the third ground truth image 23. If necessary, noise may be removed based on information about imaging and information about development by inputting information about imaging and information about development to the machine learning model together with the first training image 21. At this time, for example, information about imaging and information about development can be input to the machine learning model by linking a feature map having information about imaging and information about development for each pixel position in the channel direction of the first training image 21. Note that the machine learning model in this embodiment can generate noise feature information 24 in the process of generating the first estimated image 25 (intermediate layer). In this embodiment, the noise feature information 24 is a plurality of signal sequences (feature map) in which signal values related to noise are spatially arranged. Furthermore, the noise feature information 24 has signal sequences, the number of which may be, for example, 64 or 128, depending on the number of pixels and resolution of the first training image 21 .
[0052] In step S304, the learning unit 201d generates an upscaled estimated image (second estimated image) 26. The second estimated image 26 is generated by upscaling the first estimated image 25 using a machine learning model (upscale machine learning model). The second estimated image 26 ideally matches the second ground truth image 22. Note that, instead of the first estimated image 25, noise feature information 24 may be used as information to be input to the machine learning model. Since the noise feature information 24 has more information than the first estimated image 25, the accuracy of the upscaling process performed by the machine learning model can be improved. By inputting more information to the machine learning model, the first machine learning model can more accurately distinguish between high-frequency components and noise components of the subject included in the image, and adverse effects occurring in the image after image processing can be reduced.
[0053] The noise feature information 24 and the first estimated image 25 may be input to a machine learning model to generate a second estimated image 26. At this time, the noise feature information 24 can be input to the machine learning model, for example, by linking the noise feature information 24 in the channel direction of the first estimated image 25. If necessary, upscaling may be performed based on the information on the imaging and the information on the development by inputting the first estimated image 25 together with the information on the imaging and the information on the development to the machine learning model. The information on the imaging and the information on the development can be input to the machine learning model, for example, by linking a feature map having the information on the imaging and the information on the development for each pixel position in the channel direction of the first estimated image 25.
[0054] In step S305, the learning unit 201d updates (learns) the weights of the machine learning model. The learning unit 201d in this embodiment updates the weights of the noise removal machine learning model based on the error between the second ground truth image 22 and the second estimated image 26. Also, the learning unit 201d updates the weights of the upscale machine learning model based on the error between the third ground truth image 23 and the first estimated image 25. At this time, for example, the L2 norm or the L1 norm can be used as the loss function.
[0055] In step S306, the learning unit 201d determines whether the weight update is complete. Completion of the update can be determined by, for example, whether the number of iterations of learning (weight update) reaches a specified value, or whether the amount of change in the weight at the time of update is smaller than a specified value. If it is determined that the weight update is not complete, the process returns to step S301, and a plurality of new first answer images 20 and first training images 21 are obtained. On the other hand, if it is determined that the weight update is complete, the learning device 201 ends the learning, and stores the weight information in the storage unit 201a.
[0056] In this embodiment, image processing using a noise removal machine learning model and an upscale machine learning model has been described, but the present invention is not limited thereto. For example, a single machine learning model capable of generating images corresponding to the first estimated image 25 and the second estimated image 26 in stages may be generated by performing the same steps as in this embodiment. Furthermore, the learning device 201 of this embodiment may have a first learning means for learning the noise removal machine learning model and a second learning means for learning the upscale machine learning model, as necessary. Also, the first learning means and the second learning means may be functions of different devices.
[0057] Next, generation of an output image using a machine learning model in this embodiment will be described with reference to Fig. 10. Fig. 10 is a flowchart related to generation of an output image in this embodiment. Each step in Fig. 10 is mainly executed by the acquisition unit 223a, the noise removal unit 223b, the processing unit 223c, and the noise addition unit 223d of the image estimation unit 223.
[0058] In this embodiment, image processing is performed using the above-described noise reduction machine learning model and upscaling machine learning model. Note that the learned weights are those stored in the storage unit 224.
[0059] First, in step S401, the acquisition unit 223a acquires a captured image (first image). In this embodiment, the first image is acquired by the imaging device 202 and stored in the storage unit 224, but this embodiment is not limited to this. If necessary, information related to imaging and information related to development corresponding to the first image may be acquired.
[0060] In step S402, the noise removal unit 223b generates a first estimated image (corresponding to the second image in the first embodiment) using the noise removal machine learning model. In this embodiment, the first estimated image is an image in which high-sensitivity noise has been removed from the first image. At this time, the noise removal machine learning model can generate noise feature information in the process of generating the first estimated image. The noise feature information in the estimation phase corresponds to the noise feature information 24 in the learning phase. Note that the weight information of the noise removal machine learning model is transmitted from the learning device 201 and stored in advance in the storage unit 224.
[0061] In step S403, the processing unit 223c generates a second estimated image (corresponding to the third image in the first embodiment) using the upscale machine learning model. The second estimated image is generated by inputting at least one of the first estimated image generated by the noise removal machine learning model and the noise feature information to the upscale machine learning model. Note that the weight information of the upscale machine learning model is transmitted from the learning device 201 and stored in advance in the storage unit 224.
[0062] In step S404, the noise adding unit 223d adds noise to the second estimated image. The method of adding noise is the same as that in step S204 in the first embodiment. If there is no instruction from the user to add noise, step S204 may not be performed and the second estimated image may be used as the output image. If there is an instruction from the user, the estimated image to which noise has been added (corresponding to the fourth image in the first embodiment) is used as the output image.
[0063] [Example 3] Next, an image processing system 300 according to a third embodiment will be described with reference to Figs. 11 and 12. In this embodiment, a machine learning model is used to learn and execute image processing for upscaling an image containing a large amount of high-sensitivity noise after removing the high-sensitivity noise. The image processing system 300 of this embodiment differs from the first embodiment in that it includes a control device 304 that acquires an image to be processed from an imaging device 302 and requests an image estimation device (image processing device) 303 to process the image to be processed. In this embodiment, if the control device 304 is a user terminal, it is possible to reduce the processing load on the user terminal. Therefore, the user side can obtain an output image with a low processing load.
[0064] 11 is a block diagram of an image processing system 300 in this embodiment. The image processing system 300 includes a learning device 301, an imaging device 302, an image estimation device 303, and a control device 304. The learning device 301 and the image estimation device 303 may be servers. The control device 304 is a user terminal such as a personal computer or a smartphone. The control device 304 is connected to the image estimation device 303 via a network 305. The image estimation device 303 is connected to the learning device 301 via a network 306. That is, the control device 304 and the image estimation device 303, and the image estimation device 303 and the learning device 301 are configured to be able to communicate with each other.
[0065] A learning device 301 and an imaging device 302 in the image processing system 300 have the same configurations as the learning device 101 and the imaging device 102, respectively, and therefore their explanations are omitted.
[0066] The image estimation device 303 includes a storage unit 303a, an acquisition unit 303b, a noise removal unit 303c, a processing unit 303d, a noise adding unit 303e, and a communication unit (receiving means) 303f. The storage unit 303a, the acquisition unit 303b, the noise removal unit 303c, the processing unit 303d, and the noise adding unit 303e are similar to the storage unit 103a, the acquisition unit 103b, the noise removal unit 103c, the processing unit 103d, and the noise adding unit 103e of the first embodiment, respectively. The communication unit 303f has a function of receiving a request transmitted from the control device 304, and a function of transmitting an output image generated by the image estimation device 303 to the control device 304.
[0067] The control device 304 has a communication unit (transmission means) 304a, a display unit 304b, an input unit 304c, a processing unit 304d, and a recording unit 304e. The communication unit 304a has a function of transmitting a request to the image estimation device 303 to cause the image estimation device 303 to execute processing on the captured image, and a function of receiving an output image processed by the image estimation device 303. The communication unit 304a may communicate with the imaging device 302. The display unit 304b has a function of displaying information. The information displayed by the display unit 304b includes, for example, a captured image to be transmitted to the image estimation device 303 and an output image received from the image estimation device 303. The input unit 304c receives an instruction to start image processing from a user. The processing unit 304d has a function of further performing image processing on the output image received from the image estimation device 303. The recording unit 304e stores the captured image (first image) acquired from the imaging device 302, the output image received from the image estimation device 303, and the like.
[0068] Next, the estimation phase in this embodiment will be described with reference to Fig. 12. Fig. 12 is a flowchart relating to the estimation phase in this embodiment.
[0069] First, the operation of the control device 304 will be described.
[0070] In step S501 (first transmission step), the control device 304 transmits a request for processing the first image to the image estimation device 303. Note that the method for transmitting the first image to be processed to the image estimation device 303 is not important. For example, the first image may be uploaded to the image estimation device 303 simultaneously with S501, or may be uploaded to the image estimation device 303 before S501. Furthermore, the first image may be an image stored on a server different from the image estimation device 303. Furthermore, in S501, the control device 304 may transmit an ID for authenticating a user, information related to imaging corresponding to the first image, information related to development, and the like together with the request for processing the first image.
[0071] In step S 502 (first receiving step), the control device 304 receives the output image generated by the image estimation device 303 .
[0072] Next, the operation of the image estimation device 303 will be described.
[0073] In step S601 (second receiving step), the communication unit 303f receives a request for processing the first image transmitted from the communication unit 304a. Upon receiving the instruction to process the first image, the image estimation device 303 executes the processes from step S602 onward.
[0074] In step S602, the acquisition unit 303b acquires a first image. In this embodiment, the first image is transmitted from the control device 304. Note that information related to imaging and information related to development transmitted together with the first image may be acquired. Note that the processes of steps S601 and S602 may be performed simultaneously. Also, since steps S603 to S605 are similar to steps S202 to S204, a description thereof will be omitted.
[0075] In step S606 (second transmission step), the image estimation device 303 transmits the output image to the control device 304. The output image transmitted by the image estimation device 303 includes at least one of the estimated image generated in step S604 or the estimated image to which noise has been added in step S605.
[0076] With the above configuration, it is possible to upscale a captured image without emphasizing noise by upscaling. Therefore, it is possible to reduce the adverse effects caused by noise occurring in a generated high-resolution image. In this embodiment, the control device 304 only requests processing for a specific image. Actual image processing is performed by the image estimation device 303. Therefore, if the control device 304 is a user terminal, it is possible to reduce the processing load on the user terminal.
[0077] (Other Examples) The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-mentioned embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.
[0078] According to each embodiment, it is possible to provide an image processing method, an image processing device, a program, and a storage medium that are capable of upscaling a captured image without emphasizing noise through upscaling.
[0079] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention.
[0080] The embodiments of the present invention include the following methods, configurations, and programs.
[0081] [Method 1] A first step of generating a second image by removing noise from a first image; and a second step of generating a third image based on the second image using an upscaled machine learning model.
[0082] [Method 2] The image processing method of method 1, wherein the second image is generated using a denoising machine learning model.
[0083] [Method 3] A first step of generating noise feature information from a first image using a denoising machine learning model; and a second step of generating a third image based on the noise feature information using an upscale machine learning model.
[0084] [Method 4] The image processing method according to Method 3, wherein the noise feature information is a plurality of signal sequences in which signal values relating to noise are spatially arranged.
[0085] [Method 5] 3. The image processing method according to method 1 or 2, wherein the noise is color noise.
[0086] [Method 6] the first image is an image acquired by imaging using an optical system and an imaging element, The second image is generated using information related to the imaging; The image processing method according to method 1 or 2, characterized in that the information regarding the imaging includes at least one of the pixel pitch of the imaging element, the ISO sensitivity of the imaging element, and the type of optical low-pass filter of the optical system.
[0087] [Method 7] the first image is an image that has been subjected to a development process; The second image is generated using information about the development; The image processing method according to any one of methods 1 and 2, characterized in that the information relating to the development includes at least one of the luminance noise reduction strength, the color noise reduction strength, the sharpness strength, and the image compression ratio during development.
[0088] [Method 8] the first image is an image acquired by imaging using an optical system and an imaging element, the third image is generated using information related to the imaging; The image processing method according to any one of Methods 1 to 7, characterized in that the information relating to the imaging includes at least one of the pixel pitch of the imaging element, the ISO sensitivity of the imaging element, and the type of optical low-pass filter of the optical system.
[0089] [Method 9] the first image is an image that has been subjected to a development process; the third image is generated using information regarding development; The image processing method according to any one of Methods 1 to 7, characterized in that the information relating to development includes at least one of luminance noise reduction strength, color noise reduction strength, sharpness strength, and image compression ratio during development.
[0090] [Method 10] generating a noise map based on pixel value differences between the first image and the second image; 3. The image processing method according to method 1 or 2, further comprising the step of generating a fourth image based on the third image and the noise map.
[0091] [Method 11] the first image is an image acquired by imaging using an optical system and an imaging element, The image processing method of method 10, wherein the fourth image is generated by taking a weighted average of the noise map and the third image based on the ISO sensitivity of the image sensor.
[0092] [Program 12] A program for causing a computer to execute the image processing method according to any one of Methods 1 to 11.
[0093] [Configuration 13] A storage medium storing the program described in Program 12.
[0094] [Configuration 14] 12. An image processing device comprising a processing means capable of executing the image processing method according to any one of Methods 1 to 11.
[0095] [Configuration 15] 15. An image processing system having a control device and an image processing device according to configuration 14, which are communicable with each other, 2. An image processing system according to claim 1, wherein the control device has a means for transmitting a request for execution of processing on the first image to the image processing device.
[0096] [Method 16] obtaining a first ground truth image and a first training image; generating a second ground truth image and a second training image by removing noise from the first ground truth image and the first training image; upscaling the second training image using a machine learning model to generate an estimated image; and updating weights of a neural network based on the estimated image and the second correct answer image.
[0097] [Configuration 17] A learning device characterized by having a learning means capable of executing the method for producing a trained model described in method 16. [Explanation of symbols]
[0098] S202 First step S203 Second process
Claims
1. a first step of generating a second image by reducing noise from a first image; and a second step of generating a third image by upscaling the second image using a first machine learning model based on the second image.
2. The image processing method of claim 1 , wherein the second image is generated using a second machine learning model.
3. A first step of generating noise feature information from a first image using a first machine learning model; and a second step of generating a third image by upscaling the second image based on the noise feature information using a second machine learning model.
4. 4. The image processing method according to claim 3, wherein the noise feature information is a signal sequence in which signal values relating to noise are spatially arranged.
5. 3. The image processing method according to claim 1, wherein the noise is color noise.
6. the first image is an image acquired by imaging using an optical system and an imaging element, the second image is generated using information related to the imaging; 3. The image processing method according to claim 1, wherein the information about the image capture includes at least one of a pixel pitch of the image sensor, an ISO sensitivity of the image sensor, and a type of an optical low-pass filter of the optical system.
7. the first image is an image that has been subjected to a development process, the second image is generated using information about development; 3. The image processing method according to claim 1, wherein the information relating to development includes at least one of a luminance noise reduction strength, a color noise reduction strength, a sharpness strength, and an image compression rate at the time of development.
8. the first image is an image acquired by imaging using an optical system and an imaging element, the third image is generated using information related to the imaging; 5. The image processing method according to claim 1, wherein the information relating to the imaging includes at least one of the pixel pitch of the imaging element, the ISO sensitivity of the imaging element, and the type of optical low-pass filter of the optical system.
9. the first image is an image that has been subjected to a development process, the third image is generated using information about development; 5. The image processing method according to claim 1, wherein the information relating to development includes at least one of a luminance noise reduction strength, a color noise reduction strength, a sharpness strength, and an image compression rate at the time of development.
10. generating a noise map based on pixel value differences between the first image and the second image; 3. The image processing method according to claim 1, further comprising the step of generating a fourth image based on the third image and the noise map.
11. the first image is an image acquired by imaging using an optical system and an imaging element, 11. The image processing method according to claim 10, wherein the fourth image is generated by taking a weighted average of the noise map and the third image based on the ISO sensitivity of the image sensor.
12. A program causing a computer to execute the image processing method according to any one of claims 1 to 4.
13. A storage medium storing the program according to claim 12.
14. 5. An image processing apparatus comprising processing means capable of executing the image processing method according to claim 1.
15. An image processing system comprising a control device and the image processing device according to claim 14, which are capable of communicating with each other, The image processing system according to claim 1, wherein the control device has means for transmitting a request for execution of processing on the first image to the image processing device.
16. obtaining a first ground truth image and a first training image; generating a second ground truth image and a second training image by removing noise from the first ground truth image and the first training image; generating an estimated image based on the second training image using a machine learning model, the estimated image being upscaled from the second training image; and updating the weights of a neural network based on the estimated image and the second correct image.
17. A learning device comprising a learning means capable of executing the trained model generation method according to claim 16.