Image processing methods, image processing devices, programs

JP7927926B2Active Publication Date: 2026-10-01CANON KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025082524
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2026-10-01
Estimated Expiration
2042-07-05

AI Technical Summary

Benefits of technology

【0007】 本開示によれば、機械学習モデルを用いた画像処理において、高解像度な出力画像を得ることができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007927926000001
    Figure 0007927926000001
  • Figure 0007927926000002
    Figure 0007927926000002
  • Figure 0007927926000003
    Figure 0007927926000003
Patent Text Reader

Abstract

To provide an image processing method for obtaining an output image with high resolution in image processing using a machine learning model.SOLUTION: An image processing method has: a step S203 of dividing a first gray scale image 21 to create a plurality of second gray scale images with lower resolution than the first gray scale image 21; and an estimation step S204 of inputting the plurality of second gray scale images 23 to create a plurality of upscaled third gray scale images 24.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to image processing using a machine learning model.

Background Art

[0002] Patent Document 1 discloses an image processing method for identifying features of a color image by converting an RGB color image into YUV and inputting information of high-frequency components of the obtained Y image (luminance image) to a machine learning model. In the image processing method of Patent Document 1, a Convolutional Neural Network (CNN), which generates an output image by convolving a filter with an input image multiple times, is used as the machine learning model. Furthermore, in the image processing method of Patent Document 1, by using a reduced luminance image as an input image, the amount of computation in the CNN is reduced, thereby achieving speeding-up of processing.

Prior Art Literature

Patent Literature

[0003]

Patent Literature 1

Summary of Invention

Problem to be Solved by the Invention

[0004] However, the input image in Patent Document 1 is a reduced luminance image, and has a lower resolution than the luminance image before reduction. Therefore, with the image processing method in Patent Document 1, it is difficult to obtain a high-resolution output image.

[0005] Accordingly, an object of the present invention is to obtain a high-resolution output image in image processing using a machine learning model.

Means for Solving the Problem

[0006] The image processing method of the present disclosure comprises: A computer-based image processing method,By dividing the first grayscale image, a resolution higher than that of the first grayscale image is obtained. low The steps include generating multiple second grayscale images and providing information on the shooting conditions corresponding to the first grayscale image. and before By concatenating multiple second grayscale images in the channel direction and inputting them into a machine learning model, The aforementioned The method comprises the step of generating a plurality of upscaled third grayscale images from a plurality of second grayscale images, wherein the information on the shooting conditions is a map showing the shooting conditions for each pixel. [Effects of the Invention]

[0007] According to this disclosure, high-resolution output images can be obtained in image processing using a machine learning model. [Brief explanation of the drawing]

[0008] [Figure 1] This is a block diagram of the image processing system in Example 1. [Figure 2] This is an external view of the image processing system in Example 1. [Figure 3] This is a conceptual diagram illustrating the method for learning the weights of the neural network in Example 1. [Figure 4] This is a flowchart illustrating the learning of neural network weights in Example 1. [Figure 5] This is a conceptual diagram illustrating the method for generating output images using a neural network in Example 1. [Figure 6] This is a flowchart illustrating the generation of output images using a neural network in Example 1. [Figure 7] This is a block diagram of the image processing system in Example 2. [Figure 8] This is an external view of the image processing system in Example 2. [Figure 9]This is a flowchart illustrating the generation of output images using a neural network in Example 2. [Figure 10] This is a block diagram of the image processing system in Example 3. [Figure 11] This is a flowchart illustrating the generation of output images using a neural network in Example 3. [Modes for carrying out the invention]

[0009] The embodiments of this disclosure will be described in detail below with reference to the drawings. In each drawing, the same reference numerals are used for the same components, and redundant descriptions are omitted.

[0010] First, before describing the specific embodiments, the gist of this embodiment will be explained. This embodiment upscales luminance images (grayscale images) using a machine learning model. In this embodiment, image processing that enlarges and increases the resolution of an image is referred to as upscaling. The machine learning model in this embodiment is generated by training using a neural network. The neural network uses a convolutional filter and a summation bias on the image, and an activation function that performs a nonlinear transformation. The filter and bias are called weights and are learned (updated) using training images and ground truth images. In this embodiment, the machine learning model is trained using grayscale images as training images and ground truth images.

[0011] The image processing method of this embodiment includes the step of generating a plurality of second grayscale images with lower resolution than the first grayscale image by dividing the first grayscale image. Furthermore, it is characterized by including the estimation step of generating a plurality of upscaled third grayscale images by inputting the plurality of second grayscale images into a machine learning model.

[0012] In this embodiment, the input image to the machine learning model is a grayscale image that is reduced in size relative to the original grayscale image obtained by reversibly dividing the original grayscale image. When generating the input image from the grayscale image, the original grayscale image can be reversibly reduced in size by being divided into multiple pieces. Therefore, since information loss accompanying the size reduction can be reduced, a highly accurate estimated image (output image) can be obtained. In addition, since the input image is obtained by reducing a grayscale image which has a smaller amount of information (number of channels) than a color image, the amount of computation in image processing can be reduced, and speeding up image processing is also one of the characteristics of this embodiment.

[0013] Note that the above image processing method is an example, and the present invention is not limited thereto. Details of other image processing methods and the like will be described in the following examples.

[0014] [Example 1] First, the image processing system 100 according to Example 1 of the present invention will be described. The image processing system 100 according to the present example trains and executes image processing for upscaling images using a machine learning model.

[0015] Figure 1 is a block diagram of the image processing system 100 according to the present example. Figure 2 is an external view of the image processing system 100. The image processing system 100 includes a learning device 101, an imaging device 102, an image estimation device 103, a display device 104, a recording medium 105, an input device 106, an output device 107, and a network 108.

[0016] The learning device 101 includes a storage unit (storage means) 101a, an acquisition unit (acquisition means) 101b, a generation unit (generation means) 101c, a division unit (division means) 101d, and a learning unit (learning means) 101e.

[0017] The imaging device 102 has an optical system 102a and an image sensor 102b. The optical system 102a collects light incident on the imaging device 102 from the subject space. The image sensor 102b receives the optical image of the subject formed through the optical system 102a and acquires the captured image 20. The image sensor 102b is a CCD (Charge Coupled Device) sensor or a CMOS (Complementary Metal-Oxide Semiconductor) sensor, etc.

[0018] The imaging device 102 transmits the obtained image to the acquisition unit 103b of the image estimation device (image processing device) 103, which will be described later. If necessary, the imaging device 102 may also transmit the shooting conditions corresponding to the captured image 20 along with the captured image 20. The shooting conditions are the conditions for imaging when acquiring the captured image 20 using the optical system 102a and the image sensor 102b. For example, these include the pixel pitch of the image sensor 102b, the type of optical low-pass filter of the optical system 102a, and the ISO sensitivity. Alternatively, the shooting conditions may also be the development conditions when acquiring the captured image 20 from an undeveloped RAW image in the imaging device 102. For example, these include noise reduction intensity, sharpness intensity, and image compression ratio. In this embodiment, development is the process of converting the RAW image into an image file such as JPEG (Joint Photographic Experts Group) or TIFF (Tag Image File Format).

[0019] Note that the storage unit for storing acquired images in the imaging device 102, the display unit for displaying them, the transmission unit for transmitting them to an external source, the output unit for saving them to an external storage medium, and the control unit for controlling each part of the imaging device 102 are not shown.

[0020] The image estimation device 103 includes a storage unit (storage means) 103a, an acquisition unit (acquisition means) 103b, a generation unit (generation means) 103c, a division unit (division means) 103d, and a processing unit (estimation means) 103e, and performs image processing on the acquired captured image 20 to generate an output image.

[0021] The acquisition unit 103b acquires the captured image 20. If necessary, the acquisition unit 103b may also acquire (receive) the shooting conditions corresponding to the captured image 20 along with the captured image 20.

[0022] The generation unit 103c extracts a Y image (luminance image) and multiple color difference images (first color difference image) by performing a YUV conversion on the acquired image 20. The luminance image is a grayscale image that represents luminance value information using only the shades of a single color. The color difference images are images that contain U and V information after the YUV conversion. Details of the YUV conversion will be described later.

[0023] The division unit 103d reduces the brightness image by dividing (deforming) the obtained brightness image.

[0024] The processing unit 103e generates an estimated image (output image) by performing image processing to enlarge and increase the resolution of the reduced brightness image (input image). The processing unit 103e may also perform image processing using the shooting conditions acquired by the acquisition unit 103b. For example, when training a machine learning model, by using the pixel pitch of the image sensor, the type of optical low-pass filter, and the image compression ratio in addition to the input image, image processing can be performed even on images acquired by any imaging device where the ground truth images corresponding to the training images are different. Details regarding image processing using shooting conditions will be described later. The captured image 20 may be an image captured by the imaging device 102 or an image stored on the recording medium 105. Furthermore, the captured image 20 may be an image that is represented in grayscale from the beginning, such as an infrared image or a depth image.

[0025] Image processing in this embodiment uses a neural network. The weight information in the neural network is learned by the learning device 101. The image estimation device 103 reads the weight information from the storage unit 101a via the network 108 and stores it in the storage unit 103a. The stored weight information may be the numerical weights themselves or in an encoded format. Details regarding weight learning and image processing using the weights will be described later. The image estimation device 103 has the function to perform development processing and other image processing as needed.

[0026] The output image is output to at least one of the display device 104, recording medium 105, and output device 107. The display device 104 is, for example, a liquid crystal display or a projector. The user can check the image in progress via the display device 104 and perform image editing work via the input device 106. The recording medium 105 is, for example, semiconductor memory, a hard disk, or a server on a network. The input device 106 is, for example, a keyboard or mouse. The output device 107 is, for example, a printer. The image estimation device 103 may also display or output the image after it has been colorized. The colorization process will be described later.

[0027] Next, with reference to Figures 3 and 4, the method for learning weights (method for manufacturing a trained model) performed by the learning device 101 in this embodiment will be described. Figure 3 is a conceptual diagram showing the learning (updating) of neural network weights. Figure 4 is a flowchart related to the learning of the neural network. In this embodiment, a Convolutional Neural Network (CNN) 30 is used as the neural network. However, this embodiment is not limited to this, and for example, a Generative Adversarial Network (GAN) or a Recurrent Neural Network (RNN) may also be used.

[0028] In this embodiment, the weights of CNN30 are trained (updated) using mini-batch learning. In mini-batch learning, the error between multiple ground truth images and their corresponding estimated images is calculated, and the weights are updated. For the loss function, for example, the L2 norm or L1 norm can be used. However, this embodiment is not limited to this, and online learning or batch learning may also be used.

[0029] The convolutional layer CN performs a filter convolution operation on the information input to CNN30, calculating the sum of the input information and the bias. Furthermore, the convolutional layer CN performs a nonlinear transformation on the result of the operation based on an activation function. The initial values ​​of each filter component and the bias are arbitrary and are determined by random numbers in this embodiment. The activation function may be, for example, ReLU (Rectified Linear Unit) or a sigmoid function. Each convolutional layer CN, except for the final layer, outputs a feature map. In this embodiment, the feature map is a 4-dimensional array with batch, length, width, and channel dimensions.

[0030] The skip connection SC synthesizes feature maps output from non-contiguous layers. In this embodiment, the feature maps are synthesized using a method that calculates the element-wise sum. Alternatively, the feature maps may be synthesized by concatenation in the channel direction.

[0031] Pixel shuffle (PS) is a method for expanding feature maps. In this embodiment, a high-resolution feature map is created by expanding a low-resolution feature map in a layer close to the output layer. Note that feature map expansion may also be performed using, for example, deconvolution (or transposed convolution).

[0032] A residual block (RB) is an element (block or module) that combines multiple convolutional layers (CN). To achieve higher accuracy in learning, a network called a residual network, which is a multi-layered network of residual blocks, may be used. In this embodiment, a residual network was used for the multi-layered network, but it is not limited to this. For example, a network may be constructed by using elements such as an inception module or a dense block in a multi-layered structure.

[0033] If necessary, the convolutional layer CN may reduce the processing load by shrinking the feature map in layers closer to the input layer and expanding it in layers closer to the output layer, thereby reducing the size of the feature map in the intermediate layers. Here, pooling and stride can be used to shrink the feature map. Deconvolution, pixel shuffling, and interpolation can be used to expand the feature map.

[0034] Next, we will explain the flowchart for neural network training. Each step in Figure 4 is mainly performed by the acquisition unit 101b, the generation unit 101c, the division unit 101d, and the learning unit 101e.

[0035] First, in step S101 (acquisition step), the acquisition unit 101b acquires the first correct patch 10 (first correct image) and the first training patch 11 (first training image). The first correct patch 10 and the first training patch 11 are grayscale images that include at least brightness information. In this embodiment, the first correct patch 10 has a larger image size and higher resolution than the first training patch 11, and depicts the same subject as the first training patch 11. A patch is an image with a predetermined number of pixels. For example, the first training patch 11 may be 128 × 128 × 1 pixel, and the corresponding first correct patch 10 may be 256 × 256 × 1 pixel. Note that the batch magnification is not limited to 2 times in both the vertical and horizontal directions; any magnification is acceptable as long as the first training patch 11 and the corresponding first correct patch 10 can be acquired. In this embodiment, the first training patch 11 and the corresponding first ground truth patch 10 are generated by numerical calculation, but the present invention is not limited thereto. For example, the first training patch 11 and the corresponding first ground truth patch 10 may be obtained by imaging the same subject with optical systems with different focal lengths and cropping the corresponding parts of the two resulting images. Alternatively, the first ground truth patch 10 may be downsampled to reduce its resolution and generate the first training patch 11. Furthermore, the first ground truth patch 10 and the first training patch 11 may each be obtained using luminance patches (grayscale images) obtained by YUV conversion of color patches. By YUV conversion of color patches, luminance patches and multiple color difference patches can be generated. The generation of luminance patches and multiple color difference patches from color patches is performed according to the following formula. However, this embodiment is not limited thereto, and other definition formulas may be used.

[0036] [Mathematics 1] Y = 0.299R + 0.587G + 0.114B U = -0.14713R - 0.28886G + 0.436B V = 0.615R - 0.54199G - 0.10001B The above formula is used to convert from the RGB color space to the YUV color space. The RGB color space is represented using three color channels: Red, Green, and Blue. On the other hand, YUV is represented using a luminance channel (Y) and two chrominance channels (U and V).

[0037] In this embodiment, the acquisition unit 101b acquires a first ground truth patch 10 and a first training patch 11, which are represented in grayscale. However, the acquisition unit 101b may also acquire a training color patch and a ground truth color patch having multiple color channels. In that case, the generation unit 101c generates the first ground truth patch 10 and the first training patch 11 from the ground truth color patch and the training color patch according to formula 1. In this embodiment, only one of the first ground truth patch 10 and the first training patch 11 may be generated from a color patch, and the other may be acquired by the acquisition unit 101b as a luminance patch.

[0038] Next, in step S102 (splitting step), the splitting unit 101d generates multiple second training patches 12 (second training images) by splitting the first training patch 11. The multiple second training patches 12 are generated by a reversible deformation that does not result in any loss of information during the splitting process. In this embodiment, the second training patches 12 are generated by arranging pixel values ​​extracted from the first training patch 11, one pixel at a time in the vertical direction and one pixel at a time in the horizontal direction, in the spatial (vertical and horizontal) directions. At this time, four second training patches 12 can be generated in the channel direction from one first training patch 11 in the channel (depth) direction. Furthermore, each second training patch 12 has a smaller vertical and horizontal size and lower resolution compared to the first training patch 11. Moreover, because the deformation is reversible, the sum of the number of pixels in the multiple second training patches 12 is equal to the number of pixels in the first training patch 11.

[0039] In this embodiment, the first training patch 11 is divided into four equal parts into four second training patches 12, each having the same number of pixels. However, the method is not limited to this, and it is sufficient that at least the first training patch 11 and the multiple second training patches 12 are reversibly transformed. For example, the multiple second training patches 12 may each have a different number of pixels, and any number of second training patches 12, two or more than four, may be generated. Furthermore, frequency components obtained by multi-resolution analysis using discrete wavelet transform may also be used.

[0040] By reversibly transforming the first training patch 11, which is represented in grayscale, and using multiple second training patches 12, which are reduced in the spatial direction, as input images to the CNN30, the computational load on the CNN30 can be reduced. Furthermore, since no information is lost in the multiple second training patches 12 due to the division of the first training patch 11, image processing can be performed with high accuracy.

[0041] In step S102, the splitting unit 101d generates multiple second ground truth patches 14 (second ground truth images) by splitting the first ground truth patch 10 in the same way as the first training patch 11. In addition, the splitting unit 101d inputs the shooting conditions along with the multiple second training patches 12 to the CNN 30, and may convert the shooting conditions acquired by the acquisition unit 103b into an image (map) with shooting conditions for each pixel.

[0042] Next, in step S103 (estimation step), the learning unit 101e uses a CNN30 (machine learning model) to process the divided second training patch 12, thereby generating multiple estimated patches 13 (estimated images). These multiple estimated patches 13 are estimated images obtained by the CNN30, and ideally, they correspond to each of the corresponding second ground truth patches 14. The learning unit 101e can also input the shooting conditions to the CNN30 by concatenating images with shooting conditions for each pixel in the channel direction of the multiple second training patches 12. When images with shooting conditions for each pixel are input to the CNN30 along with the second training patches 12, the learning unit 101e performs image processing based on the shooting conditions in addition to upscaling to generate multiple estimated patches 13.

[0043] Next, in step S104 (update step), the learning unit 101e updates the weights of the CNN30 based on the error (loss) between the estimated patch 13 and the second correct patch 14. Here, the weights include the filter components and biases of each layer. In this embodiment, backpropagation is used to update the weights, but it is not limited to this. For example, gradient descent may be used.

[0044] Next, in step S105, the learning unit 101e determines whether or not the weight learning is complete. Completion can be determined by whether the number of iterations of learning (weight update) has reached a predetermined value, or whether the amount of change in weight during the update is less than a predetermined value. If it is determined that the learning is incomplete, the process returns to step S101 and a new first training patch 11 and the corresponding first correct answer patch 10 are obtained. On the other hand, if it is determined that the learning is complete, the learning device 101 terminates the learning process and stores the weight information in the storage unit 101a.

[0045] Next, the generation of the output image in this embodiment will be described with reference to Figures 5 and 6. Figure 5 is a conceptual diagram showing the generation of the output image of the neural network. Figure 6 is a flowchart showing the generation of the output image using the neural network. Each step in Figure 6 is mainly performed by the acquisition unit 103b, generation unit 103c, division unit 103d, and processing unit 103e of the image estimation device (image processing device) 103.

[0046] First, in step S201 (acquisition step), the acquisition unit 103b acquires the captured image 20 (first color image). The captured image 20 is an image that includes at least brightness information, similar to the learning process. In this embodiment, the captured image 20 is a color image transmitted from the imaging device 102, but the present invention is not limited to this. For example, it may be an image stored in the storage unit 103a, or it may be a grayscale image in which only brightness information is represented by the shades of a single color. Furthermore, the shooting conditions corresponding to the captured image 20 may be acquired along with the captured image 20 and used in subsequent steps.

[0047] Next, in step S202 (generation step), the generation unit 103c extracts a Y image (luminance image) and multiple color difference images (first color difference images) by performing a YUV conversion on the acquired captured image 20. The luminance image is a first grayscale image 21 that represents only the luminance information of the captured image 20 using only the shades of a single color. The multiple color difference images are multiple color difference images 22 (first color difference images) that have information about the color differences of the captured image 20. The generation of the Y image and multiple color difference images from the captured image 20 can be performed according to equation 1.

[0048] Next, in step S203 (splitting step), the splitting unit 103d divides the first grayscale image 21 into a plurality of second grayscale images 23. At this time, the plurality of second grayscale images 23 are generated by a reversible splitting process in which no information is lost. Therefore, each second grayscale image 23 is smaller in at least one of its width and height dimensions and has a lower resolution compared to the first grayscale image 21. Furthermore, because the transformation is reversible, the sum of the number of pixels in the second grayscale images 23 is equal to the number of pixels in the first grayscale image 21. It is also preferable that the plurality of second grayscale images have the same number of pixels (resolution). When the plurality of second grayscale images have the same number of pixels, the amount of computation for each of the plurality of second grayscale images becomes the same, which makes the calculations in the estimation step described later more efficient. Note that the method for splitting the first grayscale image 21 is the same as the method for transforming the first training patch 11 in step S102, so the explanation is omitted.

[0049] Next, in step S204 (estimation step), the processing unit 103e generates multiple first estimated images 24 (third grayscale images) from multiple second grayscale images 23 by performing image processing using the CNN 30. The weight information used to generate the multiple first estimated images 24 is transmitted from the learning device 101 and stored in the memory unit 103a, and is a neural network similar to that shown in Figure 3.

[0050] If necessary, in step S205 (combination step), the processing unit 103e may perform further image processing on the multiple first estimated images 24. For example, a second estimated image 25 (fourth grayscale image) can be generated by combining (combining) the multiple first estimated images 24. In this case, the second estimated image 25 is generated from the multiple first estimated images 24 by the reverse operation of the method used in step S203 to transform the first grayscale image 21 into multiple second grayscale images 23. In other words, in this embodiment, the processing unit 103e can generate the second estimated image 25 by adding the multiple first estimated images 24 in the spatial direction. In this case, the number of pixels in the second estimated image 25 is equal to the sum of the number of pixels in the multiple first estimated images 24. Therefore, by using multiple second grayscale images 23, which are obtained by reducing the first grayscale image 21, as input images to a machine learning model, the amount of computation can be reduced compared to using a color image as the input image when upscaling to the same magnification. Furthermore, if a second estimated image 25 is generated by combining multiple first estimated images 24, the image estimation device 103 may use the second estimated image 25 as the output image.

[0051] Furthermore, in step S206 (colorization step), the processing unit 103e may perform image processing to colorize the second estimated image 25. At this time, colorization is performed based on the second estimated image 25 and the multiple color difference images 22 generated in step S202 to generate an estimated color image 26 (second color image). The estimated color image 26 is an upscaled image of the captured image 20. In this embodiment, the colorization of the luminance image is performed according to equation 2.

[0052] [Math 2] R = Y + 1.13983V G=Y-0.39465U-0.58060V B = Y + 2.03211U The above formula is used for converting from the YUV color space to the RGB color space. Formula 2 is the inverse operation of the conversion from the RGB color space to the YUV color space performed according to Formula 1. If other definition formulas are used as the method for generating a luminance image from a color image, the inverse operation must be used as the method for generating a color image from a luminance image. If a colorized estimated color image 26 is generated from the second estimated image 25, the image estimation device 103 may use the estimated color image 26 as the output image.

[0053] Furthermore, the processing unit 103e may use multiple color difference interpolation images 27 (second color difference images) to generate the estimated color image 26. The multiple color difference interpolation images 27 are generated by interpolating each of the multiple color difference images 22 in order to increase the resolution (interpolation step). Note that the method for generating the color difference interpolation images 27 from the color difference images 22 is not limited to this, and may be performed using, for example, the binary method and the bicubic method, or a method using a machine learning model. In this case, it is preferable that each of the multiple color difference interpolation images 27 has the same resolution (number of pixels) as the second estimated image 25. By colorizing the second estimated image 25 using multiple color difference interpolation images 27 with the same resolution as the second estimated image 25, noise due to colorization can be reduced, and a more accurate estimated color image 26 can be obtained.

[0054] In this embodiment, a method was described in which an image 20 is acquired in step S201, and a first grayscale image 21 is generated from the image 20 in S202, and an output image is generated from step S203 onwards. However, if the acquisition unit 103b acquires an image that is represented in grayscale from the beginning (for example, an infrared image or a depth image) in step S201, steps S201 and S202 can be omitted, and steps S203 onwards can be executed. In that case, since there is no information regarding the color difference of the image 20, the second estimated image 25 cannot be colorized.

[0055] In this embodiment, the learning device 101 and the image estimation device 103 were described as separate devices, but the present invention is not limited thereto. The learning device 101 and the image estimation device 103 may be integrated. In other words, the learning process and the estimation process may be performed within a single device.

[0056] With the above configuration, this embodiment provides an image processing system that can obtain a high-resolution output image by using a grayscale image reduced by reversible partitioning as the input image in image processing using a machine learning model.

[0057] [Example 2] Next, the image processing system 200 in Embodiment 2 of the present invention will be described. The image processing system 200 in this embodiment learns and executes image processing that upscales images using a machine learning model.

[0058] The image processing system 200 of this embodiment differs from that of Embodiment 1 in that the imaging device 202 acquires the captured image 20 and performs image processing.

[0059] Figure 7 is a block diagram of the image processing system 200 in this embodiment. Figure 8 is an external view of the image processing system 200. The image processing system 200 has a learning device 201 and an imaging device 202 connected via a network 203. Note that the learning device 201 and the imaging device 202 do not need to be constantly connected via the network 203.

[0060] The learning device 201 includes a storage unit (storage means) 211, an acquisition unit (acquisition means) 212, a generation unit (generation means) 213, a division unit (division means) 214, and a learning unit (learning means) 215. These are used to learn (update) the weights of the neural network in order to upscale the captured image 20. The information on the neural network weights is pre-learned by the learning device 201 and stored in the storage unit 211. The method for learning (updating) the neural network weights performed by the learning device 201 is the same as in Example 1, so the explanation is omitted.

[0061] The imaging device 202 includes an optical system 221, an image sensor 222, an image estimation unit 223, a storage unit 224, a recording medium 225a, a display unit 225b, an input unit 226, and a system controller 227. The imaging device 202 captures the subject space to acquire an image 20 and generates an output image. The optical system 221 and the image sensor 222 in the imaging device 202 are the same as in Embodiment 1, so their description is omitted. The imaging device 202 also reads weight information from the storage unit 211 via the network 203 and stores it in the storage unit 224.

[0062] The image estimation unit 223 includes an acquisition unit 223a, a generation unit 223b, a splitting unit 223c, and a processing unit 223d. The acquisition unit 223a acquires the captured image 20 and the shooting conditions corresponding to the captured image 20 from the imaging device 202. The generation unit 223b and the splitting unit 223c are the same as the generation unit 103c and the splitting unit 103d in Embodiment 1. The image processing of the captured image 20 acquired by the acquisition unit 223a is performed based on the weight information stored in the storage unit 224 to generate an output image. In this embodiment, the processing unit 223d uses the shooting conditions corresponding to the captured image 20 for image processing.

[0063] The recording medium 225a stores the output image. When the user issues a command to display the estimated image via the input unit 226, the stored output image is read and displayed on the display unit 225b. The image estimation unit 223 may also read the captured image 20 and shooting conditions stored on the recording medium 225a and perform the process of generating the output image. The system controller 227 controls the processing performed by the imaging device 202.

[0064] Next, the generation of the output image in this embodiment will be described. Figure 9 is a flowchart of the generation of the output image using a neural network in this embodiment. Each step in the generation of the second estimated image 25 is mainly performed by the acquisition unit 223a (acquisition means), generation unit (generation means) 223b, division unit (division means) 223c, and processing unit (estimation means) 223d of the image estimation unit 223.

[0065] First, in step S301 (acquisition step), the acquisition unit 223a acquires the captured image 20 and the shooting conditions corresponding to the captured image 20. In this embodiment, the captured image 20 is a color image, acquired by the imaging device 202 and stored in the storage unit 224. Steps S302 (generation step) and S303 (division step) are the same as steps S202 and S203 of Embodiment 1, so their explanation is omitted.

[0066] Next, in step S304 (estimation step), the processing unit 223d generates multiple estimated images (third grayscale images) 24 from multiple second grayscale images 23 by performing image processing using a neural network. The weight information used to generate the estimated images is transmitted from the learning device 101 and stored in the storage unit 103a, and is a neural network similar to that in Figure 3. In this embodiment, in addition to the multiple first estimated images 24, the processing unit 223d performs image processing using ISO sensitivity as a shooting condition. ISO sensitivity is a shooting condition that represents how sensitive the sensor is to light, and when the ISO sensitivity is high, noise tends to appear in the image. By using ISO sensitivity as a shooting condition, image processing can be performed so as not to excessively emphasize noise when upscaling the captured image 20 with high ISO sensitivity.

[0067] Furthermore, the shooting conditions are not limited to ISO sensitivity; for example, noise reduction intensity may be used as a shooting condition. If the noise reduction intensity of the captured image 20 is weak (the captured image 20 contains many high-frequency components), image processing is performed to reduce the high-frequency components in the output image. Additionally, sharpness intensity may be used as a shooting condition. If the sharpness intensity of the captured image 20 is strong (the captured image 20 contains many high-frequency components), image processing is performed to prevent excessive high-frequency components in the output image. Moreover, image compression ratio may be used as a shooting condition. If the image compression ratio of the captured image 20 is high (high-frequency components of the captured image 20 are lost), image processing is performed to compensate for the high-frequency components in the output image.

[0068] Next, in step S305 (processing step), the processing unit 223d combines and colorizes the multiple first estimated images 24 to generate an output image. The method of combining and colorizing is the same as in Example 1, so a description is omitted.

[0069] With the above configuration, this embodiment provides an image processing system that obtains a high-resolution output image by using a grayscale image reduced by reversible deformation as the input image in image processing using a machine learning model. Furthermore, in this embodiment, image processing can be performed with even higher accuracy by inputting the shooting conditions along with the reduced grayscale image into the machine learning model.

[0070] [Example 3] Next, the image processing system 300 in Embodiment 3 of the present invention will be described. The image processing system 300 in this embodiment learns and executes image processing that upscales images using a machine learning model.

[0071] The image processing system 300 of this embodiment differs from Embodiment 1 in that it has a control device 304 that acquires an image captured image 20 from an imaging device 302 and makes requests to an image estimation device (image processing device) 303 regarding image processing of the image captured image 20.

[0072] Figure 10 is a block diagram of the image processing system 300 in this embodiment. The image processing system 300 includes a learning device 301, an imaging device 302, an image estimation device 303, and a control device 304. In this embodiment, the learning device 301 and the image estimation device 303 may be servers. The control device 304 is a user terminal such as a personal computer or a smartphone. The control device 304 is connected to the image estimation device 303 via a network 305. The image estimation device 303 is connected to the learning device 301 via a network 306. In other words, the control device 304 and the image estimation device 303, as well as the image estimation device 303 and the learning device 301, are configured to communicate with each other.

[0073] The learning device 301 and imaging device 302 in the image processing system 300 have the same configuration as the learning device 101 and imaging device 102, respectively, so their descriptions are omitted.

[0074] The image estimation device 303 includes a storage unit 303a, an acquisition unit (acquisition means) 303b, a generation unit (generation means) 303c, a division unit (division means) 303d, a processing unit (estimation means) 303e, and a communication unit (receiving means) 303f. The storage unit 303a, acquisition unit 303b, generation unit 303c, division unit 303d, and processing unit 303e in the image estimation device 303 are the same as the storage unit 103a, acquisition unit 103b, generation unit 103c, division unit 103d, and processing unit 103e, respectively.

[0075] The control device 304 includes a communication unit (transmission means) 304a, a display unit (display means) 304b, an input unit (input means) 304c, a processing unit (processing means) 304d, and a recording unit 304e. The communication unit 304a can transmit a request to the image estimation device 303 to perform processing on the captured image 20. It can also receive the output image processed by the image estimation device 303. The communication unit 304a may also communicate with the imaging device 302. The display unit 304b displays various information. The various information displayed by the display unit 304b includes, for example, the captured image 20 transmitted to the image estimation device 303 or the output image received from the image estimation device 303. The input unit 304c can receive instructions from the user to start image processing. The processing unit 304d can perform image processing, including colorization, on the output image received from the image estimation device 303. The recording unit 304e stores the captured image 20 acquired from the imaging device 302 and the output image received from the image estimation device 303.

[0076] The method by which the captured image 20 to be processed is transmitted to the image estimation device 303 is not limited; for example, the captured image 20 may be uploaded to the image estimation device 303 at the same time as S401, or it may be uploaded to the image estimation device 403 before S401. Also, the captured image 20 may be an image stored on a server different from the image estimation device 303.

[0077] Next, the generation of the output image in this embodiment will be described. Figure 11 is a flowchart of the generation of the output image using a neural network in this embodiment.

[0078] The operation of the control device 304 will now be described. In this embodiment, image processing is initiated by the user via the control device 304 when an instruction to start image processing is given.

[0079] First, in step S401 (the first transmission step), the communication unit 304a transmits a request for processing of the captured image 20 to the image estimation device 303. In step S401, the control device 304 may also transmit, along with the request for processing of the captured image 20, an ID for user authentication and shooting conditions corresponding to the captured image 20.

[0080] Next, in step S402 (first receiving step), the communication unit 304a receives the output image generated by the estimation device 303.

[0081] Next, the operation of the image estimation device 403 will be described. First, in step S501, the communication unit 303f receives a request for processing the captured image 20 transmitted from the communication unit 304a. Upon receiving the instruction to process the captured image 20, the image estimation device 303 executes the processing from step S502 onward.

[0082] Next, in step S502, the acquisition unit 303b acquires the captured image 20. In this embodiment, the captured image 20 is transmitted from the control device 304. At this time, the shooting conditions corresponding to the captured image 20 may be acquired along with the captured image 20. Note that the processing in steps S501 and S502 may be performed simultaneously. Steps S503 to S505 are the same as steps S202 to S204, so their explanation is omitted.

[0083] Next, in step S506, the image estimation device 303 transmits the output image to the control device 304. The output image transmitted by the image estimation device 303 includes one of the following: a plurality of first estimated images 24, a second estimated image 25 generated from the plurality of first estimated images 24, or an estimated color image 26.

[0084] With the above configuration, this embodiment provides an image processing system that can obtain a high-resolution output image by using a grayscale image, which has been reduced in size by reversible deformation in image processing using a machine learning model, as the input image. In this embodiment, the control device 304 only requests processing for a specific image. The actual image processing is performed by the image estimation device 303. Therefore, if the control device 304 is used as a user terminal, the processing load on the user terminal can be reduced. Consequently, the user can obtain the output image with a low processing load.

[0085] [Other examples] This embodiment can also be implemented by supplying a program that implements one or more of the functions of the above-described embodiment to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be implemented by a circuit (e.g., an ASIC) that implements one or more functions.

[0086] According to each embodiment, an image processing method, an image processing device, a program, and a storage medium can be provided that can obtain high-resolution output images in image processing using a machine learning model. The image processing device only needs to have the image processing functions of this embodiment and can be implemented in the form of an imaging device or a personal computer.

[0087] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of its gist. [Explanation of Symbols]

[0088] 21. First grayscale image 23. Second grayscale image 24. Third grayscale image

Claims

1. An image processing method performed by a computer, The steps include: dividing a first grayscale image to generate a plurality of second grayscale images with lower resolution than the first grayscale image; The process includes the step of generating a plurality of third grayscale images which are upscaled relative to the plurality of second grayscale images, by concatenating information on the shooting conditions corresponding to the first grayscale image and the plurality of second grayscale images in the channel direction and inputting this information into a machine learning model. The image processing method is characterized in that the information on the shooting conditions is a map showing the shooting conditions for each pixel.

2. The image processing method according to claim 1, characterized in that the resolution of each of the plurality of second grayscale images is the same as that of the others.

3. The image processing method according to claim 1, further comprising the step of generating a fourth grayscale image by combining the plurality of third grayscale images.

4. The image processing method according to claim 3, characterized in that the number of pixels of the fourth grayscale image is equal to the sum of the number of pixels of the plurality of third grayscale images.

5. A step of generating a first grayscale image and a plurality of first color difference images from a first color image, The image processing method according to claim 3, further comprising the step of generating a second color image based on the fourth grayscale image and the plurality of first color difference images.

6. A step of generating a first grayscale image and a plurality of first color difference images from a first color image, The steps include generating a plurality of second color difference images by upscaling the plurality of first color difference images, The image processing method according to claim 3, further comprising the step of generating a second color image based on the fourth grayscale image and the plurality of second color difference images.

7. The image processing method according to claim 6, characterized in that the resolution of each of the plurality of second color difference images is the same as the resolution of the fourth grayscale image.

8. The image processing method according to any one of claims 1 to 7, characterized in that the first grayscale image is a luminance image.

9. The image processing method according to claim 1, characterized in that the shooting conditions include at least one of the following: the pixel pitch of the image sensor, the type of optical low-pass filter of the optical system, and the ISO sensitivity.

10. The image processing method according to claim 1, characterized in that the aforementioned shooting conditions include at least one of noise reduction intensity, sharpness intensity, and image compression ratio.

11. A program characterized by causing a computer to execute the image processing method described in any one of claims 1 to 7.

12. A storage medium characterized by storing the program described in claim 11.

13. A means for generating a plurality of second grayscale images with lower resolution than the first grayscale image by dividing the first grayscale image, The system includes means for generating a plurality of third grayscale images that are upscaled to the plurality of second grayscale images by concatenating information on the shooting conditions corresponding to the first grayscale image and the plurality of second grayscale images in the channel direction and inputting this information into a machine learning model, The image processing apparatus is characterized in that the information regarding the shooting conditions is a map showing the shooting conditions for each pixel.

14. An image processing system comprising an image processing device according to claim 13 and a control device capable of communicating with the image processing device, The control device has means for transmitting a request to the image processing device to perform processing on the captured image, The image processing device is characterized by having a receiving means for receiving the request and a means for generating a plurality of third grayscale images in response to the request.

15. An acquisition unit that acquires a first training image and a first correct image, A splitting unit that splits the first training image and the first ground truth image to generate a plurality of second training images with lower resolution than the first training image and a plurality of second ground truth images with lower resolution than the first ground truth image, A processing unit that generates a plurality of estimated images upscaled to the plurality of second training images by concatenating information on the shooting conditions corresponding to the first training image and the plurality of second training images in the channel direction and inputting it to a machine learning model, The system includes a learning unit that updates the weights of a machine learning model based on the aforementioned plurality of estimated images and the aforementioned plurality of second ground truth images, The learning device is characterized in that the information on the shooting conditions is a map showing the shooting conditions for each pixel.

16. A method for generating a trained model that is executed by a computer, The steps include obtaining a first training image and a first ground truth image, The steps include generating a plurality of second training images with lower resolution than the first training image and a plurality of second ground truth images with lower resolution than the first ground truth image by splitting the first training image and the first ground truth image, The steps include: generating a plurality of estimated images upscaled to the plurality of second training images by concatenating the information on the shooting conditions corresponding to the first training image and the plurality of second training images in the channel direction and inputting it into a machine learning model; The process includes the step of updating the weights of the neural network based on the plurality of estimated images and the plurality of second ground truth images. A method for generating a trained model, characterized in that the information on the shooting conditions is a map indicating the shooting conditions for each pixel.

17. A program characterized by causing a computer to execute the method for generating a trained model described in claim 16.

Citation Information

Patent Citations

  • Image processing apparatus, control method, program, and storage medium

    JP2008167388A

  • Recognition device, recognition method, program, and data generation device

    JP2019175107A

  • Image processing method, image processing device, image processing system, creating method of learned weight, and program

    JP2020201540A

  • Image processing method and image reception device

    WO2018193333A1

  • Image processing system and program

    WO2020166596A1