Image processing method, image processing apparatus, image processing system, imaging apparatus, learning method, learning device, and program
By generating multiple upscaled images through sequential processing with smaller-scale models, the method addresses the high computational load issue in DL models, enabling efficient production of high-resolution images with varied effects.
Patent Information
- Application Number
- JP2024176390
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-08
- Filing Date
- 2024-10-08
- Publication Date
- 2025-05-20
AI Technical Summary
Existing deep learning (DL) models for generating upscaled images with different effects face high computational loads due to the large number of filter convolutions required for processing feature maps with many channels, making it impractical to generate multiple upscaled images efficiently.
The method involves generating a first image using a first machine learning model, followed by a second image using a different model, and then combining these images to create a third image, where each image has more pixels than the input and the second image has fewer high-frequency components than the first, thereby reducing computational load by using smaller-scale models for the second image processing.
This approach allows for the generation of two upscaled images with different effects while significantly reducing the computational burden, achieving efficient image processing without compromising image quality.
Smart Images

Figure 2025078591000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an image processing method, an image processing device, an image processing system, an imaging device, a learning method, a learning device, and a program. [Background technology]
[0002] Patent Document 1 discloses a method for generating two upscaled images with different effects using a DL (Deep Learning) model. Specifically, in Patent Document 1, a feature map is generated from an input low-resolution image using a DL model, and first and second high-resolution intermediate images (upscaled images) are generated from the feature map. Here, the feature map is intermediate data obtained by image processing using a DL model, and is obtained by linking multiple images in the channel (depth) direction. In DL, a feature map generally has more channels than an image; for example, a color image has three RGB channels, while a feature map has 64 or 128 channels. By using the method disclosed in Patent Document 1, two upscaled images with different effects can be generated using a DL model. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2022-046219 A Summary of the Invention [Problem to be solved by the invention]
[0004] In the method disclosed in Patent Document 1, an image is generated from a feature map with a large number of channels (large size), so the number of filter convolutions increases, resulting in a large computational load. For this reason, it is not possible to generate two upscaled images with different effects using a DL model with a small computational load. [Means for solving the problem]
[0005] An image processing method as one aspect of the present invention includes a step of generating a first image by inputting an input image or an image based on the input image into a first machine learning model, a step of generating a second image by inputting the first image into a second machine learning model different from the first machine learning model, and a step of generating a third image using the first image and the second image, wherein the first image, the second image, and the third image each have a larger number of pixels than the input image, and the second image has fewer high-frequency components than the first image.
[0006] Other objects and features of the present invention are illustrated in the following examples. Effect of the Invention
[0007] According to the present invention, it is possible to provide an image processing method capable of generating an upscaled image with a small calculation load. [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram showing a learning flow of a machine learning model in a first embodiment. [Diagram 2] 1 is a block diagram of an image processing system according to a first embodiment. [Diagram 3] 1 is an external view of an image processing system according to a first embodiment. [Figure 4] 1 is a flowchart relating to learning of a machine learning model in the first embodiment. [Diagram 5] 1 is a flowchart showing generation of an output image using a machine learning model in the first embodiment. [Figure 6] FIG. 11 is a block diagram of an image processing system according to a second embodiment. [Figure 7] FIG. 11 is an external view of an image processing system according to a second embodiment. [Figure 8] FIG. 11 is a block diagram of an image processing system according to a third embodiment. [Figure 9]13 is a flowchart for generating an output image using a machine learning model in the third embodiment. [Figure 10] FIG. 2 is an explanatory diagram of an overview of each embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the drawings, the same components are designated by the same reference numerals, and duplicated descriptions will be omitted.
[0010] First, before describing the specific embodiments, the gist of each embodiment will be described. Each embodiment uses a DL (Deep Learning) model (machine learning model) to generate two upscaled images with different effects. In each embodiment, image processing using the DL model uses a convolutional neural network that repeatedly convolves a filter with an input image, adds a bias, and performs nonlinear conversion to obtain an output image with a desired effect. Here, the DL model is composed of weights including the filter and bias of the convolutional neural network and a network configuration. In each embodiment, the calculation load of image processing using the DL model is mainly determined by the filter size (length, width, number of channels, number) used for convolution, the size of the other side to convolve the filter (length, width, number of channels of the image or feature map), and the number of convolutional layers. In addition, the feature map will be described later.
[0011] In each embodiment, upscaling is an image enlargement process that generates a sharp, high-resolution image with a large number of pixels by estimating high-frequency components that cannot be expressed in a low-resolution image with a small number of pixels. In each embodiment, an effective upscaled image is a high-resolution image that contains many high-frequency components (high resolution), and an ineffective upscaled image is a high-resolution image that contains fewer high-frequency components (low resolution).
[0012] Next, an overview of each embodiment will be described with reference to Fig. 10. Fig. 10 is an explanatory diagram of the overview of each embodiment. In each embodiment, a low-resolution image (input image) is first input to a DL model (first machine learning model) to generate a high-resolution upscaled image (first image).
[0013] Next, the first image is input to a DL model (a second machine learning model different from the first machine learning model) to generate a less effective high-resolution upscaled image (second image). In this embodiment, the number of pixels of the first image and the number of pixels of the second image are each greater than the number of pixels of the input image. Preferably, the number of pixels of the first image and the number of pixels of the second image are the same as each other. In addition, the first image and the second image contain more high-frequency components (higher resolution) than the interpolated image of the captured image. In addition, the first image contains more high-frequency components than the second image. Here, the interpolated image is a bicubic interpolated image or a bilinear interpolated image, but is not limited thereto.
[0014] Finally, a high-resolution upscaled image (third image) is generated by taking a weighted average of the first image and the second image. In this embodiment, the number of pixels of the third image is greater than the number of pixels of the captured image (input image). Preferably, the number of pixels of the third image is equal to the number of pixels of the first image and the second image. In addition, the third image contains more high-frequency components (higher resolution) than the interpolated image of the captured image, and contains high-frequency components intermediate between the first image and the second image.
[0015] In each embodiment, a first image with a high effect is input to a DL model to generate a second image with a lower effect. This reduces the number of times the filter is convolved, and the computational load can be reduced, since an estimated image is generated from an image with a small number of channels. In addition, in DL, the computational load required to generate an image with a low effect from an image with a high effect (reducing high-frequency components to blur) is generally less than the computational load required to generate an image with a high effect from an image with a low effect (increasing high-frequency components to increase the sense of resolution). Therefore, according to each embodiment, two upscaled images with different effects can be generated with a low computational load using the DL model.
[0016] In each embodiment, the third image is preferably generated by taking a weighted average of the first image and the second image, which have different effects from each other. This allows the effect of the third image to be fine-tuned by adjusting the weight of the weighted average.
[0017] In addition, the effect of the upscaled image obtained can be adjusted by taking a weighted average of the first image, which has a high effect, and the bicubic interpolated image of the captured image, and adjusting the weight of the weighted average. However, since the bicubic interpolated image of the captured image does not contain high-frequency components, the weighted average reduces the high-frequency components contained in the first image, which is the effect of upscaling. Therefore, it is desirable to take a weighted average with the second image, which contains higher frequency components than the interpolated image, as in each embodiment.
[0018] Also, it is possible to blur a first image with a high effect to generate a second image with a low effect without using a DL model. However, for the reasons described above, in order to reduce the high frequency components contained in the first image, which is the effect of upscaling, it is desirable to generate a second image that contains higher frequency components than the interpolated image using a machine learning model, as in each embodiment.
[0019] The image processing method described above is merely an example, and each embodiment is not limited to this. The image processing method of each embodiment will be described in detail below.
[0020] [Example 1] First, an image processing system 100 according to a first embodiment of the present invention will be described. In this embodiment, a DL model is used to learn and execute image processing for generating two upscaled images with different effects with a small calculation load.
[0021] Fig. 2 is a block diagram of image processing system 100. Fig. 3 is an external view of image processing system 100. Image processing system 100 includes a learning device 101, an imaging device 102, an image estimation device (image processing device) 103, a display device 104, a recording medium 105, an input device 106, an output device 107, and a network 108.
[0022] The learning device 101 includes a storage unit 101a, an image acquisition unit 101b, an image generation unit 101c, and a learning unit 101d.
[0023] The imaging device 102 has an optical system 102a and an imaging element 102b. The optical system 102a collects light incident from the subject space to the imaging device 102. The imaging element 102b receives an optical image of the subject formed via the optical system 102a to obtain a captured image (low-resolution image). The imaging element 102b is a charge coupled device (CCD) sensor or a complementary metal-oxide semiconductor (CMOS) sensor. Information on the shooting conditions of the captured image (pixel pitch of the imaging element 102b, type of optical low-pass filter, ISO sensitivity, etc.) can be obtained together with the image. In addition, the development conditions of the captured image (noise reduction strength, sharpness strength, image compression rate, etc.) can also be obtained together with the image. In addition, these image information obtained together with the image can be transmitted together with the image to an image acquisition unit 103b of the image estimation device 103 described later. In addition, a storage unit that stores the acquired image, a display unit that displays it, a transmission unit that transmits it to the outside, an output unit that stores it in an external storage medium, etc. are not shown. A control unit that controls each unit of the imaging device 102 is also not shown.
[0024] The image estimation device 103 includes a storage unit 103a, an image acquisition unit 103b, an upscale image processing unit (first generation unit) 103c, a blurred image processing unit (second generation unit) 103d, and an image synthesis unit (third generation unit) 103e. The image estimation device 103 performs processing (image processing) on a low-resolution image (captured image, input image). More specifically, first, the image acquisition unit 103b acquires a low-resolution image. Then, the upscale image processing unit 103c or the blurred image processing unit 103d generates two high-resolution images with different upscaled effects. After that, the image synthesis unit 103e synthesizes the two high-resolution images to generate a high-resolution image.
[0025] The low-resolution image may be an image captured by the imaging device 102, or may be an image stored in the recording medium 105. The upscale image processing unit 103c, the blurred image processing unit 103d, and the image synthesis unit 103e may be configured to be executable using an integrated DL model. Alternatively, the upscale image processing unit 103c and the blurred image processing unit 103d may be configured to be executable using an integrated DL model. Alternatively, the upscale image processing unit 103c and the blurred image processing unit 103d may be configured to be executable using separate DL models.
[0026] A DL model is mainly used for image processing, and the weight information is read from the storage unit 103a. The weights are obtained by learning in the learning device 101, and the image estimation device 103 reads the weight information from the storage unit 101a via the network 108 in advance and stores it in the storage unit 103a. The stored weight information may be the weight numerical value itself or may be in an encoded format.
[0027] The upscaled image is output to at least one of a display device 104, a recording medium 105, and an output device 107. The display device 104 is, for example, a liquid crystal display or a projector. A user can check an image being processed via the display device 104 and perform image editing work or the like via the input device 106. The recording medium 105 is, for example, a semiconductor memory, a hard disk, a server on a network, etc. The input device 106 is, for example, a keyboard or a mouse, etc. The output device 107 is, for example, a printer, etc.
[0028] Next, a DL model learning method (a method for manufacturing a trained model) executed by the learning device 101 in this embodiment will be described with reference to Fig. 1 and Fig. 4. Fig. 1 is a diagram showing the flow of learning the DL model. Fig. 4 is a flowchart related to learning the DL model. Each step in Fig. 4 is mainly executed by the image acquisition unit 101b or the learning unit 101d.
[0029] First, in step S101, the image acquisition unit 101b acquires a low-resolution training patch (first training image) 201 and a high-resolution training patch (second training image, correct answer image) 200 corresponding to the low-resolution training patch. The high-resolution training patch 200 and the low-resolution training patch 201 correspond to the symbols in FIG. 1, respectively. In this embodiment, a patch is an image having a predetermined number of pixels. For example, the low-resolution training patch is 128×128×3 pixels, and the corresponding high-resolution training patch is 256×256×3 pixels. In this case, since the vertical and horizontal sizes are doubled, the upscaling factor is 2 (enlarged by 4 times in terms of the number of pixels). The upscaling factor is not limited to 2, and may be any factor as long as the low-resolution training patch and the corresponding high-resolution training patch can be acquired.
[0030] In this embodiment, the low-resolution training patch and the high-resolution training patch are each a three-channel color image having RGB information, but are not limited thereto. For example, they may be a one-channel monochrome image having luminance information extracted from a color image using the image generating unit 101c. Alternatively, they may be an image in which a one-channel monochrome image is decomposed into multiple small patches using the image generating unit 101c and the patches are connected in the depth (channel) direction.
[0031] In this embodiment, the high-resolution training patch corresponding to the low-resolution training patch is an image of the same subject captured in the same scene with a different resolution (number of pixels). Such a set of images may be obtained by capturing images of the same subject in the same scene using optical systems with different focal lengths, and cutting out corresponding portions of the two images obtained. Alternatively, an equivalent low-resolution training patch captured by the imaging device 102 and a corresponding high-resolution training patch that is less affected by blurring (aberration and diffraction) due to the optical system 102a may be generated by numerical calculation. In this embodiment, the high-resolution training patch corresponding to the low-resolution training patch is generated by numerical calculation, but is not limited to this. In addition, as described later, the high-resolution training patch is an image corresponding to a ground truth image, which is the target of an upscaled patch output from the DL model by inputting the low-resolution training patch. Therefore, the high-resolution training patch contains more high-frequency components than the low-resolution training patch (higher resolution).
[0032] Next, in step S102, the learning unit 101d inputs the low-resolution training patch (first training image) 201 to the DL model (first machine learning model) and outputs (generates) an upscaled patch 202 with a high upscaled effect. Information on the shooting conditions and development conditions may be input to the DL model together with the low-resolution training patch. For example, an image having the ISO sensitivity as a pixel value may be linked in the depth (channel) direction of the low-resolution training patch and input to the DL model. This enables upscaling according to the shooting conditions and development conditions of the low-resolution training patch. Also, the low-resolution training patch may be input to the DL model after being enlarged to the same number of pixels as the high-resolution training patch by interpolation. In this case, the DL model is learned in a step described later so as to estimate high-frequency components from a patch obtained by interpolating the low-resolution training patch.
[0033] Next, in step S103, the learning unit 101d inputs the effective upscaled patch (first upscaled patch) 202 generated in step S102 to a DL model (second machine learning model). Then, the learning unit 101d outputs (generates) an upscaled patch (second upscaled patch) 203 that is less effective than the upscaled patch 202. Note that both the upscaled patch 202 and the upscaled patch 203 are estimates of the high-resolution training patch (second training image, ground truth image) 200. How the two upscaled images are differentiated in effectiveness through learning will be described later.
[0034] Next, in step S104, the learning unit 101d updates the weights of the DL models (the first machine learning model and the second machine learning model). That is, the learning unit 101d updates the weights of the DL models by using an error (loss function) based on the high-resolution training patch 200 and the upscaled patch 202 or the upscaled patch 203.
[0035] In this embodiment, the error (loss) between the high-resolution training patch 200 and the upscaled patch 202 is calculated using an adversarial loss, i.e., a generative adversarial network (GAN). Here, the generative adversarial network is a method in which the upscaled patch 202 is input, and a separately prepared image recognition DL model (classifier) is made to distinguish whether it is a real high-resolution image or a fake created by a DL model, and a large penalty (error) is added if it is detected as a fake. By learning to reduce the error based on the generative adversarial network, the finally obtained upscaled patch 202 becomes an image that contains many high-frequency components (highly effective) and is difficult to distinguish from the high-resolution training patch 200.
[0036] In this embodiment, the error (loss) between the high-resolution training patch 200 and the upscaled patch 203 is calculated using the mean square error (MSE). Here, the mean square error is the mean square value of the difference between the values of corresponding pixels in the high-resolution training patch 200 and the upscaled patch 203. In general, even if learning is performed to reduce the error based on the mean square error, the ultimately obtained upscaled patch 203 will be a blurrier (less effective) image than the high-resolution training patch 200.
[0037] In this embodiment, the method of calculating the error is a generative adversarial network and a mean square error, but is not limited to this. Other methods may be used as long as two upscaled patches 202 and 203 with different effects can be obtained.
[0038] In this embodiment, the weights of the DL models (first machine learning model, second machine learning model) are updated by the backpropagation method so that the weighted average of the errors calculated by the generative adversarial network and the mean squared error is small. However, this embodiment is not limited to this. After learning the weights of the DL model (first machine learning model) based on the errors calculated by the generative adversarial network, the learned weights may be fixed, and then the weights of the DL model (second machine learning model) may be sequentially learned based on the errors calculated by the mean squared error. That is, in this embodiment, the first machine learning model and the second machine learning model may be learned simultaneously (as a whole), or the first machine learning model and the second machine learning model may be learned separately.
[0039] Next, in step S105, the learning unit 101d judges whether the learning of the DL model is completed. Completion can be judged by whether the number of iterations of learning (updating the weights) reaches a specified value, or whether the amount of change in the weights at the time of updating is smaller than a specified value. If it is judged to be incomplete, the process returns to step S101, and multiple new low-resolution training patches (first training images) 201 and corresponding high-resolution training patches (correct images) 200 are obtained. On the other hand, if it is judged to be completed, the weight information is stored in the storage unit 101a.
[0040] In this embodiment, the convolutional neural network configuration shown in FIG. 1 is used as the DL model, but the DL model is not limited to this.
[0041] CN in FIG. 1 represents a convolution layer. In CN, the sum of the input, the convolution of the filter, and the bias is calculated, and the result is nonlinearly transformed by the activation function. The initial values of each component of the filter and the bias are arbitrary, and in this embodiment, they are determined by random numbers. For example, ReLU (Rectified Linear Unit), a sigmoid function, or a hyperbolic tangent function (Tanh) can be used as the activation function. The multidimensional array output in each layer except the output layer of the DL model is a feature map. In general, a feature map is a four-dimensional array with dimensions of batch, length, width, and channel. The skip connection 204 synthesizes feature maps output from discontinuous layers. The synthesis of feature maps may be performed by taking the sum of each element, or by concatenation in the channel direction. In this embodiment, the sum of each element is adopted.
[0042] The elements (blocks or modules) within the dotted frame in FIG. 1 represent residual blocks. A network in which residual blocks are multi-layered is called a residual network, and is widely used in DL image processing. However, this embodiment is not limited to this. In this embodiment, other elements such as an inception module or a dense block having dense skip connections may be multi-layered to form a network. Here, the inception module is a module in which convolution layers having different convolution filter sizes are juxtaposed, and the resulting feature maps are integrated to form a final feature map.
[0043] Also, the computational load (~convolution operations) can be reduced by shrinking the feature map in the layer close to the input and expanding the feature map in the layer close to the output, and reducing the size of the feature map in the intermediate layer. Here, pooling and stride can be used to shrink the feature map. Deconvolution (or transposed convolution), depth to space, interpolation, etc. can be used to expand the feature map.
[0044] In addition, the low-resolution feature map is enlarged to a high-resolution feature map and input to the output layer. In this embodiment, the depth-to-space (D2S in the figure) method, which rearranges the channel information of the feature map in the spatial (vertical and horizontal) direction, is used as a method for upsampling the feature map, but is not limited to this. In addition, when the low-resolution patch is enlarged to the same number of pixels as the high-resolution patch by interpolation in step S102 and then enlarged to the DL model, the enlargement operation of the feature map before the output layer is not necessarily required.
[0045] In this embodiment, the high-resolution feature map upsampled by depth-to-space is processed in a convolution layer (output layer (one convolution layer) in FIG. 1) to generate the upscaled patch 202. In Patent Document 1, the high-resolution feature map is processed in a separately prepared convolution layer (output layer not shown) to generate the upscaled patch 203. Therefore, in Patent Document 1, the upscaled patch 203 is generated from a feature map with a large number of channels (large size), so the number of times the filter is convolved increases, resulting in a large calculation load. On the other hand, in this embodiment, the upscaled patch 202 is processed in a convolution layer (second machine learning model) to generate the upscaled patch 203. Therefore, in this embodiment, the upscaled image is generated from an image with a small number of channels (small size), so the number of times the filter is convolved decreases, and the calculation load can be reduced.
[0046] In addition, the scale (model scale) of the second machine learning model in this embodiment is smaller than the model scale of the corresponding machine learning model in Patent Document 1. Here, the corresponding machine learning model is a separately prepared convolution layer (output layer, not shown) that processes the above-mentioned high-resolution feature map, and is a model with the same scale as the output layer in FIG. 1 in this embodiment. That is, the scale (model scale) of the second machine learning model is smaller than the scale (model scale) of the output layer of the first machine learning model. In addition, in this embodiment, the overall scale of the first machine learning model is larger than the model scale of its output layer. Therefore, the model scale of the second machine learning model is smaller than the model scale of the output layer of the first machine learning model. Preferably, the model scale of the second machine learning model is smaller than 1 / 2 the model scale of the output layer of the first machine learning model. More preferably, the model scale of the second machine learning model is smaller than 1 / 5 the model scale of the output layer of the first machine learning model. Even more preferably, the model scale of the second machine learning model is smaller than 1 / 10 the model scale of the output layer of the first machine learning model.
[0047] In this embodiment, the model scale of the machine learning models (the scale of the output layer of the first machine learning model and the scale of the second machine learning model) is expressed by the following formula (1).
[0048]
number
[0049] Here, L is the number of convolutional layers, k l is the kernel size of the lth layer (l=1~L), c l is the number of channels in the lth layer of convolution filters, n l is the number of convolution filters in the l-th layer. The number of layers L of the output layer of the first machine learning model is 1 (L = 1). Note that in formula (1), the convolution filters are assumed to be square, with no stride or dilation, and normal 2D convolution with no separation in the depth direction.
[0050] Next, the estimation (DL upscaling) performed by the image estimation device 103 in this embodiment will be described with reference to Fig. 5. Fig. 5 is a flowchart showing a process of generating an upscaled image (output image, third image) from a low-resolution image (input image) using DL models (first machine learning model, second machine learning model).
[0051] Each step in FIG. 5 is mainly executed by the image acquisition unit 103b, the upscale image processing unit 103c, the blurred image processing unit 103d, or the image synthesis unit 103e of the image estimation device 103. As described above, the upscale image processing unit 103c, the blurred image processing unit 103d, and the image synthesis unit 103e may be configured as an integrated DL model. Alternatively, the upscale image processing unit 103c and the blurred image processing unit 103d may be configured as an integrated DL model, and the image synthesis unit 103e may be configured as a separate image processing unit. Alternatively, the upscale image processing unit 103c and the blurred image processing unit 103d may be configured as separate DL models, and the image synthesis unit 103e may also be configured as a separate image processing unit.
[0052] First, in step S201, the image acquisition unit 103b acquires a captured image (input image). The captured image is a low-resolution image similar to the learning image. In this embodiment, the captured image is transmitted from the imaging device 102, but is not limited to this. For example, a captured image stored in the storage unit 103a may be used.
[0053] Next, in step S202, the upscaled image processing unit 103c generates an effective upscaled image (first image) by inputting the captured image to a DL model (first machine learning model). To generate an effective upscaled image, a convolutional neural network similar to the configuration shown in FIG. 1 is used. In this embodiment, the captured image is input to the DL model and an upscaled image is output. However, the shooting conditions and development conditions may be input to the DL model together with the captured image using the method described in step S102. In addition, in the case where the low-resolution patch is enlarged to the same number of pixels as the high-resolution patch by interpolation in step S102 and then input to the DL model for learning, the captured image is similarly enlarged by interpolation and then input to the DL model. Weight information of the DL model is transmitted from the learning device 101 and stored in the storage unit 103a.
[0054] Next, in step S203, the blurred image processing unit 103d generates a less effective upscaled image (second image) by inputting the more effective upscaled image (first image) to a DL model (second machine learning model). To generate the less effective upscaled image, a convolutional neural network similar to the configuration shown in Fig. 1 is used. Weight information of the DL model is transmitted from the learning device 101 and stored in the storage unit 103a.
[0055] In this embodiment, an effective upscaled image means an image with a high resolution, that is, an image with many high-frequency components. That is, in this embodiment, the second image has fewer high-frequency components than the first image. Preferably, an image with many high-frequency components is an image that contains many frequency components higher than the frequency components of the captured image (input image).
[0056] In this embodiment, the high frequency components of an image are expressed by the sum of the absolute values of LH, HL, and HH corresponding to the high frequency components among the four coefficients LL, LH, HL, and HH obtained by performing a one-level discrete wavelet transform on the image. Here, the discrete wavelet transform is a frequency analysis method that uses a wavelet function as a basis function to decompose the original signal into high frequency components and low frequency components. Note that, when a two-dimensional signal, an image, is subjected to a one-level discrete wavelet transform, high frequency components are obtained for each of the vertical (LH), horizontal (HL), and diagonal (HH) directions. In this embodiment, the Haar wavelet is used as the wavelet function. In addition, in this embodiment, a lifting scheme is used as a method for scaling (enlarging and reducing) the wavelet function. Here, the lifting scheme is a scheme in which a signal is divided into even elements and odd elements, the odd elements are predicted by the even elements, the deviation from the prediction is treated as a high frequency component, and the residual is treated as a low frequency component. In this case, the discrete wavelet coefficients of a one-dimensional signal x are given by the following formula (2).
[0057]
number
[0058] Here, D is the high frequency component of the one-dimensional signal X, and C is the low frequency component of the one-dimensional signal X. Note that X[0::2] represents the odd-numbered elements of the one-dimensional signal X, and X[1::2] represents the even-numbered elements of the one-dimensional signal X.
[0059] In addition, the captured image (input image) may be enlarged by BICUBIC interpolation to the same number of pixels as the upscaled image, and then high-frequency components may be calculated and compared with the upscaled image.
[0060] In this embodiment, the high frequency components of the second upscaled image are less than 1 / 4 of the high frequency components of the first upscaled image, more preferably, the high frequency components of the second upscaled image are less than 1 / 3 of the high frequency components of the first upscaled image, and even more preferably, the high frequency components of the second upscaled image are less than 1 / 2 of the high frequency components of the first upscaled image.
[0061] Next, in step S204, the image synthesis unit 103e generates a final upscaled image (third image) by taking a weighted average of the effective upscaled image and the less effective upscaled image. In this embodiment, a default weight for the weighted average is used, but a weight for the weighted average specified by the user via the input device 106 may also be used.
[0062] When inputting a captured image into a DL model, it is not necessary to cut the image to the same size as the low-resolution training patch used during learning, but it may be processed after being decomposed into multiple overlapping patches. In this case, the resulting patches can be synthesized to create an upscaled image.
[0063] In this embodiment, the learning device 101 and the image estimation device 103 are separate devices, but the present invention is not limited to this. The learning device 101 and the image estimation device 103 may be integrated. In other words, learning (the process shown in FIG. 4) and estimation (the process shown in FIG. 5) may be performed in an integrated device.
[0064] According to this embodiment, it is possible to provide an image processing method and an image processing device that use a DL model to generate two upscaled images with different effects with a small calculation load.
[0065] [Example 2] Next, an image processing system 300 according to a second embodiment of the present invention will be described. In this embodiment, an image processing DL model is learned and executed in the same manner as in the first embodiment. This embodiment differs from the first embodiment in that an imaging device acquires a captured image (low-resolution image) and processes the image.
[0066] Fig. 6 is a block diagram of image processing system 300. Fig. 7 is an external view of image processing system 300. Image processing system 300 includes a learning device 301 and an imaging device 302 connected via a network 303. Note that learning device 301 and imaging device 302 do not need to be constantly connected via network 303.
[0067] The learning device 301 includes a storage unit 311, an image acquisition unit 312, an image generation unit 313, and a learning unit 314. The learning device 301 uses these units to learn the weights of the DL model.
[0068] The imaging device 302 captures an image of a subject space, acquires a captured image (low-resolution image), and generates an image by upscaling the captured image. Details of image processing performed by the imaging device 302 will be described later. The imaging device 302 has an optical system 321 and an image sensor 322. The image estimation unit 323 has an image acquisition unit 323a, an upscale image processing unit (first generation unit) 323b, a blurred image processing unit (second generation unit) 323c, and an image synthesis unit (third generation unit) 323d. Note that, similarly to the first embodiment, the upscale image processing unit 323b, the blurred image processing unit 323c, and the image synthesis unit 323d may be configured as an integrated DL model. Alternatively, the upscale image processing unit 323b and the blurred image processing unit 323c may be configured as an integrated DL model, and the image synthesis unit 323d may be configured as a separate image processing unit. Alternatively, the upscale image processing unit 323b and the blurred image processing unit 323c may be configured as separate DL models, and the image synthesis unit 323d may also be configured as a separate image processing unit.
[0069] In this embodiment, the learning of the DL model executed by the learning device 301 is similar to the flowchart regarding the learning of the DL model described in the first embodiment with reference to FIG. 4, and therefore the description thereof will be omitted.
[0070] Details of image processing executed by the imaging device 302 will be described. The weight information of the DL model is learned in advance by the learning device 301 and stored in the storage unit 311. The imaging device 302 reads out the weight information from the storage unit 311 via the network 303 and stores it in the storage unit 324. The image estimation unit 323 uses the weight information of the DL model stored in the storage unit 324 and the captured image acquired by the image acquisition unit 323a to generate an image by upscaling the captured image in the upscale image processing unit 323b, the blurred image processing unit 323c, and the image synthesis unit 323d. The generated upscaled image is stored in the recording medium 325a. When an instruction on displaying the upscaled image is issued from the user via the input unit 326, the stored image is read out and displayed on the display unit 325b. Note that the captured image stored in the recording medium 325a and its image information may be read out and upscaled by the image estimation unit 323. The above series of controls are performed by the system controller 327.
[0071] Next, the upscale image processing using the DL model executed by the image estimation unit 323 in this embodiment will be described. The procedure of the image processing is substantially the same as the flowchart explained in the first embodiment with reference to Fig. 5, and therefore will be explained with reference to Fig. 5. Each step of the image processing is mainly executed by the image acquisition unit 323a, the upscale image processing unit 323b, the blurred image processing unit 323c, or the image synthesis unit 323d of the image estimation unit 323.
[0072] First, in step S201, the image acquisition unit 323a acquires a captured image (low-resolution image, input image). In this embodiment, the captured image is acquired by the imaging device 302 and stored in the storage unit 324, but is not limited to this.
[0073] Next, in step S202, the upscaled image processing unit 323b generates an effective upscaled image (first image) by inputting the captured image to a DL model (first machine learning model). To generate an effective upscaled image, a convolutional neural network similar to the configuration shown in FIG. 1 is used. In this embodiment, the captured image is input to the DL model and an upscaled image is output, but the shooting conditions and development conditions may be input to the DL model together with the captured image using the method described in step S102. The weight information of the DL model is transmitted from the learning device 301 and stored in the storage unit 324.
[0074] Next, in step S203, the blurred image processing unit 323c generates an upscaled image (second image) having a lower effect than the effective upscaled image by inputting the effective upscaled image (first image) to a DL model (second machine learning model). A convolutional neural network similar to the configuration shown in Fig. 1 is used to generate the less effective upscaled image. Weight information of the DL model is transmitted from the learning device 301 and stored in the storage unit 324.
[0075] Next, in step S204, the image synthesis unit 323d generates a final upscaled image (third image) by taking a weighted average of the effective upscaled image and the less effective upscaled image. In this embodiment, a default weight for the weighted average is used, but a weight for the weighted average specified by the user via the input unit 326 may also be used.
[0076] When inputting a captured image into a DL model, it is not necessary to cut the image to the same size as the low-resolution training patch used during learning, but it may be processed after being decomposed into multiple overlapping patches. In this case, the resulting patches can be synthesized to create an upscaled image.
[0077] According to this embodiment, it is possible to provide an imaging device that generates two upscaled images with different effects using a DL model with a small calculation load.
[0078] [Example 3] Next, an image processing system 400 according to a third embodiment of the present invention will be described. This embodiment differs from the first and second embodiments in that it includes a processing device (computer) that transmits a captured image (low-resolution image) to be subjected to image processing to an image estimation device (image processing device) and receives a processed output image (upscaled image) from the image estimation device.
[0079] 8 is a block diagram of an image processing system 400. The image processing system 400 includes a learning device 401, an imaging device 402, an image estimation device (image processing device) 403, and a computer (control device) 404. The learning device 401 and the image estimation device 403 are, for example, servers. The computer 404 is, for example, a user terminal (personal computer, smartphone, or camera). The computer 404 is connected to the image estimation device 403 via a network 405. The image estimation device 403 is connected to the learning device 401 via a network 406.
[0080] That is, the computer 404 and the image estimation device 403 are configured to be able to communicate with each other, and the image estimation device 403 and the learning device 401 are configured to be able to communicate with each other. The learning device 401 corresponds to a third device, the computer 404 corresponds to a fourth device, and the image estimation device 403 corresponds to a fifth device.
[0081] The configuration of the learning device 401 is similar to that of the learning device 101 in the first embodiment, and therefore a description thereof will be omitted. Also, the configuration of the imaging device 402 is similar to that of the imaging device 102 in the first embodiment, and therefore a description thereof will be omitted.
[0082] The image estimation device 403 includes a storage unit 403a, an image acquisition unit 403b, an upscale image processing unit (first generation unit) 403c, a blurred image processing unit (second generation unit) 403d, an image synthesis unit (third generation unit) 403e, and a communication unit (reception unit) 403f. The storage unit 403a, the image acquisition unit 403b, the upscale image processing unit 403c, the blurred image processing unit 403d, and the image synthesis unit 403e are similar to the storage unit 103a, the image acquisition unit 103b, the upscale image processing unit 103c, the blurred image processing unit 103d, and the image synthesis unit 103e. Note that, similar to the first embodiment, the upscale image processing unit 403c, the blurred image processing unit 403d, and the image synthesis unit 403e may be configured as an integrated DL model. Alternatively, the upscale image processing unit 403c and the blurred image processing unit 403d may be configured as an integrated DL model, and the image synthesis unit 403e may be configured as a separate image processing unit. Alternatively, the upscale image processing unit 403c and the blurred image processing unit 403d may be configured as separate DL models, and the image synthesis unit 403e may also be configured as a separate image processing unit. The communication unit 403f has a function of receiving a request transmitted from the computer 404, and a function of transmitting an output image (third image) generated by the image estimation device 403 to the computer 404.
[0083] The computer 404 has a communication unit (transmission unit) 404a, a display unit 404b, an input unit 404c, a processing unit 404d, and a recording unit 404e. The communication unit 404a has a function of transmitting a request (a request for executing processing on an input image) to the image estimation device 403 to cause the image estimation device 403 to execute processing on a captured image (input image, low-resolution image). The communication unit 404a also has a function of receiving an output image (third image) processed by the image estimation device 403.
[0084] The display unit 404b has a function of displaying various information. The information displayed by the display unit 404b includes, for example, a captured image (low resolution image) to be transmitted to the image estimation device 403 and an output image (third image) received from the image estimation device 403. The input unit 404c receives an instruction to start image processing from a user. The processing unit 404d has a function of performing image processing including noise reduction and sharpness on the output image (third image) received from the image estimation device 403. The recording unit 404e stores the captured image acquired from the imaging device 402, the output image received from the image estimation device 403, and the like.
[0085] The learning of the DL model executed by the learning device 401 is similar to the flowchart relating to the learning of the DL model shown in FIG. 4 in the first embodiment, and therefore a description thereof will be omitted.
[0086] Next, image processing in this embodiment will be described with reference to Fig. 9. Fig. 9 is a flowchart related to generation of an output image using machine learning models (first machine learning model, second machine learning model) in this embodiment. The image processing shown in Fig. 9 is started when a user issues an instruction to start image processing via computer 404.
[0087] First, the operation of the computer 404 will be described. In step S401, the computer 404 transmits a request for processing a captured image (low resolution image) to the image estimation device 403. Note that the method of transmitting the captured image to be processed and its image information to the image estimation device 403 does not matter. For example, the captured image and its image information may be uploaded to the image estimation device 403 simultaneously with step S401, or may be uploaded to the image estimation device 403 before step S401. The captured image may be an image stored on a server different from the image estimation device 403. In addition, in step S401, the computer 404 may transmit an ID for authenticating a user together with the request for processing the captured image.
[0088] Then, in step S 402 , the computer 404 receives the output image (upscaled image, third image) generated in the image estimation device 403 .
[0089] Next, a description will be given of the operation of the image estimation device 403. First, in step S501, the image estimation device 403 receives a request for processing a captured image (input image, low-resolution image) transmitted from the computer 404. The image estimation device 403 determines that processing for the captured image has been instructed, and executes the processes from step S502 onward.
[0090] Next, in step S502, the image acquisition unit 403b acquires a captured image. In this embodiment, the captured image is transmitted from the computer 404. Note that the photographing conditions and development conditions may also be acquired together with the captured image and used in the steps described below.
[0091] Next, in step S503, the upscaled image processing unit 403c generates an effective upscaled image (first image) by inputting the captured image to a DL model (first machine learning model). In this embodiment, the captured image is input to the DL model to output an effective upscaled image, but the shooting conditions and development conditions may be input to the DL model together with the captured image using the method described in step S102.
[0092] Next, in step S504, the blurred image processing unit 403d inputs the effective upscaled image (first image) to a DL model (second machine learning model) to generate an upscaled image (second image) that is less effective than the effective upscaled image.
[0093] Next, in step S505, the image synthesis unit 403e performs a weighted average of the effective upscaled image and the less effective upscaled image to generate a final upscaled image (third image). In this embodiment, a default weight for the weighted average is used, but a weight for the weighted average specified by the user via the input unit 404c may also be used.
[0094] Next, in step S506, the image estimation device 403 transmits the output image (upscaled image, third image) to the computer 404.
[0095] According to this embodiment, it is possible to provide an image processing system that uses a DL model to generate two upscaled images with different effects with a small calculation load.
[0096] In each embodiment, the first image is generated by inputting an input image (captured image) to the first machine learning model, but the present invention is not limited to this. For example, an image (image based on the input image) obtained by enlarging the input image to the same number of pixels as the first image by interpolation may be input to the first machine learning model. Similarly, instead of learning the first machine learning model based on the first training image, the first machine learning model may be learned based on an image (image based on the first training image) obtained by enlarging the first training image by interpolation.
[0097] (Other Examples) The present invention can also be realized by supplying a program for implementing one or more of the functions of the above-mentioned embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.
[0098] According to each embodiment, it is possible to provide an image processing method, an image processing device, an image processing system, an imaging device, a learning method, a learning device, and a program capable of generating an upscaled image with a small calculation load. Note that the image processing device may be any device having the image processing function of each embodiment, and may be realized in the form of an imaging device or a personal computer.
[0099] The disclosure of each embodiment includes the following methods and compositions: (Method 1) generating a first image by inputting an input image or an image based on the input image into a first machine learning model; generating a second image by inputting the first image to a second machine learning model different from the first machine learning model; generating a third image using the first image and the second image; the first image, the second image, and the third image each have a larger number of pixels than the input image; The image processing method according to claim 1, wherein the second image has fewer high frequency components than the first image. (Method 2) The image processing method described in method 1, characterized in that the scale of the second machine learning model is smaller than the scale of the output layer of the first machine learning model. (Method 3) The image processing method of method 2, characterized in that the scale of the second machine learning model is less than half the overall scale of the first machine learning model. (Method 4) The output layer of the first machine learning model is a single convolutional layer; Let the number of convolution layers be L, and the kernel size of the convolution filter in the lth layer (l=1~L) be k. l , the number of channels of the lth layer convolution filter is c l , the number of convolution filters in the lth layer is n l Then, the scale of each of the output layer of the first machine learning model and the second machine learning model is JPEG2025078591000004.jpg3578
[0100] 4. The image processing method according to method 2 or 3, wherein the image is represented by the formula: (Method 5) 5. The image processing method according to any one of Methods 1 to 4, wherein the high frequency components are higher frequency components than the frequency components of the input image. (Method 6) 6. The image processing method according to any one of Methods 1 to 5, wherein the number of pixels of the first image, the number of pixels of the second image, and the number of pixels of the third image are the same as each other. (Method 7) 7. The image processing method according to any one of methods 1 to 6, wherein the third image is generated by taking a weighted average of the first image and the second image. (Configuration 1) a first generation unit that generates a first image by inputting an input image or an image based on the input image to a first machine learning model; a second generation unit that generates a second image by inputting the first image to a second machine learning model different from the first machine learning model; a third generation unit that generates a third image by using the first image and the second image, the first image, the second image, and the third image each have a larger number of pixels than the input image; The image processing device according to claim 1, wherein the second image has fewer high frequency components than the first image. (Configuration 2) An imaging device comprising the image processing device according to configuration 1 and an imaging element. (Configuration 3) A program for causing a computer to execute the image processing method according to any one of Methods 1 to 7. (Method 8) obtaining a first training image having a low resolution or an image based on the first training image, and a second training image having a high resolution corresponding to the first training image; and learning a first machine learning model and a second machine learning model based on the first training image or the image based on the first training image and the second training image; A learning method, characterized in that a method for calculating loss when training the first machine learning model and a method for calculating loss when training the second machine learning model are different from each other. (Method 9) A loss during training of the first machine learning model is calculated using an adversarial loss based on a first upscaled patch generated by inputting the first training image to the first machine learning model and the second training image; The learning method of method 8, characterized in that a loss during training of the second machine learning model is calculated using a mean square error based on a second upscaled patch generated by inputting the first upscaled patch to the second machine learning model. (Method 10) The learning method described in method 8 or 9, characterized in that the first machine learning model and the second machine learning model are trained simultaneously. (Configuration 8) an image acquisition unit that acquires a low-resolution first training image or an image based on the first training image, and a high-resolution second training image corresponding to the first training image; a learning unit that learns a first machine learning model and a second machine learning model based on the first training image or the image based on the first training image and the second training image; A learning device, characterized in that a method for calculating loss when training the first machine learning model and a method for calculating loss when training the second machine learning model are different from each other. (Configuration 9) A program for causing a computer to execute the learning method according to any one of Methods 8 to 10. (Configuration 10) An image processing system having the image processing device according to configuration 1 and a control device capable of communicating with the image processing device, the control device has a transmission unit that transmits a request for execution of processing on an input image or an image based on the input image to the image processing device; The image processing system is characterized in that the image processing device generates the third image by executing the processing on the input image in response to the request.
[0101] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. [Explanation of symbols]
[0102] 103 Image Estimation Device (Image Processing Device) 103c upscale image processing unit (first generation unit) 103d blurred image processing unit (second generation unit) 103e Image synthesis section (3rd generation section)
Claims
1. generating a first image by inputting an input image or an image based on the input image into a first machine learning model; generating a second image by inputting the first image to a second machine learning model different from the first machine learning model; generating a third image using the first image and the second image; the first image, the second image, and the third image each have a larger number of pixels than the input image; The image processing method according to claim 1, wherein the second image has fewer high frequency components than the first image.
2. The image processing method according to claim 1 , wherein the scale of the second machine learning model is smaller than the scale of an output layer of the first machine learning model.
3. The image processing method of claim 2 , wherein the scale of the second machine learning model is less than half the overall scale of the first machine learning model.
4. The output layer of the first machine learning model is a single convolutional layer; Let the number of convolution layers be L, and the kernel size of the convolution filter in the lth layer (l = 1 to L) be k l , the number of channels of the lth layer convolution filter is c l , the number of convolution filters in the lth layer is n l Then, the scale of each of the output layer of the first machine learning model and the second machine learning model is 3. The image processing method according to claim 2, wherein the image processing method is expressed by the following formula:
5. 5. The image processing method according to claim 1, wherein the high frequency components are higher in frequency than the frequency components of the input image.
6. 5. The image processing method according to claim 1, wherein the number of pixels of the first image, the number of pixels of the second image, and the number of pixels of the third image are the same.
7. 5. The image processing method according to claim 1, wherein the third image is generated by taking a weighted average of the first image and the second image.
8. a first generation unit that generates a first image by inputting an input image or an image based on the input image to a first machine learning model; a second generation unit that generates a second image by inputting the first image to a second machine learning model different from the first machine learning model; a third generation unit that generates a third image by using the first image and the second image, the first image, the second image, and the third image each have a larger number of pixels than the input image; The image processing device according to claim 1, wherein the second image has fewer high frequency components than the first image.
9. An imaging device comprising: the image processing device according to claim 8; and an imaging element.
10. A program for causing a computer to execute the image processing method according to any one of claims 1 to 4.
11. obtaining a first training image having a low resolution or an image based on the first training image, and a second training image having a high resolution corresponding to the first training image; and learning a first machine learning model and a second machine learning model based on the first training image or the image based on the first training image and the second training image; A learning method, characterized in that a method for calculating loss when training the first machine learning model and a method for calculating loss when training the second machine learning model are different from each other.
12. A loss during training of the first machine learning model is calculated using an adversarial loss based on a first upscaled patch generated by inputting the first training image to the first machine learning model and the second training image; The method of claim 11, wherein a loss during training of the second machine learning model is calculated using a mean square error based on a second upscaled patch generated by inputting the first upscaled patch to the second machine learning model.
13. The learning method according to claim 11 or 12, characterized in that the first machine learning model and the second machine learning model are trained simultaneously.
14. an image acquisition unit that acquires a low-resolution first training image or an image based on the first training image, and a high-resolution second training image corresponding to the first training image; a learning unit that learns a first machine learning model and a second machine learning model based on the first training image or the image based on the first training image and the second training image; A learning device, characterized in that a method for calculating loss when training the first machine learning model and a method for calculating loss when training the second machine learning model are different from each other.
15. A program for causing a computer to execute the learning method according to claim 11 or 12.
16. 9. An image processing system comprising the image processing device according to claim 8 and a control device capable of communicating with the image processing device, the control device has a transmission unit that transmits a request for execution of processing on an input image or an image based on the input image to the image processing device; The image processing system according to claim 1, wherein the image processing device generates the third image by executing the processing on the input image in response to the request.
Citation Information
Patent Citations
Production of high tensile strength resistance welded tube
JP1992006219A