Training method, training apparatus, image processing method, method of generating trained model, and program

A learning method for a DL model adjusts sharpness in specific image regions by using training images and interpolation, enabling efficient upscaling with varying sharpness levels without additional processing.

JP2025094329APending Publication Date: 2025-06-25CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023209778
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-25

AI Technical Summary

Technical Problem

Existing deep learning (DL) models cannot perform upscaling with different sharpness for each region of an image.

Method used

A learning method that generates a machine learning model by using low-resolution and high-resolution training images, creating interpolated images, and adjusting sharpness in specific regions through interpolation and replacement, allowing the model to learn and apply different sharpness levels during upscaling.

Benefits of technology

Enables upscaling with different sharpness levels for each region of an image using a DL model without the need for post-processing, reducing computational load on edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025094329000001_ABST
    Figure 2025094329000001_ABST
Patent Text Reader

Abstract

To provide a training method for training a machine learning model that upscales images with different sharpness for each region.SOLUTION: A training method includes the steps of: acquiring a first training image with a low resolution and a second training image corresponding to the first training image (S101); generating a third training image by enlarging the first training image by interpolation (S102); generating a fourth training image with different sharpness for each region based on the second training image, the third training image, and at least one of a high luminance region and an edge region in the first training image (S103); and training a machine learning model based on the first training image and the fourth training image (S104 to S106).SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learning method, a learning device, an image processing method, a method for generating a learned model, and a program.

Background Art

[0002] Patent Document 1 discloses a method for correcting a blurred image caused by aberration and diffraction of an optical system using a DL (Deep Learning) model. Patent Document 2 discloses a method for colorizing a monochrome image using a DL model.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the methods disclosed in Patent Document 1 and Patent Document 2, it is not possible to perform upscaling with different sharpness for each region of an image using a DL model.

[0005] Therefore, the present invention provides a learning method for learning a machine learning model for performing upscaling with different sharpness for each region of an image.

Means for Solving the Problems

[0006] As one aspect of the present invention, a learning method includes steps of: obtaining a low-resolution first training image and a high-resolution second training image corresponding to the first training image; generating a third training image by enlarging the first training image through interpolation; generating a fourth training image with different sharpness for each region based on at least one of a high-brightness region and an edge region in the first training image, the second training image, and the third training image; and performing learning of a machine learning model based on the first training image and the fourth training image.

[0007] Other objects and features of the present invention will be described in the following embodiments.

Effect of the Invention

[0008] According to the present invention, it is possible to provide a learning method for performing learning of a machine learning model for performing upscaling with different sharpness for each region of an image.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Mode for Carrying Out the Invention

[0010] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In each figure, the same members are denoted by the same reference numerals, and overlapping descriptions are omitted.

[0011] First, before giving a specific description of each embodiment, the gist of each embodiment will be explained. Each embodiment performs upscaling with different sharpness (the amount of high-frequency components included) for each region (sub-region) of an image using a DL (Deep Learning) model (machine learning model). Here, upscaling is an image enlargement process that estimates high-frequency components that cannot be represented in a low-resolution image with a small number of pixels and generates a sharp high-resolution image with a large number of pixels.

[0012] In image processing by DL, a convolutional neural network is used as the machine learning model. In a convolutional neural network, a filter for convolving an image, a bias to be added, and an activation function for non-linear transformation are used. The filter and the bias are called weights and are generated by learning using training images. For example, in the case of DL upscaling, a low-resolution image with a small number of pixels and a sharp high-resolution image with a large number of pixels corresponding thereto are used as training images. Note that since each embodiment aims to enlarge an image captured by a camera by DL upscaling, a corresponding image captured by the camera is used as the low-resolution image among the training images.

[0013] Next, referring to FIG. 10, an overview of each embodiment will be described. In each embodiment, first, a low-resolution training image (first training image) and a corresponding high-resolution training image (second training image) are obtained. Here, the low-resolution training image is a corresponding image captured by a camera, and the high-resolution training image is a sharp image (correct image) that is the target of the upscaled image output by the DL model by inputting the low-resolution training image.

[0014] Next, the low-resolution training image is enlarged by interpolation to generate an interpolated image (third training image) having the same number of pixels as the high-resolution training image. Here, since high-frequency components cannot be created in the image even by enlarging by interpolation, the interpolated image is blurry compared to the high-resolution training image. That is, the interpolated image is an image with more blur than the high-resolution training image. Also, the interpolated image is not limited to having the same number of pixels as the high-resolution training image, and may have substantially the same number of pixels. For example, the ratio of the number of pixels of the interpolated image to the high-resolution training image may be 0.9 or more and 1.1 or less.

[0015] Next, an image (fourth training image) is generated by replacing a part of the high-resolution training image with the corresponding part of the interpolated image. Here, the part where the image is replaced corresponds to the high-luminance region of the low-resolution training image, and the determination of the high-luminance region is performed by comparing the threshold value with the luminance value. For example, when the luminance value is in the range of 0 to 255, a region having a luminance value of 220 or more is regarded as a high-luminance region. Thereby, a high-resolution training image (new correct image) with different sharpness for each region is obtained. That is, a new correct image composed of a blurry interpolated image at one location and a sharp high-resolution training image at another location is obtained.

[0016] Finally, based on the low-resolution training image and the high-resolution training images with different sharpness for each region, the weights of the DL model are learned. Although the details of the learning will be described later, the learning in this embodiment is a procedure of generating weights so that the error between the upscaled image output by the DL model by inputting the low-resolution training image and the high-resolution training images with different sharpness for each region is reduced.

[0017] In each embodiment, it is characteristic that high-resolution training images with different sharpness levels for each region are used for learning as new ground truth images. By using the obtained weights, DL upscaling with different sharpness levels for each region of the image can be performed. For example, when using the weights described in the gist of each embodiment, an upscaled image with a blurred effect close to the interpolated image in the portion corresponding to the high-brightness region of the low-resolution captured image and a sharp effect close to the ground truth image in other portions can be obtained.

[0018] Note that, as in the prior art, an upscaled image with a similar effect can also be obtained by post-processing, such as replacing the high-brightness region with the interpolated image after DL upscaling. However, the image size of the upscaled image is large, and it takes time for post-processing on edge devices with low computing power. Therefore, as in each embodiment, it is required to perform upscaling with different sharpness levels for each region of the image using only the DL model without post-processing.

[0019] Also, the above-described image processing method is an example, and the present invention is not limited thereto. Details of other image processing methods will be described in the following embodiments.

[0020] [Embodiment 1] First, an image processing system according to Embodiment 1 of the present invention will be described. In this embodiment, image processing is executed to generate an upscaled image with different sharpness levels for each region from a low-resolution image using a DL model obtained by learning.

[0021] In image processing using a DL model, a filter is convolved with the input image, a bias is added, and non-linear transformation is repeated to obtain an output image with a desired effect. Note that, in this embodiment, a convolutional neural network having weights including a filter and a bias is used as the DL model.

[0022] FIG. 2 is a block diagram of the image processing system 100 in the present embodiment. FIG. 3 is an external view of the image processing system 100. The image processing system 100 includes a learning device 101, an imaging device 102, an image estimation device 103, a display device 104, a recording medium 105, an input device 106, an output device 107, and a network 108.

[0023] The learning device 101 includes a storage unit 101a, an image acquisition unit 101b, an image generation unit (first image generation unit, second image generation unit) 101c, and a learning unit 101d.

[0024] The imaging device 102 includes an optical system 102a and an imaging element 102b. The optical system 102a condenses light incident on the imaging device 102 from the subject space. The imaging element 102b receives the optical image of the subject formed through the optical system 102a and acquires a captured image (low-resolution image). The imaging element 102b is a CCD (Charge Coupled Device) sensor, a CMOS (Complementary Metal-Oxide Semiconductor) sensor, or the like.

[0025] Note that information regarding the shooting conditions of the captured image (such as the pixel pitch of the imaging element 102b, the type of optical low-pass filter, and the ISO sensitivity) can be acquired together with the image. Also, the development conditions of the captured image (such as noise reduction intensity, sharpness intensity, and image compression ratio) can be acquired together with the image. Further, these image information acquired together with the image can be transmitted to the image acquisition unit 103b of the image estimation device 103 described later together with the image. Also, a storage unit for storing the acquired image, a display unit for displaying the image, a transmission unit for transmitting the image to the outside, an output unit for storing the image in an external storage medium, etc. are not shown. Also, a control unit for controlling each part of the imaging device 102 is not shown.

[0026] The image estimation device 103 includes a storage unit 103a, an image acquisition unit 103b, and an image processing unit (estimation unit) 103c. The image estimation device 103 performs image processing to generate a high-resolution image (output) obtained by upscaling a low-resolution image (captured image) acquired by the image acquisition unit 103b using the image processing unit 103c. The low-resolution image may be an image captured by the imaging device 102 or an image stored in the recording medium 105.

[0027] A DL model is used for the image processing, and the weight information thereof is read from the storage unit 103a. The weights are those learned (generated) by the learning device 101, and the image estimation device 103 reads the weight information from the storage unit 101a via the network 108 in advance and stores it in the storage unit 103a. The weight information to be stored may be the numerical values of the weights themselves or an encoded format.

[0028] The upscaled image is output to at least one of the display device 104, the recording medium 105, and the output device 107. The display device 104 is, for example, a liquid crystal display or a projector. The user can check the image during processing via the display device 104 and perform image editing operations or the like via the input device 106. The recording medium 105 is, for example, a semiconductor memory, a hard disk, a server on a network, etc. The input device 106 is, for example, a keyboard or a mouse. The output device 107 is, for example, a printer.

[0029] Next, with reference to FIGS. 1 and 4, a method for learning weights (a method for generating a learned model) executed by the learning device 101 in this embodiment will be described. FIG. 1 is a diagram showing the flow of learning the weights of the DL model. FIG. 4 is a flowchart regarding the learning of weights. Each step in FIG. 4 is mainly executed by the image acquisition unit 101b, the image generation unit 101c, or the learning unit 101d.

[0030] First, in step S101, the image acquisition unit 101b acquires a low-resolution training patch (first training image) 201 and a high-resolution training patch (second training image, ground truth image) 200 corresponding to the low-resolution training patch. Note that the reference numerals 200 and 201 of the patches correspond to FIG. 1. In this embodiment, a patch is an image having a predetermined number of pixels. For example, the low-resolution training patch is 128×128×3 pixels, and the corresponding high-resolution training patch is 256×256×3 pixels. In this case, since the vertical and horizontal sizes of the high-resolution training patch 200 are twice those of the low-resolution training patch 201, the upscaling ratio is 2 times (4 times in terms of the number of pixels). Note that the upscaling ratio is not limited to 2 times, and other ratios may be used as long as the low-resolution training patch and the corresponding high-resolution training patch can be acquired.

[0031] In this embodiment, the low-resolution training patch and the high-resolution training patch are each a three-channel color image having RGB information, but are not limited thereto. For example, the low-resolution training patch and the high-resolution training patch may each be a one-channel monochrome image having luminance information. Also, the low-resolution training patch and the corresponding high-resolution training patch have different resolutions (number of pixels) and are images of the same subject photographed in the same scene. Such a set of images may be obtained by photographing the same subject in the same scene with optical systems having different focal lengths and cutting out corresponding portions of the two obtained images. Note that the first training image and the second training image are not limited to captured images, and may be images generated by image processing (numerical calculation) of images corresponding to captured images. That is, the second training image may be any image that includes the same subject in the same scene as the first training image.

[0032] Alternatively, a low-resolution training patch corresponding to the patch acquired by the imaging device 102 and a corresponding high-resolution training patch with little influence of blur (aberration / diffraction) by the optical system 102a may be generated by numerical calculation. In this embodiment, the low-resolution training patch and the corresponding high-resolution training patch are each generated by numerical calculation, but the present invention is not limited to this. Further, as will be described later, an image corresponding to the ground truth image, which is the target of the upscaled patch output by the DL model by inputting the low-resolution training patch, is the high-resolution training patch. Therefore, the high-resolution training patch is rich in high-frequency components (sharp).

[0033] Subsequently, in step S102, the image generation unit (first image generation unit) 101c generates an interpolated training patch (third training image) 202 obtained by interpolating and enlarging the low-resolution training patch (first training image) 201 to the same number of pixels as the high-resolution training patch (second training image) 200. In this embodiment, bicubic interpolation is used as the interpolation method, but the present invention is not limited to this. As the interpolation method, other methods such as nearest neighbor interpolation or bilinear interpolation may be used. Note that since high-frequency components cannot be created in the image even by interpolation and enlargement, the interpolated training patch is blurrier (less sharp) than the high-resolution training patch. Further, the interpolated training patch is not limited to having the same number of pixels as the high-resolution training patch, and may have substantially the same number of pixels. For example, the ratio of the number of pixels of the interpolated training patch to the high-resolution training patch may be 0.9 or more and 1.1 or less.

[0034] Subsequently, in step S103, the image generation unit (second image generation unit) 101c generates a new high-resolution training patch (fourth training image, new ground truth image) 203 with different sharpness for each region based on the high-resolution training patch 200 and the interpolated training patch 202. In this embodiment, the image generation unit 101c generates a new high-resolution training patch 203 with different sharpness for each region from the high-resolution training patch 200 and the interpolated training patch 202 based on the luminance region (high luminance region) equal to or higher than a predetermined value in the low-resolution training patch.

[0035] In this embodiment, a new high-resolution training patch 203 is generated by replacing a portion (corresponding portion) corresponding to the high-brightness region of the low-resolution training patch (first training image) 201 in the high-resolution training patch 200 with the interpolation training patch 202. Note that the determination of the high-brightness region of the low-resolution training patch 201 is performed by comparing a threshold value with the luminance value. For example, when the luminance value of the image ranges from 0 to 255, a region of the image having a luminance value of 220 or more may be regarded as the high-brightness region. Further, instead of replacing a part of the high-resolution training patch 200 with the interpolation training patch 202, a new high-resolution training patch 203 may be generated by weighted-averaging the two.

[0036] In this embodiment, a new high-resolution training patch 203 is generated based on the high-resolution training patch 200 and the interpolation training patch 202 based on the portion corresponding to the high-brightness region of the low-resolution training patch 201, but it is not limited thereto. Note that other methods will be described in other embodiments.

[0037] Subsequently, in step S104, the learning unit 101d inputs the low-resolution training patch (first training image) 201 and outputs (generates) an upscaled patch 204 upscaled by the DL model. Note that the upscaled patch 204 is an estimation of the new high-resolution training patch (fourth training image, new correct image) 203, and ideally, the two match. Here, since the new high-resolution training patch 203 has different sharpness for each region, the generated upscaled patch 204 is also an image having different sharpness for each region. Specifically, an upscaled patch is generated in which the portion corresponding to the high-brightness region of the low-resolution training patch 201 has a blurred effect close to the interpolation training patch 202, and the other portions have a sharp effect close to the high-resolution training patch 200.

[0038] Also, together with the low-resolution training patch 201, information regarding the shooting conditions or development conditions may be input into the DL model. For example, in the depth (channel) direction of the low-resolution training patch 201, an image having its ISO sensitivity as pixel values may be concatenated and input into the DL model. As a result, upscaling according to the shooting conditions or development conditions of the low-resolution training patch 201 becomes possible.

[0039] Subsequently, in step S105, the learning unit 101d updates the weights of the DL model based on the error between the new high-resolution training patch (the fourth training image, the new correct image) 203 and the upscaled patch 204 which is the estimation result thereof. Here, the weights include the filters and biases of each layer of the convolutional neural network. In the present embodiment, the weights are updated using the backpropagation method, but are not limited thereto. In mini-batch learning, the error between the new high-resolution training patch and the corresponding upscaled patch is obtained, and the weights are updated. For the loss function, for example, the L2 norm or the L1 norm may be used. The method of updating the weights (learning method) is not limited to mini-batch learning, and may be batch learning or online learning.

[0040] Subsequently, in step S106, the learning unit 101d determines whether the learning has been completed. Completion of the learning can be determined by, for example, whether the number of iterations of the learning (weight update) has reached a specified value, or whether the amount of change in the weights during the update is smaller than the specified value. If it is determined that the learning has not been completed, the process returns to step S101, and a plurality of new low-resolution training patches (the first training image) 201 and corresponding high-resolution training patches (correct images) 200 are acquired. On the other hand, if it is determined that the learning of the weights has been completed, the weight information is stored in the storage unit 101a. In this way, the learning unit 101d performs learning of the machine learning model based on the first training image and the fourth training image.

[0041] In this embodiment, as the DL model, the configuration of the convolutional neural network shown in FIG. 1 is used, but it is not limited thereto. In FIG. 1, CN represents a convolutional layer. In CN, the convolution of the input and the filter, and the sum with the bias are calculated, and the result is non-linearly transformed by the activation function. The initial values of each component of the filter and the bias are arbitrary and are determined by random numbers in this embodiment. As the activation function, for example, ReLU (Rectified Linear Unit) or sigmoid function can be used. The multi-dimensional array output by each layer except the final layer is a feature map. Generally, the feature map is a 4D array having dimensions of batch, vertical and horizontal, and channel. The skip connection 205 synthesizes the feature maps output from non-consecutive layers. The synthesis of the feature maps may be the sum of each element or concatenation in the channel direction. In this embodiment, the sum of each element is adopted.

[0042] The elements (blocks or modules) within the dotted line frame in FIG. 1 represent a residual block. A network with multiple layers of residual blocks is called a residual network and is widely used in image processing by DL.

[0043] However, this embodiment is not limited thereto, and other elements may be layered to form a network. For example, an Inception Module can be used in which convolutional layers having different convolutional filter sizes are juxtaposed and a plurality of obtained feature maps are integrated into a final feature map. Alternatively, a Dense Block having a dense skip connection may be used.

[0044] Also, the feature map can be reduced in the layer closer to the input, enlarged in the layer closer to the output, and the size of the feature map in the intermediate layer can be made smaller to reduce the processing load (convolution operation). Here, for reducing the feature map, pooling, stride, etc. can be used. Also, for enlarging the feature map, deconvolution (or transposed convolution), pixel shuffle, interpolation, etc. can be used.

[0045] Also, the low-resolution feature map is enlarged in the layer closer to the output to obtain a high-resolution feature map. In this embodiment, pixel shuffle (PS) is used as a method for upsampling the feature map, but it is not limited to this.

[0046] Next, referring to FIG. 5, the estimation (DL upscale) executed by the image estimation device 103 in this embodiment will be described. FIG. 5 is a flowchart for generating an image obtained by upscaling a low-resolution image using a DL model. Each step in FIG. 5 is mainly executed by the image acquisition unit 103b and the image processing unit 103c of the image estimation device 103.

[0047] First, in step S201, the image acquisition unit 103b acquires a captured image. The captured image is a low-resolution image as in the learning. In this embodiment, it is transmitted from the imaging device 102, but the present invention is not limited to this. For example, a captured image stored in the storage unit 103a may be used.

[0048] Subsequently, in step S202, the image processing unit 103c generates an upscaled image for the captured image using the DL model. To generate the upscaled image, a convolutional neural network similar to the configuration shown in FIG. 1 is used. In this embodiment, the captured image is input to the DL model to output the upscaled image. However, the imaging conditions and development conditions may be input to the DL model together with the captured image using the method described in step S104. The weight information of the DL model is transmitted from the learning device 101 and stored in the storage unit 103a. When inputting the captured image to the DL model, it is not necessary to cut it out to the same size as the low-resolution training patch used during learning. However, it may be processed after decomposing it into a plurality of overlapping patches. In this case, the patches obtained after processing may be combined to form the upscaled image.

[0049] In this embodiment, the case where the learning device 101 and the image estimation device 103 are separate entities has been described as an example. However, the present invention is not limited to this. The learning device 101 and the image estimation device 103 may be integrated. That is, learning (the process shown in FIG. 4) and estimation (the process shown in FIG. 5) may be performed within an integrated device.

[0050] With the above configuration, according to this embodiment, it is possible to provide a learning method for learning a machine learning model for performing upscaling with different sharpness for each region of an image. Further, according to this embodiment, it is possible to provide an image processing method for performing upscaling with different sharpness for each region of an image using a machine learning model.

[0051] [Embodiment 2] Next, an image processing system according to Embodiment 2 of the present invention will be described. Also in this embodiment, image processing is executed to generate an upscaled image with different sharpness for each region from a low-resolution image using the DL model obtained by learning. In this embodiment, it is different from Embodiment 1 in that the imaging device acquires a captured image (low-resolution image) and performs image processing.

[0052] FIG. 6 is a block diagram of the image processing system 300 in this embodiment. FIG. 7 is an external view of the image processing system 300. The image processing system 300 includes a learning device 301 and an imaging device 302 connected via a network 303. Note that the learning device 301 corresponds to the first device, and the imaging device 302 corresponds to the second device. Also, the learning device 301 and the imaging device 302 do not necessarily need to be constantly connected via the network 303.

[0053] The learning device 301 has a storage unit 311, an image acquisition unit 312, an image generation unit (first image generation unit, second image generation unit) 313, and a learning unit 314. Using these, learning of the weights of the DL model is performed in order to perform image processing for generating an upscaled image with different sharpness for each region from a low-resolution image.

[0054] The imaging device 302 captures the subject space to acquire a captured image (low-resolution image), and generates an image obtained by upscaling the captured image. Details regarding the image processing executed by the imaging device 302 will be described later. The imaging device 302 has an optical system 321 and an image sensor 322. The image estimation unit 323 has an image acquisition unit 323a and an image processing unit 323b.

[0055] Among the learning of the weights of the DL model executed by the learning device 301, the differences from the flowchart regarding the weight learning shown in FIG. 4 of the first embodiment are only steps S103 and S104, and thus only those points will be described below.

[0056] In this embodiment, in step S103, the image generation unit 313 generates a new high-resolution training patch (the fourth training image, the new ground truth image) with different sharpness for each region from the high-resolution training patch (the second training image) and the interpolation training patch (the third training image). In this embodiment, among the high-resolution training patches, a region corresponding to the strong edge region (for example, a luminance change region with a predetermined change rate or more) of the low-resolution training patch (the first training image) is replaced with the interpolation training patch to generate a new high-resolution training patch. Note that the determination of the strength of the edge region of the low-resolution training patch is performed by comparing a threshold value with the edge intensity. For example, the edge intensity may be calculated by a differential (luminance value gradient) filter such as a Laplacian filter. Further, instead of replacing a part of the high-resolution training patch with the interpolation training patch, a new high-resolution training patch may be generated by weighted-averaging the two.

[0057] In this embodiment, a new high-resolution training patch is generated based on the high-resolution training patch and the interpolation training patch based on a portion corresponding to the strong edge region of the low-resolution training patch, but the present invention is not limited thereto. Other methods will be described in another embodiment.

[0058] Subsequently, in step S104, the learning unit 314 inputs the low-resolution training patch (the first training image) and outputs (generates) a patch upscaled by the DL model. Note that the upscaled patch is an estimation of the new high-resolution training patch (the fourth training image, the new ground truth image), and ideally, the two match. Here, since the new high-resolution training patch has different sharpness for each region, the generated upscaled patch also becomes an image with different sharpness for each region. Specifically, a upscaled patch is generated in which a portion corresponding to the strong edge region of the low-resolution training patch has a blurred effect close to the interpolation training patch, and the other portions have a sharp effect close to the high-resolution training patch.

[0059] Next, details regarding the image processing executed by the imaging device 302 will be described.

[0060] The weight information of the DL model is pre-learned by the learning device 301 and stored in the storage unit 311. The imaging device 302 reads the weight information from the storage unit 311 via the network 303 and stores it in the storage unit 324. The image estimation unit 323 uses the weight information of the DL model stored in the storage unit 324 and the captured image acquired by the image acquisition unit 323a to generate an upscaled image of the captured image in the image processing unit 323b. The generated upscaled image is stored in the recording medium 325a. When an instruction regarding the display of the upscaled image is issued from the user via the input unit 326, the stored image is read out and displayed on the display unit 325b. Note that the captured image stored in the recording medium 325a and its image information may be read out and upscaled by the image estimation unit 323. The above series of controls is performed by the system controller 327.

[0061] Next, the upscaled image processing using the DL model executed by the image estimation unit 323 in this embodiment will be described. Each step of the image processing is mainly executed by the image acquisition unit 323a and the image processing unit (estimation unit) 323b of the image estimation unit 323.

[0062] First, in step S201, the image acquisition unit 323a acquires a captured image (low-resolution image). In this embodiment, the captured image is acquired by the imaging device 302 and stored in the storage unit 324, but the present invention is not limited thereto.

[0063] Subsequently, in step S202, the image processing unit 323b generates an upscaled image from the captured image using the DL model. For generating the upscaled image, a convolutional neural network similar to the configuration shown in FIG. 1 is used.

[0064] In this embodiment, the captured image is input to the DL model to output an upscaled image. However, using the method described in step S104, the shooting conditions and development conditions together with the captured image may be input to the DL model. The weight information of the DL model is transmitted from the learning device 301 and stored in the storage unit 324. When inputting the captured image to the DL model, it is not necessary to cut it out to the same size as the low-resolution training patch used during learning, but it may be processed after decomposing it into a plurality of overlapping patches. In this case, the patches obtained after processing may be combined to form an upscaled image.

[0065] With the above configuration, according to this embodiment, it is possible to provide a learning method for learning a machine learning model for performing upscaling with different sharpness for each region of an image. Also, according to this embodiment, it is possible to provide an image processing method for performing upscaling with different sharpness for each region of an image using a machine learning model.

[0066] [Embodiment 3] Next, the image processing system in Embodiment 3 of the present invention will be described. Also in this embodiment, image processing is executed to generate an upscaled image with different sharpness for each region from a low-resolution image using the DL model obtained by learning. This embodiment is different from Embodiments 1 and 2 in that it has a processing device (computer) that transmits a captured image (low-resolution image) that is the target of image processing to the image estimation device and receives the processed output image (upscaled image) from the image estimation device.

[0067] FIG. 8 is a block diagram of the image processing system 400 in this embodiment. The image processing system 400 includes a learning device 401, an imaging device 402, an image estimation device 403, and a processing device (computer) 404. The learning device 401 and the image estimation device 403 are, for example, servers. The computer 404 is, for example, a user terminal (personal computer, smartphone, or camera). The computer 404 is connected to the image estimation device 403 via a network 405. The image estimation device 403 is connected to the learning device 401 via a network 406.

[0068] That is, the computer 404 and the image estimation device 403 are configured to be communicable, and the image estimation device 403 and the learning device 401 are configured to be communicable. The learning device 401 corresponds to the third device, the computer 404 corresponds to the fourth device, and the image estimation device 403 corresponds to the fifth device.

[0069] Since the configuration of the learning device 401 is the same as that of the learning device 101 in the first embodiment, the description thereof is omitted. Since the configuration of the imaging device 402 is the same as that of the imaging device 102 in the first embodiment, the description thereof is omitted.

[0070] The image estimation device 403 includes a storage unit 403a, an image acquisition unit 403b, an image processing unit (estimation unit) 403c, and a communication unit (reception unit) 403d. Each of the storage unit 403a, the image acquisition unit 403b, and the image processing unit 403c is the same as the storage unit 103a, the image acquisition unit 103b, and the image processing unit 103c of the image estimation device 103 in the first embodiment. The communication unit 403d has a function of receiving a request transmitted from the computer 404 and a function of transmitting an output image (upscaled image) generated by the image estimation device 403 to the computer 404.

[0071] The computer 404 includes a communication unit (transmission unit) 404a, a display unit 404b, an input unit 404c, a processing unit 404d, and a recording unit 404e. The communication unit 404a has a function of transmitting a request for causing the image estimation device 403 to execute processing on a captured image (low-resolution image) to the image estimation device 403, and a function of receiving an output image (upscaled image) processed by the image estimation device 403. The display unit 404b has a function of displaying various information. The information displayed by the display unit 404b includes, for example, a captured image (low-resolution image) to be transmitted to the image estimation device 403 and an output image (upscaled image) received from the image estimation device 403. The input unit 404c receives an instruction to start image processing or the like from the user. The processing unit 404d has a function of performing image processing including noise reduction, sharpness, etc. on the output image (upscaled image) received from the image estimation device 403. The recording unit 404e stores a captured image acquired from the imaging device 402, an output image received from the image estimation device 403, and the like.

[0072] Among the learning of the weights of the DL model executed by the learning device 401, the differences from the flowchart regarding the weight learning shown in FIG. 4 of Example 1 are only in steps S103 and S104, so only those points will be described below.

[0073] In this embodiment, in step S103, the image generation unit 401c generates a new high-resolution training patch (fourth training image, new ground truth image) with different sharpness for each region from the high-resolution training patch (second training image) and the interpolation training patch (third training image). In this embodiment, a new high-resolution training patch is generated by replacing the portions of the high-resolution training patch corresponding to the high-brightness regions and strong edge regions of the low-resolution training patch (first training image) with the interpolation training patch. Note that the determination of the high-brightness regions and edge regions of the low-resolution training patch is performed in the same manner as in Example 1 and Example 2, respectively.

[0074] Subsequently, in step S104, the learning unit 401d inputs a low-resolution training patch (the first training image) and outputs (generates) a patch upscaled by the DL model. Note that the upscaled patch is an estimation of a new high-resolution training patch (the fourth training image, a new ground truth image), and ideally, the two should match. Here, since the new high-resolution training patch has different sharpness levels for each region, the generated upscaled patch will also be an image with different sharpness levels for each region. Specifically, the part corresponding to the high-brightness region and the strong edge region of the low-resolution training patch will have a blurred effect similar to the interpolation training patch, and the other parts will generate an upscaled patch with a sharp effect similar to the high-resolution training patch.

[0075] Next, the image processing in this embodiment will be described. The image processing shown in FIG. 9 starts when an instruction to start image processing is given by the user via the computer 404. First, the operation in the computer 404 will be described.

[0076] First, in step S401, the computer 404 sends a request for processing the captured image (low-resolution image) to the image estimation device 403. Note that the method of sending the captured image to be processed and its image information to the image estimation device 403 is not limited. For example, the captured image and its image information may be uploaded to the image estimation device 403 simultaneously with S401, or may have been uploaded to the image estimation device 403 before S401. Also, the captured image may be an image stored on a server different from the image estimation device 403. Further, in S401, the computer 404 may send an ID for authenticating the user together with the request for processing the captured image.

[0077] Subsequently, in step S402, the computer 404 receives the output image (upscaled image) generated within the image estimation device 403.

[0078] Next, the operation of the image estimation device 403 will be described.

[0079] First, in step S501, the image estimation device 403 receives a processing request for the captured image (low-resolution image) transmitted from the computer 404. The image estimation device 403 determines that processing of the captured image is instructed, and executes the processing after step S502.

[0080] Subsequently, in step S502, the image acquisition unit 403b acquires the captured image. In this embodiment, the captured image is transmitted from the computer 404. Note that the shooting conditions and development conditions may also be acquired together with the captured image and used in the steps described later.

[0081] Subsequently, in step S503, the image processing unit 403c uses the DL model to generate an upscaled image with different sharpness for each region of the captured image. In this embodiment, the captured image is input to the DL model to output an upscaled image. However, the shooting conditions and development conditions may be input to the DL model together with the captured image using the method described in step S104.

[0082] Subsequently, in step S504, the image estimation device 403 transmits the output image (upscaled image) to the computer 404.

[0083] With the above configuration, according to this embodiment, it is possible to provide a learning method for learning a machine learning model for performing upscaling with different sharpness for each region of an image. Further, according to this embodiment, it is possible to provide an image processing method for performing upscaling with different sharpness for each region of an image using a machine learning model.

[0084] (Other Embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. Further, it can also be realized by a circuit (for example, ASIC) that realizes one or more functions.

[0085] According to each embodiment, it is possible to provide a learning method for learning a machine learning model for performing upscaling with different sharpness for each region of an image, a method for generating a learned model, a learning device (image processing device), and a program. Further, according to each embodiment, it is possible to provide an image processing method, an image processing device, and a program for performing upscaling with different sharpness for each region of an image using a machine learning model. The image processing device may be any device having the image processing function of each embodiment, and can be realized, for example, in the form of an imaging device or a personal computer.

[0086] The disclosure of each embodiment includes the following configurations and methods. (Method 1) A step of acquiring a low-resolution first training image and a high-resolution second training image corresponding to the first training image; A step of generating a third training image obtained by enlarging the first training image by interpolation; A step of generating a fourth training image with different sharpness for each region based on at least one of a high-brightness region and an edge region in the first training image, the second training image, and the third training image; A learning method characterized by having a step of learning a machine learning model based on the first training image and the fourth training image. (Method 2) The learning method according to Method 1, wherein the second training image has more pixels than the first training image and includes the same subject in the same scene as the first training image. (Method 3) The learning method according to Method 1 or 2, wherein the third training image has the same number of pixels as the second training image and has more blur than the second training image. (Method 4) The learning method according to any one of Methods 1 to 3, wherein the fourth training image is generated by replacing a part of the second training image with a corresponding part of the third training image. (Method 5) The learning method according to any one of Methods 1 to 3, wherein the fourth training image is generated by performing a weighted average of a part of the second training image and a corresponding part of the third training image. (Method 6) The learning method according to any one of Methods 1 to 5, wherein the fourth training image has a different sharpness from other regions in a region corresponding to at least one of the high-luminance region and the edge region of the first training image. (Method 7) The high-luminance region is a luminance region of a predetermined value or more, The learning method according to any one of Methods 1 to 6, wherein the edge region is a luminance change region of a predetermined change rate or more. (Configuration 1) An image acquisition unit that acquires a low-resolution first training image and a high-resolution second training image corresponding to the first training image, A first image generation unit that generates a third training image obtained by enlarging the first training image by interpolation, A second image generation unit that generates a fourth training image having different sharpness for each region from the second training image and the third training image based on at least one of a luminance region of a predetermined value or more and a luminance change region of a predetermined change rate or more in the first training image, A learning apparatus comprising: a learning unit that performs learning of a machine learning model based on the first training image and the fourth training image. (Configuration 2) A program for causing a computer to execute the learning method according to any one of Methods 1 to 8. (Method 8) A step of acquiring a captured image, An image processing method comprising: a step of performing upscaling with different sharpness for each region of the captured image using the machine learning model obtained by the learning method according to any one of Methods 1 to 7. (Configuration 3) A program for causing a computer to execute the image processing method according to Method 10. (Method 9) Obtaining a low-resolution first training image and a high-resolution second training image corresponding to the first training image; Generating a third training image by interpolating and enlarging the first training image; Generating a fourth training image with different sharpness for each region based on at least one of the high-brightness region and the edge region in the first training image, the second training image, and the third training image; A method for generating a learned model, comprising performing learning of a machine learning model based on the first training image and the fourth training image.

[0087] The preferred embodiments of the present invention have been described above. However, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist thereof.

Explanation of Reference Numerals

[0088] 101 Learning device 101b Image acquisition unit 101c Image generation unit (first image generation unit, second image generation unit) 101d Learning unit

Claims

1. A step (S101) of obtaining a low-resolution first training image and a high-resolution second training image corresponding to the first training image; A step (S102) of generating a third training image obtained by enlarging the first training image by interpolation; A step (S103) of generating a fourth training image with different sharpness for each region based on at least one of a high-brightness region and an edge region in the first training image, the second training image, and the third training image; A learning method characterized by having a step (S104 to S106) of learning a machine learning model based on the first training image and the fourth training image.

2. The learning method according to claim 1, wherein the second training image has more pixels than the first training image and includes the same subject in the same scene as the first training image.

3. The learning method according to claim 1, wherein the third training image has the same number of pixels as the second training image and has more blur than the second training image.

4. The learning method according to any one of claims 1 to 3, wherein the fourth training image is generated by replacing a part of the second training image with a corresponding part of the third training image.

5. The learning method according to any one of claims 1 to 3, wherein the fourth training image is generated by performing a weighted average of a part of the second training image and a corresponding part of the third training image.

6. The learning method according to any one of claims 1 to 3, wherein the fourth training image has different sharpness from other regions in a region corresponding to at least one of the high-brightness region and the edge region of the first training image.

7. The high-brightness region is a luminance region of a predetermined value or more, The learning method according to any one of claims 1 to 3, wherein the edge region is a luminance change region with a predetermined change rate or more.

8. An image acquisition unit that acquires a low-resolution first training image and a high-resolution second training image corresponding to the first training image; A first image generation unit that generates a third training image obtained by enlarging the first training image by interpolation; A second image generation unit that generates a fourth training image with different sharpness for each region from the second training image and the third training image based on at least one of a luminance region of a predetermined value or more and a luminance change region with a predetermined change rate or more in the first training image; A learning device comprising a learning unit that performs learning of a machine learning model based on the first training image and the fourth training image.

9. A program characterized by causing a computer to execute the learning method according to any one of Claims 1 to 3.

10. A step of acquiring a captured image, An image processing method comprising: performing upscaling with different sharpness for each region of the captured image using the machine learning model obtained by the learning method according to any one of Claims 1 to 3.

11. A program characterized by causing a computer to execute the image processing method according to Claim 10.

12. A step of acquiring a low-resolution first training image and a high-resolution second training image corresponding to the first training image, A step of generating a third training image obtained by enlarging the first training image by interpolation, A step of generating a fourth training image with different sharpness for each region based on at least one of a high-luminance region and an edge region in the first training image, the second training image, and the third training image, A method for generating a learned model, comprising: performing learning of a machine learning model based on the first training image and the fourth training image.

Citation Information

Patent Citations

  • Method for producing learning data, learning method, device for producing learning data, learning device, and program

    JP2021140758A

  • Use of a saliency map to train a colorization ANN

    US11403485B2