Image processing method, program, and image processing apparatus
By generating and combining intermediate high-resolution images from both discriminator-trained and untrained generators, the method effectively addresses image quality degradation issues in super-resolution image processing, achieving high-quality and controlled resolution and pseudo structure appearance.
Patent Information
- Application Number
- JP2024186516
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2040-09-10
AI Technical Summary
The generator used in network interpolation for super-resolution image processing may cause image quality degradation, such as edge multiplication and color changes, leading to double subjects and unwanted color alterations in the generated high-resolution images.
An image processing method that generates two intermediate high-resolution images using a first generator trained without a discriminator and a second generator trained with a discriminator, and then combines these images to produce an estimated high-resolution image, adjusting the resolution and pseudo structure appearance while suppressing image quality degradation.
This method provides a high-quality super-resolved image by controlling the balance between pseudo structure appearance and resolution sense, while reducing computational load and image quality deterioration.
Smart Images

Figure 0007697124000001 
Figure 0007697124000002 
Figure 0007697124000003
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for super-resolving an image using a machine learning model.
Background Art
[0002] Patent Document 1 discloses a method for super-resolving an image using a machine learning model called a generative adversarial network (GAN). This method is called SRGAN (Super Resolution GAN). SRGAN performs learning using a generator that generates a high-resolution image and a discriminator that discriminates whether the input image is an image generated by the generator or an actual high-resolution image. Here, the actual high-resolution image means a high-resolution image that is not generated by the generator.
[0003] The generator learns the weights so that the discriminator cannot distinguish between the generated high-resolution image and the actual high-resolution image. As a result, a high-resolution image with a more natural appearance having a high-resolution texture can be generated. However, at the same time, as a problem, there is a possibility that a subjectively uncomfortable pseudo structure may appear.
[0004] On the other hand, Non-Patent Document 1 discloses a method for controlling the appearance of pseudo structures and the sense of resolution. In Non-Patent Document 1, a weighted average is taken between the weights of a generator that has learned super-resolution without using a discriminator (with few pseudo structures but low resolution) and the weights of a generator that has learned using a discriminator (equivalent to SRGAN, with high resolution but possible pseudo structures). A high-resolution image is generated by a generator using this weighted average weight. This method is called network interpolation. By changing the weight of the weighted average, the balance between the appearance of pseudo structures and the sense of resolution can be controlled.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Non - Patent Literature
[0006]
Non - Patent Literature 1
Summary of the Invention
Problems to be Solved by the Invention
[0007] However, through the study of the present inventor, it has been found that the generator that performs network interpolation in Non - Patent Literature 1 may cause image quality degradation such as the multiplication of edges where the subject appears double and color changes in the generated high - resolution image.
Means for Solving the Problems
[0008] An image processing method according to an embodiment of the present invention includes a step of generating a first intermediate image with a higher resolution than the first image based on the first image using a first generator, a step of generating a second intermediate image with a higher resolution than the first image based on the first image using a second generator, and a step of generating an estimated image with a higher resolution than the first image based on the first intermediate image and the second intermediate image, wherein the first generator is obtained by learning without using a discriminator, and the second generator is obtained by learning using a discriminator.
[0009] Further, an image processing method according to an embodiment of the present invention includes a step of converting a first image into a first feature map by inputting the first image into a generator, a step of generating a first intermediate image with a higher resolution than the first image and a second intermediate image with a higher resolution than the first image based on the first feature map, and a step of generating an estimated image with a higher resolution than the first image by adjusting the resolution feeling based on the first intermediate image and the second intermediate image.
Effect of the Invention
[0010] According to the present invention, in the super-resolution of an image using a machine learning model, a high-quality image can be provided.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Mode for Carrying Out the Invention
[0012] Hereinafter, a system including the processing device of the present invention will be described based on the accompanying drawings. In each figure, the same components are denoted by the same reference numerals, and redundant descriptions are omitted.
[0013] [Example 1] First, before detailing Example 1 which is an example of the present invention, its gist will be explained. Example 1 described below is one of the preferred embodiments of the present invention, and not all of it is necessary for the realization of the present invention.
[0014] In this example, a generator which is a machine learning model converts a low-resolution image (first image) into a feature map (first feature map), and further generates two intermediate images (first intermediate image and second intermediate image) that are higher in resolution than the low-resolution image from the first feature map. Hereinafter, the first intermediate image and the second intermediate image are also referred to as the first intermediate high-resolution image and the second intermediate high-resolution image, respectively.
[0015] The generator is trained using different loss functions for these two intermediate high-resolution images. The loss function has a first loss based on the difference between the intermediate high-resolution image and the high-resolution image which is the correct answer (correct answer image), and a second loss defined based on the discrimination output of a discriminator that discriminates whether the input image is an image generated by the generator.
[0016] The first intermediate high-resolution image is generated using the result of training with a loss function in which the weight of the second loss with respect to the first loss is smaller than that of the second intermediate high-resolution image. For example, the training for generating the first intermediate high-resolution image uses only the first loss as the loss function.
[0017] The second intermediate high-resolution image may be generated using the result of learning with the weighted sum of the first loss and the second loss as a loss function. As a result, the second intermediate high-resolution image has a high-resolution texture equivalent to that when learned by SRGAN, but may also be an image in which pseudo structures appear.
[0018] On the other hand, since the learning method for generating the first intermediate high-resolution image is not GAN (or its contribution is weaker than that of the second intermediate high-resolution image), the image has both high-resolution texture and reduced pseudo structures. By combining these two intermediate high-resolution images (for example, weighted average, etc.), a high-resolution image (estimated image) with adjusted resolution and appearance of pseudo structures can be generated. The estimated image is higher resolution than the low-resolution image. Hereinafter, the estimated image is also referred to as an estimated high-resolution image.
[0019] This method combines the high-resolution images that are the targets of the loss function instead of the weights of the generator like network interpolation, so that image quality degradation such as edge multiplexing and color change can be suppressed. In addition, since two intermediate high-resolution images are generated simultaneously by one generator, an increase in calculation time can also be suppressed.
[0020] Hereinafter, the stage of determining the weights of the generator and the discriminator, which are machine learning models, based on the learning dataset is called learning, and the stage of generating an estimated high-resolution image from a low-resolution image by the generator using the learned weights is called estimation. Machine learning models include, for example, neural networks, genetic programming, Bayesian networks, etc. Neural networks include CNN (Convolutional Neural Network), GAN (Generative Adversarial Network), RNN (Recurrent Neural Network), etc.
[0021] Next, the image processing system in Example 1 will be described.
[0022] FIG. 2 and FIG. 3 are respectively a block diagram and an external view of the image processing system 100.
[0023] The image processing system 100 includes a learning device 101, a resolution enhancement device 102, and a control device 103 connected to each other by a wired or wireless network.
[0024] The control device 103 includes a storage unit 131, a communication unit 132, and a display unit 133, and transmits a request to execute resolution enhancement for a low-resolution image to the resolution enhancement device 102 according to a user's instruction.
[0025] The resolution enhancement device 102 includes a storage unit 121, a communication unit 122, an acquisition unit 123, and a resolution enhancement unit 124. Using a generator that is a learned machine learning model, the resolution enhancement device 102 performs resolution enhancement processing on a low-resolution image to generate an estimated high-resolution image. The acquisition unit 123 and the resolution enhancement unit 124 can implement their functions by one or more processors (processing means) such as a CPU. The resolution enhancement device 102 acquires information on the weights of the generator from the learning device 101 and stores it in the storage unit 121.
[0026] The learning device 101 includes a storage unit 111, an acquisition unit 112, a calculation unit 113, and an update unit 114, and learns the weights of the generator. The acquisition unit 112, the calculation unit 113, and the update unit 114 can implement their functions by one or more processors (processing means) such as a CPU.
[0027] In the image processing system 100 configured as described above, the control device 103 acquires the estimated high-resolution image generated by the resolution enhancement device 102 and presents the result to the user via the display unit 133.
[0028] Next, regarding the learning of the weights executed by the learning device 101, it will be described with reference to the flowchart of FIG. 4.
[0029] The learning of Example 1 is composed of two stages: a first learning that does not use a discriminator and a second learning (GAN) that uses a discriminator. The learning device 101 has a storage unit 111, an acquisition unit 112, a calculation unit 113, and an update unit 114, and each step is executed by any of these. Note that the functions (methods) of each flowchart described below can also be realized as a program that causes a computer to execute the functions (methods).
[0030] In step S101, the acquisition unit 112 acquires one or more sets of high-resolution images and low-resolution images from the storage unit 111. The storage unit 111 stores a learning dataset composed of a plurality of high-resolution images and low-resolution images. In the corresponding low-resolution image and high-resolution image, the same object (subject) exists within the image. Note that the low-resolution image may be generated by downsampling the high-resolution image. In Example 1, the number of pixels of the high-resolution image is 16 times that of the low-resolution image (4 times horizontally and vertically). However, the relationship of the number of pixels is not limited to this. Also, the image may be either color or grayscale. Further, deterioration other than downsampling (such as JPEG compression noise) may be added to the low-resolution image. Thereby, in addition to the high-resolution function, the generator can be given a function of correcting image deterioration.
[0031] In step S102, the calculation unit 113 inputs the low-resolution image to the generator and generates first and second intermediate high-resolution images. The generator is, for example, a CNN (Convolutional Neural Network), and in Example 1, a model having the configuration shown in FIG. 1 is used. However, the invention is not limited to this.
[0032] The initial value of the weights of the generator may be generated using random numbers or the like. In the generator shown in FIG. 1, the input low-resolution image 201 is processed by the sub-network (sub-network) 211 and converted into a first feature map 202. The sub-network 211 has one or more linear sum layers. The linear sum layer has a function of taking the linear sum of the input to the linear sum layer and the weights of the linear sum layer. The linear sum layer is, for example, a convolutional layer, a deconvolutional layer, a fully connected layer, or the like.
[0033] In addition, the subnet 211 has one or more activation functions that are non-linear transformations. The activation functions are, for example, ReLU (Rectified Linear Unit), sigmoid function, hyperbolic tangent function, etc.
[0034] In Example 1, the low-resolution image 201 has fewer pixels than the corresponding high-resolution image. Therefore, in Example 1, the subnet 211 has an upsampling layer that increases the number of pixels in the horizontal and vertical directions. That is, the upsampling layer is a layer that has a function of performing upsampling on the input to the upsampling layer. The upsampling layer in Example 1 has a function of performing sub-pixel convolution (also called Pixel Shuffler). The upsampling layer is not limited to this, and an upsampling function may be realized by transposed convolution, bilinear interpolation, nearest-neighbor interpolation, etc. However, when using sub-pixel convolution, the influence of zero-padding can be reduced compared to other methods, and the degree of freedom in convolution with weights increases, so the effect of final high-resolution can be enhanced.
[0035] In Example 1, the subnet 211 has the configuration shown in Fig. 5(A).
[0036] "conv." represents convolution, "sum" represents the sum for each pixel, and "sub-pixel conv." represents sub-pixel convolution.
[0037] Also, "residual block" represents a residual block. The configuration of the residual block in Example 1 is shown in Fig. 5(B).
[0038] A residual block is a block composed of a plurality of linear sum layers and activation functions. The residual block is configured to take the sum of the input to the residual block and the result of a series of operations within the residual block and output it. "concatenation" indicates performing concatenation in the channel direction.
[0039] In Example 1, the subnet 211 has 8 residual blocks. However, the number of residual blocks is not limited to this. If you want to further enhance the performance of the generator, it is advisable to increase the number of residual blocks.
[0040] In Example 1, there are multiple upsampling layers (sub-pixel convolutions). In Example 1, in order to upsample the number of pixels of the low-resolution image by 16 times, the sub-pixel convolution that upsamples by 4 times is executed twice. If high-magnification upsampling is performed with a single upsampling layer, a grid-like pattern or the like is likely to occur in the upsampled image. Therefore, it is desirable to perform low-magnification upsampling multiple times in this way. Note that Example 1 describes an example where the upsampling layer is in the subnet 211, but the present invention is not limited to this. The upsampling layer may be present not in the subnet 211 but in the subnet 212 and the subnet 213.
[0041] The first feature map 202 is input to the subnet 212, and the first residual component 203 is generated. Also, the first feature map 202 is input to the subnet 213, and the second residual component 204 is generated.
[0042] Each of the subnet 212 and the subnet 213 has one or more linear combination layers. In Example 1, the subnet 212 and the subnet 213 are each composed of one convolutional layer. Note that the subnet 212 and the subnet 213 can also be combined into one linear combination layer. For example, if the number of filters in the convolutional layer is doubled, an output having twice the number of channels (3 if color) of the low-resolution image 201 can be obtained. By splitting this output into two in the channel direction, it may be used as the first residual component 203 and the second residual component 204.
[0043] The first residual component 203 is added to the low-resolution image 201 to generate a first intermediate high-resolution image 205. Also, the second residual component 204 is added to the low-resolution image 201 to generate a second intermediate high-resolution image 206. The low-resolution image 201 is upsampled before addition so that the number of pixels of the first residual component 203 and the second residual component 204 match. This upsampling may be bilinear interpolation, bicubic interpolation, or may use a transposed convolution layer or the like. By estimating the residual component instead of the high-resolution image itself, it becomes difficult for image quality degradation such as color change to occur from the low-resolution image 201.
[0044] Note that the low-resolution image 201 may be upsampled in advance by bicubic interpolation or the like so that the number of pixels matches the high-resolution image and input to the generator. In this case, an upsampling layer is not required in the generator. However, when the number of pixels in the horizontal and vertical directions of the low-resolution image 201 increases, the number of times of linear addition increases and the computational load becomes large. Therefore, it is desirable to input the low-resolution image 201 to the generator without upsampling it as in the first embodiment and perform upsampling inside the generator.
[0045] In step S103, the update unit 114 updates the weights of the generator based on the first loss. The first loss in the first embodiment is a loss defined based on the difference between the high-resolution image (correct image) corresponding to the low-resolution image 201 and the intermediate high-resolution image. In the first embodiment, MSE (Mean Square Error) is used, but MAE (Mean Absolute Error) or the like may also be used.
[0046] In the first embodiment, the sum of the MSE between the first intermediate high-resolution image 205 and the high-resolution image and the MSE between the second intermediate high-resolution image 206 and the high-resolution image is used as the loss function, and the weights of the generator are updated by the error backpropagation method.
[0047] In step S104, the update unit 114 determines whether the first learning is completed. Completion can be determined by whether the number of iterations of learning (weight update) reaches a predetermined number, whether the amount of change in the weight during update is smaller than a predetermined value, and so on. If it is determined in step S104 that the weight learning is not completed, the process returns to step S101, and the acquisition unit 112 acquires one or more new low-resolution images 201 and high-resolution images. On the other hand, if it is determined that the first learning is completed, the process proceeds to step S105 to start the second learning.
[0048] In step S105, the acquisition unit 112 acquires one or more high-resolution images and low-resolution images 201 from the storage unit 111.
[0049] In step S106, the arithmetic unit 113 inputs the low-resolution image 201 to the generator to generate a first intermediate high-resolution image 205 and a second intermediate high-resolution image 206.
[0050] In step S107, the arithmetic unit 113 inputs the second intermediate high-resolution image 206 and the high-resolution image to the discriminator respectively to generate a discrimination output. The discriminator discriminates whether the input image is a high-resolution image generated by the generator or an actual high-resolution image. The discriminator may use a CNN or the like. The initial value of the weight of the discriminator is determined by a random number or the like. Note that the high-resolution image input to the discriminator may be any actual high-resolution image and does not need to be an image corresponding to the low-resolution image 201.
[0051] In step S108, the update unit 114 updates the weight of the discriminator based on the discrimination output and the correct label. In Example 1, the correct label for the second intermediate high-resolution image 206 is 0, and the correct label for the actual high-resolution image is 1. The sigmoid cross entropy is used as the loss function, but other functions may also be used.
[0052] In step S109, the update unit 114 updates the weights of the generator based on the first loss and the second loss. For the first intermediate high-resolution image 205, only the first loss (MSE with the corresponding high-resolution image) is taken. For the second intermediate high-resolution image 206, a weighted sum of the first loss and the second loss is taken. The second loss is the sigmoid cross-entropy between the discrimination output when the second intermediate high-resolution image 206 is input to the discriminator and the correct label 1. Since the generator wants to learn so that the discriminator misjudges the second intermediate high-resolution image 206 as an actual high-resolution image, the correct label is set to 1 (corresponding to the actual high-resolution image). The sum of the losses of the first intermediate high-resolution image 205 and the second intermediate high-resolution image 206 respectively is used as the loss function of the generator.
[0053] By repeating the weight update with this loss function, on the side of the second intermediate high-resolution image 206, a natural-looking high-resolution image having a high-resolution texture that causes the discriminator to make a misjudgment will be generated. However, as a side effect, a false structure may appear. On the other hand, on the side of the first intermediate high-resolution image 205, both the high-resolution texture and the false structure are suppressed, and a high-resolution image with fewer high-frequency components than the second intermediate high-resolution image 206 is output.
[0054] In the first embodiment, only the first loss is used for the first intermediate high-resolution image 205, but the second loss may be used in combination. In this case, in the first intermediate high-resolution image 205, the weight of the second loss in the first loss may be made smaller than that of the second intermediate high-resolution image 206. Note that the order of step S108 and step S109 may be reversed.
[0055] In step S110, the update unit 114 determines whether the second learning has been completed. If it is determined that it is not completed, the process returns to step S105 to obtain one or more new sets of low-resolution images 201 and high-resolution images. If it is completed, the weight information is stored in the storage unit 111. Note that since only the generator is used at the time of estimation, only the weights of the generator may be stored.
[0056] Next, regarding the estimation (generation of an estimated high-resolution image) executed by the high-resolution device 102 and the control device 103, it will be described using the flowchart of FIG. 6. The high-resolution device 102 includes a storage unit 121, a communication unit 122, an acquisition unit 123, and a high-resolution unit 124, and the control device 103 includes a storage unit 131, a communication unit 132, and a display unit 133. Each step is executed by any of these.
[0057] In step S201, the communication unit 132 of the control device 103 transmits a request for executing high-resolution to the high-resolution device 102. The execution request also includes information specifying the low-resolution image 201 for which high-resolution is to be performed. Alternatively, the low-resolution image 201 itself to be high-resolved may be transmitted together with the execution request for processing.
[0058] In step S202, the communication unit 122 of the high-resolution device 102 acquires the execution request transmitted from the control device 103.
[0059] In step S203, the acquisition unit 123 acquires the generator weight information and the low-resolution image 201 for which high-resolution is to be performed from the storage unit 121. The low-resolution image 201 may be acquired from another storage device connected via wire or wirelessly.
[0060] In step S204, the high-resolution unit 124 generates a first intermediate high-resolution image 205 and a second intermediate high-resolution image 206 from the low-resolution image 201 using the generator shown in FIG. 1. The second intermediate high-resolution image 206 is a high-resolution image with a natural appearance having a high-resolution texture, but a pseudo structure may appear. On the other hand, the first intermediate high-resolution image 205 is a high-resolution image in which both the high-resolution texture and the pseudo structure are suppressed. The first intermediate high-resolution image 205 has fewer high-frequency components than the second intermediate high-resolution image 206.
[0061] In step S205, the high-resolution unit 124 generates an estimated high-resolution image 207 based on the first intermediate high-resolution image 205 and the second intermediate high-resolution image 206. In the first embodiment, the estimated high-resolution image 207 is generated by weighted averaging the first intermediate high-resolution image 205 and the second intermediate high-resolution image 206. Note that the generation of the high-resolution image 207 is not limited to weighted averaging of the first intermediate high-resolution image 205 and the second intermediate high-resolution image 206, and may be generated by replacing a partial region of the second intermediate high-resolution image 206 with the first intermediate high-resolution image 205 or the like.
[0062] In step S206, the communication unit 122 transmits the estimated high-resolution image 207 to the control device 103.
[0063] In step S207, the communication unit 132 of the control device 103 acquires the estimated high-resolution image 207. The acquired estimated high-resolution image 207 is stored in the storage unit 131 or displayed on the display unit 133. Alternatively, it may be stored in another storage device connected via wire or wirelessly from the control device 103 or the high-resolution device 102. Further, the control device 103 may be configured to acquire the first intermediate high-resolution image 205 and the second intermediate high-resolution image 206, and the control device 103 may generate the estimated high-resolution image 207. In this case, the user can adjust the resolution feeling and the appearance of the false structure while checking the actual image on the display unit 133.
[0064] Next, desirable conditions for obtaining the effects of the present invention will be described.
[0065] It is desirable that the number of linear summation layers in the generator is smaller on the output side than on the input side of the upsampling layer that is present on the most output side. This is because when upsampling is performed early in the operations of the generator, the number of times of taking the subsequent linear summation increases, and the calculation load increases. In the first embodiment, there are 40 or more linear summation layers on the input side of the upsampling layer on the most output side, but there is only one linear summation layer for each of the first residual component 203 and the second residual component 204 on the output side.
[0066] In addition, the generator is configured to include a plurality of linear summation layers, and it is desirable that the output of at least one layer among the plurality of linear summation layers be configured to be concatenated with the input to the linear summation layer in the channel direction. This means concatenation represented by "concatenation" in FIG. 5(b), for example. As a result, the generator can transmit more feature maps after the layer, and the accuracy of the first intermediate high-resolution image 205 and the second intermediate high-resolution image 206 is improved.
[0067] Furthermore, it is desirable that at least half of the plurality of linear summation layers included in the generator be concatenated with the input to the linear summation layer in the channel direction. By doing so, the accuracy of further upscaling can be further improved.
[0068] In addition, each of the plurality of residual blocks included in the generator preferably has three or more linear summation layers. By doing so, the accuracy of upscaling can be improved. Furthermore, it is desirable that the residual block has two or more activation functions. This increases the non-linear effect and improves the accuracy of upscaling.
[0069] Also, among the plurality of residual blocks included in the generator, it is desirable that the number of blocks having a batch normalization layer for performing batch normalization be less than half. Different from the recognition task, in the regression task of estimating an image from an image, the effect of improving accuracy by batch normalization is small. Therefore, in order to suppress the computational load, it is desirable to reduce the number of batch normalizations. Furthermore, if it is desired to further suppress the computational load, the generator may be configured not to have a batch normalization layer.
[0070] In addition, in the learning of the generator, it is advisable to use a pre-trained feature extractor that converts an image into a feature map. The feature extractor converts the correct high-resolution image corresponding to the low-resolution image 201 into a second feature map and converts the second intermediate high-resolution image 206 into a third feature map. It is advisable to add a third loss based on the difference (e.g., MSE) between the second feature map and the third feature map to the loss function and learn the generator. As a result, since more abstract features are also taken into account by the loss function, the appearance of the upscaled image becomes natural.
[0071] Also, the second loss is desirably based on a comparison between a value based on each of the discrimination outputs of the discriminator for a plurality of actual high-resolution images and the discrimination output of the first intermediate high-resolution image 205 or the second intermediate high-resolution image 206. This is a method called Relativistic GAN. As the value based on each of the discrimination outputs of the discriminator for a plurality of actual high-resolution images, an average value, a median value, or the like of the plurality of discrimination outputs may be used. For example, sigmoid cross-entropy is taken so that the difference between the discrimination output of the first intermediate high-resolution image 205 or the second intermediate high-resolution image 206 and the average value of the discrimination outputs of the actual high-resolution images indicates the correct label (here, 1). This enables learning from a relative perspective of whether the high-resolution images generated by the generator look more real with respect to the set of actual high-resolution images. In the conventional GAN, there has sometimes been a problem that, during learning, the actual high-resolution images are ignored and the authenticity is learned only from the high-resolution images generated by the generator. However, this problem can be avoided by Relativistic GAN, and the stability of learning can be improved.
[0072] With the above configuration, in the upscaling of an image using a machine learning model, it is possible to provide an image processing system capable of controlling the sense of resolution and the appearance of pseudo structures while suppressing an increase in computational load and deterioration in image quality. That is, in the upscaling of an image using a machine learning model, a high-quality image can be provided.
[0073] [Example 2] An image processing system according to Example 2 of the present invention will be described.
[0074] FIG. 7 and FIG. 8 are a block diagram and an external view of the image processing system 300, respectively. The image processing system 300 includes a learning device 301 and an imaging device 302. The imaging device 302 includes an optical system 321, an imaging element 322, an image processing unit 323, a storage unit 324, a communication unit 325, a display unit 326, and a system controller 327. The optical system 321 condenses the light incident from the subject space to form a subject image. The optical system 321 has functions such as zooming, aperture adjustment, and autofocus as required. The imaging element 322 converts the subject image into an electrical signal by photoelectric conversion to generate a captured image. The imaging element 322 is, for example, a CCD (Charge Coupled Device) sensor or a CMOS (Complementary Metal-Oxide Semiconductor) sensor. The captured image is a live view of the subject space before shooting or is acquired when the user presses the release button, and is subjected to predetermined processing in the image processing unit 323 and then displayed on the display unit 326.
[0075] At the time of shooting, when the user instructs digital zoom and presses the release button, the captured image (low-resolution image) is upscaled by a generator, which is a machine learning model, in the image processing unit 323. At this time, the weights learned by the learning device 301 are used. Information on the weights is read out from the learning device 301 via the communication unit 325 in advance and stored in the storage unit 324. Details regarding the learning and estimation of the generator will be described later.
[0076] Note that, in the live view when the user instructs digital zoom, an image upsampled by a high-speed method such as bilinear interpolation is displayed on the display unit 326. The captured image (estimated high-resolution image) upscaled by the generator is stored in the storage unit 324 and displayed on the display unit 326. The above operations are controlled by the system controller 327.
[0077] FIG. 8 shows a so-called single-lens reflex camera as the imaging device 302, but the imaging device 302 may be a device such as a smartphone.
[0078] Next, regarding the learning of the weights of the generator executed by the learning device 301, it will be described using the flowchart of FIG. 9. The learning device 301 includes a storage unit 311, an acquisition unit 312, a calculation unit 313, and an update unit 314, and each step is executed by any of these.
[0079] In step S301, the acquisition unit 312 acquires one or more low-resolution images and high-resolution images from the storage unit 311. In the second embodiment, the number of pixels of the high-resolution image is 16 times that of the low-resolution image, but it is not limited to this.
[0080] In step S302, the calculation unit 313 inputs the low-resolution image into the generator and generates a first intermediate high-resolution image and a second intermediate high-resolution image. In the second embodiment, the generator has the configuration shown in FIG. 10. The sub-network 411 converts the low-resolution image 401 into the first feature map 402, and from the first feature map 402, the sub-network 412 generates the first intermediate high-resolution image 403, and the sub-network 413 generates the second intermediate high-resolution image 404. The sub-network 411 has the configuration shown in FIG. 5(A), and the residual block has the configuration shown in FIG. 11. The sub-network 411 has four residual blocks. The sub-network 412 and the sub-network 413 are each composed of one convolutional layer. However, the configuration of each sub-network is not limited to this.
[0081] In step S303, the calculation unit 313 inputs the high-resolution image and the second intermediate high-resolution image 404 into the discriminator respectively and generates a discrimination output.
[0082] In step S304, the update unit 314 updates the weights of the discriminator based on the discrimination output and the correct label.
[0083] In step S305, the update unit 314 updates the weights of the generator based on the first loss and the second loss. The first loss and the second loss are the same as those described in the first embodiment.
[0084] In step S306, the update unit 314 determines whether the learning of the generator has been completed. If it is determined that the learning of the weights has not been completed, the process returns to step S301. If it is determined that the learning has been completed, the learning is terminated and the weight information is stored in the storage unit 311.
[0085] Next, with reference to the flowchart of FIG. 12, the upscaling of the digitally zoomed captured image executed by the image processing unit 323 will be described. The image processing unit 323 includes an acquisition unit 323a, an upscaling unit 323b, and an arithmetic unit 323c, and each step is executed by one of these units.
[0086] In step S401, the acquisition unit 323a extracts a partial region (low-resolution image 401) from the captured image. Since the captured image has information on all pixels acquired by the image sensor 322, only the partial region necessary for digital zooming is extracted.
[0087] In step S402, the acquisition unit 323a acquires the weight information of the generator from the storage unit 324. Note that the order of step S401 and step S402 does not matter.
[0088] In step S403, the upscaling unit 323b inputs the partial region (low-resolution image 401) of the captured image into the generator, and generates a first intermediate high-resolution image 403 and a second intermediate high-resolution image 404.
[0089] In step S404, the arithmetic unit 323c generates an estimated high-resolution image 405 by weighted averaging of the first intermediate high-resolution image 403 and the second intermediate high-resolution image 404.
[0090] In step S405, the arithmetic unit 323c scales (upsamples or downsamples) the estimated high-resolution image 405 to a specified number of pixels. Since the generator is trained to upsample the number of pixels by 16 times, it is necessary to match the magnification of the digital zoom specified by the user. For downsampling, bicubic interpolation or the like may be used, and anti-aliasing processing may be performed as necessary. When digital zoom with a magnification exceeding 4 times in one dimension is specified, the estimated high-resolution image 405 is upsampled by bicubic interpolation or the like. Alternatively, the estimated high-resolution image 405 may be input to the generator as a new low-resolution image 401.
[0091] With the above configuration, in the super-resolution of an image using a machine learning model, it is possible to provide an image processing system that can control the resolution feeling and the appearance of false structures while suppressing an increase in computational load and deterioration of image quality. That is, in the super-resolution of an image using a machine learning model, a high-quality image can be provided.
[0092] (Other embodiments) In addition, in each of the above embodiments, an example in which MSE or MAE is used as the first loss and the discrimination result by the discriminator is used as the second loss has been described, but the present invention is not limited to this. The effect of the present invention can be obtained by generating a first intermediate image and a second intermediate image having different characteristics from each other from a first feature map based on a low-resolution image, and generating an estimated image based on the first intermediate image and the second intermediate image. This is because it is possible to cover the adverse effects that occur in one of the first intermediate image and the second intermediate image with the other.
[0093] Further, the present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device reading and executing the program. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.
[0094] According to each embodiment, it is possible to provide an image processing apparatus, an imaging apparatus, an image processing method, an image processing program, and a storage medium that can generate a high-quality image in the upscaling of an image using a machine learning model.
[0095] As described above, the preferred embodiments of the present invention have been described. However, the present invention is not limited to these embodiments, and various combinations, modifications, and changes are possible within the scope of the gist thereof. For example, in the above-described embodiments, an example of setting body information or motion information as target information has been described, but the present invention is not limited thereto. As the target information, the amount of study or the amount of reading may be set.
Explanation of Reference Numerals
[0096] 200 Low-resolution image 202 First feature map 205 First intermediate image 206 Second intermediate image 207 Estimated image
Claims
1. generating a first intermediate image based on the first image using a first generator, the first intermediate image having a higher resolution than the first image; generating a second intermediate image based on the first image using a second generator, the second intermediate image having a higher resolution than the first image; generating an estimated image having a higher resolution than the first image based on the first intermediate image and the second intermediate image; The first generator is obtained by learning without using a discriminator; The image processing method according to claim 1, wherein the second generator is obtained by learning using a classifier.
2. The image processing method according to claim 1 , wherein the classifier identifies whether an image input in the learning is an image generated by a generator.
3. 3. The method of claim 1, wherein the first generator and the second generator comprise a common sub-network.
4. The method of claim 3 , wherein the sub-network generates a first feature map based on the first image.
5. The first generator generates the first intermediate image by taking a sum of a first residual component generated based on the first feature map and the first image; The image processing method according to claim 4 , wherein the second generator generates the second intermediate image by taking a sum of a second residual component generated based on the first feature map and the first image.
6. converting a first image into a first feature map by inputting the first image into a generator; generating a first intermediate image having a higher resolution than the first image and a second intermediate image having a higher resolution than the first image based on the first feature map; and generating an estimated image having a higher resolution than the first image by adjusting a sense of resolution based on the first intermediate image and the second intermediate image.
7. The generator includes: generating the first intermediate image by taking a sum of a first residual component generated based on the first feature map and the first image; The image processing method according to claim 6, further comprising the step of generating the second intermediate image by taking a sum of a second residual component generated based on the first feature map and the first image.
8. 8. The image processing method according to claim 1, wherein the first intermediate image and the second intermediate image have a larger number of pixels than the first image.
9. 8. The image processing method according to claim 5, wherein the first image is upsampled before summing so that the number of pixels of the first residual component and the number of pixels of the second residual component match.
10. 10. The image processing method according to claim 1, wherein the estimated image is generated by a weighted average of the first intermediate image and the second intermediate image.
11. obtaining a plurality of training images and a plurality of ground truth images; a step of converting the training image into a first feature map by inputting the training image into a generator, and generating a first intermediate image having a higher resolution than the training image and a second intermediate image having a higher resolution than the training image based on the first feature map; A step of identifying whether an image input to a classifier is an image generated by the generator; learning the generator based on a first loss based on a difference between the correct image corresponding to the training image and the first intermediate image or the second intermediate image, and a second loss based on a classification output of the classifier when the first intermediate image or the second intermediate image is input; and generating an estimated image by inputting a first image to the generator.
12. A program causing a computer to execute the image processing method according to any one of claims 1 to 11.
13. means for generating a first intermediate image having a higher resolution than the first image based on the first image using a first generator; means for generating a second intermediate image based on the first image using a second generator, the second intermediate image having a higher resolution than the first image; means for generating an estimated image having a higher resolution than the first image based on the first intermediate image and the second intermediate image; The first generator is obtained by learning without using a discriminator; The image processing device according to claim 1, wherein the second generator is obtained by learning using a classifier.
14. means for converting a first image into a first feature map by inputting the first image into a generator; means for generating a first intermediate image having a higher resolution than the first image and a second intermediate image having a higher resolution than the first image based on the first feature map; and means for generating an estimated image having a higher resolution than the first image by adjusting a sense of resolution based on the first intermediate image and the second intermediate image.
Citation Information
Patent Citations
Super resolution using a generative adversarial network
US20180075581A1