Method, apparatus and computer program product for generating super-resolution images
A cascaded network architecture iteratively generates higher resolution images by calculating residual images, addressing blurring issues in existing super-resolution methods and enhancing image quality and fidelity.
Patent Information
- Application Number
- CN202410051905.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-12
- Publication Date
- 2025-07-15
AI Technical Summary
Existing image/video super-resolution methods tend to lead to blur effects when magnifying images, especially in areas where high data fidelity is required, such as medicine and architecture, it is difficult to effectively improve resolution without degrading image quality.
The cascading reverse projection network architecture is adopted, and super-resolution images are generated through iterative learning of multiple reverse projection networks. The model selection module, LIIF module and image combination module are used, and the image resolution is gradually improved and data fidelity is maintained.
While improving image resolution, it effectively reduces the signal-to-noise ratio, avoids blur effects, improves image quality, and is suitable for real-time reconstruction of edge devices and clouds.
Smart Images

Figure CN120318069A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and more particularly, to a method, apparatus, and computer program product for generating super-resolution images. Background Art
[0002] Nowadays, digital display devices have made great progress. Many display devices such as document screens, billboards, etc. can support 4K or even 8K videos / images. With the rapid development of hardware, effective tools for digital signal compression, transmission, and restoration are desired. For large devices given for capturing megapixel or even gigapixel images, it is desired to effectively transmit and store these image pixels with lossless or lossy restoration.
[0003] In this context, image / video super-resolution schemes have been proposed, which are used to upsample source data to a much larger resolution or FPS. Image / video super-resolution schemes can relieve the burden of data storage and transmission, thus reducing the risk of data corruption and loss. However, existing image / video super-resolution methods are usually based on fixed upsampling, which causes a blurring effect when magnifying the source image data to a larger scale. This is a huge problem, especially for fields such as medicine and architecture that require high data fidelity. Summary of the Invention
[0004] Embodiments of the present disclosure provide a method, apparatus, and computer program product for generating super-resolution images.
[0005] In one aspect of the present disclosure, a method for generating a super-resolution image is provided. The method includes: generating, by a first network, a second image with a first super-resolution based on a first image with a first resolution; determining, by the first network, a first residual image based on the first image and the second image; and generating, by a second network, a third image with a second super-resolution based on the first residual image and the second image, where the first super-resolution is greater than the first resolution, and the second super-resolution is greater than the first super-resolution.
[0006] In another aspect of the present disclosure, an electronic device is provided. The device includes at least one processor and a memory, where the memory is coupled to the at least one processor and has instructions stored thereon. When executed by the at least one processor, the instructions cause the electronic device to perform the following actions: generating, by a first network, a second image with a first super-resolution based on a first image with a first resolution; determining, by the first network, a first residual image based on the first image and the second image; and generating, by a second network, a third image with a second super-resolution based on the first residual image and the second image, where the first super-resolution is greater than the first resolution, and the second super-resolution is greater than the first super-resolution.
[0007] In yet another aspect of the present disclosure, a computer program product is provided, which is tangibly stored on a non-transitory computer-readable medium and includes machine-executable instructions. When executed, the machine-executable instructions cause the machine to perform the following actions: generating, by a first network, a second image with a first super-resolution based on a first image with a first resolution; determining, by the first network, a first residual image based on the first image and the second image; and generating, by a second network, a third image with a second super-resolution based on the first residual image and the second image, where the first super-resolution is greater than the first resolution, and the second super-resolution is greater than the first super-resolution.
[0008] It should be understood that the content described in the Summary of the Invention section is not intended to limit the key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0010] Figure 1 A schematic diagram of a network architecture for generating a super-resolution image according to some embodiments of the present disclosure is shown;
[0011] Figure 2 A flowchart of a method for generating a super-resolution image according to some embodiments of the present disclosure is shown;
[0012] Figure 3 A flowchart of another method for generating a super-resolution image according to some embodiments of the present disclosure is shown;
[0013] Figure 4 A schematic diagram of a cascaded back-projection network according to some embodiments of the present disclosure is shown;
[0014] Figure 5 Shows a schematic diagram of a model selection module of a cascaded back-projection network according to some embodiments of the present disclosure;
[0015] Figure 6 Shows a schematic diagram of a local implicit image function (LIIF) module and an image ensemble module of a cascaded back-projection network according to some embodiments of the present disclosure;
[0016] Figure 7 Shows a schematic diagram of a back projection process according to some embodiments of the present disclosure;
[0017] Figure 8 Shows a schematic diagram of an image generated by a network architecture for generating a super-resolution image according to some embodiments of the present disclosure;
[0018] Figure 9 Shows a schematic diagram of the classification score of a model selection module according to some embodiments of the present disclosure; and
[0019] Figure 10 Shows a block diagram of a device that can implement multiple embodiments of the present disclosure. Detailed implementation manners
[0020] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0021] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter. Hereinafter, the term "image" may refer to various static images or video frames extracted from dynamic videos, and thus the present disclosure may be applicable to image processing or video processing. Hereinafter, the term "super resolution (SR)" image may refer to an image obtained by upsampling an image and having a resolution higher than that of the source image, and may be used interchangeably with a high resolution (HR) image. Hereinafter, the source image has the lowest resolution compared with other images, and thus may be referred to as a low resolution (LR) image. Hereinafter, the terms "feature map", "depth feature", and "feature vector" are all related to the features of an image and may be used interchangeably.
[0022] Figure 1 FIG. shows a schematic diagram of a network architecture 100 for generating a super-resolution image according to some embodiments of the present disclosure. As an example, the network architecture 100 may include a plurality of cascaded backprojection networks CBPN1104a to CBPNN 104n arranged in a cascade manner, such that the cascaded backprojection network architecture may be used for continuous learning, where n is any positive integer. It should be understood that CBPN is provided only as an example here, and the present disclosure is not limited thereto. In other words, in the cascaded backprojection network architecture 100 according to the present disclosure, the new image data obtained based on each image data may be used to fine-tune the architecture to automatically improve the image quality. In Figure 1 , n residual images 108a - 108n may be generated through CBPN1 104a to CBPNN104n, and CBPN1 104a to CBPNN 104n may correspond to n servers 110a - 110n. It should be understood that the number of networks, residual images, and servers here is provided only as an example, and the present disclosure is not limited thereto.
[0023] As Figure 1As shown, the network architecture 100 may include CBPN1 104a to CBPNN 104n arranged in series. CBPN1 104a may generate a second image 106a with a first super-resolution based on a first image 102 with a first resolution, and may determine a first residual image 108a based on the first image 102 and the second image 106a, where the first super-resolution is greater than the first resolution. CBPN2 104b may generate a third image 106b with a second super-resolution based on the second image 106a and the first residual image 108a, where the second super-resolution is greater than the first super-resolution. In some examples, the first image 102 may correspond to a low-resolution image, the second image 106a generates a first super-resolution image, and the third image 106b may correspond to a second super-resolution image.
[0024] In some examples, the first image 102 includes a static image or an image frame extracted from a dynamic video, and the first network 104a is disposed on an edge device or in the cloud.
[0025] In some embodiments, CBPN2 104b may also determine a second residual image 108b based on the third image 106b and the first image 102, and CBPN3 104c generates a fourth image 106c with a third super-resolution based on the second residual image 108b and the third image 106b, where the third super-resolution is greater than the second super-resolution. In some embodiments, subsequent CNPBs may iteratively perform the above process until an image with a desired resolution is generated.
[0026] By having the first network CBPN1 104a generate a second image 106a with a first super-resolution higher than the first resolution based on the first image 102 with the first resolution, and determining the first residual image 108a based on the first image 102 and the second image 106a, and having the second network CBPN 104b generate a third image 106b with a second super-resolution higher than the first super-resolution based on the second image 106a and the first residual image 108a, the network architecture 100 can maintain data fidelity while achieving super-resolution upsampling, that is, minimizing the signal-to-noise ratio (SNR), thereby greatly improving the quality of the generated super-resolution image. Moreover, the network architecture 100 is a scalable framework for implementing continuous learning and can provide more flexibility.
[0027] As Figure 1As shown, CBPN1 104a to CBPNN 104n can be respectively set in servers 110a - 110n. Servers 110a - 110n can all be set on edge devices or all be set on the cloud. In some embodiments, servers 110a - 110n can be individually set on edge devices or on the cloud to iteratively improve image quality without complex model training and setup. That is, some servers are set on edge devices while some other servers are set on the cloud, such that each CBPN can be independently deployed and trained.
[0028] It can be understood that CBPN1 104a to CBPNN 104n can be the same CBPN or can be different CBPNs. That is, CBPN1 104a to CBPNN 104n can include the same modules or different modules, as long as CBPN1 104a to CBPNN 104n can generate an SR image, determine a residual image between the SR image and the LR image, and generate the next SR image based on the SR image and the residual image.
[0029] Figure 2 A flowchart of a method 200 for generating a super - resolution image according to some embodiments of the present disclosure is shown. Figure 2 The method 200 shown can be implemented in Figure 1 the network architecture 100 shown. The method 200 includes: at 202, generating a second image with a first super - resolution based on a first image with a first resolution by a first network; at 204, determining a first residual image based on the first image and the second image by the first network; and at 206, generating a third image with a second super - resolution based on the first residual image and the second image by a second network, where the first super - resolution is greater than the first resolution, and the second super - resolution is greater than the first super - resolution.
[0030] By generating an SR image based on an LR image, determining a residual image based on the SR image and the LR image, and generating the next SR image based on the SR image and the residual image, the method 200 can maintain data fidelity while increasing the image resolution, thereby reducing the signal - to - noise ratio, avoiding the blurring effect, and improving the image quality.
[0031] Figure 3 A flowchart of another method 300 for generating a super - resolution image according to some embodiments of the present disclosure is shown. Figure 3 The method 300 shown can be in Figure 1It is implemented in the network architecture 100 shown. The method 300 includes: at 302, generating, by a first network, a second image with a first super-resolution based on a first image with a first resolution, where the first super-resolution is greater than the first resolution; at 304, determining, by the first network, a first residual image based on the first image and the second image; at 306, generating, by a second network, a third image with a second super-resolution based on the first residual image and the second image, where the second super-resolution is greater than the first super-resolution; at 308, determining, by the second network, a second residual image based on the third image and the first image; at 310, generating, by a third network, a fourth image with a third super-resolution based on the second residual image and the third image, where the third super-resolution is greater than the second super-resolution; at 312, determining whether the resolution of the fourth image meets the desired resolution; if so, at 314, stop iterating; if not, at 316, continue to iteratively execute steps 308 - 312 until the output of 312 is yes, that is, until the resolution of the generated super-resolution image meets the desired resolution.
[0032] By iteratively generating the i-th super-resolution image (i.e., the super-resolution image) based on the LR image (i.e., the first image), determining the i-th residual image between the i-th super-resolution image and the LR image, and generating the (i + 1)-th super-resolution image based on the i-th super-resolution image (i.e., the super-resolution image) and the i-th residual image, the resolution of the image can be further improved and the residual can be reduced, thereby further improving the data fidelity.
[0033] In some embodiments, it may be determined whether the last generated super-resolution image meets the desired resolution after generating any number of super-resolution images. For example, it may be determined whether the super-resolution image meets the desired resolution after generating two super-resolution images, or it may be determined whether the super-resolution image meets the desired resolution after each generation of a super-resolution image, or it may be determined whether the last generated super-resolution image meets the desired resolution after generating n (n is any positive integer) super-resolution images, which is not limited herein.
[0034] In some embodiments, determining whether the resolution of the generated super-resolution image meets the desired resolution may be determined intuitively by the human eye, or by comparing the resolution of the generated super-resolution image with a threshold resolution by means of machine detection, which is not limited herein.
[0035] Figure 4 A schematic diagram of a cascaded back-projection network CBPN 400 according to some embodiments of the present disclosure is shown. Figure 4 The CBPN 400 shown may correspond to Figure 1 any one of CBPN1 - CBPNN in Figure 4As shown, the CBPN 400 may include a model selection module 404, an LIIF module 406, an image combination module 408, and a back-projection module 410. The model selection module 404 may be used to select an appropriate model to process the LR image 402 to obtain a feature map. The LIIF module 406 may be used to generate an image patch based on the feature map and query coordinates. The image combination module 408 may be used to generate an SR image based on the image patch. The back-projection module 410 may be used to generate a residual image based on the ST image and the LR image. Below will be described with respect to Figure 6 the LIIF module 406 and the image combination module 407 in detail, and with respect to Figure 7 the back-projection module 410 will be described in detail.
[0036] As Figure 4 shown, in the model selection module 404 of the CBPN 400, a specific model is selected based on the category of the first image 402 to process the first image to obtain a feature map, which includes the depth features of the first image. In the LIIF module 406 of the CBPN 400, an image patch is generated based on the feature map and a query coordinate grid. In the image combination module 408 of the CBPN 400, a second image of the first super-resolution is generated based on the image patch. In some embodiments, in the back-projection module 410 of the CBPN 400, the second image is downsampled and a first residual image between the downsampled second image and the first image is determined. Then, the first residual image is added to the second image via the model selection module and LIIF of the next CBPN to obtain a third image of the second super-resolution. It should be understood that the model selection module 404 of the CBPN 400 and the model selection module of the next CBPN may be the same or different, may include the same or different models, and may select the same or different models to process the images of the corresponding super-resolution. In some embodiments, any number of such iterations may be experienced to obtain the final super-resolution image 412. The super-resolution image 412 may have a desired resolution.
[0037] In some examples, the LR image X corresponding to the above-mentioned first image 402 is input into the CBPN 400. In the model selection module 404 of the CBPN 400, an appropriate model for feature extraction is selected according to the category of the LR image, and a feature map f(X) is obtained via this model. In some embodiments, in the LIIF module 406 of the CBPN 400, image patches are generated based on the feature map f(X) and a query coordinate grid. In some embodiments, the feature map f(X) is upsampled in the LIIF module 406. In some embodiments, in the image combination module 408 of the CBPN 400, the i-th SR image Yi’ is generated based on the image patches. In some embodiments, the image combination module 408 is used to learn an overlapping scheme based on the image patches for complete image reconstruction, thereby obtaining the i-th SR image Yi’. In some embodiments, in the back-projection module 410 of the CBPN 400, the i-th SR image Yi’ is downsampled and a first residual image between the downsampled SR image Yi’ and the LR image X is determined. In some embodiments, downsampling the SR image Yi’ can obtain an estimated LR image Xi’. In some embodiments, the estimated LR image Xi’ is fed back by the back-projection module 410 to the model selection module of the next CBPN for feature extraction again. Then, in the LIIF module of the next CBPN, the residual between the estimated SR image Xi’ and the LR image X is added back to the i-th SR image Yi’ to obtain the (i + 1)-th SR image Yi+1’.
[0038] In some embodiments, generating image patches based on a feature map and a query coordinate grid includes: in the multi-layer perceptron MLP layer of the LIIF module, using a continuous function to generate the image patches based on the feature map and the query grid, where the image patches are overlapping. In some embodiments, the continuous function can be s = f(z, x q - v), where z is a latent feature vector, x q is the query coordinate value of the second image, and v is the spatial coordinate of the first image.
[0039] Figure 5FIG. 0 shows a schematic diagram of a model selection module 500 of a cascaded back-projection network according to some embodiments of the present disclosure. The model selection module 500 may identify the content of a first image to determine the category of the first image; for each corresponding category in the category, train the models in the model selection module respectively; and for the first image, select a specific model to process the first image based on the performance of the corresponding model in the models. The model selection module 500 may also be referred to as a model bank. The model bank is a collection of multiple pre-trained super-resolution models, which may form a domain of experts (DoE) for optimal feature extraction. The model selection module proposed following the idea of DoE is a multi-modal selection scheme, which not only realizes the automatic optimization of the model, but also provides the user with the freedom of performance inspection.
[0040] In some embodiments, in the model selection module 500, a specific model is selected to process the first image based on the category of the first image to obtain a feature map for subsequent processing by the LIIF module 406. In some embodiments, selecting a specific model to process the first image includes: identifying the content of the first image to determine the category of the first image; for each corresponding category in the category, train the models in the model selection module respectively; and for the first image, select a specific model to process the first image based on the performance of the corresponding model in the models.
[0041] The content of an image / video is domain-specific. For example, based on the content of an image / video, an image can be classified as a text image from a document file containing a large amount of text and numbers, a screenshot from an online meeting containing a large amount of digital synthetic patterns, a natural image from a scenery containing a wide color distribution and texture, a cartoon image from a comic containing a large amount of drawing styles, and so on. Each image domain has a unique pattern that requires dedicated super-resolution. In some embodiments, the model selection module 500 may identify the content of the LR image to determine the LR image category and select a model suitable for a specific LR image category from all candidate models, thereby enabling the use of domain-specific data to train a specific model.
[0042] In some embodiments, the model selection module 500 may utilize a multi-classifier to identify the content of the first image based on a classification score. The classification score is an image evaluation method that determines the category of the input LR image based on the image content, such as a natural image, a cartoon image, a text image, etc. The classification score is the output of a multi-classifier, which gives the probability of each category, and the higher the probability, the greater the likelihood that the image belongs to that category. In some embodiments, the classification score can be used to select the super-resolution model most suitable for the image content from the model selection module based on the image content.
[0043] In some embodiments, the model selection module 500 includes a multi-classifier for identifying the content of the input image / video. As shown in Figure 5 , six image categories can be considered: text images (not shown), screenshots 502, natural images 504, cartoon images 506, low-contrast images 508, old photos 510, etc. In some embodiments, more image categories can be considered. In some embodiments, the model selection module 500 can create an autoencoder with skip connections between the encoder and the decoder for feature sharing. In some embodiments, the output layer of the model selection module 500 can be a softmax function for learning multi-class prediction. Data can be scripted from the Internet for training. As shown in Figure 5 , images of different categories can have different features, and the softmax function can be used to classify the input LR images. Hereinafter, the performance of the multi-classifier on a small dataset will be described with reference to Figure 9 .
[0044] The model selection module 500 can include various models, such as Super-Resolution Convolutional Neural Network (SRCNN), Very Deep Super-Resolution (VDSR), Enhanced Deep Super-Resolution (EDSR), Robust Corner Detection Network (RCDN), Swin Transformer-based Image Restoration (SwinIR), Single Image Super-Resolution (SAN), Laplacian Pyramid Super-Resolution Network (LapSRN), Super-Resolution Residual Network (SRResNet), Generative Adversarial Network (GAN), Super-Resolution-based Generative Adversarial Network (SRGAN), Visual Geometry Group (VGG), and Variational Autoencoder (VAE), etc. Please refer back to Figure 4 , where the model selection module 404 can include EDSR, RCDN, VDSR, SwinIR, SRGAN, SAN, etc. We use their public code to train the models on each dataset and obtain the corresponding model files. For the N super-resolution models trained on 6 different datasets, we obtain a total of 6N models to form the model selection module.
[0045] In some embodiments, the model selection module 500 may determine the performance of the corresponding model in the model based on non-reference metrics for perceptual estimation. The model selection module 500 may measure the performance of each model on unknown LR images. Without knowing the ground truth, non-reference metrics may be used for perceptual estimation. The non-reference metrics identify the richness of texture and color based on natural image patterns and may include both conventional metrics and deep learning-based metrics. Examples of non-reference metrics may include the PI metric, the NIQE (Natural Image Quality Evaluator) metric, and the FID (Fréchet Inception Distance) metric. These metrics are used to learn a final score for image quality assessment. In some embodiments, weights may be assigned to all metrics, or the weights may be adjusted to select one of the metrics.
[0046] A non-reference metric is an image quality assessment method that does not require a reference image. It can estimate the visual quality of an image based on features such as the naturalness, sharpness, and contrast of the image. In some embodiments, non-reference metrics are used to select the most suitable super-resolution model for an input low-resolution image in the absence of a ground truth high-resolution image. The non-reference metric generates a score regarding the image quality, and a lower score indicates higher image quality. In some embodiments, non-reference metrics can be used to select the most suitable super-resolution model for the image content from the model selection module based on the image content.
[0047] Figure 6 A schematic diagram of the LIIF module and the image combination module of the cascaded back-projection network labeled 600 according to some embodiments of the present disclosure is shown. To perform LIIF, a feature vector or a feature map is extracted from the LR image 602, the image patch 608 is generated by the LIIF module based on the feature vector and the query coordinate grid, and the SR image 610 is generated by the image combination module based on the image patch 608. In some embodiments, the image combination module may tile the image patch 608 to generate the SR image 610. In some embodiments, the image patch 608 is an overlapping image patch
[0048] As Figure 6 shown, LIIF is built on implicit arbitrary image super-resolution. LIIF is used to learn feature-based interpolation using a multilayer perceptron (MLP) layer. Given a query coordinate and an LR feature vector, a continuous function f shared by all images can be learned, such as s = f(z, x q -v), where z is the latent feature vector obtained by the encoder, and xq where \(x\) is the sampled query coordinate value and \(v\) is the spatial coordinate of the LR image. In some embodiments, in LIIF, the image is split into the spatial coordinate \(X_{hr}\) and the value range \(S_{hr}\). \(X_{hr}\) can refer to the coordinate matrix of the image, and \(S_{hr}\) can refer to the RGB values corresponding to each coordinate of the image, that is, the pixel values at the corresponding positions. Here, \(X_{hr}\) and \(S_{hr}\) correspond to the coordinate matrix of the SR image and the RGB values of each pixel. It should be understood that for the above function \(f\), \(s\) belongs to \(S_{hr}\), and \(x\) q belongs to \(X_{hr}\). In some embodiments, the LR image is downsampled as the input. Based on LIIF, \(f\) is used to predict the RGB value corresponding to each pixel of \(X_{hr}\), that is, to obtain \(S_{hr}\). In some embodiments, the function \(f\) can represent the potential feature vector (i.e., \(z\)) corresponding to the spatial coordinate \(v\) that is closest to the query coordinate value \(x\) q (the distance of the coordinates is measured using the Euclidean distance) and the relative coordinate \((x\) q - \(v\)) to predict the RGB value (i.e., \(s\)) of the continuous image at the coordinate \(x\) q .
[0049] In some embodiments, the LIIF module can be used to generate a high - resolution (HR) image. First, the LIIF module can include an encoder that converts the LR image into depth features. The LIIF module also includes an MLP, which can take the coordinates and depth features as inputs and output the corresponding RGB values. The MLP can achieve the continuous representation of the image.
[0050] To generate a high - resolution image, it is necessary to query the output of LIIF on a finer grid. This grid is the query grid, and its size is related to the target resolution. For example, if the image needs to be magnified by two times, then the size of the query grid is four times that of the original image. Each point on the query grid is a query coordinate, which represents the pixel position on the high - resolution image.
[0051] However, directly using the query coordinates and the entire depth features as the inputs of the MLP may be unreasonable because this will lead to a large amount of computation and ignore the local information of the image. Therefore, it is necessary to divide the query coordinates so that each query coordinate only uses a small piece of depth features as the input. This small piece of depth features is the image patch, and its size is related to the neighborhood of the query coordinate. For example, if a \(3\times3\) neighborhood is desired, then the size of the image patch is a \(3\times3\) depth feature.
[0052] To ensure the alignment between the query coordinates and the image patches, it is necessary to perform a certain offset on the query coordinates so that each query coordinate is located at the center of the image patch. In this way, we can use each query coordinate and the corresponding image patch as the inputs of the MLP to obtain the RGB values on the high - resolution image.
[0053] To improve the generalization ability of the MLP, it is also necessary to overlap the query coordinates to a certain extent so that each query coordinate can use multiple different image patches as input. In this way, the diversity of the images can be utilized to increase the training data of the MLP. This degree of overlap can be represented by a parameter. For example, if each query coordinate is desired to use 4 different image patches, then the degree of overlap is 50%.
[0054] In some embodiments, Figure 6 The LIIF in [reference] can include: inputting the LR image and extracting depth features with an encoder; querying the output of the LIIF on a finer grid to obtain the query grid and query coordinates; partitioning the query coordinates to obtain image patches associated with the depth features; offsetting the query coordinates to align them with the image patches; overlapping the query coordinates so that they use multiple image patches as input; and using each query coordinate and the corresponding image patch as the input to the MLP to obtain the RGB values on the high-resolution image. In some embodiments, the query grid is partitioned and overlapped to obtain image patches, and these image patches can be overlapping because each query coordinate can use multiple image patches as input.
[0055] In some embodiments, the spatial grid refers to the spatial coordinate grid of the LR image, which can be generated by the model. The query grid refers to the query coordinate grid of the HR image or SR image, which is obtained by random sampling. In some embodiments, Figure 6 The grid sampling in [reference] can input the feature vector (including 2D LR features and the spatial grid) of the LR image extracted from the spatial grid and the coordinates in the query grid into the MLP layer to learn the continuous function f for interpolation and upsampling.
[0056] In some embodiments, to achieve upsampling of large and irregular images, the query coordinates of the pixels can be randomly sampled into the LIIF for upsampling. The conventional method simply passes all the query coordinates to the LIIF and combines all the RGB pixels to form the final image. In the embodiments of the present disclosure, an overlapping scheme based on image patches is used, where the query coordinates are partitioned into overlapping image patches, and these overlapping image patches are fed into the LIIF module to obtain the overlapping HR image patches 608. Next, these HR image patches 608 are tiled together with overlap by the image combination module to obtain the final HR image. As shown in Figure 6 the last stage shown, the overlapping area of the image patches will be the average of all the image patches. In this way, the technical solution of the present disclosure can not only achieve better visual quality but also avoid the blocking effect on the boundaries of the image patches.
[0057] Figure 7A schematic diagram of a back-projection process 700 according to some embodiments of the present disclosure is shown. In some embodiments, the back-projection process 700 may be executed in a back-projection module. In some embodiments, the back-projection process 700 may be jointly executed by the back-projection module of the first network, the model selection module of the second network, and the LIIF model. In other examples, the back-projection process 700 may also be jointly executed by the back-projection module of the first network, the model selection module of the second network, the LIIF model, and the image combination module.
[0058] As Figure 7 shown, the back-projection process 700 may include determining a residual image based on the LR image 702 and the SR image 704 generated based on the LR image. For example, the SR image 704 may be downsampled and the residual image 706 between the downsampled SR image and the LR image may be determined. Additionally, a residual block may be obtained based on the residual image 706, and the residual block may be deconvolved (e.g., performing Conv2Dtranspose) to upsample the residual image 704 to obtain the upsampled residual image 708. Further, the SR image 704 is added to the upsampled residual image 708 to obtain a refined SR image 710. It should be understood that when the LR image 702 corresponds to the first image, the refined SR image 710 may correspond to the third image of the second super-resolution. When the LR image 702 corresponds to the LR image X, the SR image 704 may correspond to the i-th SR image Yi’, and the refined SR image 710 may correspond to the (i + 1)-th SR image Yi+1’.
[0059] In some embodiments, determining the first residual image based on the first image and the second image may include downsampling the second image in the back-projection module of the first network and determining the first residual image between the downsampled second image and the first image.
[0060] In some embodiments, generating the third image based on the first residual image and the second image includes: in the model selection module, the LIIF module, and the image combination module of the second network, upsampling the first residual image and adding the upsampled first residual image to the second image to generate the third image.
[0061] In some embodiments, the back-projection module may downsample the SR image and generate a residual based on the downsampled SR image and the LR image. This residual represents the error between the SR image and the LR image, that is, if the SR image is perfect, then the residual should be zero. Therefore, it is desirable to minimize the residual to improve the quality of the SR image.
[0062] In some embodiments, the residual is fed into the second CBPN, and through the model selection module, LIIF module, and image combination module of the second CBPN, the residual is added to the SR image to generate a second SR image. It should be noted that the model selection module of the second CBPN does not directly use the residual as input. Instead, the residual is first upsampled to the same resolution as the SR image, and then the model selection module is used to extract features so that the residual and the SR image are in the same feature space for easy addition. Additionally, the model selection module of the second CBPN is not necessarily the same as that of the first CBPN, and it can be adjusted and optimized according to the dataset and device. In some embodiments, Figure 7 The generation of the residual image 704 and the steps before it can be executed in the first CBPN, and the steps after generating the residual image 704 can be executed in the second CBPN. In the CBPN architecture, the second CBPN is arranged after the CBPN. In some embodiments, more CBPNs can be designed as needed to further improve the quality of the SR image.
[0063] In some embodiments, the back-projection module 700 is based on the assumption that the ideal SR image should have the same corresponding estimated LR image as the original LR image. As Figure 7 shown, it is desired to minimize the residual between the LR image and the downsampled SR image corresponding to the estimated LR image, so that the SR image is closer to the ground truth. As described above, it is desired to maximize the data fidelity in the super-resolution image. That is, by minimizing the distance between the LR and the downsampled SR image, the SR image is forced to be constrained in the same space as the LR image.
[0064] In the CBPN according to the present disclosure, a back-projection mechanism is used to estimate the i-th residual between the i-th SR and the LR. The residual is added back to the (i + 1)-th SR image through another model selection module, LIIF, and image combination. In practice, the number of back-projection stages to be performed and trained independently for each stage of back-projection can be determined according to actual needs. We highlight the importance of this design. Each stage of back-projection is independent of the final state. At the edge device, the previous CBPN network can be inherited for super-resolution. The edge device can also develop its own back-projection module to train its own dataset. The advantage is that it can adjust the model bias of a specific dataset, thereby achieving a better SNR. At the same time, the back-projection module can be transferred to or inherited by any system or device, which makes it scalable and cost-effective.
[0065] Figure 8 A schematic diagram of an image generated by a network architecture for generating super-resolution images according to some embodiments of the present disclosure is shown.Figure 8 shows the images obtained after four iterations. The upper row shows the corresponding SR images obtained after the iteration, and the lower row shows the corresponding residual images. It can be understood that the number of iterations corresponds to the number of CBPNs and the number of stages of the back-projection module. In Figure 8 , the first iteration 802 corresponds to the following stages: generating a second image of the first super-resolution for the first image based on the first resolution, determining a first residual image based on the first image and the second image, and generating a third image of the second super-resolution based on the first residual image and the second image. It should be understood that the upper image corresponds to the second image of the first super-resolution, and the lower image corresponds to the first residual image. Similarly, the second iteration 804 corresponds to the third image of the second super-resolution and the second residual image, the third iteration 806 corresponds to the fourth image of the third super-resolution and the third residual image, and the fourth iteration 808 corresponds to the fifth image of the fourth super-resolution and the fourth residual image. In Figure 8 's example, a four-stage CBPN for generating super-resolution images is proposed, that is, a four-stage back-projection network. Figure 8 shows the effect of using multi-stage back-projection for super-resolution. It can be seen that as the number of stages of the back-projection module increases, the image quality gets better and better. The image quality obtained using the four-stage back-projection module (808) is much better than that obtained using only the one-stage back-projection module (802).
[0066] Figure 9 shows a schematic diagram of the classification scores of the model selection module according to some embodiments of the present disclosure. Specifically, Figure 9 the classification scores in correspond to the outputs of the multi-classifiers of the model selection module. The inventors tested the performance of the multi-classifiers for image content recognition. The inventors manually labeled 500 images from the DIV2K and Flcikr2K datasets for training and selected 60 images for testing, with 10 images in each category. As described above, the classification scores can give the probability of each category, and the higher the probability, the greater the possibility that the image belongs to that category. In Figure 9 , for image 902, the classification scores of the multi-classifier of the model selection module of the present disclosure show that image 902 corresponds to a screenshot; for image 904, the classification scores show that it corresponds to a cartoon image; for image 906, the classification scores show that it corresponds to a natural image; for image 908, the classification scores show that it corresponds to a low-contrast image; and for image 910, the classification scores show that it corresponds to an old photo. It can be seen that the classifier of the model selection module according to the embodiments of the present disclosure can correctly identify the content of the image and determine the image category according to the image content.
[0067] It should be understood that the method, device, and computer program product for generating super-resolution images provided by the present disclosure can maintain data fidelity, reduce the signal-to-noise ratio, and improve image quality when generating high-resolution images by generating a second image with a first super-resolution higher than the first resolution based on a first image of the first resolution by a first network, determining a first residual image based on the first image and the second image, and generating a third image with a second super-resolution higher than the first super-resolution based on the second image 106a and the first residual image by a second network, thereby improving the user experience. The architecture of the present disclosure sets up a back-projection scheme in a cascaded manner, enabling continuous learning and contributing to compressing and restoring data with a fixed data storage space. Moreover, the architecture of the present disclosure can also be applied to the edge or cloud for real-time reconstruction. In addition, in the present disclosure, following the idea of DoE, a multi-modal selection scheme is proposed to form a unique model selection module. This can not only automatically optimize the model but also provide users with the freedom to perform performance checks.
[0068] Figure 10 FIG. shows a schematic block diagram of an exemplary device 1000 that can be used to implement embodiments of the present disclosure. As shown, the device 1000 includes a determination unit 1001, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 1002 or computer program instructions loaded from a storage unit 10010 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The determination unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0069] Multiple components in the device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disc, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0070] The determination unit 1001 can be various general-purpose and / or special-purpose processing components with processing and determination capabilities. Some examples of the determination unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) determination chips, various determination units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The determination unit 1001 executes the various methods and processes described above, such as method 300. For example, in some embodiments, method 300 may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the determination unit 1001, one or more steps of the method 300 described above may be executed. Alternatively, in other embodiments, the determination unit 1001 may be configured to execute method 300 in any other suitable manner (e.g., by means of firmware).
[0071] The functions described above herein can be performed at least in part by one or more hardware logic components. By way of example and not limitation, the exemplary classes of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0072] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on a remote machine or server.
[0073] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. Further, although the operations are depicted in a particular order, this should be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of the present disclosure. Certain features that are described in the context of separate embodiments can also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation can also be implemented separately or in any suitable sub-combination in multiple implementations.
[0074] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for generating a super-resolution image, comprising: generating, by a first network, a second image with a first super-resolution based on a first image with a first resolution; determining, by the first network, a first residual image based on the first image and the second image; and generating, by a second network, a third image with a second super-resolution based on the first residual image and the second image, where the first super-resolution is greater than the first resolution, and the second super-resolution is greater than the first super-resolution.
2. The method according to claim 1, further comprising: determining, by the second network, a second residual image based on the third image and the first image; generating, by a third network, a fourth image with a third super-resolution based on the second residual image and the third image, where the third super-resolution is greater than the second super-resolution.
3. The method according to claim 2 further comprises: determining whether the resolution of the fourth image meets a desired resolution, and if the resolution of the fourth image meets the desired resolution, stopping the iteration.
4. The method according to claim 1, wherein generating the second image based on the first image comprises: in a model selection module of the first network, selecting a specific model based on the category of the first image to process the first image to obtain a feature map; in a local implicit image function (LIIF) module of the first network, generating image patches based on the feature map and a query coordinate grid; and in an image combination module of the first network, generating the second image based on the image patches.
5. The method according to claim 4, wherein selecting the specific model to process the first image comprises: identifying the content of the first image to determine the category of the first image; and for corresponding categories in the category, respectively training the models in the model selection module; for the first image, selecting the specific model to process the first image based on the performance of the corresponding model in the models.
6. The method according to claim 5, wherein using a multi-classifier of the model selection module, based on a classification score, identifying the content of the first image to determine the category of the first image; and wherein determining the performance of the corresponding model in the models based on a non-reference metric for perceptual estimation.
7. The method according to claim 4, wherein generating the image patches based on the feature map and the query coordinate grid comprises: in a multi-layer perceptron (MLP) layer of the LIIF module, generating the image patches based on the feature map and the query grid using a continuous function, where the image patches are overlapping.
8. The method according to claim 1, wherein determining the first residual image based on the first image and the second image comprises: in a back-projection module of the first network, downsampling the second image; and determining the first residual image between the downsampled second image and the first image.
9. The method according to claim 8, wherein generating the third image based on the first residual image and the second image comprises: In the model selection module, LIIF module, and image combination module of the second network, perform upsampling on the first residual image; and Add the upsampled first residual image to the second image to generate the third image.
10. The method according to claim 1, wherein the first image includes a static image or an image frame extracted from a dynamic video, and wherein the first network is disposed on an edge device or in the cloud.
11. An electronic device, comprising: At least one processor; and A memory coupled to the at least one processor and having instructions stored thereon, the instructions, when executed by the at least one processor, cause the electronic device to perform actions, the actions including: Generate a second image of the first super-resolution based on a first image of a first resolution by a first network; Determine a first residual image based on the first image and the second image by the first network; and Generate a third image of the second super-resolution based on the first residual image and the second image by a second network, wherein the first super-resolution is greater than the first resolution, and the second super-resolution is greater than the first super-resolution.
12. The electronic device according to claim 11, the actions further including: Determine a second residual image based on the third image and the first image by the second network; Generate a fourth image of the third super-resolution based on the second residual image and the third image by a third network, wherein the third super-resolution is greater than the second super-resolution.
13. The electronic device according to claim 12, wherein the action further comprises: Determine whether the resolution of the fourth image meets the desired resolution, and if the resolution of the fourth image meets the desired resolution, stop the iteration.
14. The electronic device according to claim 11, wherein generating the second image based on the first image includes: In the model selection module of the first network, select a specific model based on the category of the first image to process the first image to obtain a feature map; In the local implicit image function (LIIF) module of the first network, generate image patches based on the feature map and a query coordinate grid; and In the image combination module of the first network, generate the second image based on the image patches.
15. The electronic device according to claim 14, selecting the specific model to process the first image for the little red flower includes: Identify the content of the first image to determine the category of the first image; and For the corresponding categories in the category, train the models in the model selection module respectively; For the first image, select the specific model to process the first image based on the performance of the corresponding model in the models.
16. The electronic device according to claim 15, wherein the model selection module includes a multi-classifier, the multi-classifier identifies the content of the first image based on a classification score to determine the category of the first image; and wherein the model selection module determines the performance of the corresponding model in the models based on a non-reference metric for perceptual estimation.
17. The electronic device according to claim 14, wherein generating the image patch based on the feature map and the query coordinate grid includes: In a multi-layer perceptron (MLP) layer of the LIIF module, using a continuous function to generate the image patch based on the feature map and the query grid, wherein the image patches are overlapping.
18. The electronic device according to claim 11, wherein determining the first residual image based on the first image and the second image includes: In a back-projection module of the first network, downsampling the second image; And Determining the first residual image between the downsampled second image and the first image.
19. The electronic device according to claim 18, wherein generating the third image based on the first residual image and the second image includes: In a model selection module, an LIIF module, and an image combination module of the second network, upsampling the first residual image; And Adding the upsampled first residual image to the second image to generate the third image.
20. A computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including machine-executable instructions that, when executed, cause a machine to perform operations, the operations including: Generating, by a first network, a second image with a first super-resolution based on a first image with a first resolution; Determining, by the first network, a first residual image based on the first image and the second image; And Generating, by a second network, a third image with a second super-resolution based on the first residual image and the second image, wherein the first super-resolution is greater than the first resolution, and the second super-resolution is greater than the first super-resolution.