Image Network System Based on Unsupervised Degradation Feature Learning and Its Super-Resolution Algorithm

Through unsupervised degradation feature learning and contrast learning methods, degradation features in real-world images are extracted, and the problem of lack of degradation features in low-resolution images in the prior art is solved, achieving efficient image super-scoring effect.

CN114612300BActive Publication Date: 2025-06-10HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210198241.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-02
Publication Date
2025-06-10
Estimated Expiration
2042-03-02

AI Technical Summary

Technical Problem

The existing single image super-scoring methods perform poorly on real-world images because the low-resolution images used in training lack degenerate features in real-world, which makes the model unable to effectively learn super-scoring methods.

Method used

An image network system based on unsupervised degradation feature learning is adopted to extract degradation features in the image through comparative learning, generate low-resolution images with real-world degradation features, and use the super-segment reconstruction module to perform image super-segment.

Benefits of technology

This method can effectively extract and utilize degraded features in real-world images to generate high-quality high-resolution images, improving the effect and generalization ability of image super-score.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114612300B_ABST
    Figure CN114612300B_ABST
Patent Text Reader

Abstract

The present invention discloses an image network system based on unsupervised degradation feature learning and its super-resolution algorithm. The system includes a degradation feature extraction module, a downsampling module, and a super-resolution reconstruction module. The present invention relates to the technical fields of deep learning and image super-resolution. Specifically, an image network system based on unsupervised degradation feature learning and its super-resolution algorithm are provided. Compared with other super-resolution algorithms for real-world images, this algorithm does not require explicit degradation estimation, but directly learns feature expressions in the representation space of images to distinguish image degradation, and also has good adaptability for the extraction of complex degradation feature images. Compared with the super-resolution algorithm for blur kernel estimation based on a generative adversarial network, this method is less affected by estimation errors and has strong accuracy in feature extraction. Finally, the super-resolution network has a small number of parameters, strong coupling, and great optimization potential.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of deep learning and image super-resolution, and specifically to an image network system based on unsupervised degradation feature learning and its super-resolution algorithm. Background Art

[0002] Image super-resolution is an important research direction in the field of computer vision in recent years. The purpose of single-image super-resolution is to perform end-to-end training on low-resolution images to obtain clearer high-resolution images. Methods such as linear interpolation, nearest-neighbor interpolation, and bicubic interpolation based on the sampling theory have emerged in the super-resolution field and are a simple and convenient super-resolution means. According to these operators, a high-resolution image can be quickly generated without occupying additional space. Recent research has shown that methods based on convolutional neural networks have achieved more significant results than traditional methods. By using the end-to-end mapping from low-resolution images to high-resolution images in deep learning, these deep networks can extract high-level features in the input images, thereby learning the complex non-linear mapping relationships between images. However, the above single-image super-resolution methods have not achieved good results when applied to real-world image super-resolution. The reason is that the low-resolution images used in the training of these methods are generally obtained by interpolating and downsampling high-resolution images and do not have the degradation features in the real world. To address this problem, how to obtain real-world low-resolution-high-resolution image pairs is the main research direction at present. One type of method is to use a generative adversarial network to generate low-resolution images with degradation features, but the generated images lack details and realism; another type of method is to model image degradation, believing that image degradation is caused by a blur kernel and noise, and then estimate the blur kernel to generate the corresponding image. The disadvantage of this type of method is that it is difficult to predict the blur kernel in the real world and there are large errors.

[0003] Unsupervised learning is a learning method in machine learning. It can find previously undetected patterns in a dataset without pre-existing labels and with minimal human supervision. In computer vision, unsupervised usually means that there are no labeled images in the training set that match the training pictures. The key feature of unsupervised learning is that the data passed to the algorithm is very rich in internal structure, while the targets and rewards for training are very scarce. Therefore, unsupervised learning aims to learn the same characteristics of similar data from a large amount of data, encode them into high-level representations, and then fine-tune the learning model according to different specific tasks to achieve excellent results. Contrastive learning is a type of unsupervised learning, and its core idea is to reduce the distance between positive samples and increase the distance between positive samples and negative samples.

[0004] During the formation, recording, and processing of images, the quality of the images is prone to degradation due to the imperfections of the imaging system, recording device, and processing method. Therefore, low-resolution images in the real world can generally be regarded as degraded images. In view of the poor performance of traditional single-image super-resolution methods on real-world images, it can be analyzed that the low-resolution images used in the training of traditional single-image super-resolution are obtained through simple downsampling operators, such as bicubic interpolation. This simple downsampling strategy makes the training data not have the degradation characteristics contained in low-resolution images in the real world, that is, the model cannot learn how to deal with the super-resolution method for images with degradation characteristics. In real-world image super-resolution, the low-resolution images in our training data all have degradation characteristics and there are no corresponding high-resolution images. Therefore, it is not feasible to directly train with supervised methods. Although there are no paired high-resolution images, we can start from uncorresponding high-resolution images, downsample them into low-resolution images with degradation characteristics, and thus form image pairs for the model to use. For the extraction of the degradation characteristics of images, due to the lack of labeled images, the traditional variational encoder method cannot be used, so it is also necessary to learn the degradation characteristics of images through unsupervised methods. For the above problems, the use of contrast learning and perceptual loss optimization algorithms can solve them one by one.

[0005] Current super-resolution algorithms based on convolutional neural networks (CNNs) are usually trained on large labeled datasets of high-quality images. However, the networks obtained in this way often have poor generalization ability for low-resolution images with blur kernels and noise in the real world, and the quality of the generated high-resolution images is not good. Since the degradation of images destroys the structure and statistical characteristics of pixels in the neighborhood, it is impossible to obtain low-resolution degraded images by common bicubic interpolation downsampling. The degradation model of images can generally be regarded as a high-resolution image being downsampled by a blur kernel and then adding noise to obtain a degraded image. Some existing methods have proposed super-resolution algorithms based on blur kernel estimation based on this model. However, this explicit degradation estimation in the pixel space not only depends on the effectiveness of the degradation model but also requires a strong prior to assist the model in learning the blur kernel. Therefore, this patent proposes an image network system based on unsupervised degradation feature learning and its super-resolution algorithm, aiming to use contrast learning to extract the implicit degradation feature representation in real-world images, so as to generate corresponding low-resolution images. Summary of the Invention

[0006] In view of the above situation, to make up for the above existing defects, the present invention provides an image network system based on unsupervised degradation feature learning and its super-resolution algorithm. Compared with other super-resolution algorithms for real-world images, this algorithm does not require explicit degradation estimation, but directly learns feature expressions in the representation space of images to distinguish image degradation, and also has good adaptability for the extraction of complex degradation feature images. Compared with the super-resolution algorithm for blur kernel estimation based on the generative adversarial network, this method is less affected by estimation errors and has strong accuracy in feature extraction. Finally, the super-resolution network has a small number of parameters, strong coupling, and great optimization potential.

[0007] The present invention provides the following technical solutions: The image network system based on unsupervised degradation feature learning proposed by the present invention includes a degradation feature extraction module, a downsampling module, and a super-resolution reconstruction module;

[0008] The degradation feature extraction module is responsible for training an encoder capable of extracting image degradation features through contrastive learning;

[0009] The downsampling module is responsible for generating a low-resolution image from a high-resolution image through a linear depth convolution downsampling network. At the same time, the trained encoder will ensure that the generated image can carry the same degradation features as the real world;

[0010] The super-resolution reconstruction module is responsible for inputting the paired training images obtained from the degradation feature extraction module and the downsampling module into the super-resolution reconstruction module, and training the super-resolution network and the downsampling network at the same time to further enhance the generation and super-resolution effects and complete the super-resolution task of real-world images.

[0011] Further, the degradation feature extraction module includes two networks, an encoder and a multi-layer perceptron. The encoder includes a feature encoder and a momentum encoder. The role of the feature encoder is to extract degradation features in the image, and its input is a picture; the role of the momentum encoder is to process high-resolution image blocks, and its input is a picture; the role of the multi-layer perceptron is to convert the degradation representation into positive and negative samples.

[0012] Further, the downsampling module generates a low-resolution image from a high-resolution image through a generator.

[0013] The super-resolution algorithm of the image network system based on unsupervised degradation feature learning specifically includes the following steps:

[0014] (1) The degradation features in the real-world images are extracted by the degradation feature extraction module based on contrastive learning. According to the degradation model of the real-world images, it can be reasonably assumed that the degradation features of the low-resolution real-world images in the same data set are the same, but different from the degradation features of the high-resolution images. Therefore, the degradation feature extraction module classifies the degradation features extracted from the real-world images as positive samples and the degradation features extracted from the high-resolution images as negative samples. First, the real-world images and non-corresponding high-resolution images in the data set are divided into two groups and input into the feature encoder and momentum encoder respectively. The feature encoder and momentum encoder adopt the Moco strategy. The network structure of the feature encoder and the momentum encoder are the same, and the initial parameters are also the same. After the training starts, the parameters of the two will no longer be the same. The output of the feature encoder is a vector representation of a feature. These features are grouped in pairs and then passed through a multi-layer perceptron. One feature in the real-world image will be used as a query sample, the other as its positive sample, and the features of the two high-resolution images will be used as negative samples. The process of contrastive learning is to shorten the distance between the query sample and the positive sample and to increase the distance between it and the negative sample. In other words, we want to make the positive samples as similar as possible and the positive and negative samples as dissimilar as possible. InfoNCE is used here to measure similarity. Previous work on contrastive learning has shown that a large dictionary containing rich negative samples is critical for learning good representations. Therefore, a queue is used here to maintain the negative samples obtained in each batch.

[0015] (2) Establish an image degradation model. Starting from the image degradation model, a linear downsampling network is used to convolve and downsample the convolution kernel of the image. A high-resolution image is input and a low-resolution image is output. A multi-layer network is used for training. During the training process, a weighted loss of color loss, adversarial loss, and feature matching loss is used as the total loss function. Among them, the feature matching loss is used to ensure that the features of the generated image obtained by the encoder are similar to the features of the real-world image obtained by the encoder. The role of color loss and adversarial loss is to ensure that the basic structural information of the image after downsampling does not change too much.

[0016] (3) Based on the paired image data sets generated by the degradation feature extraction module and the downsampling module, a super-resolution reconstruction module is used to super-resolve the generated low-resolution images, and the corresponding high-resolution images are used for training. The super-resolution reconstruction module is a deep convolutional neural network that uses a residual network to reduce the problem of network overfitting and gradient disappearance.

[0017] The beneficial effects achieved by the present invention using the above structure are as follows: The image network system and its super-resolution algorithm based on unsupervised degradation feature learning proposed by the present invention have two main tasks, one is to generate paired low-resolution and high-resolution image tasks, and the other is the image super-resolution task; for the first task, contrastive learning can be used to shorten the distance between low-resolution image features (positive samples) and distance from them to negative samples (high-resolution image features), thereby obtaining an implicit expression of degradation features. For the second task, the super-resolution module is mainly based on a residual feedforward neural network, and different loss functions are introduced to ensure the image super-resolution quality and the restoration of texture details; the main steps of the method can be summarized as first training a degradation feature encoder by contrastive learning, and the feature encoder can extract its implicit degradation features from the image. Using the extracted degradation features, the model can make the low-resolution images obtained through the downsampling network also have the same degradation features, thereby generating paired training data for the super-resolution network.

[0018] It is worth mentioning that, unlike the general unsupervised real-world image super-resolution algorithm, this method learns implicit representations based on the degradation differences between high-resolution and low-resolution images, and does not require the introduction of some special prior knowledge for explicit estimation, so it can better adapt to the super-resolution of different data sets. In addition, the super-resolution module used in this method has a small number of parameters and a moderate network depth, which ensures good super-resolution effects while reducing computational overhead and time consumption. Moreover, this module can be arbitrarily replaced by other excellent single-image super-resolution networks, and has strong coupling. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0020] Figure 1 This is a schematic diagram of the overall framework structure of the image network system based on unsupervised degradation feature learning and its super-resolution algorithm proposed in the present invention;

[0021] Figure 2 A schematic diagram of the image network system based on unsupervised degradation feature learning and the downsampling module of the super-resolution algorithm thereof based on contrastive learning for extracting degradation features proposed by the present invention;

[0022] Figure 3 This is a network structure diagram of the encoder and generator of the image network system based on unsupervised degradation feature learning and its super-resolution algorithm proposed in the present invention.

[0023] Figure 4Test results of the image network system and its super-resolution algorithm based on unsupervised degradation feature learning proposed by the present invention in two evaluation methods: peak signal-to-noise ratio (PSNR) and structural similarity (SSIM).

[0024] Figure 5 Image experimental results of the image network system and its super-resolution algorithm based on unsupervised degradation feature learning proposed by the present invention. Detailed implementation manners

[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0026] It should be noted that the terms "front", "rear", "left", "right", "up" and "down" used in the following description refer to the directions in the drawings, and the terms "inner" and "outer" respectively refer to the directions towards or away from the geometric center of a specific component.

[0027] The image network system based on unsupervised degradation feature learning proposed by the present invention includes a degradation feature extraction module, a downsampling module and a super-resolution reconstruction module; the degradation feature extraction module is responsible for training an encoder capable of extracting image degradation features through contrastive learning; the downsampling module is responsible for generating a low-resolution image from a high-resolution image through a linear depth convolutional downsampling network, and the trained encoder will ensure that the generated image can carry the same degradation features as the real world; the super-resolution reconstruction module is responsible for inputting the paired training images obtained by the degradation feature extraction module and the downsampling module into the super-resolution reconstruction module, and training the super-resolution network and the downsampling network at the same time to further enhance the generation and super-resolution effects and complete the super-resolution task of real-world images.

[0028] Furthermore, the degradation feature extraction module includes two networks, an encoder and a multi-layer perceptron. The encoder includes a feature encoder and a momentum encoder. The function of the feature encoder is to extract the degradation features in the image, and its input is a picture; the function of the momentum encoder is to process high-resolution image patches, and its input is a picture; the function of the multi-layer perceptron is to convert the degradation representations into positive and negative samples.

[0029] Furthermore, the downsampling module generates a low-resolution image from a high-resolution image through a generator.

[0030] The super-resolution algorithm of the image network system based on unsupervised degradation feature learning specifically includes the following steps:

[0031] (1) The degradation features in the real-world images are extracted by the degradation feature extraction module based on contrastive learning. According to the degradation model of the real-world images, it can be reasonably assumed that the degradation features of the low-resolution real-world images in the same data set are the same, but different from the degradation features of the high-resolution images. Therefore, the degradation feature extraction module classifies the degradation features extracted from the real-world images as positive samples and the degradation features extracted from the high-resolution images as negative samples. First, the real-world images and non-corresponding high-resolution images in the data set are divided into two groups and input into the feature encoder and momentum encoder respectively. The feature encoder and momentum encoder adopt the Moco strategy. The network structure of the feature encoder and the momentum encoder are the same, and the initial parameters are also the same. After the training starts, the parameters of the two will no longer be the same. The output of the feature encoder is a vector representation of a feature. These features are grouped in pairs and then passed through a multi-layer perceptron. One feature in the real-world image will be used as a query sample, the other as its positive sample, and the features of the two high-resolution images will be used as negative samples. The process of contrastive learning is to shorten the distance between the query sample and the positive sample and to increase the distance between it and the negative sample. In other words, we want to make the positive samples as similar as possible and the positive and negative samples as dissimilar as possible. InfoNCE is used here to measure similarity. Previous work on contrastive learning has shown that a large dictionary containing rich negative samples is critical for learning good representations. Therefore, a queue is used here to maintain the negative samples obtained in each batch.

[0032] The overall process is as follows, assuming that the batch size is N and the queue size is K (K>N):

[0033] Step S1: First, N real-world images and N high-resolution images are passed through the encoder and dynamic encoder to obtain image features F q and F k , the dimension is N×C;

[0034] Step S2: transform the image feature F q and F k Send it to the multi-layer perceptron to get the query sample q and positive sample k + and negative samples k + ;

[0035] Step S3: Use InfoNCE to measure the similarity between positive and negative samples. During training, the InfoNCE Loss should be minimized, that is,

[0036] Step S4: Perform gradient backpropagation according to Loss to update encoder parameters, and the parameters of the momentum encoder are updated using momentum;

[0037] Step S5: Update the queue, delete the earliest batch, and add a new batch;

[0038] (2) Establish an image degradation model. The image degradation model can usually be defined in the following form:

[0039]

[0040] Among them, y and x represent the low-resolution image and the high-resolution image respectively. represents the convolution operation between the image and the blur kernel,↓ s represents the downsampling scaling factor, and n represents the noise.

[0041] Starting from the image degradation model, because the convolution operation is a linear operation, the use of a nonlinear downsampling network is obviously inconsistent with the model theory, so a linear downsampling network is used here, that is, there is no activation function. The convolution kernels of the first three layers of the network are 7×7, 5×5 and 3×3, and the convolution kernels of the next three layers are 1×1. The last layer is composed of a hidden layer with a step size of 2. The number of layers here depends on the downsampling scaling factor. The entire network is equivalent to the image convolving and downsampling a 13×13 convolution kernel, which can be regarded as a 13×13 single-layer linear network, which meets the definition of the model. Since single-layer networks are difficult to optimize and converge slowly, a multi-layer network is used for training here. For the specific network structure, see Figure 2 . During the training process, the weighted loss of color loss, adversarial loss and feature matching loss is used as the total loss function. The feature matching loss function is used to ensure that the features of the generated image obtained by the encoder are similar to the features of the real-world image obtained by the encoder. The role of color loss and adversarial loss is to ensure that the basic structural information of the image after downsampling will not change too much;

[0042] (3) Based on the paired image data sets generated by the degradation feature extraction module and the downsampling module, a super-resolution reconstruction module is used to super-resolve the generated low-resolution images, and the corresponding high-resolution images are used for training. The super-resolution reconstruction module is a deep convolutional neural network that uses a residual network to reduce the problem of network overfitting and gradient disappearance;

[0043] Each residual block first consists of two convolutional layers with a kernel size of 3×3 and 64 channels, followed by a batch normalization layer and ReLU as the activation function, and then two Pixelshuffler layers are used to enlarge the size of the features, and finally a 3×3 convolution outputs a 3-channel image. 1 and l per As the loss function, l 1 The loss is defined as:

[0044]

[0045] Among them, s is the scaling factor, and W and H represent the width and height of the scaled image. represents the pixel value of the image, and F is the super-resolution reconstruction network. This loss function is the most common super-resolution loss function, and the distance between pixels is measured by the L 1 norm. If only a single l 1 loss is used, the super-resolved image often lacks high-frequency parts, and the details and textures of the image cannot be well displayed. Therefore, another perceptual loss, l per loss is defined as:

[0046]

[0047] Among them, φ i,j represents the feature map of the j-th convolutional layer before the i-th maximum pooling layer of the VGG network. The perceptual loss can compare the features of the super-resolved high-resolution image with the features of the label image, making the content and global structure between the images close. The super-resolved image has more edge, color, and detail information, and the visual effect is good.

[0048] Figure 1 is the schematic diagram of the overall architecture of the present invention. The super-resolution algorithm of the image network system based on unsupervised degradation feature learning of the present invention is completed by three modules in total. First, an encoder capable of extracting image degradation features is trained by contrastive learning. Then, a high-resolution image is used to generate a low-resolution image through a linear depth convolutional downsampling network. At the same time, the trained encoder will ensure that the generated image has the same degradation features as the real world. Finally, the paired training images obtained from the first two modules are input into the super-resolution reconstruction module, and the super-resolution network and the downsampling network are trained simultaneously to further enhance the generation and super-resolution effects, thereby completing the super-resolution task of real-world images.

[0049] In this process, the algorithm of the present invention focuses on two tasks: generating real low-resolution images and super-resolving real images. As Figure 1As shown in the figure, the process marked by the first row of the training stage is the process of contrastive learning. It is necessary to first send the two groups of pictures to the encoder to obtain their respective features, and then pass them through a multi-layer perceptron to obtain positive and negative samples. The goal of contrastive learning is to shorten the distance between positive samples and increase the distance between positive and negative samples. The blue arrow in the second row of the training stage is the process of image downsampling. A high-resolution image is input and a low-resolution image is output. After the output image and the real-world image are passed through the trained encoder to obtain their respective features, feature matching loss is used to make the two as similar as possible. The third row of the training stage is the process of super-resolution reconstruction. A low-resolution image generated by the downsampling network is input, and a high-resolution image is obtained after passing through the super-resolution network. The high-resolution image in the data set is used as a label. The super-resolution network is trained to learn the mapping relationship between low-resolution images and high-resolution images, and a real-world image super-resolution device can be obtained.

[0050] Figure 2 Shown is a schematic diagram of degradation feature extraction through contrastive learning.

[0051] Figure 3 The network structure diagram of the encoder and generator is shown; each rectangle in the figure represents a layer in the network, Conv in the rectangle represents the convolution layer, BN represents the batch normalization layer, and LeakyReLU is the activation function. These three parts constitute a hidden block in the encoder. k3n64s1 indicates that the convolution kernel size of the convolution layer is 3, the number of channels is 64, and the step size is 1. Figure 3 (a) You can see that the encoder ends with two fully connected layers of multilayer perceptrons. The generator is a linear convolutional network, so the figure only contains convolutional layers without other activation functions. The first 6 layers are fixed, and the number of convolutional layers with a stride of 2 depends on the size of the scaling factor.

[0052] Figure 4 The test results of the algorithm of the present invention on two evaluation methods, peak signal-to-noise ratio (PSNR) and structural similarity (SSIM), are shown. The peak signal-to-noise ratio is based on the error between corresponding pixels, that is, an error-sensitive image quality evaluation. The structural similarity is a full-reference image quality evaluation index, which measures image similarity from three aspects: brightness, contrast, and structure. As can be seen from the table above, the algorithm is higher than the existing super-resolution algorithms in both indicators, indicating that this invention can better complete the super-resolution task of real-world images.

[0053] Figure 5The experimental results of the algorithm of the present invention are shown. In the validation set of NTIRE2020 Track1, the present invention successfully performs super-resolution on real-world images and compares them with other existing algorithms. It can be seen that the super-resolution images obtained by our algorithm are relatively clear, can be close to the true high-resolution images in terms of details, have distinct edges, few blurred artifacts, and have good visual effects.

[0054] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, material or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, material or device.

[0055] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An image network system based on unsupervised degradation feature learning, characterized in that, it includes a degradation feature extraction module, a downsampling module and a super-resolution reconstruction module; the degradation feature extraction module is responsible for training an encoder that can extract image degradation features through contrastive learning; the downsampling module is responsible for generating a low-resolution image from a high-resolution image through a linear depth convolutional downsampling network. At the same time, the trained encoder will ensure that the generated image can carry the same degradation features as the real world; the super-resolution reconstruction module is responsible for inputting the paired training images obtained by the degradation feature extraction module and the downsampling module into the super-resolution reconstruction module, and training the super-resolution network and the downsampling network at the same time to further enhance the generation and super-resolution effects and complete the super-resolution task of real-world images; the degradation feature extraction module includes two networks, an encoder and a multi-layer perceptron. The encoder includes a feature encoder and a momentum encoder. The role of the feature encoder is to extract the degradation features in the image, and its input is a picture; the role of the momentum encoder is to process high-resolution image patches, and its input is a picture; the role of the multi-layer perceptron is to convert the degradation representation into positive and negative samples; the super-resolution reconstruction module is a deep convolutional neural network, and a residual network is used to reduce the problems of overfitting and gradient disappearance in the network; Each residual block is first composed of two convolutional layers with a kernel size of 3×3 and 64 channels, followed by a batch normalization layer and ReLU as the activation function. Then, two Pixelshuffler layers are used to enlarge the size of the features. Finally, a 3×3 convolution outputs an image with 3 channels. During training, l 1 and l per are used as loss functions. l 1 The loss is defined as: where s is the scaling factor, and W and H represent the width and height of the scaled image, represents the pixel values of the image, F is the super-resolution reconstruction network, and another perceptual loss, l per loss is defined as: Among them, φ i,j represents the feature map of the j-th convolutional layer before the i-th max pooling layer of the VGG network.

2. The image network system based on unsupervised degradation feature learning according to claim 1, characterized in that, the downsampling module generates a low-resolution image from a high-resolution image through a generator.

3. The super-resolution algorithm of the image network system based on unsupervised degradation feature learning according to any one of claims 1 to 2, characterized in that, it specifically includes the following steps: (1) Extract the degradation features in the real-world image through the degradation feature extraction module based on contrastive learning. According to the degradation model of the real-world image, it can be reasonably assumed that the degradation features of the low-resolution real-world images in the same dataset are the same, while different from the degradation features of the high-resolution images; Therefore, the degradation feature extraction module classifies the degradation features extracted from the real-world image as positive samples, and the degradation features extracted from the high-resolution image as negative samples; first, the real-world images and non-corresponding high-resolution images in the data set are divided into two groups and input into the feature encoder and momentum encoder respectively. The feature encoder and momentum encoder adopt the Moco strategy. The network structure of the feature encoder and the momentum encoder are the same, and the initial parameters are also the same. After the training starts, the parameters of the two will no longer be the same; the output of the feature encoder is a vector representation of a feature, and these features are grouped in pairs and then passed through a multi-layer perceptron, where one feature in the real-world image will be used as a query sample, the other as its positive sample, and the features of the two high-resolution images will be used as negative samples; the process of contrastive learning is to shorten the distance between the query sample and the positive sample and to increase the distance between it and the negative sample. In other words, it is to make the positive samples similar and the positive and negative samples dissimilar. Here, InfoNCE is used to measure similarity; previous work on contrastive learning shows that a large dictionary containing rich negative samples is critical to learning good representations, so a queue is used here to maintain the negative samples obtained in each batch; (2) Establish an image degradation model. Starting from the image degradation model, a linear downsampling network is used to convolve and downsample the convolution kernel of the image. A high-resolution image is input and a low-resolution image is output. A multi-layer network is used for training. During the training process, a weighted combination of color loss, adversarial loss, and feature matching loss is used as the total loss function. The feature matching loss is used to ensure that the features of the generated image obtained by the encoder are similar to the features of the real-world image obtained by the encoder. The role of color loss and adversarial loss is to ensure that the basic structural information of the image after downsampling does not change significantly. (3) Based on the paired image data sets generated by the degradation feature extraction module and the downsampling module, a super-resolution reconstruction module is used to super-resolve the generated low-resolution images, and the corresponding high-resolution images are used for training. The super-resolution reconstruction module is a deep convolutional neural network that uses a residual network to reduce the problem of network overfitting and gradient disappearance.

Citation Information

Patent Citations

  • Image degradation processing method and device, storage medium and electronic equipment

    CN112419151A

  • HSI super-resolution reconstruction method based on unsupervised learning and related equipment

    CN113269677A