Small sample image super-resolution reconstruction method combining self-supervised learning and generative adversarial network

By combining self-supervised learning and generating adversarial network methods, image super-resolution reconstruction is used to use label-free data to solve the image reconstruction problem under small sample data sets, and high-quality and efficient image super-resolution reconstruction is achieved.

CN120278885APending Publication Date: 2025-07-08NORTHEASTERN UNIV CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510437169.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art cannot effectively utilize labelless data in small sample data sets, resulting in insufficient generalization ability of image super-resolution reconstruction, and relying on a large amount of labeled data leads to poor performance in data scarce environments.

Method used

Combining self-supervised learning and generative adversarial networks, through dual-path self-supervised pre-training, multi-scale deformable upsampling modules and self-consistent adversarial training systems, the model's performance under small sample conditions is improved by using label-free data.

Benefits of technology

Significantly improve image reconstruction quality and generalization capabilities under small sample conditions, reduce dependence on labeled data, and achieve efficient and accurate image super-resolution reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278885A_ABST
    Figure CN120278885A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and particularly relates to a small sample image super-resolution reconstruction method combining self-supervised learning and a generative adversarial network. By introducing a self-supervised learning mechanism, unlabeled data can be effectively utilized, so that the expression of the model under a small sample condition is improved, and the demand on labeled data is reduced. On the basis, through the innovative generator and discriminator architecture design, the image reconstruction quality and the detail recovery capability are improved. Through the technical innovation, the method not only breaks through the limitation of a traditional method in a small sample environment, but also realizes more efficient and more accurate image super-resolution reconstruction, and provides a feasible solution for a data scarce scene in practical application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a small-sample image super-resolution reconstruction method combining self-supervised learning and generative adversarial networks. Background Art

[0002] In recent years, significant progress has been made in image super-resolution reconstruction methods based on deep learning. For example, SRCNN (Super-Resolution Convolutional Neural Network) directly maps from low-resolution images to high-resolution images through a convolutional neural network, significantly improving the reconstruction quality. Subsequently, EDSR (Enhanced Deep Super-Resolution Network) further improves the performance by removing redundant batch normalization layers and increasing the network depth. However, these methods usually require a large number of paired low-resolution and high-resolution images for training and are difficult to apply in small-sample scenarios. Generative adversarial networks (GANs) have been widely applied to the image super-resolution reconstruction task. For example, SRGAN (Super-Resolution Generative Adversarial Network) can generate high-resolution images with more realistic visual effects by introducing adversarial loss and perceptual loss. However, SRGAN still relies on a large amount of labeled data and is prone to overfitting problems in small-sample scenarios, resulting in the loss of details or distortion of the reconstructed images.

[0003] The following are the related research and applications of two existing technologies:

[0004] CN106600533B proposes a single-image super-resolution reconstruction method. This method includes preprocessing the original image to obtain the corresponding low-resolution image; then dividing the low-resolution image into multiple groups and performing adaptive dictionary learning on each group to calculate the adaptive learning dictionary for each group. Based on the adaptive learning dictionaries of each group, calculate its sparse coding, and then restore and reconstruct the image patches of each group. Finally, calculate the average of the restored images of all groups to obtain the complete high-resolution image. This method has high requirements for hardware resources during the calculation process, and when dealing with super-resolution reconstruction, its adaptability to small-sample datasets may be poor, resulting in relatively limited performance in low-sample cases. The proposed method relies on adaptive dictionary learning and sparse coding techniques, which can improve the image quality to a certain extent when the computational resources requirements are high, but it performs poorly when dealing with small-sample datasets. This method overly relies on the sparse representation of the dataset, has weak generalization ability in small-sample scenarios, and cannot effectively utilize unlabeled data, resulting in limited reconstruction quality. Therefore, this method cannot fully address the challenges in data-scarce and low-sample situations and cannot provide a more efficient solution for small-sample super-resolution reconstruction.

[0005] CN114782247A proposes a super-resolution model based on PUGAN-Charbon (SRPUGAN-Charbon). This model includes a generator network for synthesizing super-resolution (SR) images and a discriminator network that is trained to distinguish SR images from real high-resolution (HR) images. However, this method relies on a large amount of labeled data for training, which poses high requirements for the diversity and scale of data during the training process. Therefore, its application on small-sample datasets may face certain challenges. The proposed SRPUGAN Charbon performs super-resolution image reconstruction through a generative adversarial network (GAN). Although it has achieved good visual effects, this method still relies on a large amount of labeled data for training. Due to the high cost of obtaining labeled data and its poor performance on small-sample datasets, SRPUGAN Charbon is prone to overfitting problems in low-sample scenarios, resulting in the loss of details or distortion of the reconstructed images, and the training process poses high requirements for the diversity and scale of data. Therefore, the applicability of this method is limited under small-sample conditions and it is difficult to be effectively promoted.

[0006] Generally speaking, the main problems faced by existing methods are: on the one hand, they cannot guarantee sufficient model generalization ability in small-sample scenarios; on the other hand, they overly rely on a large amount of labeled data for training and fail to fully utilize unlabeled data, resulting in poor performance in data-scarce environments. Especially in the super-resolution reconstruction task, the existing technologies have not solved the problem of how to achieve high-quality image reconstruction in a low-sample data environment.

[0007] By introducing a self-supervised learning mechanism, the present invention can effectively exploit the potential of unlabeled data under small-sample conditions, significantly improving the quality and generalization ability of image reconstruction. In this way, the present invention can overcome the defects of the prior art, especially in low-sample data sets, breaking through the limitation of the existing methods that rely on a large amount of labeled data, and providing a new solution for image super-resolution reconstruction. Summary of the Invention

[0008] The object of the present invention is to propose a small-sample image super-resolution reconstruction method combining self-supervised learning and generative adversarial networks, aiming to overcome the problems of insufficient generalization ability and over-reliance on a large amount of labeled data in the prior art under small-sample data sets. Specifically, the present invention is applied to technical fields such as image processing, computer vision, and deep learning. Especially in a low-sample data environment, self-supervised learning and generative adversarial networks are used to improve the image resolution and detail restoration ability.

[0009] By introducing a self-supervised learning mechanism, the present invention can effectively utilize unlabeled data, thereby improving the performance of the model under small-sample conditions and reducing the need for labeled data. On this basis, the present invention also improves the quality of image reconstruction and the detail restoration ability through innovative generator and discriminator architecture designs. Through these technological innovations, the present invention not only breaks through the limitations of traditional methods in small-sample environments but also realizes more efficient and accurate image super-resolution reconstruction, providing a practical solution for data-scarce scenarios in practical applications.

[0010] The technical solution of the present invention is as follows: A small-sample image super-resolution reconstruction method combining self-supervised learning and generative adversarial networks, which selects multiple pairs of high-resolution images and low-resolution images; the high-resolution images are used for dual-path self-supervised pre-training of the encoder of training path A and the encoder of path B. The encoder of path A is used to extract the low-level features of the image, and the encoder of path B is used to obtain the semantic information of the image. The two jointly provide high-quality feature representations for the subsequent generation network; the low-resolution images are generated into new high-resolution images by the generator network G-Net, and the generator network G-Net is continuously optimized and updated during the training process;

[0011] The discriminator network D-Net is used to evaluate the quality of the new high-resolution images generated by the generator network G-Net and optimize the generation ability of the generator network G-Net through adversarial training; among them, at the initial stage of training, the generator network G-Net is fixed, and only the parameters of the discriminator network D-Net itself are trained; in the subsequent stage of training, the generator network G-Net and the discriminator network D-Net are alternately trained to jointly optimize the two.

[0012] The initial training stage is the first 50 training epochs: fix the generator network and train the discriminator network; use the Adam optimizer with a learning rate of 1e-4, β1 = 0.5, and β2 = 0.999; the training loss function uses an improved Hinge Loss function and adds a gradient penalty term with λ = 10; the subsequent training stage is from the 51st to the 200th training epoch, a total of 150 epochs: alternately train the generator network and the discriminator network, and continuously optimize their performance through adversarial training;

[0013] L Hinge = E[max(0, 1 - D(I real ))] + E[max(0, 1 - D(I fake ))]

[0014] where E represents the expected value over all training samples, I real represents the real high-resolution image, and I fake represents the high-resolution image generated by the generator network G-Net;

[0015] The subsequent training stage is the subsequent 150 training epochs after the first 50 training epochs: alternately train the generator network and the discriminator network; the learning rate of the generator network is set to 5e-5, and the learning rate of the discriminator network is 2e-5; the combination of training loss functions:

[0016] L total = 0.7L adv + 0.2L perceptual + 0.1L pixel

[0017] where L total is the total loss function, L adv is the adversarial loss, which measures the adversarial difference between the generated image and the real image, and L perceptual is the perceptual loss, which extracts features based on a pre-trained deep network and measures the difference between the generated image and the real image at the perceptual level, and L pixel is the pixel-level loss, which directly measures the difference between the generated image and the real image at the pixel level and usually uses the L2 norm for calculation.

[0018] A dynamic learning mechanism is added in the subsequent training stage, specifically:

[0019] The input image size gradually increases; it is set as follows: from the 0th to the 30th training epoch, the input image size is 64×64; from the 31st to the 60th training epoch, the input image size is 128×128; from the 61st to the 200th training epoch, the size of the input low-resolution image is 256×256;

[0020] Noise injection: Gaussian noise is added to the input data of the discriminator network, and the noise intensity σ linearly decays from 0.1 to 0.01;

[0021] The confidence of the generated images is evaluated by calculating the variance of the output of the discriminator network; according to the results of the confidence evaluation, the generated images with a confidence greater than 0.85 are retained; the filtered high-quality images are added to the training set as new incremental samples for further training of the generator network and the discriminator network.

[0022] The low-resolution images are generated by bicubic downsampling;

[0023] Random geometric transformations and color jitters are applied to the high-resolution images to generate a large number of high-resolution images.

[0024] The dual-path self-supervised pre-training includes path A self-supervised pre-training for pixel-level reconstruction and path B self-supervised pre-training for semantic-level prediction;

[0025] Path A is as follows:

[0026] Input processing: Randomly erase 32×32 pixel blocks in the high-resolution images, and the erasure rate is dynamically adjusted;

[0027] The encoder uses a 12-layer Vision Transformer encoder;

[0028] The L2 loss function is used to calculate the pixel-level error between the predicted image and the real image, where, is the predicted image, is the real image, and N is the number of pixels in the image:

[0029]

[0030] Path B is as follows:

[0031] Input processing: The high-resolution image is divided into 4×4 blocks, randomly shuffled and then input into the Transformer encoder that shares the parameters of path A, and a classification layer is connected at the output end; the cross-entropy loss is used to calculate the difference between the predicted block order and the actual order, where, y i is the real block position label, and p i is the predicted probability:

[0032]

[0033] Path A and Path B share the Transformer encoder parameters at the 4th and 8th layers. Both adopt the same encoder architecture in the network structure design, so the corresponding layers have the same hierarchical structure and parameter dimensions. The sharing condition is that the feature similarity threshold θ = 0.75, that is, when the similarity of the feature maps output by Path A and Path B at this layer is greater than the threshold, the parameter sharing mechanism is enabled to improve the consistency of feature representation and the network training efficiency.

[0034] The dynamic adjustment benchmark of the erasure rate is 40%, and the absolute value of the random fluctuation is 15%. The number of heads in the multi-head attention mechanism of each layer of the VisionTransformer encoder is 8, and the dimension of each head is 64. The classification layer includes a fully connected layer and Softmax.

[0035] The generator network G-Net adopts an improved U-Net structure, with the input being a low-resolution image I LR ∈R H×W×3 and the output being the predicted high-resolution image I HR ∈R 4H×4W×3 ; The generator network G-Net includes a shallow feature extraction module, a residual dense downsampling module RRDB, a multi-scale deformable upsampling module MSFF, a detail enhancement unit, and an output layer.

[0036] The shallow feature extraction module includes three layers of convolution operations, and the specific structure and parameter settings of each layer of convolution are as follows:

[0037] The first layer of convolution: uses a 5×5 convolution kernel, with a stride of 1, and the number of output channels is 64; the activation function after the convolution operation is LeakyReLU, and the leakage coefficient α = 0.2;

[0038] The second layer of convolution: uses a 3×3 convolution kernel, with a stride of 1, and the number of output channels is 64; the activation function after the convolution operation is LeakyReLU, and the leakage coefficient α = 0.2;

[0039] The third layer of convolution: uses a 3×3 convolution kernel, with a stride of 1, and the number of output channels is 64; the activation function after the convolution operation is LeakyReLU, and the leakage coefficient α = 0.2;

[0040] Each layer of convolution operation performs a non-linear transformation through the LeakyReLU activation function, and the number of output channels of each layer is 64;

[0041] Output: Shallow feature map F shallow ∈R H×W×64 ;

[0042] The residual dense downsampling module RRDB includes 4 levels of residual dense blocks. Each level of residual dense block sequentially includes 3 dense connection layers, a residual connection layer, and a downsampling layer. After processing the input shallow feature map through the residual dense downsampling module, a deep feature map is output.

[0043] The multi-scale deformable upsampling module includes a dynamic convolution kernel generation network, a deformable upsampling layer, and a cross-scale skip connection. The shallow feature map and the deep feature map are input into the multi-scale deformable upsampling module. Features are extracted from intermediate layers of different scales, and the features are fused into the final output feature map through skip connections.

[0044] The dynamic convolution kernel generation network KernelNet is used to generate dynamic convolution kernels. By processing the deep features, convolution kernels of different sizes and shapes are generated. Among them:

[0045] The feature dimension of the input layer is 512, and the output dimension is 18, corresponding to the x / y offsets of a 3×3 convolution kernel. Each convolution kernel contains 9 offset values, which are used to dynamically adjust the position and shape of the convolution kernel.

[0046] After being processed, the convolution kernels generated by the dynamic convolution kernel generation network are used to perform deformable convolution operations on the feature map, so as to perform a more refined upsampling on the image.

[0047] The deformable upsampling layer uses the convolution kernels generated by the dynamic convolution kernel generation network to perform upsampling operations on the input feature map. The deformable upsampling layer dynamically adjusts the convolution operation according to the offsets of the convolution kernels.

[0048] The cross-scale skip connection combines image information of different scales through cross-scale feature fusion connections to enhance the structural consistency and detail restoration ability of the generated image.

[0049] The specific operation of the cross-scale skip connection is as follows:

[0050] Feature extraction and downsampling: In the generator network, the input image first undergoes feature extraction through the shallow feature extraction module, and then is downsampled through the residual dense downsampling module to obtain multi-scale deep feature maps.

[0051] Cross-scale feature fusion: Let F1, F2, and F3 be feature maps of three different scales. The low-resolution feature maps F3 and F2 are upsampled to the same spatial size as F1 using upsampling operations. Then, the features of these three scales are fused element-wise with weights.

[0052] Final upsampling and output: The fused feature F fuse is sent to the detail enhancement unit for texture restoration and a super-resolution image is generated through the final output layer.

[0053] The detailed enhancement unit includes a local texture branch and a global structure branch, and the outputs of the two branches are dynamically fused:

[0054] F enhanced = αF local + (1 - α)F global

[0055] Among them, F enhanced represents the finally enhanced feature representation, which is used to generate a high-quality super-resolution image; F local represents the local texture feature, which is extracted through a 3×3 standard convolutional layer and is used to enhance local details; F global represents the global structure feature, which is obtained through dilated convolution and is used to maintain the overall structural consistency; α is the adaptive fusion weight, which is used to control the proportion of the local texture feature and the global structure feature.

[0056] The discriminator network D-Net adopts a multi-scale PatchGAN structure, and the specific design is as follows:

[0057] Three discriminators are in parallel, with the same structure, and they process image patches of different scales respectively;

[0058] Spectral normalization is performed on the convolutional weights of each layer, and the calculation formula is as follows, where W is the convolutional kernel and W normal is the normalized convolutional kernel:

[0059]

[0060] The perceptual loss is calculated using the features of the conv5-4 layer of the pre-trained VGG19 network, and the calculation formula is as follows:

[0061]

[0062] Among them, Φ j is the feature extraction function of the j-th layer of the pre-trained VGG19 network, is the predicted image, is the real image, and N is the number of pixels.

[0063] The beneficial effects of the present invention: The present invention proposes a small-sample image super-resolution reconstruction method combining self-supervised learning and generative adversarial networks. The core innovation lies in solving the problem that traditional methods are difficult to achieve high-quality image reconstruction under small-sample conditions through a dual-path self-supervised pre-training mechanism, a multi-scale deformable upsampling module, and a self-consistent adversarial training system.

[0064] 1. Dual-path self-supervised pre-training mechanism: The present invention first creates a dual-task collaborative pre-training framework of pixel-level reconstruction and semantic-level prediction. Through random block erasing reconstruction (the erasing rate is dynamically adjusted at 40% ± 15%) in Path A and adversarial jigsaw recombination (the block perturbation intensity is adaptively adjusted) in Path B, it achieves the pre-training effect of 1000 pairs of samples in traditional supervised learning with only 100 pairs of training samples (verified by the ImageNet subset); the cross-path parameter sharing strategy (shared layer dynamic selection algorithm: based on the feature similarity threshold θ = 0.75) further improves the pre-training efficiency and significantly reduces the model's dependence on large-scale labeled data.

[0065] 2. Multi-scale deformable upsampling module: The present invention designs a dynamic convolution kernel generation network (Kernel-Net), generates 3×3 deformable convolution kernel parameters through the input feature map, and combines the spatial attention guidance mechanism (coordinate attention map adjusts the deformation offset) to achieve more accurate feature alignment and detail reconstruction. In the ×4 super-resolution task, compared with traditional bicubic interpolation, this module reduces 37% of the artifact phenomenon (quantitative analysis based on the LIVE1 dataset) and significantly improves the quality of the reconstructed image.

[0066] 3. Self-consistent adversarial training system: The present invention proposes a pseudo-label self-enhancement strategy (automatically screens generated samples based on the discriminator confidence threshold > 0.85) and curriculum noise injection (the noise intensity σ = 0.1 → 0.01 linearly decays during training), combined with multi-granularity discriminators (a 3-level PatchGAN architecture integrates VGG19 deep semantic constraints), effectively solving the mode collapse problem in small-sample training. Under the training condition of 200 samples, the FID index of the generated image reaches 18.7 (better than 23.5 of ESRGAN), significantly improving the training stability and the authenticity of the generated image.

[0067] In summary, through innovative self-supervised pre-training strategies, dynamic feature fusion mechanisms, and adaptive training techniques, the present invention has achieved a performance breakthrough in the small-sample image super-resolution reconstruction task, with significant technical advantages and application values. Brief Description of the Drawings

[0068] Figure 1 is the logic flow chart of this method;

[0069] Figure 2 is the structural diagram of the multi-scale deformable upsampling module;

[0070] Figure 3 is the structural diagram of the generative adversarial network model proposed by this method. Detailed Embodiments

[0071] The present invention proposes a few-shot image super-resolution reconstruction method combining self-supervised learning and generative adversarial networks. This method effectively utilizes unlabeled data through a self-supervised learning mechanism, and at the same time combines generative adversarial networks to perform super-resolution reconstruction on images, significantly improving the quality of the reconstructed images. Select multiple pairs of high-resolution images and low-resolution images; the high-resolution images are used for dual-path self-supervised pre-training, training the encoder of path A and the encoder of path B. The encoder of path A is used to extract the low-level features of the image, and the encoder of path B is used to obtain the semantic information of the image. Both jointly provide high-quality feature representations for the subsequent generative network; the low-resolution images are used to generate new high-resolution images via the generator network G-Net, and the generator network G-Net is continuously optimized and updated during the training process;

[0072] The discriminator network D-Net is used to evaluate the quality of the new high-resolution images generated by the generator network G-Net, and optimize the generation ability of the generator network G-Net through adversarial training; among them, at the initial stage of training, the generator network G-Net is fixed, and only the parameters of the discriminator network D-Net itself are trained; in the subsequent stages of training, the generator network G-Net and the discriminator network D-Net are alternately trained to jointly optimize the two.

[0073] The following is a detailed technical implementation step of an embodiment of the present invention, and the specific implementation steps are as follows:

[0074] I. Data preprocessing and self-supervised pre-training

[0075] 1. Construction of a few-shot dataset

[0076] Data source: Select 100 - 300 pairs of high-resolution images (HR) and low-resolution images (LR), and LR is generated by bicubic downsampling (scaling factor ×4).

[0077] Data augmentation: Apply random geometric transformations (rotation ±15°, translation ±10%) and color jitter (brightness ±0.1, contrast ±0.1) to the HR images to generate diverse samples.

[0078] 2. Dual-path self-supervised pre-training

[0079] Path A (pixel-level reconstruction):

[0080] (1) Input processing: Randomly erase a 32×32 pixel block in the HR image, and the erasure rate is dynamically adjusted (benchmark 40%, random fluctuation ±15%).

[0081] (2) Model structure: Adopt a 12-layer Vision Transformer encoder, with a multi-head attention mechanism in each layer (number of heads = 8, dimension of each head = 64).

[0082] (3) Loss function: The L2 loss function is used to calculate the pixel-level error between the reconstructed image and the original image, where, is the predicted image, is the real image, and N is the number of pixels in the image:

[0083]

[0084] Path B (semantic-level prediction):

[0085] (1) Input processing: The HR image is divided into 4×4 blocks (16 blocks in total), randomly shuffled and then input into the model.

[0086] (2) Model structure: Share the Transformer encoder of Path A, and connect a classification layer (fully connected layer + Softmax) at the output end.

[0087] (3) Loss function: Use the cross-entropy loss to calculate the difference between the predicted block order and the actual order, where, y i is the real block position label, and p i is the predicted probability:

[0088]

[0089] Parameter sharing mechanism: The two paths share the Transformer encoder parameters at the 4th and 8th layers, and the sharing condition is that the feature similarity threshold θ = 0.75;

[0090] II. Generator network architecture (G-Net)

[0091] The generator (G-Net) adopts an improved U-Net structure, with the input being the low-resolution image I LR ∈R H×W×3 , and the output being the high-resolution image I HR ∈R 4H×4W×3 . The overall structure of the generator includes the following core modules:

[0092] 1. Shallow feature extraction module

[0093] Input: Low-resolution image I LR such as (256×256×3)

[0094] Structure: 3-layer convolution sequence:

[0095] Conv1: kernel = 5×5, stride = 1, channels = 64 → LeakyReLU(α = 0.2),

[0096] Conv2:kernel=3×3, stride=1, channels=64→LeakyReLU(α=0.2),

[0097] Conv3:kernel=3×3, stride=1, channels=64→LeakyReLU(α=0.2),

[0098] Output: shallow feature map F shallow ∈R H×W×64 ;

[0099] 2. Residual Dense Downsampling Module (RRDB)

[0100] Structure: 4-level residual dense block (RRDB), each level contains the following operations:

[0101] Level 1 (RRDB1):

[0102] Input: shallow feature map F shallow ∈R H×W×64 ;

[0103] operate:

[0104] 1.Dense Block: 3 densely connected layers (64 output channels per layer);

[0105] 2. Residual connection;

[0106] 3. Downsampling: MaxPooling (kernel = 2 × 2, stride = 2);

[0107] Output: Feature map size

[0108] Subsequent levels (RRDB2-RRDB4):

[0109] The number of channels doubles step by step (128→256→512), and the structure is the same as RRDB1;

[0110] The final output is deep features

[0111] 3. Multi-scale Deformable Upsampling Module (MSFF)

[0112] Step 1: Dynamic deformable convolution kernel generation;

[0113] Input: deep features F deep ;

[0114] Operation: KernelNet is a 3-layer fully connected network (input 512 dimensions → output 18 dimensions, corresponding to the x / y offset of the 3×3 convolution kernel) △ = KernelNet (Fdeep );

[0115] Step 2: Deformable Upsampling;

[0116] Output: The size of the upsampled feature map is

[0117] Step 3: Cross-scale Skip Connection;

[0118] 4. Detail Enhancement Unit

[0119] Structure:

[0120] Branch 1 (Local Texture):

[0121] 3×3 Convolution → Number of Channels 64 → LeakyReLU (α = 0.2);

[0122] Branch 2 (Global Structure):

[0123] Atrous Convolution (kernel = 3×3, dilation = 2) → Number of Channels 64 → LeakyReLU (α = 0.2);

[0124] Dynamic Fusion:

[0125] F enhanced = αF local + (1 - α)F global

[0126] 5. Output Layer

[0127] Input: The size of the last upsampled feature is 4H×4W×64;

[0128] Operation:

[0129] Conv4: kernel = 3×3, stride = 1, channels = 64 → LeakyReLU (α = 0.2);

[0130] Conv5: kernel = 3×3, stride = 1, channels = 3 → Tanh Activation;

[0131] Output: High-resolution image I HR ∈R 4H×4W×3 ;

[0132] III. Discriminator Network Architecture Construction (D-Net)

[0133] The discriminator adopts a multi-scale PatchGAN structure, and the specific design is as follows:

[0134] 1. Basic Architecture: Three discriminators (D1 - D3) in parallel, respectively processing image patches of different scales:

[0135] D1: 70×70 pixel block, 5-layer convolution (number of channels: 64→128→256→512→1024);

[0136] D2: 140×140 pixel block, with the same structure as D1;

[0137] D3: 280×280 pixel block, with the same structure as D1;

[0138] Convolution parameters: kernel = 4×4, stride = 2, activation function is LeakyReLU (α = 0.2).

[0139] 2. Spectral normalization: Perform spectral normalization on the convolutional weights of each layer to improve the stability of training. The calculation formula is as follows, where W is the convolutional kernel and W normal is the normalized convolutional kernel:

[0140]

[0141] 3. Semantic consistency constraint: Calculate the perceptual loss using the features of the conv5-4 layer of the pre-trained VGG19 network. The calculation formula is as follows:

[0142]

[0143] where, Φ j is the feature extraction function of the j-th layer of VGG19, is the predicted image, is the real image, and N is the number of pixels.

[0144] IV. Joint training

[0145] 1. Adopt a two-stage training mechanism, specifically:

[0146] Stage 1 (the first 50 epochs): Fix the generator and train the discriminator. Use the Adam optimizer with a learning rate of 1e-4, β1 = 0.5, and β2 = 0.999. The loss function uses the improved Hinge Loss and adds a gradient penalty term (λ = 10):

[0147] L Hinge = E[max(0,1 - D(I real ))] + E[max(0,1 - D(I fake ))]

[0148] Stage 2 (the subsequent 150 epochs): Alternately train the generator and the discriminator. The learning rate of the generator is set to 5e-5, and the learning rate of the discriminator is 2e-5. Loss function combination:

[0149] Ltotal = 0.7L adv + 0.2L perceptual + 0.1L pixel

[0150] 2. Add a dynamic learning mechanism:

[0151] The input image size gradually increases: 64×64 (0 - 30 epochs) → 128×128 (31 - 60 epochs) → 256×256 (61 - 200 epochs).

[0152] Noise injection: Gaussian noise is added to the discriminator input, and the noise intensity σ linearly decays from 0.1 to 0.01.

[0153] V. Adaptive pseudo-label generation

[0154] 1. Confidence evaluation: The confidence of the generated images is evaluated by calculating the variance of the discriminator output. Specifically, the discriminator outputs a credibility measure based on the comparison between the generated images and the real images. By calculating the variance of this output, the quality and authenticity of the generated images can be judged. High variance indicates that the model is uncertain, while low variance indicates that the model is more confident.

[0155] 2. Sample screening: According to the results of the confidence evaluation, retain those generated images with a confidence greater than 0.85. These images are considered to be of higher quality and can effectively increase the diversity and effectiveness of the training set. The screened high-quality images will be added to the training set as new incremental samples for further training of the generator and discriminator.

[0156] VI. Image reconstruction

[0157] After completing the self-supervised pre-training and the training of the generative adversarial network, the generator model has been trained to be able to recover high-resolution images from low-resolution images. The process of image reconstruction is as follows:

[0158] S1: Input low-resolution image

[0159] Input a low-resolution image, which may be generated by methods such as bicubic downsampling. Usually, its resolution is 1 / 4 of the high-resolution image. At this time, the detail information of the image is relatively blurred and needs to be restored.

[0160] S2: The generator performs feature extraction and restoration

[0161] The generator network first extracts image features from the low-resolution image and gradually restores the lost details and textures. The generator processes the low-resolution image through the following steps:

[0162] 1. Feature extraction: The generator extracts the features of the image through multiple layers of convolution, especially the key texture and structure information.

[0163] 2. Upsampling and detail restoration: The generator gradually restores the details of the high-resolution image through multi-level upsampling modules and deformable convolutions. The detail enhancement unit further supplements local details and global structures.

[0164] S3: Output high-resolution image

[0165] After a series of convolutions, upsamplings, and detail enhancements, the generator outputs the reconstructed high-resolution image. This high-resolution image is the same size as the input low-resolution image, but has significant improvements in details, texture, and clarity.

[0166] S4: Final output

[0167] The trained generator can accurately reconstruct the high-resolution image from the low-resolution image, restore the lost details, and improve the quality of the image. The finally output image is visually close to or even exceeds the real high-resolution image.

[0168] The experimental results of the present invention on the standard test sets Set5 and Set14 show that its PSNR index reaches 30.85dB and the SSIM index is 0.882, which is significantly improved compared with the traditional SRGAN method (PSNR 29.45dB, SSIM 0.847). At the same time, the number of model parameters is only 9.3M, which is 44% less than 16.7M of ESRGAN, realizing the lightweight of the model while ensuring the performance. In addition, the adaptive pseudo-label generation method and domain adaptation transfer scheme of the present invention under small sample conditions further expand its application scenarios, and are particularly suitable for fields with scarce labeled data such as medical images and satellite images.

Claims

1. A small-sample image super-resolution reconstruction method combining self-supervised learning and generative adversarial networks, characterized in that Select multiple pairs of high-resolution images and low-resolution images; the high-resolution images are used for dual-path self-supervised pre-training of the encoder of training path A and the encoder of path B. The encoder of path A is used to extract low-level features of the image, and the encoder of path B is used to obtain semantic information of the image. The two jointly provide high-quality feature representations for the subsequent generation network; the low-resolution images are used to generate new high-resolution images via the generator network G-Net, and the generator network G-Net is continuously optimized and updated during the training process; The discriminator network D-Net is used to evaluate the quality of the new high-resolution images generated by the generator network G-Net and optimize the generation ability of the generator network G-Net through adversarial training; among them, at the initial stage of training, the generator network G-Net is fixed, and only the parameters of the discriminator network D-Net are trained; in the subsequent stage of training, the generator network G-Net and the discriminator network D-Net are alternately trained to jointly optimize the two.

2. The few-shot image super-resolution reconstruction method combining self-supervised learning and generative adversarial network according to claim 1, characterized in that The initial training stage is the first 50 training epochs: fix the generator network and train the discriminator network; use the Adam optimizer with a learning rate of 1e-4, β1 = 0.5, and β2 = 0.999; The training loss function uses an improved HingeLoss loss function and adds a gradient penalty term with λ = 10; the subsequent training stage is from the 51st to the 200th training epoch, a total of 150 epochs: alternately train the generator network and the discriminator network, and continuously optimize the performance of both through adversarial training; L Hinge = E[max(0, 1 - D(I real ))] + E[max(0, 1 - D(I fake ))] where E represents the expected value over all training samples, I real represents the true high-resolution image, and I fake represents the high-resolution image generated by the generator network G-Net; The subsequent training stage is the subsequent 150 training epochs after the first 50 training epochs: alternately train the generator network and the discriminator network; the learning rate of the generator network is set to 5e-5, and the learning rate of the discriminator network is 2e-5; the combination of training loss functions: L total = 0.7L adv + 0.2L perceptual + 0.1L pixel Among them, L total Total loss function, L adv is the adversarial loss, which is used to measure the adversarial difference between the generated image and the real image. L perceptual is the perceptual loss. Based on the features extracted by a pre-trained deep network, it measures the difference between the generated image and the real image at the perceptual level. L pixel is the pixel-level loss, which directly measures the difference between the generated image and the real image at the pixel level. Usually, the L2 norm is used for calculation.

3. The few-shot image super-resolution reconstruction method combining self-supervised learning and generative adversarial network according to claim 2, characterized in that, A dynamic learning mechanism is added in the subsequent training stage, specifically: The input image size gradually increases; it is set as follows: from the 0th to the 30th training epoch, the input image size is 64×64; from the 31st to the 60th training epoch, the input image size is 128×128; from the 61st to the 200th training epoch, the size of the input low-resolution image is 256×256; Noise injection: Gaussian noise is added to the input data of the discriminator network, and the noise intensity σ linearly decays from 0.1 to 0.01; The confidence of the generated images is evaluated by calculating the variance of the output of the discriminator network; according to the results of the confidence evaluation, the generated images with a confidence greater than 0.85 are retained; the screened high-quality images are added to the training set as new incremental samples for further training of the generator network and the discriminator network.

4. The few-shot image super-resolution reconstruction method combining self-supervised learning and generative adversarial network according to claim 1, characterized in that The low-resolution images are generated by bicubic downsampling; Random geometric transformations and color jitters are applied to the high-resolution images to generate a large number of high-resolution images.

5. The few-shot image super-resolution reconstruction method combining self-supervised learning and generative adversarial network according to claim 1, characterized in that The dual-path self-supervised pre-training includes self-supervised pre-training of path A for pixel-level reconstruction and self-supervised pre-training of path B for semantic-level prediction; Path A is specifically as follows: Input processing: Randomly erase 32×32 pixel blocks in the high-resolution image, and the erasure rate is dynamically adjusted; The encoder uses a 12-layer Vision Transformer encoder; Calculate the pixel-level error between the predicted image and the ground truth image using the L2 loss function, where, is the predicted image, is the ground truth image, and N is the number of pixels in the image: The specific path of Path B is as follows: Input processing: The high-resolution image is segmented into 4×4 blocks, randomly shuffled, and then input into the Transformer encoder with the parameters of shared path A. A classification layer is connected to the output end. The cross-entropy loss is used to calculate the difference between the predicted block order and the actual order, where y i is the true block position label, and p i is the predicted probability: Path A and Path B share the Transformer encoder parameters at the 4th and 8th layers. They adopt the same encoder architecture in the network structure design, so the corresponding layers have the same hierarchical structure and parameter dimensions. The sharing condition is that the feature similarity threshold θ = 0.75, that is, when the similarity of the feature maps output by Path A and Path B at this layer is greater than the threshold, the parameter sharing mechanism is enabled to improve the consistency of feature representation and the network training efficiency.

6. The few-shot image super-resolution reconstruction method combining self-supervised learning and generative adversarial network according to claim 5, characterized in that The dynamic adjustment benchmark of the erasure rate is 40%, and the absolute value of the random fluctuation is 15%. The number of heads in the multi-head attention mechanism of each layer of the Vision Transformer encoder = 8, and the dimension of each head is 64. The classification layer includes a fully connected layer and Softmax.

7. The few-shot image super-resolution reconstruction method combining self-supervised learning and generative adversarial network according to claim 1, characterized in that The generator network G-Net adopts an improved U-Net structure, with the input being a low-resolution image I LR ∈R H×W×3 , and the output being the predicted high-resolution image I HR ∈R 4H×4W×3 ; the generator network G-Net includes a shallow feature extraction module, a residual dense downsampling module RRDB, a multi-scale deformable upsampling module MSFF, a detail enhancement unit, and an output layer; The shallow feature extraction module includes three layers of convolution operations. The specific structure and parameter settings of each layer of convolution are as follows: The first layer of convolution: uses a 5×5 convolution kernel, the stride is 1, and the number of output channels is 64; the activation function after the convolution operation is LeakyReLU, and the leakage coefficient α = 0.2; The second layer of convolution: uses a 3×3 convolution kernel, the stride is 1, and the number of output channels is 64; the activation function after the convolution operation is LeakyReLU, and the leakage coefficient α = 0.2; The third layer of convolution: uses a 3×3 convolution kernel, the stride is 1, and the number of output channels is 64; the activation function after the convolution operation is LeakyReLU, and the leakage coefficient α = 0.2; Each layer of convolution operation performs non-linear transformation through the LeakyReLU activation function, and the number of output channels of each layer is 64; Output: Shallow feature map F shallow ∈R H×W×64 ; The residual dense downsampling module RRDB includes 4 levels of residual dense blocks. Each level of residual dense block sequentially includes 3 dense connection layers, a residual connection layer, and a downsampling layer. After processing the input shallow feature map through the residual dense downsampling module, a deep feature map is output The multi-scale deformable upsampling module includes a dynamic convolution kernel generation network, a deformable upsampling layer, and cross-scale skip connections; the shallow feature map and the deep feature map are input into the multi-scale deformable upsampling module; Extract features from intermediate layers of different scales, and fuse the features into the final output feature map through skip connections; The dynamic convolution kernel generation network KernelNet is used to generate dynamic convolution kernels. By processing the deep features, convolution kernels of different sizes and shapes are generated; among them: The feature dimension of the input layer is 512, and the output dimension is 18, corresponding to the x / y offsets of the 3×3 convolution kernel; each convolution kernel contains 9 offset values, which are used to dynamically adjust the position and shape of the convolution kernel; After being processed, the convolution kernels generated by the dynamic convolution kernel generation network are used to perform deformable convolution operations on the feature map, so as to perform more refined upsampling on the image; The deformable upsampling layer uses the convolution kernels generated by the dynamic convolution kernel generation network to perform upsampling operations on the input feature map; the deformable upsampling layer dynamically adjusts the convolution operation according to the offsets of the convolution kernels; The cross-scale skip connection combines image information of different scales through cross-scale feature fusion connections to enhance the structural consistency and detail recovery ability of the generated image.

8. The few-shot image super-resolution reconstruction method combining self-supervised learning and generative adversarial network according to claim 7, characterized in that The specific operation of the cross-scale skip connection is as follows: Feature extraction and downsampling: In the generator network, the input image first undergoes feature extraction through a shallow feature extraction module, and then is downsampled through a residual dense downsampling module to obtain multi-scale deep feature maps; Cross-scale feature fusion: Let F1, F2, and F3 be feature maps of three different scales. The low-resolution feature maps F3 and F2 are upsampled to the same spatial size as F1 using an upsampling operation. Then, the features of these three scales are fused element-wise with weights; Final upsampling and output: The fused feature F fuse is fed into the detail enhancement unit for texture restoration and a super-resolution image is generated through the final output layer.

9. The few-shot image super-resolution reconstruction method combining self-supervised learning and generative adversarial network according to claim 7, wherein The detail enhancement unit includes a local texture branch and a global structure branch, and the outputs of the two branches are dynamically fused: F enhanced = αF local + (1 - α)F global Among them, F enhanced represents the finally enhanced feature representation, which is used to generate high-quality super-resolution images; F local represents the local texture feature, which is extracted by a 3×3 standard convolutional layer and is used to enhance local details; F global represents the global structural feature, which is obtained through dilated convolution and is used to maintain the overall structural consistency; α is the adaptive fusion weight, which is used to control the proportion of the local texture feature and the global structural feature.

10. The few-shot image super-resolution reconstruction method combining self-supervised learning and generative adversarial network according to claim 1, wherein The discriminator network D-Net adopts a multi-scale PatchGAN structure, and the specific design is as follows: Three discriminators are in parallel, with the same structure, and process image patches of different scales respectively; Perform spectral normalization on the convolutional weights of each layer. The calculation formula is as follows, where W is the convolutional kernel, and W normal is the normalized convolutional kernel: The perceptual loss is calculated using the features of the conv5-4 layer of the pre-trained VGG19 network, and the calculation formula is as follows: Among them, Φ j is the feature extraction function of the j-th layer of the pre-trained VGG19 network, is the predicted image, is the real image, and N is the number of pixels.

Citation Information

Patent Citations

  • Single Image Super-Resolution Reconstruction Methods

    CN106600533B