An efficient super-high-resolution image restoration method based on dynamic dataset distillation

By employing a dynamic dataset distillation algorithm, utilizing ViT and CLIP encoders to calculate image complexity scores, and selecting high-quality samples for fine-tuning, the high cost problem in deep learning image restoration methods is solved, achieving efficient ultra-high resolution image restoration.

CN120765511BActive Publication Date: 2025-11-04SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511240626.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-11-04
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

Existing deep learning image restoration methods are costly to construct large-scale datasets, and existing dataset distillation techniques are difficult to apply directly to image restoration tasks, resulting in a decline in image restoration quality.

Method used

The TripleD algorithm is used to synthesize a simplified subset from a large-scale image restoration dataset. The variance of the feature maps is extracted by ViT and CLIP encoders to calculate the complexity score. A dynamic sample selection mechanism is used to screen high-quality samples. A convolutional neural network is used for fine-tuning to construct the data subset and train the image restoration model.

Benefits of technology

It achieves near-full dataset training performance with only 2%-5% of the data, reducing training costs and improving image restoration quality. It adapts to image datasets from different sources and with different characteristics, improving training efficiency and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765511B_ABST
    Figure CN120765511B_ABST
Patent Text Reader

Abstract

The application provides an efficient super-high-resolution image restoration method based on dynamic dataset distillation, comprising the following steps: obtaining an original degraded image dataset and an original clear image dataset, and constructing a complete dataset; after down-sampling the original degraded image dataset and the original clear image dataset, respectively sending them into a ViT encoder and a CLIP encoder for feature extraction, calculating the variance of the first feature map extracted by the ViT encoder as a first complexity score, and calculating the variance of the second feature map extracted by the CLIP encoder as a second complexity score; based on a dynamic sample selection mechanism, assigning a complexity weight and an uncertainty weight to each image data, and calculating a selection score for each image data; selecting image data with a high selection score to construct a data subset, and fine-tuning the data subset based on a convolutional neural network; and training an image restoration model based on the data subset. The application solves the high cost problem of current deep learning in image restoration.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image restoration, and in particular to an efficient super-high resolution image restoration method based on dynamic dataset distillation. BACKGROUND

[0002] In the field of computer vision, image restoration is a key technology that can provide high-quality image information for downstream vision tasks such as denoising, deblurring, and deraining. In recent years, deep learning technology has been widely applied in image restoration due to its powerful feature extraction and modeling capabilities, significantly improving the effectiveness of image restoration and achieving leading results in various benchmark tests.

[0003] Although deep learning algorithms perform exceptionally well in image restoration tasks, their performance is highly dependent on large-scale datasets. To pursue better restoration results, constructing large-scale benchmark datasets has become the current trend. However, this process is accompanied by many challenges. Building large-scale datasets not only requires a large amount of manpower, material resources, and time costs for data collection, organization, and annotation, but also requires exponential growth in computing resources during the model training phase, resulting in extremely high training costs, which to some extent limits the further promotion and application of deep learning in the field of image restoration.

[0004] To alleviate this problem, dataset distillation technology has been proposed as a potential solution. The core of this technology is to synthesize a representative subset from the original large-scale dataset, and the data volume of this subset is much smaller than the original dataset. Through carefully designed distillation strategies, the model can achieve similar performance when trained on this subset as when trained on the large-scale dataset, thereby reducing the data volume while preserving as much key information as possible from the original data.

[0005] However, the dataset distillation technology is still in its infancy in the field of image restoration. Image restoration is essentially a fine regression operation that requires accurate prediction and restoration of image pixel values, which is significantly different from classification tasks. In classification tasks, the probability values output by the model fluctuate within a certain range, which can still represent the same class, and there is a certain tolerance. However, image restoration tasks require high accuracy of the prediction results, and any minor error can significantly degrade the quality of the restored image. Therefore, existing dataset distillation methods cannot be directly applied to image restoration tasks, and there is an urgent need for an efficient dataset distillation-based method specifically for image restoration to address the high cost problem currently faced by deep learning in image restoration. SUMMARY

[0006] The present application aims to overcome the above technical deficiencies, and provides an efficient super-high resolution image restoration method based on dynamic dataset distillation, which provides a TripleD algorithm that can efficiently synthesize a compact subset from a large-scale image restoration dataset. The TripleD algorithm uses only 2%-5% of the data, and can achieve a performance close to that of full-dataset training, reaching 90%-95%.

[0007] To achieve the above technical purpose, in a first aspect, the technical solution of the present application provides an efficient super-high resolution image restoration method based on dynamic dataset distillation, comprising the steps of:

[0008] Obtaining an original degraded image dataset and an original clear image dataset, and constructing a complete dataset;

[0009] Downsampling the original degraded image dataset and the original clear image dataset, and feeding them into a ViT encoder and a CLIP encoder for feature extraction, respectively. The variance of the first feature map extracted by the ViT encoder is calculated as the first complexity score, and the variance of the second feature map extracted by the CLIP encoder is calculated as the second complexity score;

[0010] Based on a dynamic sample selection mechanism, a complexity weight and an uncertainty weight are assigned to the first complexity score and the second complexity score of each image data, respectively, and a selection score is calculated for each image data based on the dynamic sample selection mechanism;

[0011] Selecting the image data with the top pre-set proportion of selection scores to construct a data subset, and fine-tuning the data subset based on a convolutional neural network;

[0012] Training an image restoration model based on the data subset.

[0013] Compared with the prior art, the present application has the following advantages:

[0014] The TripleD algorithm of the present application extracts features from down-sampled images (all input images are down-sampled to 128x128 resolution) using a visual Transformer (ViT) and CLIP to represent high-level semantic information of the images. The key reason why the present application selects to perform sample screening based on high-level semantic features (rather than low-level features) is that, with the help of GPU shaders, ViT and CLIP can run efficiently when processing down-sampled images, while still retaining the semantic information of the original resolution images. Subsequently, the present application calculates the variance of these features to measure the complexity of the samples. In this process, the inventors of the present application have also tried measurement methods such as cosine distance, standard deviation, KL divergence and entropy, but they are not good enough in evaluating image complexity. The contributions of this work are as follows: (1) the present application first designs a data set distillation framework (TripleD) for image restoration tasks, which reduces the cost of training image restoration models on large-scale data sets. (2) the present application develops an efficient dynamic distillation method that gradually increases the complexity of samples to select high-quality sample clusters for distillation, thereby obtaining a sub-data set of 2% of the whole data set. (3) the method of the present application can also achieve data set distillation on a single GPU within a limited time when processing ultra-high-definition data sets. A large number of experimental results show that the method of the present application can help the model to maintain 90% of the performance at a lower cost.

[0015] According to some embodiments of the present application, down-sampling the original degraded image data set, the original clear image data set comprises the steps of:

[0016] Adjusting the original degraded image data set, the original clear image data set to the scale of 128x128 through bilinear interpolation.

[0017] According to some embodiments of the present application, calculating the variance of the first feature map extracted by the ViT encoder as the first complexity score comprises the steps of:

[0018] For an image , the first complexity score is obtained by calculating the variance of the feature vector extracted by the ViT encoder, denoted as , wherein is the number of image blocks, is the feature dimension of each image block, and the calculation formula of the first complexity score is:

[0019] ,

[0020] wherein, denotes the The average eigenvalue of each dimension of the variance is calculated by the following formula:

[0021] ,

[0022] wherein, is the average eigenvalue of the dimension of all image blocks, denotes the dimension of all blocks, and the calculation formula of is as follows:

[0023] .

[0024] According to some embodiments of the present application, the second feature map extracted by the CLIP encoder is calculated for the variance as a second complexity score, comprising the steps of:

[0025] For an image , the second complexity score is obtained by calculating the variance of the feature vector extracted by the CLIP encoder, and let denote the feature map, wherein is the number of image blocks, is the feature dimension of each image block, and the calculation formula of the second complexity score is as follows:

[0026] ,

[0027] wherein, denotes the eigenvalue of the dimension of all image blocks, and the variance of each dimension is calculated by the following formula:

[0028] ,

[0029] wherein, is the average eigenvalue of the dimension of all image blocks, denotes the dimension of all blocks, and the calculation formula of is as follows:

[0030] .

[0031] According to some embodiments of the present application, the first complexity score and the second complexity score of each image data are respectively assigned with a complexity weight and an uncertainty weight based on a dynamic sample selection mechanism, and a selection score of each image data is calculated based on the dynamic sample selection mechanism, comprising the steps of:

[0032] The selection score , the calculation formula is:

[0033] ,

[0034] wherein, alpha and beta are dynamic weights, which are adjusted throughout the training process in order to emphasize complexity or uncertainty according to the training stage.

[0035] According to some embodiments of the application, alpha and beta are set to (0.1, 0.1) in the early training stage; alpha and beta are set to (0.5, 0.5) in the middle training stage; and alpha and beta are set to (0.7, 0.7) in the later stage.

[0036] According to some embodiments of the application, the data subset is constructed by selecting the image data with the top preset proportion of the selection score, comprising the steps of:

[0037] The data subset is constructed by selecting the image data with the top 2%-5% of the selection score.

[0038] According to some embodiments of the application, the data subset is fine-tuned based on a convolutional neural network, comprising the steps of:

[0039] An 8-layer convolutional neural network (CNN) is constructed, with a convolution kernel size of 3x3 and a step size of 1, for fine-tuning the distribution of the data subset. The fine-tuned data subset is used to train an image restoration model, and the convolutional neural network is updated by means of the loss of the image restoration model. L1 loss is used to provide pixel-level supervision, and the calculation formula is:

[0040] ,

[0041] wherein, L1 norm, the data subset, denotes the output of the image restoration model.

[0042] In a second aspect, the technical scheme of the application provides an efficient super-high resolution image restoration system based on dynamic data set distillation, comprising:

[0043] A data set construction module is used to obtain an original degraded image data set and an original clear image data set, and to construct a complete data set.

[0044] A complexity calculation module is used to downsample the original degraded image data set and the original clear image data set, respectively, and then send them into a ViT encoder and a CLIP encoder for feature extraction. The first feature map extracted by the ViT encoder is used to calculate the variance as a first complexity score, and the second feature map extracted by the CLIP encoder is used to calculate the variance as a second complexity score.

[0045] The selection score calculation module assigns a complexity weight and an uncertainty weight to the first complexity score and the second complexity score of each image data respectively based on the dynamic sample selection mechanism, and calculates a selection score for each image data based on the dynamic sample selection mechanism;

[0046] The subset construction module is configured to select the image data with a top pre-set proportion of selection scores to construct a data subset, and fine-tune the data subset based on a convolutional neural network.

[0047] The model training module is configured to train the data subset to an image restoration model.

[0048] In a third aspect, the present application provides a computer readable storage medium storing computer executable instructions for causing a computer to execute the efficient super-high resolution image restoration method based on dynamic dataset distillation according to any one of the first aspect.

[0049] Additional aspects and advantages of the present application will be made apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0050] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description, taken in conjunction with the accompanying drawings, in which:

[0051] Figure 1 A flowchart of the efficient super-high resolution image restoration method based on dynamic dataset distillation provided by an embodiment of the present application is shown in FIG. 1.

[0052] Figure 2 A flowchart of the efficient super-high resolution image restoration method based on dynamic dataset distillation provided by an embodiment of the present application is shown in FIG. 1. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0054] It should be noted that although the functional modules are divided in the system schematic diagram, and the logical order is shown in the flowchart, in some cases, the steps shown or described can be performed in a manner different from the module division in the system or the order in the flowchart. The terms "first", "second", etc. in the specification and claims and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.

[0055] Referring to Figure 1 and Figure 2 , Figure 1 The flowchart of the efficient super-high resolution image restoration method based on dynamic dataset distillation provided by an embodiment of the present application;

[0056] Figure 2 The flowchart of the efficient super-high resolution image restoration method based on dynamic dataset distillation provided by an embodiment of the present application. The efficient super-high resolution image restoration method based on dynamic dataset distillation includes but is not limited to the following steps:

[0057] Step S110, obtaining an original degraded image dataset and an original clear image dataset, and constructing a complete dataset;

[0058] Step S120, after down-sampling the original degraded image dataset and the original clear image dataset, respectively sending them into a ViT encoder and a CLIP encoder for feature extraction, calculating the variance of the first feature map extracted by the ViT encoder as a first complexity score, and calculating the variance of the second feature map extracted by the CLIP encoder as a second complexity score;

[0059] Step S130, based on a dynamic sample selection mechanism, assigning complexity weights and uncertainty weights to the first complexity score and the second complexity score of each image data, respectively, and calculating a selection score for each image data based on the dynamic sample selection mechanism;

[0060] Step S140, selecting image data with a selection score in the top preset proportion to construct a data subset, and fine-tuning the data subset based on a convolutional neural network;

[0061] Step S150, training the image restoration model with the data subset.

[0062] In an embodiment, the efficient super-high-resolution image restoration method based on dynamic dataset distillation comprises the steps of: obtaining an original degraded image dataset and an original clear image dataset, and constructing a complete dataset; after downsampling the original degraded image dataset and the original clear image dataset, respectively sending them into a ViT encoder and a CLIP encoder for feature extraction, calculating the variance of the first feature map extracted by the ViT encoder as a first complexity score, and calculating the variance of the second feature map extracted by the CLIP encoder as a second complexity score; based on a dynamic sample selection mechanism, assigning a complexity weight and an uncertainty weight to the first complexity score and the second complexity score of each image data respectively, and calculating a selection score for each image data based on the dynamic sample selection mechanism; selecting a preset proportion of image data with the top selection scores to construct a data subset, and fine-tuning the data subset based on a convolutional neural network; and training an image restoration model based on the data subset.

[0063] The present application filters representative image data from the original large-scale complete dataset to perform subsequent model training by constructing a data subset. This avoids the high cost of directly using a large-scale dataset, effectively controls the subsequent consumption of computing resources, and solves the problem of resource waste caused by excessive reliance on large-scale datasets in traditional deep learning image restoration methods.

[0064] The present application uses a ViT encoder and a CLIP encoder to extract features and calculate corresponding complexity scores, which can mine data in the dataset that is most critical and most representative of image restoration task-related features. Compared with simple data subset construction methods such as random sampling, this feature-based screening can better ensure the quality of the selected data subset, so that a limited amount of data can cover more abundant and important image restoration-related features, improving data utilization efficiency. The present application assigns a complexity weight and an uncertainty weight to each image data based on a dynamic sample selection mechanism, and calculates a selection score to select a data subset, meaning that the data included in the training is carefully selected and takes into account multiple factors, and is more in line with the actual needs of image restoration model training. This allows the model to focus on these high-quality, high-value data during training, avoiding interference from irrelevant or low-quality data, which helps to improve the relevance of model training, speed up the convergence process, and improve training efficiency.

[0065] Although the data subset used by the present application has a much smaller data volume than the original data set, it is still able to promote the model to learn sufficient effective image restoration knowledge and patterns during the training of the image restoration model, so that it can achieve more than 90% performance when facing the super high resolution image restoration task, ensuring that the model can output high-quality restoration results in downstream visual tasks such as denoising, deblurring, and deraining. The application of the dynamic sample selection mechanism of the present application enables the method to flexibly cope with original degraded image data sets and original clear image data sets of different sources and different characteristics. Regardless of the differences in image content, resolution, degradation degree, etc. of the data set itself, the mechanism can filter out suitable subsets for model training, has strong universality and adaptability, and can be widely applied to super high resolution image restoration work in various scenes.

[0066] With the further improvement of subsequent requirements for image restoration tasks or the emergence of new image feature dimensions, the method can relatively easily optimize and adjust the dynamic sample selection mechanism and other links on the basis of the existing framework to adapt to new changes, such as incorporating more feature evaluation indicators or adjusting the weight distribution strategy, etc. to facilitate technical iteration and expand the application range. The present application selects data subsets for training, greatly reducing the computational load of the convolutional neural network during the training process, reducing the dependence on long-term high-intensity operation of hardware resources such as GPU, saving a large amount of power, hardware loss, and other costs. At the same time, the time period of model training is shortened, indirectly improving the overall efficiency of research, production, and other links, so that the image restoration technology can be applied in more resource-constrained practical environments.

[0067] Overall process summary: TripleD first extracts features from the complete data set using CLIP (c1) and ViT (c2) models. Then, the selection mechanism calculates a score for each image based on complexity (α×c1 + β×c2). In each cycle of the training phase, a light-weight CNN is used to fine-tune the data distribution of the top 2% images. Finally, the image restoration model is trained using these 2% fine-tuned data.

[0068] The TripleD method includes image complexity calculation, data selection mechanism, and data distribution fine-tuning.

[0069] Image complexity calculation:

[0070] Given a data set X, the method extracts features from two pre-trained encoders (ViT and CLIP) to evaluate the complexity of the image distribution.

[0071] The effectiveness of the ViT-based complexity estimation is reflected in the variance of the features extracted by the ViT. Since the input images are down-sampled to 128x128 by bilinear interpolation, the ViT is able to amplify the differences between low-resolution images. Here, the main purpose of down-sampling the images in this invention is to improve the efficiency of the data distillation algorithm. Mathematically, for an image , the first complexity score is obtained by taking the variance of the feature vectors extracted by the ViT encoder, where denotes the feature map, and is the number of image patches, is the feature dimension of each image patch, and the formula for calculating the first complexity score is as follows:

[0072] ,

[0073] where denotes the feature value in the th dimension of all image patches, and the variance of each dimension is calculated by:

[0074] ,

[0075] where is the average feature value in the th dimension of all image patches, denotes the th dimension among all patches, and the formula for calculating

[0076] .

[0077] To more clearly illustrate the effectiveness of this method, the invention visualizes and analyzes some degenerate images. The image with uniform rain stripes shows the lowest variance, indicating the simplest complexity. In contrast, the mixed image containing simple features and more diverse features has moderate complexity. Finally, the image containing various objects and textures exhibits the highest variance, reflecting the greatest complexity.

[0078] Since the ViT is only trained on visual images, the invention attempts to use CLIP to supplement the image semantics that it may have overlooked. Specifically, CLIP consists of a visual encoder and a text encoder, and the invention only uses the visual encoder (which implicitly contains rich text representation semantics) and adjusts the input image to a scale of 128x128 by bilinear interpolation. Subsequently, CLIP calculates the variance in the same way as the ViT.

[0079] calculating variance of the second feature map extracted by the CLIP encoder as the second complexity score, comprising the steps of:

[0080] For an image , the second complexity score is obtained by calculating the variance of the feature vector extracted by the CLIP encoder, and let denote the feature map, where is the number of image blocks, is the feature dimension of each image block, and the calculation formula of the second complexity score is as follows:

[0081] ,

[0082] wherein denotes the feature value of the th dimension in all image blocks, and the variance of each dimension is calculated by the following formula:

[0083] ,

[0084] wherein is the average feature value of the th dimension in all image blocks, denotes the th dimension in all blocks, and the calculation formula is as follows:

[0085] .

[0086] Dynamic sample selection mechanism:

[0087] So far, four complexity scores have been obtained, namely the variance of the degraded image and the variance of the GT image. In order to accurately construct the subset, a dynamic sample selection mechanism is introduced. According to the combination of the ViT-based complexity and the CLIP-based complexity of the image , a selection score is assigned to each image, and the calculation formula is as follows:

[0088] ,

[0089] Here, α and β are dynamic weights that are adjusted throughout the training process to emphasize complexity or uncertainty depending on the training stage. Throughout training, the complexity weight α gradually increases, while the uncertainty weight β also gradually increases, causing the model's focus to gradually shift from simpler examples to more uncertain and challenging samples. Specifically, in the early stages of training, α and β are set to (0.1, 0.1); in the middle stages, they are set to (0.5, 0.5); and in the later stages, they are set to (0.7, 0.7). Within each training epoch, only the score is selected. The top 2% of samples were used for training. This fixed threshold ensured computational efficiency during training while providing the model with diverse samples at different stages. The 2% threshold was determined experimentally, offering the optimal balance between computational cost and model performance. This invention experimented with other thresholds (e.g., 1% and 5%), but found that 2% achieved the best trade-off, allowing the model to achieve 90%-95% of the full dataset training performance while significantly reducing computational requirements. Finally, a distillation dataset B was obtained.

[0090] Fine-tuning a subset of data using a convolutional neural network includes the following steps:

[0091] An 8-layer convolutional neural network (CNN) with 3×3 kernels and a stride of 1 is constructed to fine-tune the distribution of a subset of data. This fine-tuned subset is then used to train the image restoration model. The CNN is updated using the loss function of the image restoration model, employing L1 loss to provide pixel-level supervision. The calculation formula is as follows:

[0092] ,

[0093] in, Describing the L1 norm, For a subset of data, This represents the output of the image restoration model. The invention also tested perceptual loss and adversarial loss, finding that they had limited effectiveness in improving visual quality and were more computationally expensive.

[0094] The effect after improvement:

[0095] The experiment mainly included three groups:

[0096] One group consists of multi-task image restoration experiments;

[0097] One group consists of integrated image restoration experiments;

[0098] The other group is an experiment on the restoration of ultra-high-definition low-light images.

[0099] This invention uses two widely accepted metrics to evaluate performance: PSNR and SSIM. PSNR and SSIM focus on evaluating the fidelity and structural integrity of the reconstructed image.

[0100] Among them, PSNR (Peak Signal-to-Noise Ratio) is a commonly used metric for measuring image quality. It is mainly used to evaluate the difference between the processed image and the original image. The higher the value, the better the image quality. The calculation formula is:

[0101] ;

[0102] ;

[0103] MSE stands for Mean Squared Error, which measures the difference between two images at the pixel level. and represents the pixel value at position (i,j) of the generated fused image and the corresponding gold standard image (ground truth), respectively; m and n represent the width and height of the image, respectively; and MAX represents the maximum value of the image pixels.

[0104] SSIM (Structural Similarity Index Measure) is an indicator used to evaluate the quality of two images. It mainly measures the similarity of images by comparing their brightness, contrast, and structural information. The closer the value is to 1, the better the image quality.

[0105] The formula for calculating SSIM is as follows:

[0106] ;

[0107] ;

[0108] ;

[0109] in, It is the fused image output by the model. The average value, Is with Corresponding gold standard The average value, It represents an image and covariance, yes variance yes The variance; L is the dynamic range of pixel values. and denotes a preset hyper-parameter, here = 0.01, = 0.03, and denotes a smoothing parameter.

[0110] Multi-task image restoration experiments:

[0111] The present application has been tested on multiple standard low-level vision tasks, including single-image rain removal, motion deblurring, defocus deblurring, and image denoising. All comparative methods are retrained on small-scale datasets extracted from the Restormer method. The experiments use the Rain100L, Rain100H, GoPro, RealBlur-J, RealBlur-R, DPDD, SIDD, and DND datasets for benchmarking. These datasets cover a wide range of image degradation types, enabling the present application to evaluate TripleD's generalization ability across different tasks.

[0112] Table 1. Multi-task image restoration experiment results

[0113]

[0114] Table 1 compares the performance of Restormer under different training methods, covering multiple image restoration tasks and using PSNR and SSIM as evaluation metrics. Restormer is a model that performs well in the field of image restoration, where Restormer (multiple GPUs) represents the version trained using multiple image processing units, which can fully utilize computing resources to achieve better performance; Restormer (single GPU) refers to the version trained using a single image processing unit, which is limited by the computing power of a single image processing unit; and Restormer (our method (single GPU)) trained using the dynamic dataset distillation method proposed in this paper, under the condition of a single image processing unit, only uses 2% of the dataset, and the goal is to reduce data requirements while achieving performance close to that of multiple image processing units.

[0115] Experiments cover multiple datasets and tasks to comprehensively evaluate the generalization ability of the model in different image restoration scenarios. In various tasks, the Restormer trained by TripleD (our method (single GPU)) still approaches the Restormer trained on multiple image processing units in terms of PSNR and SSIM, although it only uses 2% of the data and is trained on a single image processing unit. For example, in the Rain100L rain removal task, the PSNR of Restormer (multiple GPUs) is 36.52, and the SSIM is 0.925, while the PSNR of Restormer trained by TripleD (our method (single GPU)) is 32.08, and the SSIM is 0.910, far exceeding the Restormer trained directly on a single image processing unit (single GPU) (PSNR is 33.22, and SSIM is 0.921). This result shows that TripleD can support efficient training with very little data under limited computing resources, not only significantly reducing training costs, but also maintaining high image restoration quality, effectively balancing computing efficiency and model performance.

[0116] Integrated image restoration experiments:

[0117] Table 2. Results of integrated image restoration experiments

[0118]

[0119] The present application evaluates the effectiveness of the TripleD method in training the well-known integrated model PromptIR under multiple degradation types. To ensure that all types of degradation (such as rain, fog, and noise) are fully represented during training, TripleD dynamically selects samples from each degradation category and balances the number of samples in each category. Table 2 shows the quantitative results of PSNR and SSIM for different image restoration tasks, including de-fogging on the SOTS dataset, rain removal on the Rain100L dataset, and de-noising on the BSD68 dataset at different noise levels (σ = 15, 25, 50), where σ represents the variance of Gaussian noise. The experimental results show that the PromptIR model trained using TripleD performs close to or slightly better than the retrained PromptIR model on multiple tasks. The hardware configuration and parameter settings of the experiment are consistent with the multi-task image restoration experiment, and all comparison methods are retrained on the small-scale dataset extracted from PromptIR.

[0120] Ultra-high-definition low-light image restoration experiments:

[0121] Table 3. Quantitative results on the UHD-LOL4K dataset

[0122]

[0123] Table 4 Quantitative results on UHD-LL dataset

[0124]

[0125] The present application evaluates the effectiveness of the TripleD method in training the famous ultra-high definition model UHDFormer against multiple low-light degradations. All the comparative methods are retrained on a small-scale dataset extracted from UHDFormer. Tables 3 and 4 show the PSNR and SSIM quantitative results of the low-light image restoration task. Experiments show that the UHDFormer model trained using TripleD has a performance close to that of the retrained UHDFormer, maintaining more than 90% of the restoration effect. This result shows that the TripleD method can still achieve excellent image restoration performance while significantly reducing the amount of training data.

[0126] In an embodiment, the efficient ultra-high resolution image restoration system based on dynamic dataset distillation comprises: a dataset construction module for obtaining an original degraded image dataset and an original clear image dataset, and constructing a complete dataset; a complexity calculation module for down-sampling the original degraded image dataset and the original clear image dataset and respectively sending them into a ViT encoder and a CLIP encoder for feature extraction, calculating the variance of the first feature map extracted by the ViT encoder as a first complexity score, and calculating the variance of the second feature map extracted by the CLIP encoder as a second complexity score; a selection score calculation module for assigning complexity weights and uncertainty weights to the first complexity score and the second complexity score of each image data based on a dynamic sample selection mechanism, and calculating a selection score for each image data based on the dynamic sample selection mechanism; a subset construction module for selecting a preset proportion of image data with a high selection score to construct a data subset, and fine-tuning the data subset based on a convolutional neural network; and a model training module for training an image restoration model based on the data subset.

[0127] Memory, as a kind of non-transient computer readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transient memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transient solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0128] The apparatus embodiments described above are only illustrative, wherein the units described as separate components can or can not be physically separated, and can be located in one place or distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment.

[0129] In addition, one embodiment of the present application also provides a computer readable storage medium, which stores computer executable instructions, and the computer executable instructions are executed by a processor or a controller, for example, a processor in the terminal embodiment described above, so that the processor executes the high-efficiency super-high-resolution image restoration method based on dynamic data set distillation in the above embodiment.

[0130] Those skilled in the art can understand that all or some steps in the method disclosed above, the system can be implemented as software, firmware, hardware and appropriate combination thereof. Some or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or as hardware, or as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer readable medium, which can include computer storage media (or non-transitory media) and communication media (or transitory media). As known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. In addition, as known to those skilled in the art, communication media generally includes computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism, and can include any information delivery medium.

[0131] The above is a specific description of the preferred embodiment of the present application, but the present application is not limited to the above-mentioned embodiments. Those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the present application, and these equivalent modifications or replacements are all included in the scope defined by the claims of the present application.

[0132] The above description of the specific embodiments of the present application is not intended to limit the scope of the present application. Any other corresponding changes and modifications made according to the technical concept of the present application should be included in the scope of protection of the claims of the present application.

Claims

1. An efficient super-high resolution image restoration method based on dynamic dataset distillation, characterized in that, Including the following steps: Obtain the original degraded image dataset and the original clear image dataset, and construct the complete dataset; After downsampling the original degraded image dataset and the original clear image dataset, they are respectively fed into the ViT encoder and the CLIP encoder for feature extraction. The variance of the first feature map extracted by the ViT encoder is calculated as the first complexity score, and the variance of the second feature map extracted by the CLIP encoder is calculated as the second complexity score. Based on the dynamic sample selection mechanism, complexity weights and uncertainty weights are assigned to the first complexity score and the second complexity score of each image data, respectively, and a selection score is calculated for each image data based on the dynamic sample selection mechanism. A subset of image data is constructed by selecting the image data with the highest selected scores from a predetermined proportion, and the subset of image data is fine-tuned based on a convolutional neural network. The image restoration model is trained using the subset of data.

2. The efficient super-high resolution image restoration method based on dynamic dataset distillation according to claim 1, characterized in that, Downsampling the original degraded image dataset and the original clear image dataset includes the following steps: The original degraded image dataset and the original clear image dataset are adjusted to a scale of 128×128 using bilinear interpolation.

3. The efficient super-high resolution image restoration method based on dynamic dataset distillation according to claim 1, characterized in that, The variance of the first feature map extracted by the ViT encoder is calculated as the first complexity score, including the following steps: For an image , the first complexity score is obtained by calculating the variance of the feature vectors extracted by the ViT encoder, denoted as , where is the number of image blocks, is the feature dimension of each image block, and the calculation formula of the first complexity score is , in, Indicates the first image patch among all image patches. The eigenvalues ​​of the dimension, the variance of each dimension is calculated by the following formula: , in, It is the first of all image patches The average eigenvalues ​​of the dimension, Indicates the first in all blocks One dimension, The calculation formula is as follows: 。 4. The efficient ultra-high resolution image restoration method based on dynamic dataset distillation according to claim 3, characterized in that, The variance of the second feature map extracted by the CLIP encoder is calculated as the second complexity score, including the following steps: For an image Second complexity score It is obtained by calculating the variance of the feature vectors extracted by the CLIP encoder. Let... Represents the feature map, where It is the number of image patches. It is the feature dimension of each image patch, and the second complexity score. The calculation formula is: , in, Indicates the first image patch among all image patches. The eigenvalues ​​of the dimension, the variance of each dimension is calculated by the following formula: , in, It is the first of all image patches The average eigenvalues ​​of the dimension, Indicates the first in all blocks One dimension, The calculation formula is as follows: 。 5. The efficient ultra-high resolution image restoration method based on dynamic dataset distillation according to claim 4, characterized in that, The first complexity score and the second complexity score of each image data are assigned complexity weights and uncertainty weights based on a dynamic sample selection mechanism, and a selection score is calculated for each image data based on the dynamic sample selection mechanism, including the following steps: Calculate the selection score The calculation formula is: , Here, α and β are dynamic weights that are adjusted throughout the training process to emphasize complexity or uncertainty depending on the training stage.

6. The efficient ultra-high resolution image restoration method based on dynamic dataset distillation according to claim 5, characterized in that, In the early stages of training, α and β were set to (0.1, 0.1); in the middle stages of training, α and β were set to (0.5, 0.5); and in the later stages, α and β were set to (0.7, 0.7).

7. The efficient ultra-high resolution image restoration method based on dynamic dataset distillation according to claim 6, characterized in that, Selecting a subset of image data with the highest selected scores from a predetermined proportion to construct a data set includes the following steps: A subset of data is constructed by selecting the image data whose selection scores are in the top 2%-5%.

8. The efficient ultra-high resolution image restoration method based on dynamic dataset distillation according to claim 7, characterized in that, Fine-tuning the data subset based on a convolutional neural network includes the following steps: An 8-layer convolutional neural network (CNN) with a kernel size of 3×3 and a stride of 1 is constructed to fine-tune the distribution of the data subset. This fine-tuned data subset is then used to train the image restoration model. The CNN is updated using the loss function of the image restoration model, employing L1 loss to provide pixel-level supervision. The calculation formula is as follows: , in, Describing the L1 norm, For the data subset, This represents the output of the image restoration model.

9. A high-efficiency ultra-high resolution image restoration system based on dynamic dataset distillation, characterized in that, include: The dataset building module is used to obtain the original degraded image dataset and the original clear image dataset, and to build the complete dataset; The complexity calculation module is used to downsample the original degraded image dataset and the original clear image dataset and then send them to the ViT encoder and CLIP encoder for feature extraction, respectively. The variance of the first feature map extracted by the ViT encoder is calculated as the first complexity score, and the variance of the second feature map extracted by the CLIP encoder is calculated as the second complexity score. The selection score calculation module assigns complexity weights and uncertainty weights to the first complexity score and the second complexity score of each image data based on a dynamic sample selection mechanism, and calculates a selection score for each image data based on the dynamic sample selection mechanism. A subset construction module is used to select a preset proportion of the image data with the highest selection scores to construct a data subset, and to fine-tune the data subset based on a convolutional neural network; The model training module is used to train the image recovery model from the subset of data.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the efficient ultra-high resolution image restoration method based on dynamic dataset distillation as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • General image restoration method for severe weather based on two-stage distillation learning

    CN118521511A

  • Domain generalization method combining vision-language pre-training and prompt learning

    CN118607591A