Method for generating super-resolution image dataset, super-resolution image model, and training method

The method addresses the challenge of generalizing image super-resolution models by using blind degradation to simulate real-world image degradations, training a model, and then using it to enhance a simpler model's ability to handle diverse degradation causes, resulting in improved training efficiency and effectiveness.

JP2025516410AInactive Publication Date: 2025-05-30WELLINK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023575571
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-04-26
Filing Date
2023-09-08
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing image super-resolution models face challenges in generalizing well to real-world images with various degradation causes, leading to low training effectiveness and slow learning speeds, especially on terminal-side models with simpler structures.

Method used

A method is proposed to generate an image super-resolution dataset using blind degradation processing, which simulates different real-world degradation causes. This dataset is used to train a first model, and the trained model is then used to infer low-resolution images to generate a secondary dataset for training a second model with a simpler structure.

Benefits of technology

The approach enhances the generalization ability of the model, achieving a relatively good training effect and fast learning speed, even with a simple model structure, thereby improving the super-resolution performance on images with various degradation causes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025516410000001_ABST
    Figure 2025516410000001_ABST
Patent Text Reader

Abstract

A method for generating a super-resolution image dataset, a super-resolution image model, and a training method. The method for generating the super-resolution image dataset includes: step S101 of constructing a high-resolution image set; step S102 of performing image blind degradation processing on all high-resolution images HR1 to obtain an LR1-HR1 dataset; step S103 of training a first model using the LR1-HR1 dataset, and obtaining and saving the model parameters of the first model after training is completed; step S104 of constructing a low-resolution image set; and step S105 of inputting all low-resolution images LR2 into the obtained first model, inferring by the first model, and obtaining an LR2-SR2 dataset. A training method for a super-resolution image model, which trains a second model using the LR2-SR2 dataset obtained by the above method. A super-resolution image model obtained by training a second model by the above method, where the second model is an ECBSR model. By using the LR2-SR2 dataset of the present invention, it is possible to train with a model having a simple structure, with a fast learning speed and a strong generalization ability of the model after training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image super-resolution, and specifically to a method for generating an image super-resolution dataset, a model training method, and an image super-resolution model.

Background Art

[0002] Image super-resolution (SR) is a technology that restores a low-resolution input to a high-resolution (HR) output to improve the display quality. Image super-resolution is divided into cloud super-resolution and terminal-side super-resolution. Cloud super-resolution performs image super-resolution on a server, and terminal-side super-resolution performs image super-resolution on a terminal device. There are various causes for reducing the quality of images in the real world. Therefore, a dataset generated by a specific degradation method often has low effectiveness when processing real pictures. In order to process the super-resolution tasks of low-resolution images caused by various causes and improve the generalization ability of the super-resolution model, it is necessary to simulate different low-resolution images with different degradation methods to train the super-resolution model. However, for the super-resolution model installed on the terminal side, due to the limitations of the terminal-side hardware, the model structure of the terminal-side super-resolution is usually simpler than that of the cloud super-resolution model. Therefore, when training with the above low-resolution images and corresponding high-resolution images as a training set, there are problems such as low training effect of the model and slow learning speed.

[0003] The method for training an image super-resolution model in Patent Document CN112488924A includes the steps of obtaining a low-resolution image, a corresponding real high-resolution image, and a real visible light image to form a training sample set; inputting the low-resolution image in the training sample set into a preset image super-resolution model to obtain a candidate high-resolution image; performing image mode conversion on the candidate high-resolution image and the real high-resolution image respectively to obtain a first visible light image and a second visible light image; constructing a loss function based on the differences between the first visible light image, the second visible light image and the real visible light image, and the differences between the candidate high-resolution image and the real high-resolution image; and performing model training on the preset image super-resolution model based on the loss function to obtain a preset image super-resolution model with completed training.

[0004] In the related art, it does not consider constructing training images with different degradation causes. When processing real-world images, it is impossible to output super-resolution images with relatively good effects for low-resolution images with different degradation causes, and the generalization ability of the model is low.

Summary of the Invention

Problems to be Solved by the Invention

[0005] The present invention provides a training set of low-resolution images and corresponding high-resolution images applicable to a model with a simple structure to solve problems such as low training effect and slow learning speed of the model during training to improve the generalization ability of the model when training on the terminal side model. The training set has a relatively good training effect, a fast learning speed, and a strong generalization ability of the trained model. In addition, the present invention provides a model training method based on the above training set and an image super-resolution model obtained from the training.

Means for Solving the Problems

[0006] In view of the existing limitations described above, the present invention proposes a method for generating an image super-resolution dataset, step S101 of constructing a high-resolution image set, performing image blind degradation processing on the high-resolution image HR1 of the high-resolution image set to obtain a corresponding low-resolution image LR1, obtaining an LR1-HR1 data pair, and performing the above operation on all the high-resolution images HR1 of the high-resolution image set to obtain an LR1-HR1 dataset in step S102; training a first model, which is an image super-resolution model, using the LR1-HR1 dataset, and obtaining and saving the model parameters of the first model after the training is completed in step S103; step S104 of constructing a low-resolution image set, inputting the low-resolution image LR2 in the low-resolution image set into the first model having the model parameters to obtain a super-resolution image SR2, thereby obtaining an LR2-SR2 data pair, and performing the above operation on all the low-resolution images LR2 in the low-resolution image set to obtain an LR2-SR2 dataset in step S105.

[0007] Furthermore, step S103 includes step S1031 of inputting the low-resolution image LR1 into the first model, and the first model outputs a super-resolution image SR1; step S1032 of calculating a loss function using the high-resolution image HR1 and the super-resolution image SR1; step S1033 of saving the model parameters when the loss function is less than a first preset threshold; repeating steps S1031 to S1033 for each LR1-HR1 data pair in the training set, and obtaining the model parameters of the first model after completion in step S1034.

[0008] Furthermore, the image blind degradation process selects and executes one or more of the following operations on the high-resolution image HR1 according to a random selection method: Fuzzy operation: According to the random selection method, select and operate one or two of Gaussian fuzzy and Sinc filter fuzzy; Zoom operation: According to the random selection method, select and operate one or more of bilinear interpolation, bicubic interpolation, and region interpolation; Noise superposition operation: According to the random selection method, select and operate one or two of Gaussian noise and Poisson noise; Picture compression operation: Compress the image, and the compression rate factor of the compression is 30% - 95%; The image blind degradation process is executed once, or the image blind degradation process is repeatedly executed twice.

[0009] Furthermore, for the random selection method, a random score between 0 and 1 is randomly given to all options. If the random score of an option is less than a second preset threshold, the operation is not performed. Normalize all random scores greater than or equal to the second preset threshold as the weight of the corresponding option, execute the operation corresponding to the option whose random score is greater than or equal to the second preset threshold, and perform weighted calculation according to the weight on the execution results of all options to obtain an output result.

[0010] Furthermore, in step S103, the first model includes an ESRGAN model, a SwinIR model, and a HAT model. The ESRGAN model, the SwinIR model, and the HAT model are respectively trained using the LR1-HR1 dataset, and the model parameters corresponding to each model are saved. In step S105, the low-resolution image LR2 is respectively input into the ESRGAN model, the SwinIR model, and the HAT model, and weighted fusion is performed on the obtained results according to preset weights to obtain the super-resolution image SR2. Regarding the preset weights, the weight of the ESRGAN model is 0.2, the weight of the SwinIR model is 0.4, and the weight of the HAT model is 0.4.

[0011] Furthermore, in the steps S101 to S102, the obtained LR1-HR1 dataset includes n types of sub-training sets. In the step S103, when performing model training on the first model, each of the n types of sub-training sets is used for training to obtain n groups of sub-model parameters. In the step S105, for the k groups of sub-model parameters, one group of sub-model parameters is sequentially selected as the model parameters of the first model, the low-resolution image LR2 is input, an output result is obtained, and weighted fusion is performed on all the output results according to the preset weights to obtain the super-resolution image SR2.

[0012] Furthermore, in the steps S101 to S102, the obtained LR1-HR1 dataset includes one base training set and k types of sub-training sets. In the step S103, the first model is trained using the base training set to obtain base model parameters. The first model with base model parameters is trained using k types of sub-training sets respectively to obtain k groups of sub-model parameters. In the step S105, for the k groups of sub-model parameters, one group of sub-model parameters is sequentially selected as the model parameters of the first model, the low-resolution image LR2 is input, an output result is obtained, and weighted fusion is performed on all the output results according to the preset weights to obtain the super-resolution image SR2.

[0013] A method for training an image super-resolution model, which uses the LR2-SR2 dataset obtained by the above method to train a second model, the second model being an image super-resolution model, and the training method being Step S201 of inputting an LR2 image into the second model and the second model outputting an SR2' image, Calculating a loss function using the super-resolution image SR2 and the super-resolution image SR2', and when the loss function is less than a third preset threshold, saving model parameters in step S202, Repeating steps S201 to S202 for each LR2-SR2 data pair in the LR2-SR2 dataset, and after completion, obtaining the model parameters of the second model in step S203.

[0014] Furthermore, the calculation method of the loss function is Calculating the L1 loss function, the GAN loss function, and the perceptual loss function respectively, Performing weighted calculation according to preset weights on the results of the above calculations to obtain the loss function.

[0015] An image super-resolution model obtained by training the second model by the above method, the second model being an ECBSR model.

Advantages of the Invention

[0016] Compared with related technologies, the present invention has the following advantages.

[0017] The method for generating an image super-resolution dataset according to one aspect of the present invention is an image blind degradation processing method. A blind degradation high-low resolution image LR1-HR1 dataset is obtained and used as the training set for the first model. By simulating the reduction of image resolution due to different causes in the real world through blind image degradation, the first model can learn a strong generalization ability, that is, it has a relatively good effect when processing low-resolution images caused by various degradation causes in the real world. Using the trained first model to infer the low-resolution image LR2 to obtain the LR2-SR2 dataset, and using the LR2-SR2 dataset to train the model to be trained. Even if the structure of the model to be trained is simple, it can quickly approximate the first model in the way of knowledge transfer, and thereby quickly learn the generalization ability of the first model.

[0018] The model training method according to one aspect of the present invention is to train the second model with the LR2-SR2 dataset obtained by inferring the low-resolution image LR2 using the first model. In the way of knowledge transfer, the second model can quickly approximate the first model, and thereby quickly learn the generalization ability of the first model.

[0019] The image super-resolution model according to one aspect of the present invention is to train the second model with the LR2-SR2 dataset obtained by inferring the low-resolution image LR2 using the first model. Even if the structure of the second model is simple, the second model can quickly approximate the first model, and thereby quickly learn the generalization ability of the first model. The trained second model has a strong generalization ability. Therefore, in the actual situation, the low-resolution problem of images caused by various reasons can be processed with a small model structure, perform restoration and super-resolution on the above pictures, and obtain an excellent super-resolution effect.

Brief Description of the Drawings

[0020]

Figure 1

Figure 2

Figure 3

Figure 4

Embodiments for Carrying Out the Invention

[0021] To make the objectives, technical solutions and advantages of the present invention more clear, the present invention will be described in more detail below. It should be understood that the descriptions herein are only for interpreting the present invention and not for limiting the scope of the present invention.

[0022] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the art. The terms used in the specification of the present invention are only for explaining specific embodiments and not for limiting the present invention. All the expression means related to this specification can refer to the related descriptions of the prior art and will not be repeatedly described herein.

[0023] To further understand the present invention, the present invention will be described in more detail below in combination with the optimal embodiments.

[0024] Example 1

[0025] As shown in FIG. 1, a method for generating an image super-resolution dataset, comprising: Step S101 of constructing a high-resolution image set; Performing image blind degradation processing on the high-resolution image HR1 of the high-resolution image set to obtain a corresponding low-resolution image LR1, obtaining an LR1-HR1 data pair, and performing the above operation on all the high-resolution images HR1 to obtain an LR1-HR1 dataset, step S102; Step S103 of training a first model, which is an image super-resolution model, using the LR1-HR1 dataset, and obtaining and saving the model parameters of the first model after the training is completed; Step S104 of constructing a low-resolution image set; Step S105 of inputting a low-resolution image LR2 in the low-resolution image set into the first model having the model parameters to obtain a super-resolution image SR2, thereby obtaining an LR2-SR2 data pair, and performing the above operation on all the low-resolution images LR2 to obtain an LR2-SR2 dataset.

[0026] When training a super-resolution model, the method of generating high- and low-resolution data pairs in the training set usually uses a method of degrading high-resolution pictures. Related degradation methods generally assume a predefined degradation process from high-resolution images to low-resolution images, but this method is difficult to hold for real images with complex degradation types. To solve the above problems, blind degradation methods have emerged. That is, an uncertain degradation process is used to complete the transformation from high-resolution images to low-resolution images. The training set generated by this blind degradation method can enable the model to learn a strong generalization ability, thereby better processing real images with complex degradation types. However, the training set generated by this blind degradation method has high requirements for the model structure and hardware computing power during training. Therefore, the first model can have a complex model structure, and thus can quickly learn a strong generalization ability from the training set by blind degradation. After learning by the first model, the actual low-resolution picture LR2 is used, and through the inference of the first model, the LR2-SR2 dataset is generated as the training set, which is used for the learning of a model with a simple structure, making the inference effect of the model with a simple structure approximate to that of a model with a complex structure and strong learning ability, and completing the knowledge transfer from a complex model to a simple model. Therefore, the training set has high training efficiency, high speed, good effect, and is applicable to the training of a model with a simple structure. The blind degradation method may be a degradation method based on the CMDSR framework, or a learning degradation method based on the adaptation of the CycleGAN framework, etc.

[0027] The first model is an image super-resolution model, including but not limited to the ESRGAN model, SwinIR model, HAT model, etc.

[0028] In the image blind degradation processing method, a blind degradation high and low resolution image LR1-HR1 dataset is obtained and used as the training set of the first model. The blind image degradation is used to simulate the reduction of image resolution caused by different real-world causes, so that the first model can learn a strong generalization ability, that is, it has a relatively good effect when processing low-resolution images caused by various real-world degradation causes. The trained first model is used to infer the low-resolution image LR2 to obtain the LR2-SR2 dataset, and the LR2-SR2 dataset is used for training the model to be trained. Even if the structure of the model to be trained is simple, it can quickly approximate the first model in the way of knowledge transfer, so that the generalization ability of the first model can be quickly learned.

[0029] Embodiment 2

[0030] As shown in FIGS. 1-2, based on Embodiment 1, further, step S103 is as follows: Step S1031 of inputting the low-resolution image LR1 into the first model and the first model outputting the super-resolution image SR1; Step S1032 of calculating a loss function using the high-resolution image HR1 and the super-resolution image SR1; Step S1033 of saving the model parameters when the loss function is less than a first preset threshold; Repeating steps S1031-S1033 for each LR1-HR1 data pair in the training set, and after completion, step S1034 of obtaining the model parameters of the first model.

[0031] The numerical range of the first preset threshold is 0.01 to 0.05.

[0032] Furthermore, the image blind degradation processing is to select and execute one or more of the following operations on the high-resolution image HR1 according to a random selection method. Fuzzy operation: According to the random selection method, one or two of Gaussian fuzzy and Sinc filter fuzzy are selected for operation. Zoom operation: According to the random selection method, one or more of bilinear interpolation, bicubic interpolation, and area interpolation are selected for operation. Noise superposition operation: According to the random selection method, one or two of Gaussian noise and Poisson noise are selected for operation. Picture compression operation: Compress the image, and the compression ratio factor of the compression is 30% - 95%. The image blind degradation process is executed once, or the image blind degradation process is repeatedly executed twice.

[0033] The Sinc filter is an ideal low-pass filter and is used to remove the high-frequency part of a signal from the spectrum. Its design is based on the Sinc function, that is, sinc(t) = sin(πt) / πt, where t represents time. The Sinc function has a very smooth frequency response in the frequency domain, but due to its infinite-length time-domain response, it cannot be directly used in actual applications.

[0034] Regarding the realization of the Sinc filter, the filter coefficients are defined in the form of the Sinc function in the frequency domain, and the time-domain response of the filter is obtained by discretizing these coefficients. In the filter design, the performance of the filter can be controlled by adjusting parameters such as the cut-off frequency and the filter size.

[0035] By increasing the Sinc filtering, the Sinc filter increases the artifacts by setting the ringing and overshoot artifact phenomena mimicked by different factors, thereby removing the vibration artifacts in the picture after training.

[0036] Gaussian fuzzing, also known as Gaussian smoothing, and the Gaussian fuzzing process of an image is to convolve the image with a normal distribution. The value of the original pixel has the maximum Gaussian distribution value and the maximum weight, and the weights of adjacent pixels decrease as the distance from the original pixel increases. In Gaussian fuzzing, the mathematical expression of the normal distribution is used.

[0037] The Sinc filter is an ideal electronic filter that removes all signal components on a predetermined bandwidth and leaves only the low-frequency signal. In the digital signal field, the normalized Sinc function is defined as follows. TIFF2025516410000002.tif16170

[0038] Bilinear interpolation refers to first using linear interpolation in one direction and then using linear interpolation in another direction to perform bilinear interpolation. Bilinear interpolation calculates the value of one pixel using four pixel points.

[0039] Bicubic interpolation calculates the value of one pixel using 16 surrounding pixel points. Compared with bilinear interpolation, bicubic interpolation has a more complex interpolation formula and performs a smoothing process on the pixels during the process. In actual use, for example, the bicubic interpolation method based on the BiCubic basis function can be selected, but it is not limited to these.

[0040] Region interpolation refers to the reintegration of data from one group of surfaces (source surfaces) to another group of surfaces (target surfaces).

[0041] Gaussian noise refers to noise whose probability density function follows a Gaussian distribution (i.e., a normal distribution), usually sensor noise caused by poor lighting and high temperature. By adding Gaussian noise to an image, the noise caused by the above reasons can be simulated.

[0042] Poisson Noise is noise whose probability density follows a Poisson distribution.

[0043] Picture compression is a method of compressing images based on picture compression algorithms, including, but not limited to, lossy compression such as JPEG and JP2, and lossless compression such as TIFF, PNG, and GIF.

[0044] The compression rate factor = the size of the file after compression / the size of the file before compression.

[0045] In each picture compression operation, a random number between 30% and 95% is randomly selected as the compression rate factor.

[0046] The fact that the image blind degradation process is repeatedly executed twice means that, as shown in FIG. 4, the result of the first image blind degradation process is input and substituted into the second blind degradation process.

[0047] In the related art, a blind degradation dataset is generated using an unknown degradation kernel. Usually, a continuously modified degradation kernel is used, and there is no fixed degradation kernel in the process, so it can be called blind degradation. However, such continuous modification of the degradation kernel causes a large amount of computation. The image blind degradation processing method used in the present invention simulates the causes of real-world degradation in a random selection manner and can realize blind degradation processing with less computation.

[0048] Furthermore, for the random selection method, a random score between 0 and 1 is randomly given to all options. If the random score of an option is less than a second preset threshold, the operation is not performed. All random scores greater than or equal to the second preset threshold are normalized to be the weights of the corresponding options. The operation corresponding to the option whose random score is greater than or equal to the second preset threshold is executed, and weighted calculation is performed according to the weights for the execution results of all options to obtain the output result.

[0049] The normalization refers to redistributing all random scores that are greater than or equal to the second preset threshold so that their sum is equal to 1. For example, if the second preset threshold is 0.7 and the scores greater than or equal to 0.7 are [0.8, 0.9], then through normalization, they are converted to [0.8 / 1.7, 0.9 / 1.7], and no operation is performed on the options corresponding to other scores less than 0.7.

[0050] Performing weighted calculation according to the weights for the execution results of all options means multiplying the output values corresponding to all options by the weights of the respective options and then superimposing them. The corresponding output values refer to the same position numbers of the output values (based on the representation method, for example, the matrix numbers of a matrix, etc.). The concepts of "weighted calculation" below are all the same.

[0051] By the above random selection method and the setting of the second preset threshold, the cause of degradation can be randomly simulated to achieve blind degradation of the image. The second preset threshold is 0.4 - 0.8.

[0052] The dataset generated by the above blind degradation can improve the generalization ability, representation ability, and image reconstruction accuracy of the first model.

[0053] Furthermore, in step S103, the first model includes the ESRGAN model, the SwinIR model, and the HAT model. The ESRGAN model, the SwinIR model, and the HAT model are respectively trained using the LR1 - HR1 dataset, and the model parameters corresponding to each model are saved. In step S105, the low - resolution image LR2 is input into the ESRGAN model, the SwinIR model, and the HAT model respectively, and weighted fusion is performed on the obtained results according to the preset weights to obtain the super - resolution image SR2. Regarding the preset weights, the weight of the ESRGAN model is 0.2, the weight of the SwinIR model is 0.4, and the weight of the HAT model is 0.4.

[0054] The weighted fusion refers to adding, after multiplying the predicted values of components with the same number (position) output from each model by the corresponding weights. The following "weighted fusions" are all the same.

[0055] The ESRGAN model has an SRResNet network structure and includes one or more residual high-density blocks (RRDB) modules.

[0056] The SwinIR model is composed of a shallow feature extraction module, a deep feature extraction module, and a reconstruction module.

[0057] The shallow feature extraction module uses one or more convolutional layers.

[0058] The deep feature extraction module includes one or more Swin Transfomer modules with residuals (RSTB). One convolutional layer is connected to the back of all RSTB modules, and one residual connection is connected to the back of the convolutional layer.

[0059] The RSTB module includes one or more STL layers (Swin Transformer Layer). One convolutional layer is connected to the back of all STL layers, and one residual connection is connected to the back of the convolutional layer.

[0060] The reconstruction module fuses the features obtained by the shallow feature extraction module and the deep feature extraction module to reconstruct the image.

[0061] The HAT model includes one or more RSTB modules and one or more groups of residual hybrid attention modules.

[0062] The residual hybrid attention module group includes one or more hybrid attention modules (HAB, hybrid attention block). One or more channel attention modules are connected to the subsequent stage of all the hybrid attention modules. One convolutional layer is connected to the subsequent stage of all the channel attention modules. One residual connection is connected to the subsequent stage of the convolutional layer.

[0063] The channel attention module includes a convolutional module group including one convolutional layer, one activation function, and one convolutional layer, and a residual module group including one global pooling layer, one convolutional layer, one activation function, and one convolutional layer.

[0064] The hybrid attention module includes those that add a channel attention module (CAB) based on the STL layer and a multi-head self-attention layer based on the shifted window (SW-MSA).

[0065] The RSTB module includes one or more Swin Transformer Layers (STL layers). One convolutional layer is connected to the subsequent stage of all the STL layers. One residual connection is connected to the subsequent stage of the convolutional layer.

[0066] A residual connection is drawn between the convolutional module group and the residual module group, and the residual connection is connected to the subsequent stage of the residual module group.

[0067] The structure of the above model can be adjusted as needed. For example, a certain layer can be increased or decreased, or the number of layers can be increased or decreased, etc. It can be adjusted within the framework of the basic model.

[0068] The ESRGAN model is mainly excellent in solving the problems of detail fuzziness and artifacts.

[0069] The SwinIR model has a stronger local representation ability, can achieve higher performance using less information, can recover high-frequency details, reduce the fuzzy effect, and generate sharp and natural edges.

[0070] The HAT model can recover more and clearer details. When there are many repetitive textures, HAT has significant advantages. In text recovery, HAT can recover clearer text edges compared to other methods.

[0071] In the above, the first model is formed by weighted fusion of one or more of the ESRGAN model, SwinIR model, and HAT model.

[0072] By fusing the above-mentioned multiple complex blind super-resolution big models, the differences in the learning ability and feature representation of different super-resolution models are utilized to complete the fusion of diverse semantic information, thereby obtaining a more accurate inference effect.

[0073] In the steps S101 - S102, the obtained LR1-HR1 dataset includes n types of sub-training sets. In the step S103, when performing model training on the first model, each of the n types of sub-training sets is used for training to obtain n groups of sub-model parameters. In the step S105, for the k groups of sub-model parameters, one group of sub-model parameters is sequentially selected as the model parameters of the first model, the low-resolution image LR2 is input, the output result is obtained, and weighted fusion is performed on all the output results according to the preset weights to obtain the super-resolution image SR2.

[0074] The n types of sub-training sets refer to training sets for certain types of features. For example, a specific training set for the processing of the face of a portrait, a specific training set for the processing of characters, etc. Depending on application needs, it is not limited to the above types. The first model is trained by the training sets for the classified specific types, and the trained first model can have a relatively good processing effect in this field. After training the first model with a specific type of training set, the model parameters corresponding to that type are obtained and are called sub-model parameters.

[0075] Therefore, in the above, after training the first model with multiple sub-training sets to obtain multiple model parameters, when the first model makes inferences, the output results corresponding to different sub-model parameters are weighted and fused to obtain the output of the first model.

[0076] If necessary, one or more types of sub-model parameters are selected, and by weighting and fusing the results of different sub-model parameters, a specific feature can be processed and has a relatively good effect. For example, after weighting and fusing the results of using sub-model parameters for training on portraits and sub-model parameters for character images, the same image with portraits and characters can be processed and has a relatively good super-resolution effect. n is the quantity of types that need to be trained, n is an integer, and n = 2 to 6.

[0077] Furthermore, In the steps S101 to S102, the obtained LR1-HR1 dataset includes one base training set and k types of sub-training sets. In the step S103, the first model is trained using the base training set to obtain base model parameters. Using k types of sub-training sets to train the first model with base model parameters respectively to obtain k groups of sub-model parameters, where k = 2 to 6.

[0078] In step S105, for the k groups of sub-model parameters, one group of sub-model parameters is selected in sequence as the model parameters of the first model, the low-resolution image LR2 is input, the output result is obtained, and weighted fusion is performed on all the output results according to the preset weights to obtain the super-resolution image SR2.

[0079] The base training set is a training set with various types of features, which can achieve a certain level of output effect when performing super-resolution inference after training the model. The sub-training set is a training set for certain types of features, for example, a specific training set for processing the faces of human portraits, a specific training set for processing characters, etc., and is not limited to the above types according to application needs. Training the first model with the training sets for the above classified specific types, the first model has a relatively good processing effect in this field.

[0080] The base model obtained after training the first model using the base training set refers to the first model having the base model parameters obtained after training with the base training set. By further training the above model using the sub-training set for a specific type, the first model can quickly learn in this field and achieve a relatively good effect, thereby shortening the learning time and improving the learning efficiency. This process is also called fine-tuning of the model, which is a training method for quickly training the model to adapt to tasks in different fields.

[0081] One or more sub-training sets can be selected as needed to train the base model respectively. After the training is completed, the corresponding sub-model parameters of multiple groups can be obtained. Furthermore, the output results of different sub-model parameters can be weighted and fused, which has a relatively good effect when processing specific tasks and can shorten the training time at the same time.

[0082] Fusing the output results of the above different models, fusing the output results during inference after training with a classification training set, and further training a specific classification based on the base model to perform fine-tuning of the model, and fusing the output results of the models (having different model parameters) obtained from the training of each classification during inference are all optimization methods of the model, which can improve the inference quality of the model and make the inferred super-resolution image closer to the actual high-resolution image. Regarding the above three optimization methods, they can be used in combination as needed, or used alone, and the accuracy of the output results of the super-resolution model can be improved.

[0083] Furthermore, the calculation method of the loss function is Calculate the L1 loss function, GAN loss function, and perceptual loss function respectively, and obtain the loss function by performing weighted calculation according to preset weights on the results of the above calculations.

[0084] The L1 loss function (MAE) is also called the mean absolute error and refers to the average value of the absolute difference between the model prediction value f(x) and the actual value y. The formula is as follows: TIFF2025516410000003.tif17170 Here, i f(x i ) and y respectively represent the prediction value and the corresponding actual value of the i-th sample, and n is the number of samples.

[0085] Perceptual loss function: Output image I and original high-resolution image I HRInput into a single differentiable function, and the formula for the perceptual loss function is as follows: TIFF2025516410000004.tif13170 The formula for the GAN loss function is as follows: TIFF2025516410000005.tif13170 Here, the GAN loss function consists of two parts: a discriminative network and a generative network. The V in V(D, G) is the symbol indicating a loss function, D indicates the discriminative network, and G indicates the generative network.

[0086] Example 3

[0087] As shown in Figure 3, based on Example 1 or 2, a model training method is provided. Using the LR2-SR2 dataset obtained by the method described in Example 1 or 2, a second model is trained. The second model is an image super-resolution model, and the training method includes: Step S201: Input an LR2 image into the second model, and the second model outputs an SR2' image; Step S202: Calculate a loss function using the super-resolution image SR2 and the super-resolution image SR2'. If the loss function is less than a third preset threshold, save the model parameters; Step S203: Repeat steps S201 - S202 for each LR2-SR2 data pair in the LR2-SR2 dataset. After completion, obtain the model parameters of the second model.

[0088] Furthermore, the calculation method of the loss function is: Calculate the L1 loss function, the GAN loss function, and the perceptual loss function respectively; Perform weighted calculation according to preset weights on the results of the above calculations to obtain the loss function.

[0089] The numerical range of the third preset threshold is 0.01 - 0.05.

[0090] The second model is trained by the LR2-SR2 dataset obtained by inferring the low-resolution image LR2 using the first model. The second model can quickly approximate the first model, thereby quickly learning the generalization ability of the first model. The second model may be an image super-resolution model with a simple structure.

[0091] Example 4

[0092] An image super-resolution model obtained by training the second model by the method of Example 3, wherein the second model is an ECBSR model.

[0093] The second model is trained by the LR2-SR2 dataset obtained by inferring the low-resolution image LR2 using the first model. Even when the structure of the second model is simple, the second model can quickly approximate the first model, thereby quickly learning the generalization ability of the first model. The trained second model has strong generalization ability. Therefore, in actual situations, the problem of low resolution of images caused by various reasons can be processed with a small model structure, perform restoration and super-resolution on the above pictures, and obtain excellent super-resolution effects.

[0094] In the embodiments of the present invention, the method for generating the image super-resolution dataset of the present invention, the model training method, and the model can be used for the training and learning of the terminal-side model. Understandably, the above method and model are not limited to the above applications and can be used in all application scenarios where it is necessary to improve the generalization ability of the model and the training efficiency of the model.

[0095] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principle of the present invention should all be included within the protection scope of the present invention.

Claims

1. A method for generating a super-resolution image dataset, comprising: a step S101 of constructing a high-resolution image set; performing an image blind degradation process on a high-resolution image HR1 of the high-resolution image set to obtain a corresponding low-resolution image LR1, obtaining an LR1-HR1 data pair, and performing the above operation on all the high-resolution images HR1 of the high-resolution image set to obtain an LR1-HR1 data set in step S102; training a first model, which is an image super-resolution model, using the LR1-HR1 data set, and obtaining and storing model parameters of the first model after the training is completed in step S103; a step S104 of constructing a low-resolution image set; inputting a low-resolution image LR2 of the low-resolution image set into the first model having the model parameters to obtain a super-resolution image SR2, thereby obtaining an LR2-SR2 data pair, and performing the above operation on all the low-resolution images LR2 of the low-resolution image set to obtain an LR2-SR2 data set in step S105. A method for generating a super-resolution image dataset, characterized by including the above steps.

2. The step S103 includes: a step S1031 of inputting a low-resolution image LR1 into the first model, and the first model outputs a super-resolution image SR1; a step S1032 of calculating a loss function using the high-resolution image HR1 and the super-resolution image SR1; a step S1033 of storing model parameters when the loss function is less than a first preset threshold; repeating steps S1031 to S1033 for each LR1-HR1 data pair of the LR1-HR1 data set, and obtaining the model parameters of the first model after completion in step S1034. The method according to claim 1, characterized by including the above steps.

3. The image blind degradation process is to select and execute any one or more of the following operations on the high-resolution image HR1 according to a random selection method: Fuzzy operation: Select and operate one or two of Gaussian fuzzy and Sinc filter fuzzy according to a random selection method; Zoom operation: Select and operate one or more of bilinear interpolation, bicubic interpolation, and area interpolation according to a random selection method. Noise superposition operation: According to the random selection method, one or two of Gaussian noise and Poisson noise are selected and operated on. Picture compression operation: Compress the image, and the compression rate factor of the compression is 30% to 95%. The method according to claim 1, characterized in that the image blind degradation process is executed once, or the image blind degradation process is repeatedly executed twice.

4. Regarding the random selection method, a random score between 0 and 1 is randomly assigned to all options. If the random score of an option is less than a second preset threshold, the operation is not performed. All random scores greater than or equal to the second preset threshold are normalized to be the weights of the corresponding options. The operation corresponding to the option whose random score is greater than or equal to the second preset threshold is executed, and weighted calculation is performed according to the weights for the execution results of all options to obtain an output result. The method according to claim 3, characterized in that.

5. In step S103, the first model includes an ESRGAN model, a SwinIR model, and a HAT model. The ESRGAN model, the SwinIR model, and the HAT model are respectively trained using the LR1-HR1 dataset, and the model parameters corresponding to each model are saved. In step S105, the low-resolution image LR2 is respectively input into the ESRGAN model, the SwinIR model, and the HAT model, and weighted fusion is performed on the obtained results according to preset weights to obtain the super-resolution image SR2. Regarding the preset weights, the weight of the ESRGAN model is 0.2, the weight of the SwinIR model is 0.4, and the weight of the HAT model is 0.

4. The method according to claim 1, characterized in that.

6. In steps S101 to S102, the obtained LR1-HR1 dataset includes n types of sub-training sets. In step S103, when performing model training on the first model, the n types of sub-training sets are respectively used for training to obtain n groups of sub-model parameters. In step S105, for the sub-model parameters of the k groups, one group of sub-model parameters is sequentially selected as the model parameters of the first model, the low-resolution image LR2 is input, an output result is obtained, and weighted fusion is performed on all the output results according to preset weights to obtain the super-resolution image SR2. The method according to claim 1, characterized in that.

7. In steps S101 to S102, the obtained LR1-HR1 dataset includes one base training set and k types of sub-training sets. In step S103, the first model is trained using the base training set to obtain base model parameters, and the first model with base model parameters is trained using k types of sub-training sets respectively to obtain k groups of sub-model parameters. In step S105, for the sub-model parameters of the k groups, one group of sub-model parameters is sequentially selected as the model parameters of the first model, the low-resolution image LR2 is input, an output result is obtained, and weighted fusion is performed on all the output results according to preset weights to obtain the super-resolution image SR2. The method according to claim 1, characterized in that.

8. A method for training an image super-resolution model, using the LR2-SR2 dataset obtained by the method according to any one of claims 1 to 7 to train a second model, wherein the second model is an image super-resolution model, and the training method is as follows: Step S201 of inputting an LR2 image into the second model and the second model outputting an SR2' image; Step S202 of calculating a loss function using the super-resolution image SR2 and the super-resolution image SR2', and saving the model parameters when the loss function is less than a third preset threshold; Repeating steps S201 to S202 for each LR2-SR2 data pair in the LR2-SR2 dataset, and after completion, obtaining the model parameters of the second model in step S203. A method for training an image super-resolution model, characterized by comprising the above steps.

9. The calculation method of the loss function is as follows: Calculating the L1 loss function, the GAN loss function, and the perceptual loss function respectively. The method according to claim 2 or 8, characterized in that a weighted calculation is performed according to a preset weight with respect to the result of the above calculation to obtain a loss function.

10. An image super-resolution model, which is obtained by training the second model by the method according to claim 8, and the second model is an ECBSR model.

Citation Information

Patent Citations

  • Method, apparatus, and computer program for improving the reconstruction of high-density super-resolution images from diffraction-limited images acquired by single-molecule localization microscopy

    JP2020529083A

  • Image processing method, processing device and processing device

    JP2021502644A

  • Image processing method and device, electronic device, and storage medium

    JP2021528742A

  • Super-resolution reconstruction method and related device

    JP2023508512A

  • Information processing device, information processing method, and program

    WO2022269963A1