Training data acquisition method and system, and super-division model training method and system
By randomly perturbing, slicing, degrading, and upsampling clear images, high- and low-resolution image pairs are generated, which solves the problem of insufficient training data diversity in existing technologies and improves the training effect and mid-to-low frequency resolution of super-resolution models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies using simulated degradation methods to prepare training data for super-resolution models suffer from significant discrepancies and insufficient diversity, particularly in terms of resolution in the low-to-mid frequency range.
By randomly perturbing the original clear image, extracting local images, and performing degradation and upsampling processes to increase high-fidelity details, high- and low-resolution image pairs are generated for training the super-resolution model.
It improves the diversity and robustness of training data, narrows the gap with real-world scenarios, and enhances the training performance and mid-to-low frequency resolution of the super-resolution model.
Smart Images

Figure CN121639489A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to a training data acquisition method and system and a super-resolution model training method and system. BACKGROUND
[0002] Real-world video super-resolution (VSR) technology aims to convert low-resolution video into high-resolution video to improve visual quality. This technology faces multiple challenges in practical applications, especially in mobile applications, mainly involving training data preparation, computational efficiency, and the complexity of real-world scenarios. The method of training data preparation mainly falls into two categories. One is the simulated degradation method, which uses clear images / videos to synthesize blurred noise images / videos close to real degradation. However, this method often uses downsampling to prepare degraded images, which results in a certain gap between the high and low quality image pairs obtained in this way and the real scene, easily leading to excessive smearing in certain frequency bands. The second is the real scene data collection method, which uses cameras, mobile phones, and other devices with zoom lenses to collect and construct real degraded and high-definition image pairs for training. However, this method has a complex process, and there are problems such as blur caused by parallax and alignment errors of clear images. Moreover, the scale of scene diversity is limited. In practical applications, the two methods are often combined. SUMMARY
[0003] The technical problem to be solved by the present application is to overcome the above-mentioned defects in the prior art based on the simulated degradation method for preparing super-resolution model training data, and to provide a training data acquisition method and system and a super-resolution model training method and system.
[0004] The present application solves the above technical problems by the following technical solutions:
[0005] A training data acquisition method, the training data being used to train a super-resolution model, the acquisition method comprising:
[0006] randomly perturbing an original clear image to obtain a perturbed image;
[0007] cutting out the perturbed image to obtain a local image;
[0008] degrading the local image to obtain a low-resolution image;
[0009] upsampling the local image and adding high-fidelity details to obtain a high-resolution image;
[0010] wherein the low-resolution image and the high-resolution image constitute an image pair in the training data.
[0011] Preferably, the random perturbation processing on the original clear image comprises:
[0012] random scale perturbation processing on the original clear image.
[0013] Preferably, the degradation processing on the local image comprises:
[0014] isotropic Gaussian blur processing and / or anisotropic Gaussian blur processing on the local image;
[0015] and / or,
[0016] adding noise to the local image, wherein the noise comprises sensor noise.
[0017] Preferably, the high-fidelity details comprise image textures, and the up-sampling processing on the local image and the increase of high-fidelity details comprise:
[0018] up-sampling processing on the local image and restoration and enhancement of image textures based on a pre-trained high-fidelity generation model.
[0019] Preferably, the high-fidelity generation model comprises a diffusion model.
[0020] A training method of a super-resolution model, the training method comprising:
[0021] training the super-resolution model using training data, wherein the training data is obtained according to any one of the above training data obtaining methods.
[0022] An obtaining system of training data, the training data being used for training a super-resolution model, the obtaining system comprising:
[0023] a perturbation module configured to perform random perturbation processing on an original clear image to obtain a perturbed image;
[0024] an extraction module configured to perform extraction processing on the perturbed image to obtain a local image;
[0025] a degradation module configured to perform degradation processing on the local image to obtain a low-resolution image;
[0026] a generation module configured to perform up-sampling processing on the local image and increase high-fidelity details to obtain a high-resolution image;
[0027] wherein the low-resolution image and the high-resolution image constitute an image pair in the training data.
[0028] A training system of a super-resolution model, the training system comprising a super-resolution model, a quantification module and the above-mentioned training data acquisition system, the training data of the super-resolution model being provided by the acquisition system, the training data comprising image pairs composed of low-resolution images and high-resolution images, for the same image pair:
[0029] The super-resolution model is used to output a super-resolution image based on the low-resolution image;
[0030] The quantification module is used to quantify the difference between the super-resolution image and the high-resolution image.
[0031] An electronic device comprising a memory, a processor and a computer program stored on the memory and executable on the processor, the processor implementing the above-mentioned any one of the training data acquisition method or the above-mentioned super-resolution model training method when executing the computer program.
[0032] A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the steps of the above-mentioned any one of the training data acquisition method or the steps of the above-mentioned super-resolution model training method.
[0033] The positive progress effect of the present application is that the present application optimizes the simulation degradation method for preparing training data of a super-resolution model in the prior art, proposes a generative training data preparation method, specifically by randomly disturbing an original clear image to obtain a disturbed image, by cutting the disturbed image to obtain a local image, and then by degrading the local image to obtain a low-resolution image and by upsampling the local image and adding high-fidelity details to obtain a high-resolution image, a high-low resolution image pair for super-resolution model training is obtained. In this way, the present application can make full use of existing large-scale public data sets and can avoid the gap between the training data and the real scene as much as possible, thereby improving the training effect and practicality of the super-resolution model and significantly improving the resolving power of the super-resolution model in the medium-low frequency band. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 A flowchart of the training data acquisition method according to embodiment 1 of the present application.
[0035] Figure 2 A module schematic diagram of the training data acquisition system according to embodiment 3 of the present application.
[0036] Figure 3 A module schematic diagram of the super-resolution model training system according to embodiment 4 of the present application. DETAILED DESCRIPTION
[0037] The present invention will be further illustrated by way of embodiments below, but the invention is not limited to the scope of the embodiments described herein.
[0038] Example 1
[0039] This embodiment provides a method for obtaining training data, wherein the training data is used to train a super-resolution model. Figure 1 A flowchart of this embodiment is shown, with reference to... Figure 1 The acquisition method in this embodiment includes:
[0040] S11. Randomly perturb the original clear image to obtain the perturbed image.
[0041] In this embodiment, the original sharp image can be taken from a pre-prepared sharp, noise-free image dataset or video dataset, such as an existing publicly available large-scale image dataset or video dataset. In this embodiment, random perturbation processing of the original sharp image aims to increase the sample size and diversity of the training data, narrow the gap between the training data and the real scene, thereby improving the robustness and generalization ability of the super-resolution model and reducing overfitting. Therefore, the specific implementation of random perturbation processing can be customized according to the actual application. For example, to narrow the scale gap between the training data and the real scene, random scale perturbation processing can be applied to the original sharp image.
[0042] Reference Figure 1 The acquisition method in this embodiment further includes:
[0043] S12. Perform image extraction on the disturbed image to obtain a local image.
[0044] In this embodiment, local images of the region of interest can be obtained by extracting the perturbed image. The region of interest and specific implementation method for extraction can be customized according to the actual application. It should be understood that obtaining local images by extracting perturbed images can further increase the sample size and diversity of the training data. For example, for an original clear image, after M types of random perturbation, M perturbed images can be obtained. If each perturbed image is processed to obtain N local images, then a total of M*N local images can be obtained. In this way, the robustness and generalization ability of the super-resolution model can be further improved, and overfitting can be reduced.
[0045] Reference Figure 1 The acquisition method in this embodiment further includes:
[0046] S13. Perform degradation processing on the local image to obtain a low-resolution image.
[0047] In this embodiment, the degradation processing of local images is intended to simulate actual degradation signals. For example, the degradation processing may include, but is not limited to, at least one of blur degradation processing and noise degradation processing. Blur degradation processing may be implemented by performing isotropic Gaussian blur processing and / or anisotropic Gaussian blur processing on local images. Noise degradation processing may be implemented by adding noise to local images. The noise may include, but is not limited to, sensor noise from the real end side, Poisson Gaussian noise, etc., to obtain a blurry, noisy, low-resolution image.
[0048] Reference Figure 1 The acquisition method in this embodiment further includes:
[0049] S14. Upsample the local image and add high-fidelity details to obtain a high-resolution image.
[0050] In this embodiment, to obtain a clean and clear high-resolution image, not only is local image upsampling processed, but high-fidelity details in the local image are also added to narrow the gap between the training data and the real scene, thereby improving the robustness and generalization ability of the super-resolution model. Therefore, the high-fidelity details to be added and the method of addition can be customized according to the actual application. For example, high-fidelity details can include, but are not limited to, image textures, and adding high-fidelity details can be implemented as restoring and enhancing image textures. Compared to images obtained by only upsampling local images, the high-resolution image obtained in this embodiment is obviously more suitable as a clear, large image for training ground truth. Thus, the low-resolution image and high-resolution image obtained in this embodiment constitute an image pair in the training data, where the high-resolution image serves as the training ground truth.
[0051] Furthermore, in this embodiment, step S14 can be implemented as upsampling the local image and adding high-fidelity details based on a pre-trained high-fidelity generation model that takes fidelity into account, to obtain a high-resolution image. More specifically, when the high-fidelity details include image texture, step S14 can be implemented as upsampling the local image and restoring and enhancing the image texture based on a pre-trained high-fidelity generation model that takes fidelity into account. The high-fidelity generation model may include, but is not limited to, a diffusion model.
[0052] In existing technologies, the general practice for preparing training data is to extract a large image from a clear image as the training ground truth, and then perform simulated degradation processing on the extracted large image. For example, the extracted large image is first subjected to compound blur degradation and Poisson Gaussian noise, and then downsampled to obtain a low-quality small image corresponding to the large image. It should be understood that the texture scale and other characteristics of the extracted large image are often limited in diversity, and there is a certain gap between the large image and the small image and the real scene, which can easily lead to over-smearing of certain frequency bands (such as mid-low frequency bands).
[0053] Based on this, this embodiment optimizes the simulated degradation method for preparing training data for super-resolution models in the prior art and proposes a generative training data preparation method. Specifically, it obtains a perturbed image by randomly perturbing the original clear image, obtains a local image by extracting the perturbed image, obtains a low-resolution image by degrading the local image, and obtains a high-resolution image by upsampling the local image and adding high-fidelity details. This results in a pair of high- and low-resolution images for super-resolution model training. In this way, this embodiment can make full use of existing large-scale public datasets and avoid the gap between training data and real scenes as much as possible, thereby improving the training effect and practicality of the super-resolution model and significantly improving the resolution of the super-resolution model in the mid-to-low frequency band.
[0054] Example 2
[0055] This embodiment provides a training method for a super-resolution model. The training method of this embodiment includes training the super-resolution model using training data. The training data is obtained according to the training data acquisition method provided in Embodiment 1. The training data obtained in Embodiment 1 includes image pairs consisting of low-resolution images and high-resolution images. The high-resolution images are used as training ground truth. The super-resolution model is iteratively trained by calculating the loss between the output image of the super-resolution model based on the input low-resolution image and the corresponding high-resolution image.
[0056] The super-resolution model training method provided in this embodiment obtains training data based on the training data acquisition method provided in Embodiment 1. This training data makes full use of existing large-scale public data and avoids the gap between training data and real scenes as much as possible. Therefore, the training effect and practicality of the super-resolution model trained based on this training data are significantly improved, and the resolution in the mid-to-low frequency band is also significantly improved, making it applicable to real-world image / video super-resolution technology.
[0057] Example 3
[0058] This embodiment provides a system for acquiring training data, wherein the training data is used to train a super-resolution model.Figure 2 A schematic diagram of the modules in this embodiment is shown, with reference to... Figure 2 The acquisition system in this embodiment includes:
[0059] The perturbation module 31 is used to perform random perturbation processing on the original clear image to obtain the perturbed image.
[0060] In this embodiment, the original sharp image can be taken from a pre-prepared sharp, noise-free image dataset or video dataset, such as an existing publicly available large-scale image dataset or video dataset. In this embodiment, random perturbation processing of the original sharp image aims to increase the sample size and diversity of the training data, narrow the gap between the training data and the real scene, thereby improving the robustness and generalization ability of the super-resolution model and reducing overfitting. Therefore, the specific implementation of random perturbation processing can be customized according to the actual application. For example, to narrow the scale gap between the training data and the real scene, random scale perturbation processing can be applied to the original sharp image.
[0061] Reference Figure 2 The acquisition system in this embodiment also includes:
[0062] The extraction module 32 is used to extract parts of the disturbed image to obtain a local image.
[0063] In this embodiment, local images of the region of interest can be obtained by extracting the perturbed image. The region of interest and specific implementation method for extraction can be customized according to the actual application. It should be understood that obtaining local images by extracting perturbed images can further increase the sample size and diversity of the training data. For example, for an original clear image, after M types of random perturbation, M perturbed images can be obtained. If each perturbed image is processed to obtain N local images, then a total of M*N local images can be obtained. In this way, the robustness and generalization ability of the super-resolution model can be further improved, and overfitting can be reduced.
[0064] Reference Figure 2 The acquisition system in this embodiment also includes:
[0065] The degradation module 33 is used to perform degradation processing on local images to obtain low-resolution images.
[0066] In this embodiment, the degradation processing of local images is intended to simulate actual degradation signals. For example, the degradation processing may include, but is not limited to, at least one of blur degradation processing and noise degradation processing. Blur degradation processing may be implemented by performing isotropic Gaussian blur processing and / or anisotropic Gaussian blur processing on local images. Noise degradation processing may be implemented by adding noise to local images. The noise may include, but is not limited to, sensor noise from the real end side, Poisson Gaussian noise, etc., to obtain a blurry, noisy, low-resolution image.
[0067] Reference Figure 2 The acquisition system in this embodiment also includes:
[0068] The generation module 34 is used to upsample the local image and add high-fidelity details to obtain a high-resolution image.
[0069] In this embodiment, to obtain a clean and clear high-resolution image, not only is local image upsampling processed, but high-fidelity details in the local image are also added to narrow the gap between the training data and the real scene, thereby improving the robustness and generalization ability of the super-resolution model. Therefore, the high-fidelity details to be added and the method of addition can be customized according to the actual application. For example, high-fidelity details can include, but are not limited to, image textures, and adding high-fidelity details can be implemented as restoring and enhancing image textures. Compared to images obtained by only upsampling local images, the high-resolution image obtained in this embodiment is obviously more suitable as a clear, large image for training ground truth. Thus, the low-resolution image and high-resolution image obtained in this embodiment constitute an image pair in the training data, where the high-resolution image serves as the training ground truth.
[0070] Furthermore, in this embodiment, the generation module 34 can be implemented to upsample the local image and add high-fidelity details based on a pre-trained high-fidelity generation model that takes fidelity into account, thereby obtaining a high-resolution image. More specifically, when the high-fidelity details include image texture, the generation module 34 can be implemented to upsample the local image and restore and enhance the image texture based on a pre-trained high-fidelity generation model that takes fidelity into account. The high-fidelity generation model may include, but is not limited to, a diffusion model.
[0071] In existing technologies, the general practice for preparing training data is to extract a large image from a clear image as the training ground truth, and then perform simulated degradation processing on the extracted large image. For example, the extracted large image is first subjected to compound blur degradation and Poisson Gaussian noise, and then downsampled to obtain a low-quality small image corresponding to the large image. It should be understood that the texture scale and other characteristics of the extracted large image are often limited in diversity, and there is a certain gap between the large image and the small image and the real scene, which can easily lead to over-smearing of certain frequency bands (such as mid-low frequency bands).
[0072] Based on this, this embodiment optimizes the simulation degradation system for preparing training data for super-resolution models in the prior art and proposes a generative training data preparation system. Specifically, it randomly perturbs the original clear image to obtain a perturbed image, performs cropping processing on the perturbed image to obtain a local image, then performs degradation processing on the local image to obtain a low-resolution image, and then performs upsampling processing on the local image to add high-fidelity details to obtain a high-resolution image. This results in a pair of high- and low-resolution images for super-resolution model training. In this way, this embodiment can make full use of existing large-scale public datasets and can avoid the gap between training data and real scenes as much as possible, thereby improving the training effect and practicality of the super-resolution model and significantly improving the resolution of the super-resolution model in the mid-to-low frequency band.
[0073] Example 4
[0074] This embodiment provides a training system for a super-resolution model. Figure 3 A schematic diagram of the modules in this embodiment is shown. (Refer to...) Figure 3 The training system of this embodiment includes a super-resolution model 41, a quantization module 42, and a training data acquisition system provided in embodiment 3. The training data of the super-resolution model 41 is provided by the acquisition system. The training data provided by the acquisition system includes image pairs consisting of low-resolution images and high-resolution images. The high-resolution images are used as training ground truth. For the same image pair, the super-resolution model 41 is used to output a super-resolution image based on the low-resolution image. The quantization module 42 is used to quantize the difference between the super-resolution image and the high-resolution image, thereby realizing iterative training of the super-resolution model.
[0075] The super-resolution model training system provided in this embodiment obtains training data based on the training data acquisition system provided in Embodiment 3. This training data makes full use of existing large-scale public data and avoids the gap between training data and real scenes as much as possible. Therefore, the training effect and practicality of the super-resolution model trained based on this training data are significantly improved, and the resolution in the low and medium frequency bands is also significantly improved. It can be applied to real-world image / video super-resolution technology.
[0076] Example 5
[0077] This embodiment provides an electronic device, which may include a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it can implement the training data acquisition method provided in Embodiment 1 or the super-resolution model training method provided in Embodiment 2.
[0078] Example 6
[0079] This embodiment provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the program implements the steps of the training data acquisition method provided in Embodiment 1 or the steps of the super-resolution model training method provided in Embodiment 2. The readable storage medium may include, but is not limited to, portable disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0080] In a possible implementation, the present invention can also be implemented as a program product comprising program code. When the program product is run on a terminal device, the program code causes the terminal device to execute the steps of the training data acquisition method provided in Embodiment 1 or the training method of the super-resolution model provided in Embodiment 2. The program code for executing the present invention can be written in any combination of one or more programming languages. The program code can be executed entirely on the user device, partially on the user device, as a standalone software package, partially on the user device and partially on a remote device, or entirely on a remote device.
[0081] While specific embodiments of the present invention have been described above, those skilled in the art should understand that these are merely illustrative examples, and the scope of protection of the present invention is defined by the appended claims. Those skilled in the art can make various changes or modifications to these embodiments without departing from the principles and essence of the present invention, but all such changes and modifications fall within the scope of protection of the present invention.
Claims
1. A method for acquiring training data, characterized in that, The training data are used for training a super-resolution model, and the obtaining method comprises: randomly perturbing an original clear image to obtain a perturbed image; performing cutout processing on the perturbed image to obtain a local image; degrading the local image to obtain a low-resolution image; performing up-sampling processing on the local image and adding high-fidelity details to obtain a high-resolution image; wherein the low-resolution image and the high-resolution image constitute an image pair in the training data.
2. The method of claim 1, wherein the training data is obtained by, The random perturbation processing on the original clear image comprises: random scale perturbation processing on the original clear image.
3. The method of claim 1, wherein the training data is obtained by, The degradation processing on the local image comprises: performing isotropic Gaussian blur processing and / or anisotropic Gaussian blur processing on the local image; and / or, adding noise to the local image, wherein the noise comprises sensor noise.
4. The method of claim 1, wherein the training data is obtained from a plurality of users. The high-fidelity details comprise image textures, and the up-sampling processing on the local image and the addition of high-fidelity details comprise: performing up-sampling processing on the local image and recovering and enhancing the image textures based on a pre-trained high-fidelity generation model.
5. The method of claim 4, wherein the training data is obtained by, The high-fidelity generation model comprises a diffusion model.
6. A method for training a super-resolution model, the method comprising: The training method comprises: training the super-resolution model using training data, wherein the training data are obtained according to the training data obtaining method of any one of claims 1-5.
7. An acquisition system of training data, characterized in that, The training data are used for training a super-resolution model, and the obtaining system comprises: a perturbation module configured to randomly perturb an original clear image to obtain a perturbed image; a cutout module configured to perform cutout processing on the perturbed image to obtain a local image; a degradation module configured to degrade the local image to obtain a low-resolution image; a generation module configured to perform up-sampling processing on the local image and add high-fidelity details to obtain a high-resolution image; wherein the low-resolution image and the high-resolution image constitute an image pair in the training data.
8. A system for training a super-resolution model, the system comprising: The training system comprises a super-resolution model, a quantization module, and the training data obtaining system of embodiment 7, wherein the training data of the super-resolution model are provided by the obtaining system, the training data comprise an image pair comprising a low-resolution image and a high-resolution image, and for the same image pair: the super-resolution model is configured to output a super-resolution image based on the low-resolution image; the quantization module is configured to quantify the difference between the super-resolution image and the high-resolution image.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the training data obtaining method of any one of claims 1-5 or the training method of the super-resolution model of claim 6.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the training data obtaining method of any one of claims 1-5 or the steps of the training method of the super-resolution model of claim 6.