Real-world noise image super-resolution data set construction method
By quantifying the noise intensity through ISO parameter control and coordinated with optical zoom, as well as SIFT feature matching and maximum correlation coefficient registration, a strictly aligned real-world noise super-resolution dataset was constructed, which solved the problems of missing noise modeling and dataset limitations in traditional methods and improved the robustness and quality of image restoration.
Patent Information
- Application Number
- CN202510701040.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-30
AI Technical Summary
Existing real-world super-resolution datasets fail to effectively simulate real sensor noise, resulting in noise amplification, artifacts and detail distortion in the reconstructed image, and the datasets do not cover image pairs with controllable noise intensity, making them unable to support the training and evaluation of noise-robust super-resolution models.
Through a noise intensity quantifiable acquisition strategy that coordinates ISO parameter control with optical zoom, combined with SIFT feature matching and maximum correlation coefficient registration, a strictly aligned and noise-intensity controllable real-world noise super-resolution dataset is constructed, including a triplet dataset of noisy low-resolution, clean low-resolution, and high-resolution image patches.
It achieves quantifiable modeling of real sensor noise, improves the robustness of image restoration, and significantly improves the quality of super-resolution reconstruction in real scenes.
Smart Images

Figure CN120725872A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer underlying vision, and in particular relates to a method for constructing a super-resolution dataset of real-world noisy images. Background Art
[0002] Super-resolution technology aims to reconstruct high-resolution images from low-resolution images. In recent years, super-resolution methods based on deep learning have made significant progress on synthetic datasets. However, due to the complex degradation process of low-resolution images in real scenes (including noise, blur, compression artifacts, etc.), models trained based on synthetic degraded data have significant performance degradation in practical applications. This phenomenon is called the "domain difference problem". To solve this problem, real-world super-resolution technology was proposed. Its core is to construct a dataset through real-world image pairs to narrow the domain gap between training and test data. Existing real-world super-resolution datasets collect multi-scale image pairs through optical zoom strategies, but they assume that the input images are noise-free or approximately clean images, ignoring the degradation effects caused by sensor noise in real imaging. In real scenes, due to insufficient lighting, high ISO settings or sensor limitations, low-resolution images often contain complex noise with non-Gaussian distribution. Existing real-world super-resolution research has the following limitations:
[0003] 1) Lack of noise modeling: Existing methods usually assume that the input low-resolution image is noise-free or contains only synthetic Gaussian noise. However, real sensor noise has signal dependence and channel correlation, which is difficult to be effectively removed by traditional denoising algorithms. As a result, the noise is amplified during the super-resolution process, and the reconstructed image has artifacts and detail distortion.
[0004] 2) Dataset Limitations: Current real-world super-resolution datasets lack image pairs where noise intensity can be quantified and controlled, making them incapable of supporting the training and evaluation of noise-robust super-resolution models. While some studies have attempted to superimpose noise on synthetic data, the resulting noise distribution differs from actual sensor noise and fails to reflect the true characteristics of imaging.
[0005] Constructing a dataset containing realistic noise, rigorously registered, and with controllable noise intensity is key to overcoming the bottleneck of real-world noisy image super-resolution technology. The RealNSR dataset proposed in this paper is the first to achieve real-world noise super-resolution data construction with quantifiable noise intensity (controlled by ISO parameters), cross-modal alignment (based on SIFT feature matching and maximum correlation coefficient registration), and color consistency correction. This provides fundamental support for the development of noise-robust super-resolution models. Summary of the Invention
[0006] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the existing technology and provide a method for constructing a super-resolution dataset of real-world noisy images. Through a noise intensity quantifiable acquisition strategy that coordinates ISO parameter control with optical zoom, combined with a cross-modal registration technology based on SIFT feature matching and maximum correlation coefficient registration fusion, the first super-resolution dataset RealNSR that covers real noise distribution, strict alignment and noise intensity grading is constructed. This provides a high-quality benchmark for joint denoising and super-resolution model training, and significantly improves the robustness of image restoration in real scenes.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] In a first aspect, the present invention provides a method for constructing a real-world noisy image super-resolution dataset, comprising the following steps:
[0009] The steps for acquiring images with different noise intensities are as follows: In the same natural scene, a camera is used to simultaneously capture low-resolution images with different noise intensities and their corresponding high-resolution images without noise by adjusting the ISO parameter and optical zoom factor.
[0010] Image registration and alignment steps: Based on SIFT feature matching and the homography matrix optimization method that maximizes the correlation coefficient, the low-resolution image and the high-resolution image are aligned at the pixel level to eliminate the displacement error caused by changes in shooting parameters;
[0011] Cross-modal data processing steps: Screen and eliminate image groups containing moving objects, and based on the characteristics of the camera's image signal processor, correct the color space differences and brightness deviations between the low-resolution and high-resolution images;
[0012] Dataset generation steps: The registered image pairs are layered and cropped according to the preset size to generate a triplet dataset containing noisy low-resolution image patches, clean low-resolution image patches, and high-resolution image patches, and are classified and stored as training sets and test sets according to noise intensity.
[0013] As a preferred technical solution, the multi-noise intensity image acquisition step is specifically as follows:
[0014] Use the camera to shoot low-resolution images at the first focal length, set the ISO parameter to 100 to generate a clean low-resolution image without noise, and set the ISO parameter to 1600 and 3200 to generate low-intensity noise and high-intensity noise low-resolution images respectively;
[0015] The focus is adjusted to the second focal length using optical zoom, and the ISO is reset to 100 to capture a noise-free high-resolution image of the same scene, forming an image pair with a 4x super-resolution magnification compared to the low-resolution image.
[0016] Repeat the above steps to collect raw data in multiple different natural scenes, covering indoor and outdoor environments and static objects.
[0017] As a preferred technical solution, the image registration and alignment step is specifically as follows:
[0018] SIFT feature points are extracted from the high-resolution image HR-ISO100 and the low-resolution images LR-ISO100, LR-ISO1600, and LR-ISO3200, and feature descriptors are calculated.
[0019] Based on the RANSAC algorithm, the matching feature point pairs are selected and the homography matrix is initialized;
[0020] Assume that the high-resolution image is I H , the low-resolution image is I L , the homography matrix τ is iteratively optimized by maximizing the correlation coefficient algorithm, and the objective function is defined as:
[0021]
[0022] Where C is a cropping operation that can transform the low-resolution image I L The field of view is cropped and the high-resolution image I H For the same field of view, α and β are brightness adjustment parameters, ||·|| p is the p-norm;
[0023] The optimized homography matrix is applied to the low-resolution image and resampled by bilinear interpolation to generate noisy low-resolution image patches that are aligned with the high-resolution image.
[0024] As a preferred technical solution, the cross-modal data processing is specifically as follows:
[0025] The aligned image I′ is obtained by maximizing the correlation coefficient algorithm L =C(τ(I L ));
[0026] The calculation of brightness adjustment parameters is as follows:
[0027] α=std(I H ) / std(I′ L ),β=mean(I H )-αmean(I′ L )
[0028] Where mean is the pixel mean and std is the pixel variance;
[0029] Brightness adjustment can ensure I' L and I H have the same pixel mean and variance.
[0030] As a preferred technical solution, the data set generation step is specifically as follows:
[0031] Perform non-overlapping sliding cropping on the registered high-resolution image according to the first pixel size to generate a high-resolution image block;
[0032] Synchronously cropping the low-resolution image to generate a low-resolution image block of a second pixel, retaining a coordinate correspondence with the high-resolution image block;
[0033] Divide the dataset into subsets based on noise intensity:
[0034] Add metadata tags to each image block, including ISO value, noise intensity level, scene category, and crop position coordinates.
[0035] As a preferred technical solution, the data set is divided into multiple subsets according to noise intensity, specifically:
[0036] Low-intensity noise training set: 40,059 groups (LR-ISO1600, LR-ISO100, HR-ISO100) image patches;
[0037] Low-intensity noise test set: 4,082 image patches;
[0038] High-intensity noise training set: 40,010 groups (LR-ISO3200, LR-ISO100, HR-ISO100) image patches;
[0039] High-intensity noise test set: 4,087 image patches.
[0040] As a preferred technical solution, the image block cropping adopts overlapping sliding windows to maximize data utilization.
[0041] In a second aspect, the present invention provides a real-world noisy image super-resolution dataset construction system, which is applied to the real-world noisy image super-resolution dataset construction method, including an image acquisition module, an image registration and alignment module, a cross-modal data processing module, and a dataset generation module;
[0042] The image acquisition module is used to synchronously capture low-resolution images containing different noise intensities and their corresponding noise-free high-resolution images in the same natural scene by using a camera by adjusting ISO parameters and optical zoom multiples;
[0043] The image registration and alignment module is used to perform pixel-level alignment of low-resolution images and high-resolution images based on SIFT feature matching and a homography matrix optimization method that maximizes the correlation coefficient, thereby eliminating displacement errors caused by changes in shooting parameters;
[0044] The cross-modal data processing module is used to filter out image groups containing moving objects and correct the color space difference and brightness deviation between the low-resolution and high-resolution images based on the characteristics of the camera image signal processor;
[0045] The dataset generation module is used to perform hierarchical cropping of the registered image pairs according to a preset size to generate a triplet dataset containing noisy low-resolution image blocks, clean low-resolution image blocks, and high-resolution image blocks, and store them as training sets and test sets according to noise intensity.
[0046] In a third aspect, the present invention provides an electronic device, comprising:
[0047] at least one processor; and,
[0048] a memory communicatively connected to the at least one processor; wherein,
[0049] The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor to enable the at least one processor to perform the method for constructing a real-world noisy image super-resolution dataset.
[0050] In a fourth aspect, the present invention provides a computer-readable storage medium storing a program, which, when executed by a processor, implements the method for constructing a real-world noisy image super-resolution dataset.
[0051] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0052] 1. This paper proposes a realistic noise modeling and controllable data acquisition strategy to achieve quantifiable modeling of real sensor noise. Existing real-world super-resolution datasets only target noise-free or synthetic noise scenarios, ignoring the true distribution characteristics of sensor noise.
[0053] 2. This paper proposes a cross-modal, high-precision registration method based on the fusion of SIFT feature matching and maximum correlation coefficient registration. Global alignment is achieved by initializing the homography matrix through SIFT feature point matching, and local sub-pixel alignment is ensured through iterative optimization based on maximum correlation coefficient. Traditional registration methods (such as single SIFT feature matching) are prone to registration errors in the presence of noise. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0055] Figure 1 This is a flowchart of a method for constructing a real-world noise image super-resolution dataset according to an embodiment of the present invention;
[0056] Figure 2 A block diagram of a system for constructing a real-world noisy image super-resolution dataset according to an embodiment of the present invention.
[0057] Figure 3 2 is a structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0058] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0059] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0060] like Figure 1 As shown, this embodiment provides a method for constructing a real-world noise image super-resolution dataset, including the following steps:
[0061] S1. Image acquisition with controllable noise intensities: In the same natural scene, a camera is used to simultaneously capture low-resolution images with different noise intensities and their corresponding noise-free high-resolution images by adjusting the ISO parameters and optical zoom ratio.
[0062] It is used to generate raw image pairs with quantifiable noise intensity, graded resolution and strict alignment, specifically:
[0063] A Sony SLR camera, mounted on a tripod, was used in both RAW and JPEG modes. The lens focal lengths were 24mm (for low-resolution capture) and 96mm (for high-resolution capture). ISO parameters were set to 100 (no noise), 1600 (low-intensity noise), and 3200 (high-intensity noise). The shutter speed was automatically adjusted to ensure consistent exposure. The following operations were performed in sequence across 580 natural scenes:
[0064] S11. Use a SLR camera with a focal length of 24mm to shoot low-resolution images. Set the ISO parameter to 100 to generate a clean low-resolution image without noise (LR-ISO100). Set the ISO parameter to 1600 and 3200 to generate low-intensity noise (LR-ISO1600) and high-intensity noise (LR-ISO3200) low-resolution images, respectively.
[0065] S12, adjusting the focal length to 96 mm through optical zoom and resetting the ISO to 100, capturing a noise-free high-resolution image (HR-ISO100) of the same scene, thereby forming an image pair having a 4x super-resolution magnification with the low-resolution image;
[0066] S13. Repeat steps S12-S12 to collect raw data in 580 different natural scenes, covering indoor and outdoor environments and static objects, such as offices, buildings, vegetation, etc.
[0067] S2. SIFT feature matching and maximum correlation coefficient fusion registration and alignment. Based on the homography matrix optimization method of SIFT feature matching and maximum correlation coefficient, low-resolution images and high-resolution images are aligned at the pixel level to eliminate the displacement error caused by changes in shooting parameters.
[0068] Step S2 is used to achieve sub-pixel alignment of the low-noise low-resolution image, the high-noise low-resolution image, the clean low-resolution image, and the clean high-resolution image, specifically:
[0069] S21. Extract SIFT feature points from the high-resolution image HR-ISO100 and the low-resolution images LR-ISO100, LR-ISO1600, and LR-ISO3200, and calculate feature descriptors.
[0070] S22, screening matching feature point pairs based on the RANSAC algorithm and initializing the homography matrix;
[0071] S23, assuming that the high-resolution image is I H , the low-resolution image is I L , the homography matrix τ is iteratively optimized by maximizing the correlation coefficient algorithm, and the objective function is defined as:
[0072]
[0073] Where C is a cropping operation that can convert the low-resolution image I L The field of view is cropped and the high-resolution image I H For the same field of view, α and β are brightness adjustment parameters, ||·|| p is the p-norm.
[0074] S24. Apply the optimized homography matrix to the low-resolution image and generate a noisy low-resolution image patch aligned with the high-resolution image through bilinear interpolation resampling.
[0075] S3. Cross-modal data processing step: Filter and eliminate image groups containing moving objects, and based on the characteristics of the camera image signal processor, correct the color space difference and brightness deviation between the low-resolution and high-resolution images.
[0076] Step S3 is used to eliminate the cross-modal color difference and brightness deviation caused by the camera ISP, specifically: obtain the aligned image I′ L =C(τ(I L )), using the calculation of brightness adjustment parameters:
[0077] α=std(I H ) / std(I′ L ),β=mean(I H )-αmean(I′ L )
[0078] Where mean is the pixel mean, std is the pixel variance, and brightness adjustment can ensure I′ L and I H have the same pixel mean and variance.
[0079] S4. Generation of real-world noisy image super-resolution dataset: The registered image pairs are layered and cropped according to the preset size to generate a triple dataset containing noisy low-resolution image patches, clean low-resolution image patches, and high-resolution image patches. The dataset is then classified and stored as a training set and a test set according to the noise intensity.
[0080] Step S4 is used to generate the image block pairs required for training and testing, specifically:
[0081] The registered HR images were cropped non-overlappingly to 192×192 pixels to generate high-resolution image blocks. The LR images were also cropped synchronously to generate low-resolution image blocks of 48×48 pixels, retaining a 4x scaling relationship. The images were classified and stored according to noise intensity. The low-intensity noise image set contained 40,059 training blocks (including LR-ISO1600 / LR-ISO100 / HR) and 4,082 test blocks. The high-intensity noise image set contained 40,010 training blocks (including LR-ISO3200 / LR-ISO100 / HR) and 4,087 test blocks.
[0082] S41, performing non-overlapping sliding cropping on the registered high-resolution image according to a size of 192×192 pixels to generate a high-resolution image block;
[0083] S42, synchronously cropping the low-resolution image to generate a low-resolution image block of 48×48 pixels, retaining the coordinate correspondence with the high-resolution image block;
[0084] S43. Divide the dataset into the following subsets according to noise intensity:
[0085] 1) Low-intensity noise training set: 40,059 groups of (LR-ISO1600, LR-ISO100, HR-ISO100) image patches;
[0086] 2) Low-intensity noise test set: 4,082 image patches;
[0087] 3) High-intensity noise training set: 40,010 groups of (LR-ISO3200, LR-ISO100, HR-ISO100) image patches;
[0088] 4) High-intensity noise test set: 4,087 image patches;
[0089] S44. Add metadata tags to each image block, including ISO value, noise intensity level, scene category, and cropping position coordinates.
[0090] The RealNSR dataset produced based on the embodiment of the present invention includes:
[0091] (a) 580 sets of original scene data, each set includes:
[0092] 1) Low-noise intensity low-resolution image (LR-ISO1600);
[0093] 2) low-resolution images with high noise intensity (LR-ISO3200);
[0094] 3) Clean low-resolution image (LR-ISO100);
[0095] 4) High-resolution images (HR-ISO100);
[0096] (b) 80,141 sets of training image patches and 8,169 sets of testing image patches are stored in the following structure:
[0097] 1) train / low_iso / : 40,059 sets of 48×48 low-intensity noise image patches and corresponding 192×192 high-resolution image patches;
[0098] 2) train / high_iso / : 40,010 sets of 48×48 high-intensity noise image patches and corresponding 192×192 high-resolution image patches;
[0099] 3) test / low_iso / : 4,082 sets of 48×48 low-intensity noise test image patches and corresponding 192×192 high-resolution test image patches;
[0100] 4)test / high_iso / : 4,087 sets of 48×48 low-intensity noise test image patches and corresponding 192×192 high-resolution test image patches.
[0101] It should be noted that, for the sake of convenience, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited to the described order of actions, because according to the present invention, certain steps can be performed in other orders or simultaneously.
[0102] Based on the same principles as the method for constructing a real-world noise image super-resolution dataset in the above-mentioned embodiment, the present invention also provides a system for constructing a real-world noise image super-resolution dataset, which can be used to implement the above-mentioned method. For ease of explanation, the structural diagram of the embodiment of the system for constructing a real-world noise image super-resolution dataset only shows the parts relevant to the embodiment of the present invention. Those skilled in the art will understand that the illustrated structure does not constitute a limitation of the device, and the device may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0103] See also Figure 2 In another embodiment of the present application, a real-world noisy image super-resolution dataset construction system 100 is provided, the system comprising an image acquisition module 101, an image registration and alignment module 102, a cross-modal data processing module 103, and a dataset generation module 104;
[0104] The image acquisition module 101 is used to synchronously capture low-resolution images containing different noise intensities and their corresponding noise-free high-resolution images in the same natural scene by using a camera by adjusting ISO parameters and optical zoom factors;
[0105] The image registration module 102 is used to perform pixel-level alignment of the low-resolution image and the high-resolution image based on SIFT feature matching and a homography matrix optimization method that maximizes the correlation coefficient, thereby eliminating displacement errors caused by changes in shooting parameters;
[0106] The cross-modal data processing module 103 is used to filter out image groups containing moving objects and correct the color space difference and brightness deviation between the low-resolution and high-resolution images based on the characteristics of the camera image signal processor;
[0107] The dataset generation module 104 is used to perform hierarchical cropping of the registered image pairs according to a preset size to generate a triple dataset containing noisy low-resolution image blocks, clean low-resolution image blocks, and high-resolution image blocks, and store them into training sets and test sets according to noise intensity.
[0108] It should be noted that the real-world noise image super-resolution dataset construction system of the present invention corresponds one-to-one to the real-world noise image super-resolution dataset construction method of the present invention. The technical features and beneficial effects described in the above-mentioned embodiment of the real-world noise image super-resolution dataset construction method are all applicable to the embodiment of the real-world noise image super-resolution dataset construction. For specific contents, please refer to the description in the embodiment of the method of the present invention. No further details will be given here. This is hereby declared.
[0109] In addition, in the implementation of the real-world noise image super-resolution dataset construction system of the above embodiment, the logical division of each program module is only an example. In actual application, the above functions can be assigned to different program modules as needed, for example, for the configuration requirements of the corresponding hardware or the convenience of software implementation. That is, the internal structure of the real-world noise image super-resolution dataset construction system is divided into different program modules to complete all or part of the functions described above.
[0110] See also Figure 3 In one embodiment, an electronic device for implementing a method for constructing a real-world noise image super-resolution dataset is provided. The electronic device 200 may include a first processor 201, a first memory 202 and a bus, and may also include a computer program stored in the first memory 202 and executable on the first processor 201, such as a real-world noise image super-resolution dataset construction program 203.
[0111] The first memory 202 includes at least one type of readable storage medium, including a flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the first memory 202 may be an internal storage unit of the electronic device 200, such as a mobile hard disk of the electronic device 200. In other embodiments, the first memory 202 may also be an external storage device of the electronic device 200, such as a plug-in mobile hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 200. Furthermore, the first memory 202 may include both an internal storage unit of the electronic device 200 and an external storage device. The first memory 202 can be used not only to store application software and various types of data installed in the electronic device 200, such as the code of the real-world noise image super-resolution dataset construction program 203, but also to temporarily store data that has been output or is about to be output.
[0112] In some embodiments, the first processor 201 may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The first processor 201 is the control core (Control Unit) of the electronic device, connecting the various components of the entire electronic device using various interfaces and lines, and executing or executing programs or modules stored in the first memory 202, as well as calling data stored in the first memory 202, to perform various functions of the electronic device 200 and process data.
[0113] Figure 3 Only the electronic device with components is shown, and it can be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the electronic device 200 , and the electronic device 200 may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0114] The real-world noise image super-resolution dataset construction program 203 stored in the first memory 202 of the electronic device 200 is a combination of multiple instructions. When executed in the first processor 201, the program can achieve:
[0115] The steps for acquiring images with different noise intensities are as follows: In the same natural scene, a camera is used to simultaneously capture low-resolution images with different noise intensities and their corresponding high-resolution images without noise by adjusting the ISO parameter and optical zoom factor.
[0116] Image registration and alignment steps: Based on SIFT feature matching and the homography matrix optimization method that maximizes the correlation coefficient, the low-resolution image and the high-resolution image are aligned at the pixel level to eliminate the displacement error caused by changes in shooting parameters;
[0117] Cross-modal data processing steps: Screen and eliminate image groups containing moving objects, and based on the characteristics of the camera's image signal processor, correct the color space differences and brightness deviations between the low-resolution and high-resolution images;
[0118] Dataset generation steps: The registered image pairs are layered and cropped according to the preset size to generate a triplet dataset containing noisy low-resolution image patches, clean low-resolution image patches, and high-resolution image patches, and are classified and stored as training sets and test sets according to noise intensity.
[0119] Furthermore, if the modules / units integrated in the electronic device 200 are implemented as software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium. The computer-readable medium may include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0120] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0121] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0122] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. A method for constructing a real-world noisy image super-resolution dataset, characterized by: The steps include: The steps for acquiring images with different noise intensities are as follows: In the same natural scene, a camera is used to simultaneously capture low-resolution images with different noise intensities and their corresponding high-resolution images without noise by adjusting the ISO parameter and optical zoom factor. Image registration and alignment steps: Based on SIFT feature matching and the homography matrix optimization method that maximizes the correlation coefficient, the low-resolution image and the high-resolution image are aligned at the pixel level to eliminate the displacement error caused by changes in shooting parameters; Cross-modal data processing steps: Screen and eliminate image groups containing moving objects, and based on the characteristics of the camera's image signal processor, correct the color space differences and brightness deviations between the low-resolution and high-resolution images; Dataset generation steps: The registered image pairs are layered and cropped according to the preset size to generate a triplet dataset containing noisy low-resolution image patches, clean low-resolution image patches, and high-resolution image patches, and are classified and stored as training sets and test sets according to noise intensity.
2. The method for constructing a real-world noise image super-resolution dataset according to claim 1, characterized in that: The multi-noise intensity image acquisition step is specifically as follows: Use the camera to shoot low-resolution images at the first focal length, set the ISO parameter to 100 to generate a clean low-resolution image without noise, and set the ISO parameter to 1600 and 3200 to generate low-intensity noise and high-intensity noise low-resolution images respectively; The focus is adjusted to the second focal length using optical zoom, and the ISO is reset to 100 to capture a noise-free high-resolution image of the same scene, forming an image pair with a 4x super-resolution magnification compared to the low-resolution image. Repeat the above steps to collect raw data in multiple different natural scenes, covering indoor and outdoor environments and static objects.
3. The method for constructing a real-world noise image super-resolution dataset according to claim 1, characterized in that: The image registration and alignment steps are specifically as follows: SIFT feature points are extracted from the high-resolution image HR-ISO100 and the low-resolution images LR-ISO100, LR-ISO1600, and LR-ISO3200, and feature descriptors are calculated. Based on the RANSAC algorithm, the matching feature point pairs are selected and the homography matrix is initialized; Assume that the high-resolution image is I H , the low-resolution image is I L , the homography matrix τ is iteratively optimized by maximizing the correlation coefficient algorithm, and the objective function is defined as: Where C is a cropping operation that can transform the low-resolution image I L The field of view is cropped and the high-resolution image I H For the same field of view, α and β are brightness adjustment parameters, ||·|| p is the p-norm; The optimized homography matrix is applied to the low-resolution image and resampled by bilinear interpolation to generate noisy low-resolution image patches that are aligned with the high-resolution image.
4. The method for constructing a real-world noise image super-resolution dataset according to claim 1, characterized in that: The cross-modal data processing is specifically as follows: The aligned image I′ is obtained by maximizing the correlation coefficient algorithm L =C(τ(I L )); The calculation of the brightness adjustment parameters is as follows: α=std(I H ) / std(I′ L ),β=mean(I H )-αmean(I′ L ) Where mean is the pixel mean and std is the pixel variance; Brightness adjustment can ensure I' L and I H have the same pixel mean and variance.
5. The method for constructing a real-world noise image super-resolution dataset according to claim 1, characterized in that: The data set generation steps are specifically as follows: Perform non-overlapping sliding cropping on the registered high-resolution image according to the first pixel size to generate a high-resolution image block; Synchronously cropping the low-resolution image to generate a low-resolution image block of a second pixel, retaining a coordinate correspondence with the high-resolution image block; Divide the dataset into subsets based on noise intensity: Add metadata tags to each image block, including ISO value, noise intensity level, scene category, and crop position coordinates.
6. The method for constructing a real-world noise image super-resolution dataset according to claim 1, characterized in that: The data set is divided into multiple subsets according to the noise intensity, specifically: Low-intensity noise training set: 40,059 groups (LR-ISO1600, LR-ISO100, HR-ISO100) image patches; Low-intensity noise test set: 4,082 image patches; High-intensity noise training set: 40,010 groups (LR-ISO3200, LR-ISO100, HR-ISO100) image patches; High-intensity noise test set: 4,087 image patches.
7. The method for constructing a real-world noise image super-resolution dataset according to claim 1, characterized in that: The image block cropping adopts overlapping sliding windows to maximize data utilization.
8. A system for constructing a real-world noise image super-resolution dataset, characterized by: A method for constructing a real-world noise image super-resolution dataset as applied to any one of claims 1-7, comprising an image acquisition module, an image registration and alignment module, a cross-modal data processing module, and a dataset generation module; The image acquisition module is used to synchronously capture low-resolution images containing different noise intensities and their corresponding noise-free high-resolution images in the same natural scene by using a camera by adjusting ISO parameters and optical zoom multiples; The image registration and alignment module is used to perform pixel-level alignment of low-resolution images and high-resolution images based on SIFT feature matching and a homography matrix optimization method that maximizes the correlation coefficient, thereby eliminating displacement errors caused by changes in shooting parameters; The cross-modal data processing module is used to filter out image groups containing moving objects and correct the color space difference and brightness deviation between the low-resolution and high-resolution images based on the characteristics of the camera image signal processor; The dataset generation module is used to perform hierarchical cropping of the registered image pairs according to a preset size to generate a triplet dataset containing noisy low-resolution image blocks, clean low-resolution image blocks, and high-resolution image blocks, and store them as training sets and test sets according to noise intensity.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the method for constructing a real-world noisy image super-resolution dataset as described in any one of claims 1-7.
10. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the method for constructing a real-world noise image super-resolution dataset as described in any one of claims 1 to 7 is implemented.