Acquisition and alignment method for real multiple degraded images

By adjusting camera parameters during image acquisition and using optical flow estimation network for pixel alignment, the problem of single degradation type and time-consuming and labor-consuming process in the prior art is solved, efficient real multi-degradation image acquisition and alignment is achieved, and high-quality image data sets are constructed.

CN120047367APending Publication Date: 2025-05-27UNIV OF ELECTRONICS SCI & TECH OF CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510062382.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing acquisition and alignment methods of real multi-degradation images have problems such as single degradation types, time-consuming and labor-intensive alignment processes, and poor results, making it difficult to build high-quality image datasets.

Method used

Low-resolution, degraded images and high-resolution, degraded images are acquired in various scenarios by using camera adjustment parameters, and pixel alignment is performed by combining pre-trained optical flow estimation networks and linear iterative algorithms, and illumination alignment is performed through channel mean correction.

Benefits of technology

It realizes efficient acquisition of real, multi-type, composite superimposed degraded images, and is well aligned with the corresponding high-quality images, and builds a high-quality image data set, improving the performance of the model in practical application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047367A_ABST
    Figure CN120047367A_ABST
Patent Text Reader

Abstract

The invention discloses a real multi-degradation image acquisition and alignment method, which comprises the following steps of: adjusting parameters by using a camera in various scenes, and acquiring low-resolution and degradation-containing images and high-resolution and non-degradation images to form image pairs; aligning image pixels by using a pre-trained optical flow estimation network and a linear iterative algorithm; and finally, performing alignment on image illumination through channel mean value correction so as to efficiently acquire real, multi-type and composite superposition degraded images, and performing good alignment on the real, multi-type and composite superposition degraded images and the corresponding high-quality images so as to construct a high-quality image data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and more specifically, relates to a method for collecting and aligning real multi-degraded images. Background Art

[0002] Image restoration refers to restoring a low-quality image containing various degradation factors (such as low resolution, noise interference, blur distortion, etc.) to a high-quality, non-degraded image. This task has important application values in fields such as medical image processing, cultural relic restoration, and video enhancement. Currently, deep learning methods have become the main technical means for image restoration tasks due to their excellent efficiency and performance. However, the effectiveness of these methods highly depends on large-scale, high-quality datasets for training to learn the mapping relationship from low-quality images to high-quality images.

[0003] In existing image restoration research, many training methods tend to use synthetic datasets, that is, by specific algorithms or models to simulate the real-world degradation process, thereby generating low-quality images and pairing them with the original high-quality images. However, there are significant differences between the degraded images generated in this way and real-world degradation. Due to the complexity, diversity, and randomness of degradation factors in the real world, synthetic datasets are difficult to fully cover the distribution characteristics of real degradation, resulting in problems such as performance degradation and insufficient generalization of the model in actual application scenarios. Especially when the model processes completely unknown types or combined degradations, its restoration effect is often unsatisfactory. Therefore, constructing a high-quality image restoration dataset closer to the real-world degradation characteristics has become a key link in promoting the development of this field.

[0004] At present, some studies have begun to attempt to use real datasets to address the differences between synthetic datasets and real degradation distributions. Such real datasets are typically constructed by collecting multi-degraded images and their corresponding high-quality images in actual scenarios. The acquisition of real image inpainting datasets mainly faces the following two types of problems. First, it is difficult to achieve precise alignment between high-quality images and multi-degraded images. Due to the shooting environment, device characteristics, operation errors, and the complexity of the degradation process, there are often perspective deviations, position offsets, or scale differences between high-quality images and degraded images. Such alignment errors will significantly affect the quality of model training and may even lead to unstable inpainting effects. In addition, the manual alignment process is both time-consuming and labor-intensive, making it difficult to construct large-scale datasets, thus restricting the widespread application of real datasets in image inpainting tasks. Second, the degradation types in existing real datasets are relatively single, making it difficult to comprehensively cover the complex and diverse degradation situations in the real world. Degradation factors in real scenarios often exhibit diversity and randomness, such as the coexistence or superposition of blur, noise, and low resolution problems. However, the degradation in existing datasets is often dominated by a single factor, lacking a comprehensive description of multi-degradation combinations, resulting in significant performance bottlenecks when the model processes complex degraded images.

[0005] In summary, the existing methods for collecting and aligning real multi-degraded images have the following disadvantages: 1. The degradation types of the collected images are single. 2. Alignment requires a large amount of manual operations, which is time-consuming and labor-intensive, and it is difficult to align a large number of images. 3. The alignment effect is not good, and there are still misalignments, inconsistent brightness, etc. between the pictures. Summary of the Invention

[0006] The purpose of the present invention is to overcome the deficiencies of the prior art and propose a method for collecting and aligning real multi-degraded images. The method proposed in this patent can efficiently collect real, multi-type, and composite superposition degraded images and align them well with the corresponding high-quality images, so as to be used to generate high-quality real images.

[0007] To achieve the above invention purpose, a method for collecting and aligning real multi-degraded images of the present invention is characterized by including the following steps:

[0008] (1) Pairing image shooting and introducing multiple degradations;

[0009] (1.1) Set the focal length to f 1 , and set the camera exposure parameters to appropriate values according to the specific scenario, specifically denoted as: the ISO value is I 1 , the exposure time is E 1 , the aperture size is A 1 , and then shoot the image to obtain a high-resolution non-degraded image GT-RAW with a length of H and a width of W;

[0010] (1.2) Set the focal length to f 2 = 0.25f 1 , keep other camera parameters and the camera position unchanged, then take an image, and crop the central part to obtain a low-resolution but non-degraded image LQ-RAW-low resolution with a length of H / 4 and a width of W / 4;

[0011] (1.3) Based on step (1.2), take an integer n greater than or equal to 5 and less than or equal to 100, and set the camera sensitivity to I 2 = nI 1 , and set the exposure time to E 2 = E 1 / n, keep other settings unchanged to ensure the same light input to the lens, then take an image, and crop the central part to obtain a low-resolution and noisy degraded image LQ-RAW-low resolution + noise with a length of H / 4 and a width of W / 4;

[0012] (1.4) Based on step (1.2), take an integer k greater than or equal to 2 and less than or equal to 5, and set the exposure time to E 3 = E 1 / k 2 , adjust the aperture size to A 3 = kA 1 , and turn the focus ring to adjust the focus position to ensure that the object to be photographed is beyond the depth of field range, thus generating defocus blur, keep other settings unchanged to ensure the same light input to the lens, then take an image, and crop the central part to obtain a low-resolution and defocus-blurred degraded image LQ-RAW-low resolution + defocus blur with a length of H / 4 and a width of W / 4;

[0013] (1.5) Set all camera parameters to be the same as in step (1.2), vibrate the camera shooting device periodically during exposure, ensure that the period is much smaller than the exposure time, and take multiple images to obtain the corresponding low-resolution and motion-blurred degraded image LQ-RAW-low resolution + motion blur;

[0014] (1.6) Obtain a set of paired images to be processed, namely GT-RAW, LR-RAW-low resolution, LR-RAW-low resolution + noise, LR-RAW-low resolution + defocus blur, LR-RAW-low resolution + motion blur;

[0015] (2) Build and pre-train an optical flow estimation network;

[0016] (2.1) Randomly select two equal-sized images to be aligned I a 、I b , and the displacement map P of their positional relationship between thema-b ;

[0017] (2.2), Input I into convolutional layers with a convolutional kernel size of 3*3, a stride of 1, and channel numbers of 32, 128, and 256 respectively. After each convolution, perform 2-fold downsampling to obtain feature maps a Input I into convolutional layers with a convolutional kernel size of 3*3, a stride of 1, and channel numbers of 32, 128, and 256 respectively. After each convolution, perform 2-fold downsampling to obtain feature maps

[0018] (2.3), Input I into convolutional layers with a convolutional kernel size of 3*3, a stride of 1, and channel numbers of 32, 128, and 256 respectively to obtain feature maps b Input I into convolutional layers with a convolutional kernel size of 3*3, a stride of 1, and channel numbers of 32, 128, and 256 respectively to obtain feature maps

[0019] (2.4), Concatenate F and F along the channel dimension to obtain feature maps a with F b in the channel dimension to obtain feature maps

[0020] (2.5), Input the feature map F into convolutional layers with a convolutional kernel size of 3*3, a stride of 1, and channel numbers of 256, 128, and 64 respectively. After each convolution, perform 2-fold upsampling to obtain the feature map F ∈ R a-b Input the feature map F into convolutional layers with a convolutional kernel size of 3*3, a stride of 1, and channel numbers of 256, 128, and 64 respectively. After each convolution, perform 2-fold upsampling to obtain the feature map F ∈ R H×W×64 ;

[0021] (2.6), Input the feature map F into convolutional layers with a convolutional kernel size of 3*3, a stride of 1, and channel numbers of 64, 16, and 2 respectively to obtain the optical flow map

[0022] (2.7), Calculate the sum of the absolute differences of the pixel values of the corresponding pixel points in the optical flow map and the displacement map P a-b , then calculate the pixel mean as the loss value loss, and then update the parameters using the gradient descent method according to the loss value loss;

[0023] (2.8), Repeat steps (2.1) to (2.7) until the optical flow estimation network converges;

[0024] (3), Optical flow estimation correction and enhancement correlation coefficient correction;

[0025] (3.1), Render the two images GT-RAW and LR-RAW-low resolution using ISP into RGB images, namely GT-RGB and LR-RGB-low resolution;

[0026] (3.2), Upsample LR-RGB-low resolution using bilinear interpolation to process it into the same size as GT-RGB;

[0027] (3.3), Input LR-RGB-low resolution and GT-RGB into the pre-trained optical flow estimation network to calculate the optical flow map

[0028] (3.4), According to the optical flow map Correct GT-RAW to obtain the image GT-RAW-low resolution that is completely aligned with LQ-RAW-low resolution;

[0029] (3.5), Initialize the affine matrix M noise ∈R 2×3 as

[0030] (3.6), Denoise LQ-RAW-low resolution + noise using a denoising algorithm to obtain the image C b1 , Denote the GT-RAW-low resolution image as C a , Then use the linear iteration method to update the affine matrix M noise , Such that C a ·M noise = C b1 ; Finally, correct the image C through the updated affine matrix M noise , To obtain the image C that is completely aligned with the image C a ; b1 a1 ;

[0031] (3.7), Denote LQ-RAW-low resolution + defocus blur image as C b2 , Use the linear iteration method to update the affine matrix M noise , Such that C a ·M noise = C b2 ; Finally, correct the image C through the updated affine matrix M noise , To obtain the image C that is completely aligned with the image C a ; b2 a2 ;

[0032] (3.8), Denote LQ-RAW-low resolution + motion blur image as C b3 , Use the linear iteration method to update the affine matrix M noise , Such that C a ·M noise = C b3 ; Finally, correct the image C through the updated affine matrix M noise , To obtain the image C that is completely aligned with the image C a ; b3 a3 ;

[0033] (4), Global illumination adjustment;

[0034] (4.1) Calculate the mean value of each channel of GT-RAW-low resolution [μ 1 , μ 2 , …, μ z , and calculate the mean value of each channel of LQ-RAW-low resolution [ν 1 , ν 2 , …, ν z , where z represents the total number of channels of the image;

[0035] (4.2) Divide the corresponding channel means of [μ 1 , μ 2 , …, μ z and [ν 1 , ν 2 , …, ν z to obtain the correction coefficients [B 1 , B 2 , …, B z ;

[0036] (4.3) Multiply the GT-RAW-low resolution with z channels by the correction coefficients [B 1 , B 2 , …, B z to obtain the GT-RAW-low resolution of the image after illumination adjustment and the LQ-RAW-low resolution of the image;

[0037] (4.4) Similarly, perform illumination adjustment on the GT-RAW-low resolution + noise image, the GT-RAW-low resolution + motion blur image, the GT-RAW-low resolution + defocus blur image and their corresponding LQ-RAW images according to steps (4.1) to (4.3) to obtain multiple pairs of aligned images.

[0038] The object of the present invention is achieved as follows:

[0039] A method for collecting and aligning real multi-degraded images of the present invention first uses a camera in various scenarios, adjusts parameters, collects low-resolution, degraded images and high-resolution, non-degraded images to form image pairs; then uses a pre-trained optical flow estimation network and a linear iterative algorithm to align the image pixels; finally, aligns the image illumination through channel mean correction, thereby efficiently collecting real, multi-type, composite and superimposed degraded images and aligning them well with the corresponding high-quality images, so as to construct a high-quality image dataset.

[0040] At the same time, a method for collecting and aligning real multi-degraded images of the present invention also has the following beneficial effects:

[0041] (1). The types of image degradation collected are diverse; compared with the current single-degradation image collection method, the present invention can collect various degraded images including noise, motion blur, defocus blur, and low resolution.

[0042] (2). The image alignment process of the present invention does not require manual operation, and the entire process automatically performs image processing by an algorithm.

[0043] (3). The alignment effect of the present invention is good; compared with the current alignment methods, the present invention can achieve good alignment at the pixel level and brightness level. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a flowchart of a method for collecting and aligning real multi-degraded images of the present invention;

[0045] Figure 2 is a flowchart of shooting degraded images;

[0046] Figure 3 is a flowchart of optical flow estimation correction and enhancement correlation coefficient correction;

[0047] Figure 4 is a flowchart of global illumination adjustment;

[0048] Figure 5 is the image after collection and alignment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] The following describes the specific embodiments of the present invention with reference to the drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.

[0050] Embodiment

[0051] In this embodiment, as Figure 1 shown, a method for collecting and aligning real multi-degraded images of the present invention includes the following steps:

[0052] S1. Pairing image shooting and introducing multiple degradations;

[0053] Denote the focal length as f, the camera sensitivity as ISO, the exposure time as E, the aperture size as A, the high-resolution non-degraded image taken as GT-RAW, and the low-resolution image with a certain type of degradation as LQ-RAW - a certain type of degradation; the length and width of the high-resolution image are H and W, and the length and width of the low-resolution image are 1 / 4 of the high-resolution image. The overall process of pairing image shooting and introducing multiple degradations is as Figure 2 shown, and the following will describe this process in detail step by step;

[0054] S1.1. Set the focal length to f 1 = 80, set the camera exposure parameters to appropriate values according to the specific scenario, specifically record: the ISO is I 1 = 160, the exposure time is E 1 = 1 / 80 s, the aperture size is A 1 = 1 / 16, then take a picture to obtain a high-resolution non-degraded image GT-RAW with a length H = 1920 mm and a width W = 1080 mm;

[0055] S1.2. Set the focal length to f 2 = 0.25f 1 , keep other camera parameters and the camera position unchanged, then take a picture, and crop the central part to obtain a low-resolution but non-degraded image LQ-RAW - low resolution with a length of H / 4 and a width of W / 4;

[0056] S1.3. On the basis of step S1.2, take an integer n greater than or equal to 5 and less than or equal to 100, set the camera ISO to I 2 = nI 1 , and set the exposure time to E 2 = E 1 / n, keep other settings unchanged to ensure the same light input to the lens, then take a picture, and crop the central part to obtain a low-resolution and noisy degraded image LQ-RAW - low resolution + noise with a length of H / 4 and a width of W / 4; In this embodiment, n can be flexibly adjusted according to requirements. The larger the value, the greater the introduced noise. In this embodiment, n = 50 is taken.

[0057] S1.4. On the basis of S1.2, take an integer k greater than or equal to 2 and less than or equal to 5, set the exposure time to E 3 = E 1 / k 2 , adjust the aperture size to A 3 = kA 1 , and turn the focus ring to adjust the focus position to ensure that the photographed object is out of the depth of field range, thus generating defocus blur. Keep other settings unchanged to ensure the same light input to the lens, then take a picture, and crop the central part to obtain a low-resolution and defocus-blurred degraded image LQ-RAW - low resolution + defocus blur with a length of H / 4 and a width of W / 4; In this embodiment, k can be flexibly adjusted according to requirements. The larger the value, the more serious the introduced defocus blur. In this embodiment, k = 4 is taken.

[0058] S1.5. Adjust all camera settings to be the same as in S1.2. During the exposure, periodically vibrate the camera shooting device to ensure that the period is much smaller than the exposure time, and take multiple images to obtain corresponding low-resolution and motion-blurred degraded images LQ-RAW - Low-resolution + Motion Blur.

[0059] S1.6. Obtain a set of paired images to be processed, namely GT-RAW, LR-RAW - Low-resolution, LR-RAW - Low-resolution + Noise, LR-RAW - Low-resolution + Defocus Blur, LR-RAW - Low-resolution + Motion Blur.

[0060] S2. Build and pre-train an optical flow estimation network;

[0061] S2.1. Randomly select two equally sized images to be aligned I a 、I b from the publicly available dataset FlyingChairs, as well as the displacement map P a-b between them;

[0062] S2.2. Input I a into a convolutional layer with a convolutional kernel size of 3*3, a stride of 1, and channel numbers of 32, 128, and 256 respectively. After each convolution, perform 2-fold downsampling to obtain the feature map

[0063] S2.3. Input I b into a convolutional layer with a convolutional kernel size of 3*3, a stride of 1, and channel numbers of 32, 128, and 256 respectively to obtain the feature map

[0064] S2.4. Concatenate F a and F b in the channel dimension to obtain the feature map

[0065] S2.5. Input the feature map F a-b into a convolutional layer with a convolutional kernel size of 3*3, a stride of 1, and channel numbers of 256, 128, and 64 respectively. After each convolution, perform 2-fold upsampling to obtain the feature map F∈R H×W×64 .

[0066] S2.6. Input the feature map F into a convolutional layer with a convolutional kernel size of 3*3, a stride of 1, and channel numbers of 64, 16, and 2 respectively to obtain the optical flow map

[0067] S2.7. Calculate the optical flow map and the displacement map P a-bThe sum of the absolute differences of the pixel values of the corresponding pixel points is calculated, and then the pixel mean is calculated as the loss value loss. Then, according to the loss value loss, the parameters are updated using the gradient descent method;

[0068] S2.8. Repeat S2.1 - S2.7 until the training converges. The trained network can calculate the corresponding relationship between the pixel points of the two images.

[0069] S3. Optical flow estimation correction and enhancement correlation coefficient correction;

[0070] The purpose of the correction is to perform pixel - level alignment on the to - be - processed paired images obtained in step S1, so as to obtain a high - quality aligned image pair. The correction process is as Figure 3 shown. The detailed correction process is as follows:

[0071] S3.1. Render the two images GT - RAW and LR - RAW - low - resolution into RGB images using ISP, that is, GT - RGB and LR - RGB - low - resolution.

[0072] S3.2. Upsample LR - RGB - low - resolution using bilinear interpolation to the same size as GT - RGB.

[0073] S3.3. Input LR - RGB - low - resolution and GT - RGB into the pre - trained optical flow estimation network to calculate the optical flow map

[0074] S3.4. According to the optical flow map correct GT - RAW to obtain the image GT - RAW - low - resolution that is completely aligned with LQ - RAW - low - resolution;

[0075] S3.5. Initialize the affine matrix M noise ∈R 2×3 as

[0076] S3.6. Denoise LQ - RAW - low - resolution + noise using a denoising algorithm to obtain the image C b1 , denote the GT - RAW - low - resolution image as C a , and then update the affine matrix M noise using the linear iteration method, so that C a ·M noise = C b1 ; finally, correct the image C noise through the updated affine matrix M a to obtain the image C b1 that is completely aligned with the image C a1 ;

[0077] S3.7. Denote the LQ-RAW low-resolution + defocus blurred image as C b2 , and update the affine matrix M using the linear iteration method noise such that C a ·M noise = C b2 ; Finally, correct the image C through the updated affine matrix M noise to obtain the image C a that is perfectly aligned with the image C b2 ; a2 ;

[0078] S3.8. Denote the LQ-RAW low-resolution + motion blurred image as C b3 , and update the affine matrix M using the linear iteration method noise such that C a ·M noise = C b3 ; Finally, correct the image C through the updated affine matrix M noise to obtain the image C a that is perfectly aligned with the image C b3 ; a3 ;

[0079] S4. Global illumination adjustment, the process is as Figure 4 shown;

[0080] S4.1. Calculate the mean values of each channel of GT-RAW low-resolution [μ 1 , μ 2 , …, μ z , and calculate the mean values of each channel of LQ-RAW low-resolution [ν 1 , ν 2 , …, ν z , where z represents the total number of channels of the image;

[0081] S4.2. Divide the corresponding channel means of [μ 1 , μ 2 , …, μ z and [ν 1 , ν 2 , …, ν z to obtain the correction coefficients [B 1 , B 2 , …, B z ;

[0082] S4.3. Multiply the GT-RAW low-resolution with z channels by the correction coefficients [B 1 , B 2 , …, B z to obtain the GT-RAW low-resolution of the image after illumination adjustment and the LQ-RAW low-resolution image;

[0083] S4.4. Similarly, for the images GT-RAW-low resolution + noise, GT-RAW-low resolution + motion blur, and GT-RAW-low resolution + defocus blur, and their corresponding LQ-RAW images, perform illumination adjustment according to steps S4.1 to S4.3 to obtain multiple pairs of aligned images.

[0084] Figure 5 These are three pairs of images after acquisition and alignment processing using the method proposed in the present invention. In Figure 5 each row, the left side is the actually acquired degraded low-quality image, and the right side is the high-quality image that is perfectly aligned with it; the degradation in the first row is low resolution and noise, the degradation in the second row is low resolution and motion blur, and the degradation in the third row is low resolution and defocus blur. By comparing the three pairs of images, it can be found that the present invention can be well aligned with the corresponding high-quality images, and thus can be used to generate high-quality real images.

[0085] Although the above describes the illustrative specific embodiments of the present invention for the understanding of those skilled in the art of the present technology, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

Claims

1. A method for collecting and aligning real multi-degraded images, characterized in that: The following steps are involved: (1) Paired image capture and introduction of multiple degradations; (1.1) Set the focal length to f1, and set the camera exposure parameters to appropriate values ​​according to the specific scene. Specifically, the sensitivity is I1, the exposure time is E1, and the aperture size is A1. Then take an image to obtain a high-resolution, non-degraded GT-RAW image with a length of H and a width of W. (1.2) Set the focal length to f2 = 0.25f1, keep other camera parameters and camera position unchanged, then take an image, and crop the central part to obtain a low-resolution but non-degraded image LQ-RAW-low resolution with a length of H / 4 and a width of W / 4; (1.3) Based on step (1.2), take an integer n greater than or equal to 5 and less than or equal to 100, set the camera sensitivity to I2 = nI1, and set the exposure time to E2 = E1 / n, keep other settings unchanged to ensure that the amount of light entering the lens is the same, then take an image, and crop the center part to obtain a low-resolution and noisy degraded image LQ-RAW-low resolution + noise with a length of H / 4 and a width of W / 4; (1.4) Based on step (1.2), take an integer k greater than or equal to 2 and less than or equal to 5, and set the exposure time to E3 = E1 / k 2 , adjust the aperture size to A3=kA1, and twist the focus ring to adjust the focus position to ensure that the photographed object is beyond the depth of field, thereby generating defocus blur. Keep other settings unchanged to ensure that the amount of light entering the lens is the same, then shoot an image and crop the center part to obtain a degraded image LQ-RAW-low resolution + defocus blur with a length of H / 4 and a width of W / 4; (1.5) Set all camera parameters to be the same as step (1.2), periodically vibrate the camera shooting device during exposure, ensure that the period is much smaller than the exposure time, and take multiple images to obtain the corresponding low-resolution and motion-blurred degraded image LQ-RAW-low resolution + motion blur; (1.6) Obtain a set of paired images to be processed, namely, GT-RAW, LR-RAW-low resolution, LR-RAW-low resolution + noise, LR-RAW-low resolution + defocus blur, LR-RAW-low resolution + motion blur; (2) Build and pre-train the optical flow estimation network; (2.1) Randomly select two images of equal size to be aligned from the public dataset FlyingChairs a ,I b , and the positional relationship displacement diagram P between them a-b ; (2.2), will I a The input convolution kernel size is 3*3, the step size is 1, and the number of channels is 32, 128, and 256 respectively. Each convolution layer is downsampled by 2 times to obtain the feature map (2.3), will I b Input convolution kernel size is 3*3, step size is 1, and convolution layer with number of channels is 32, 128, and 256 respectively, and get feature map (2.4), F a With F b Connect in the channel dimension to get the feature map (2.5) The feature map F a-b The input convolution kernel size is 3*3, the step size is 1, and the number of channels is 256, 128, and 64 respectively. After each convolution layer, it is upsampled by 2 times to obtain the feature map F∈R H×W×64 ; (2.6) Input the feature map F into the convolution layer with a convolution kernel size of 3*3, a step size of 1, and a number of channels of 64, 16, and 2 respectively, to obtain the optical flow map (2.7) Calculate the optical flow map With displacement map P a-b The sum of the absolute differences of the pixel values ​​of the corresponding pixels in the image is calculated, and then the pixel mean is calculated as the loss value loss. Then, according to the loss value loss, the gradient descent method is used to update the parameters. (2.8) Repeat steps (2.1) to (2.7) until the optical flow estimation network converges; (3) Correction of optical flow estimation and enhancement correlation coefficient; (3.1) Use ISP to render the GT-RAW and LR-RAW-low resolution images into RGB images, namely GT-RGB and LR-RGB-low resolution; (3.2) Use bilinear upsampling to process the LR-RGB-low resolution to the same size as GT-RGB; (3.3) Input LR-RGB-low resolution and GT-RGB into the pre-trained optical flow estimation network to calculate the optical flow map (3.4), according to the optical flow graph Correct GT-RAW to get an image GT-RAW-low resolution that is completely aligned with LQ-RAW-low resolution; (3.5), Initialize the affine matrix M noise ∈R 2×3 for (3.6) De-noise LQ-RAW-low resolution + noise using denoising algorithm to obtain image C b1 , let GT-RAW-low resolution image be C a , and then use the linear iteration method to update the affine matrix M noise , so that C a ·M noise =C b1 ; Finally, the updated affine matrix M noise Corrected image C a , get the same as image C b1 Completely aligned image C a1 ; (3.7), let LQ-RAW-low resolution + defocused blur image be C b2 , update the affine matrix M using linear iteration noise , so that C a ·M noise =C b2 ; Finally, the updated affine matrix M noise Corrected image C a , get the same as image C b2 Completely aligned image C a2 ; (3.8), let LQ-RAW-low resolution + motion blurred image be C b3 , update the affine matrix M using linear iteration noise , so that C a ·M noise =C b3 ;Finally, the updated affine matrix M noise Corrected image C a , and get the same as image C b3 Completely aligned image C a3 ; (4) Global illumination adjustment; (4.1), calculate the mean of each channel of GT-RAW-low resolution [μ 1 ,μ 2 ,…,μ z ], calculate the mean value of each channel of LQ-RAW-low resolution [ν 1 ,ν 2 ,…,ν z ], where z represents the total number of channels of the image; (4.2),[μ 1 ,μ 2 ,…,μ z ] and [ν 1 ,ν 2 ,…,ν z ] is divided by the corresponding channel mean to obtain the correction coefficient [B 1 ,B 2 ,…,B z ]; (4.3), the GT-RAW-low resolution with z channels and the correction coefficient [B 1 ,B 2 ,…,B z ] to obtain the illumination-adjusted image GT-RAW-low resolution and the image LQ-RAW-low resolution; (4.4) Similarly, perform illumination adjustment on the image GT-RAW-low resolution + noise, the image GT-RAW-low resolution + motion blur, the image GT-RAW-low resolution + defocus blur and their corresponding images LQ-RAW according to steps (4.1) to (4.3) to obtain multiple pairs of aligned images.