Noise Estimation

By using the method of combining two training models in the image processing device, the noise removal problem in the prior art is solved and the image quality is significantly improved.

CN113168671BActive Publication Date: 2025-05-13HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN201980076434.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-03-21
Publication Date
2025-05-13
Estimated Expiration
2039-03-21

AI Technical Summary

Technical Problem

The prior art is difficult to effectively estimate and remove noise in images, especially in images captured under low light conditions, where noise removal is crucial to improving image quality.

Method used

An image processing device and method are used to perform noise estimation through a combination of two training models. The first model is used to detect random noise, and the second model is used to detect pixel extreme values, and the two combine to form a aggregate noise estimation. This method is suitable for detecting Poisson and Gaussian noise, and improves the accuracy of noise estimation by adapting the model.

Benefits of technology

Accurate estimation and removal of noise in the image is achieved, significantly improving the quality of captured images under low light conditions, and improving the fidelity and clarity of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113168671B_ABST
    Figure CN113168671B_ABST
Patent Text Reader

Abstract

An image processing device includes a processor, wherein the processor is used to estimate noise in an image by the following steps, wherein the image is represented by a set of pixels and each pixel has a value associated with the pixel on each of one or more channels: processing data derived from the image through a first training model for detecting random noise to form a first noise estimate; processing data derived from the image through a second training model for detecting pixel extreme values ​​to form a second noise estimate; and combining the first noise estimate and the second noise estimate to form an aggregated noise estimate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to estimating noise in an image. Background Art

[0002] When a camera captures an image, noise may be present in the captured image. Noise reduces the fidelity of the captured image. When the camera used is a digital camera, image noise can originate from a number of sources. First, it can come from random processes associated with image capture. For example, this random noise can be generated by noise in the image sensor. Random noise varies from image to image and is detected by a particular camera. Second, noise can originate from deterministic sources, such as failed pixels in the image sensor. This noise is usually consistent for a particular camera unless additional pixels fail or are likely to recover.

[0003] Certain situations are particularly prone to image noise. For example, when a digital camera captures an image in low-light conditions, the camera can increase the gain of its sensor to amplify the brightness of the captured data. However, this also increases random noise. The downside of increasing sensor gain is that it amplifies noise. Therefore, noise removal can significantly improve the quality of images captured in low-light conditions.

[0004] It is desirable to be able to process an image after it has been captured to reduce the presence of noise. This way, the image quality will be higher. One approach is to use a filter to blur the image. This reduces high frequency noise but makes the image less sharp. Another approach is to estimate the noise in the original captured image and then attempt to remove the estimated noise from the original captured image so as to form an adjusted denoised image.

[0005] Estimating noise in an image is difficult. Some prior art methods are described in Chatterjee, Priyam et al., "Noise suppression in low-light images through joint denoising and demosaicing", CVPR 2011. IEEE, 2011, and Remez, Tal et al., "Deep convolutional denoising of low-light images", arXiv preprint arXiv:1701.01687 (2017). Rudin, Leonid I., Stanley Osher, and Emad Fatemi, "Nonlinear total variation based noise removal algorithms", Physics D: Nonlinear Phenomena 60.1-4 (1992): 259-268, propose a nonlinear variational algorithm to remove noise based on partial differential equations that implement a total variation loss.

[0006] An improved method is needed to estimate noise in images. Summary of the invention

[0007] Embodiments of the invention are defined by the features of the attached independent claims. Further advantageous implementations of these embodiments are further defined by the features of the dependent claims.

[0008] According to a first aspect, there is provided an image processing device, comprising a processor, wherein the processor is used to estimate noise in an image by the following steps, wherein the image is represented by a set of pixels and each pixel has a value associated with the pixel on each of one or more channels: processing data derived from the image by a first training model for detecting random noise to form a first noise estimate; processing data derived from the image by a second training model for detecting pixel extreme values ​​to form a second noise estimate; and combining the first noise estimate and the second noise estimate to form an aggregated noise estimate.

[0009] According to a second aspect, a method for training an image processing model is provided, comprising: (a) receiving a plurality of pairs of images, wherein each pair of images represents a common scene and the first image in each pair of images includes more noise than the second image in each pair of images; (b) for each pair of images: (i) processing data derived from the first image in each pair of images by a first model for estimating random noise in the images to form a first noise estimate; (ii) processing data derived from the first image in each pair of images by a second model for detecting pixel extrema to form a second noise estimate; (iii) combining the first noise estimate and the second noise estimate to form an aggregate noise estimate; (iv) estimating the difference between (A) the second image in each pair of images and (B) the first image in each pair of images denoised according to the aggregate noise estimate; and (v) adapting the first model and the second model according to the estimated difference.

[0010] The first training model may be suitable and / or adapted to detect Poisson noise and / or Gaussian noise. Such types of noise may appear in digitally captured images.

[0011] The first training model may be more accurate at detecting Poisson noise and / or Gaussian noise than the second training model. In this regard, the functions of the models may be different, which may result in the overall system performing better after training on different aspects of the models.

[0012] The second training model may have a higher accuracy in detecting bad pixel noise than the first training model. The second model may be suitable and / or adapted to detect bad pixel noise. It may be used to detect independent pixel extreme values.

[0013] The apparatus may be configured to subtract the aggregate noise estimate from the image to form a denoised image. The denoised image may appear better to a viewer.

[0014] The device may be used to process the image and the aggregated noise point estimate through a third trained model to form a denoised image. This may improve the perceived result of denoising the original image.

[0015] The first training model and the second training model may include a processing architecture for: (a) processing data derived from the image to gradually reduce resolution through a first series of stages to form intermediate data; (b) processing the intermediate data to gradually increase resolution through a second series of stages to form corresponding noise estimates, wherein there are jump connections between corresponding stages of the first series and the second series to provide feedthrough. This may be an effective way to configure the model to achieve good trainability and applicability.

[0016] The first series of stages of the second training model may include: (a) a first stage for processing data derived from the image to reduce resolution and increase data depth to form second intermediate data, (b) a second stage for processing the second intermediate data to reduce resolution without increasing data depth to form third intermediate data. This may be particularly suitable for processing bad pixel noise.

[0017] The first level may be a space-depth level. By increasing the data depth at the retained pixel point to avoid data loss, the spatial resolution may be reduced while retaining data.

[0018] The second level may be a maximum pooling level. None of the first series of levels of the first training model may increase data depth. This may reduce resolution in a lossy manner, thereby facilitating more efficient subsequent processing.

[0019] The first model and the second model may include a processing architecture, which is used to: (a) process data derived from the first image in each pair of images to gradually reduce the resolution through a first series of levels to form intermediate data; (b) process the intermediate data to gradually increase the resolution through a second series of levels to form corresponding noise estimates, wherein there are jump connections providing feedthrough between corresponding levels of the first series and the second series; the first series of levels of the second training model may include: (a) a first stage for processing data derived from the first image in each pair of images to reduce the resolution and increase the data depth to form second intermediate data; (b) a second stage for processing the second intermediate data to reduce the resolution without increasing the data depth to form third intermediate data.

[0020] According to a third aspect, there is provided an image processing model adapted by the above method. The model may be stored on a data carrier. The model may be stored in a non-transient form. The model may include neural network weights. The model may include a neural network.

[0021] According to a fourth aspect, there is provided an image processing device comprising a processor and a memory, wherein the memory stores instructions executed by the processor in a non-transitory form to implement the image processing model described above.

[0022] The details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The present invention will now be described by way of examples with reference to the accompanying drawings.

[0024] In the attached picture:

[0025] Figure 1 The overall architecture of a device for implementing the present invention is shown;

[0026] Figure 2 A noisy image and a clearer image are shown;

[0027] Figure 3 shows the arrangement of pixels in a RAW format image;

[0028] Figure 4 The results for bad pixels are shown;

[0029] Figure 5 A first processing architecture is shown;

[0030] Figure 6 An example of a trainable subnet architecture is shown;

[0031] Figure 7 A second processing architecture is shown;

[0032] Figure 8 A comparison between denoised images is shown;

[0033] Fig. 9 A second comparison between the denoised images is shown. DETAILED DESCRIPTION

[0034] Figure 1 The overall architecture of a device for implementing the present invention is shown. The device has a processor 1, a memory 2, a camera 3, a display 4, a keyboard 5 and a data interface 6. The memory stores instruction codes executable by the processor in a non-transient manner so that the device performs the functions described herein. The camera can be controlled using the keyboard 5 to capture images. The images are then stored as data in the memory 2. The keyboard can be integrated with the display in the touch screen. The data interface allows the device to send and receive data with a remote location via the Internet or the like. For example, the device can be a mobile phone or a cellular phone. The device can be a server computer, in which case it may not have a camera 3, a display 4 and a keyboard 5. The device can be a dedicated camera, in which case it may not have a data interface 6. It can also be other physical structures.

[0035] In operation, the memory 2 can store images captured by the camera 3 or received via the data interface 6. The processor can then execute the stored code to reduce noise in the image by the method discussed below to form a de-noised image. The processor can then cause the display to display the de-noised image, or can transmit the de-noised image to another location via the interface 6.

[0036] Figure 2An example of denoising is shown. An observed noisy image is passed through a denoiser. The denoiser outputs a denoised or "cleaned" image. A denoised image will not typically have all noise removed.

[0037] Digital images are composed of pixels. Different interpretations of the term "pixel" are common. In one interpretation, a single pixel can encode multiple primary colors. In this interpretation, a pixel can include separate red ("R"), green ("G"), and blue ("B") values ​​or areas. In another interpretation, a pixel can encode a single color. For example, an image in RAW format is divided into multiple 2×2 blocks, each block includes four elements (see Figure 3 ). Each element encodes a single color, such as Figure 3 The elements in each block can be considered as pixels. In a monochrome image, a pixel can encode only grayscale data.

[0038] Noise in digital camera images comes from three main sources:

[0039] 1. Read noise: Generated by the electronics in the imaging sensor. This is random noise that can usually be modeled with a Gaussian distribution.

[0040] 2. Shot noise: related to the quantization of the number of photons reaching the sensor, and therefore depends on the brightness. This is random noise and can usually be modeled with a Poisson distribution.

[0041] 3. Bad pixel noise: Individual pixels on a camera sensor may have different sensitivities to incident light. This can be a result of the sensor manufacturing process or due to random failures during service. While the incidence of these effects can change gradually over time, the change is usually slow. Bad pixels in pixels can cause impulse noise effects: that is, causing the pixel to become saturated (therefore reporting maximum brightness) or not capture light (therefore reporting minimum brightness). It can also cause pixels to report incorrect intermediate levels. Figure 4 An example of bad pixel noise is shown.

[0042] In the method described below, an artificial intelligence model is trained to detect noise in an image. This can be done by training the model on a dataset that includes multiple pairs of images of the same scene, one image in each pair showing considerable noise and the other image being substantially noise-free. For each pair of images, a relatively noisy image is input to the model in its current state. The model estimates the noise in the image. A hypothetical denoised image is formed by removing the estimated noise from the relatively noisy image. The hypothetical denoised image is then compared to the relatively noise-free image. Based on the results of the comparison, the model is adapted (e.g., by changing the weights in a neural network included in the model) to strengthen the model's accurate noise estimation. Once the model is trained in this manner on multiple training image pairs, it can be run on other images to estimate noise in other images. Conveniently, the model can be trained on one device, such as a computer server, and then loaded into other devices, such as a camera or a mobile / cellular phone that includes a camera. These other devices can then run the model to help remove noise from images captured by the device. Although Figure 1 A device of the type shown in can be used to train a model, but the device running the training model need not have the ability to train the model itself. The training model can be provided to the device in the form of executable code, or, if the device already has executable code suitable for the general model, the training model can be provided to such device in the form of weights or values ​​to be used in the model. The training model can be provided to the device dynamically (e.g., as a download) at the time of manufacture or in non-volatile memory on a storage device. The training model can be stored in a suitable memory in a non-transient form.

[0043] In addition to training the model based on the appearance of the training images, the ISO or sensitivity at which the images were taken can also be used as input for training. It can also be used as input when running the trained model.

[0044] Preferably, the system uses the appearance of the image and the ISO used (known when the photo was taken) to estimate the noise in the image. The system decomposes the noise into (i) random noise estimates, preferably including Gaussian noise and Poisson noise and (ii) deterministic noise, preferably including bad pixel noise. A separate subnet can be used to estimate each of the above noise. It has been found that using separate trainable subnets or training subnets for these two types of noise can produce improved results in terms of the accuracy of estimating noise. The network can be beneficially trained in a multi-task setting. One task is to estimate random noise and the other task is to estimate deterministic noise. In a convenient embodiment, there can be separate subnets for these two estimates. The random noise (i.e., Gaussian+Poisson noise) and the deterministic noise (i.e., bad pixel noise) are combined to form an overall noise estimate. The original image is then denoised using the overall noise estimate. This can be achieved by subtracting the noise estimate from the original image or using another trained denoising network.

[0045] Each subnetwork can be pre-trained individually using synthetic noise. Thus, the above training process can be implemented separately on each network, in each case with relatively noisy input images consisting mostly of the type of noise being trained (random or deterministic / bad pixel noise). The two pre-trained models can then be combined into a complete architecture. If desired, the complete architecture can be trained using images that include both types of noise: for example non-synthetic or real image data.

[0046] To train the subnetworks individually, noise can be synthesized and applied to non-noisy training images to form relatively noisy training images. To synthesize random noise, a series (e.g., 12) images captured by a static camera in a low-light environment can be used. These frames can be averaged to obtain an average image, which is used as a relatively low-noise training image. Then, a variance image can be calculated for each pixel in the sequence. Poisson noise in the image is intensity-dependent. Therefore, a linear equation can be fit to the noise variance as a function of intensity / brightness using the least squares method and / or RANSAC. A random (Gaussian + Poisson) noise model can be used to characterize the noise using this linear equation. Any pixel in the image in the sequence that shows noise that is inconsistent with the model can be considered a bad pixel noise. Conveniently, any pixel whose intensity is outside the 99% confidence interval of the estimated random noise distribution can be considered a bad pixel. In this way, for each image in the sequence, the estimate can be formed by (a) random noise and (b) deterministic noise. Then, for each image in the sequence, two images can be formed: one containing only random noise, and one containing only deterministic noise. Each of these can be paired with a relatively low-noise image and used to train the corresponding part of the model.

[0047] RAW images are a lossless image format. Unlike RGB images, RAW images are composed of a Bayer array. Each pixel block in the Bayer array can be BGGR, RGGB, GRGB, or RGGB. At each pixel, there is only red, green, or blue. During training and runtime, it is beneficial to use RAW images as input to the model. Using RAW images can preserve more details than images with dynamic range compression applied. The Bayer array structure of RAW images (see Figure 3 ) makes it difficult to perform image enhancement using some traditional methods. Using the above split model method, each model can be trained to denoise one of the random noise and the impulse noise respectively.

[0048] The system provides a decomposition network to estimate noise decomposed into (i) random / Gaussian+Poisson noise and (i) deterministic / impulsive noise. Figure 5Such a network is shown schematically. The overall structure is a network consisting of two subnetworks. Each subnetwork is a neural network that operates independently of the other. The weights in one subnetwork do not affect the operation of the other. One of the subnetworks estimates random noise. The other estimates deterministic noise. The subnetworks can have any suitable form, but it is common to say that they can be networks that each regress the noise value at each pixel in the input RAW data.

[0049] exist Figure 5 In the architecture of , the following steps are performed on a single input image. In this example, it is assumed that the input image is in RAW format, but it can be in any suitable format.

[0050] 1. First pack the input image into four channels (R, G1, G2, B), corresponding to the Bayer array used in RAW encoding. For other forms of encoding, the packing method may be different. The resolution of these channels in width and height units is half of the original input image. One advantage of packing in this way is that pixels of the same color are grouped together.

[0051] 2. Provide the ISO setting of the captured image as input on the fifth channel.

[0052] 3. The network can use "shared layer" blocks to extract features. This can be done using convolutions and rectified linear units (ReLU) or other methods. Shared layers extract features that are common to the tasks of the two subnets. Shared layers can also be called coupling layers. These layers can be used to assemble features at relatively high resolution to establish a connection between the two subnets. After pre-training each subnet separately, these coupling layers enable the overall architecture to update the weights of each subnet based on training data that includes two types of noise.

[0053] 4. The upper subnet branch (“Subnet A”) estimates random noise points.

[0054] 5. The lower subnetwork branch (“Subnetwork B”) estimates the deterministic noise.

[0055] 6. Add the noise estimates to form an aggregate noise estimate. The aggregate noise estimate can be subtracted from the original image to produce a denoised result.

[0056] Each subnetwork A or B may be independently formed using any suitable trainable network architecture. Examples include neural networks, such as Unet. (See Ronneberger, Olaf, Philipp Fischer, and Thomas Brox, "Unet: Convolutional Networks for Biomedical Image Segmentation," in International Conference on Medical Image Computing and Computer-Assisted Intervention, Cham, Springer, 2015.) Conveniently, Unet may be modified, such as Figure 6 shown. Figure 6 The steps of downsampling and then upsampling the input image that can be performed in the example of subnetwork A or B are shown. Figure 6 The image input is provided in the upper left corner of the process shown. This input is located at Figure 6 The top level of the data flow. Figure 6 In , the rectangles represent the convolution and leaky rectified linear unit (ReLU) steps. Figure 6 In the downsampling path shown on the left, the input is converted to data for the second layer using the first spatial-depth transform. The data for the second layer is converted to data for the third layer using the second spatial-depth transform. The data for the third layer is converted to data for the fourth layer using the first maximum pooling transform. The data for the fourth layer is converted to data for the fifth layer using the second maximum pooling transform. Figure 6The upsampling path is shown on the right. A jump connection is provided between the data in each layer of the downsampling path and the data in the corresponding layer of the upsampling path. The upsampling path starts with the fifth layer of upsampled data derived from the jump connection of the fifth layer of upsampled data. The upsampled data of the fifth layer is converted to the upsampled data of the fourth layer using a first deconvolution transform. The upsampled data of the fourth layer is converted to the upsampled data of the third layer using a second deconvolution transform. The upsampled data of the third layer is converted to the upsampled data of the second layer using a first depth-space transform. The upsampled data of the second layer is converted to the upsampled data of the top layer using a second depth-space transform. The depth-space transform reduces the resolution of the input data of these transforms, but increases the data depth of the value at the resulting pixel position according to the input data. Therefore, these transforms reduce the resolution of the data without losing information proportional to the reduction in resolution. Preferably, all data is retained by storing data defining pixels, wherein the pixels are lost in the process of reducing the resolution in the increased data depth at each remaining pixel. In this way, the space-depth transform can be lossless. The space-depth transform performs the opposite operation. The max pooling transform reduces the resolution by selecting only the pixels with the maximum value in the multi-pixel region to be reduced to a single pixel. The data depth is maintained. In this way, the max pooling transform is lossy. The deconvolution transform can form an estimate from a low-resolution image to a high-resolution image. It has been found that while the max pooling operation is relatively effective as part of the process of identifying bad pixels (because bad pixels usually produce extreme values), the spatial-depth operation has an advantage in the process of identifying random noise, which may usually produce intermediate noise values.

[0057] Figure 7 An alternative architecture is shown. Figure 7 The architecture and Figure 5 The architecture is similar to that of Figure 7 Some key points to note about the architecture are:

[0058] 1. Use different networks in the subnetworks to detect random noise and deterministic noise. The subnetwork used to detect random noise ignores the maximum pooling. The subnetwork used to detect deterministic noise retains the maximum pooling (e.g., Figure 6 This can improve the efficiency and accuracy of detecting bad pixels and noise.

[0059] 2. The noise estimate formed by the two subnetworks is concatenated with the original image to a common dataset, and all three are passed to the denoising network. The denoising network applies a trained neural network model to form an adapted noise estimate. The adapted noise estimate is then subtracted from the original image in a subtraction step to form an adjusted image. The performance of the system can be improved compared to the alternative approach of summing the noise estimates from the two subnetworks and subtracting the noise estimate from the input image. The denoising network can be any suitable network trained for this purpose. The noise estimates from each subnetwork can be packed into multiple color-specific channels (e.g., R, G, G, B in the case of RAW images) for input to the cascade block.

[0060] Figure 8 A comparison is shown between an example denoised image by the present system (right) and an image denoised using the well-known method DnCNN. (See Zhang, Kai et al., "Beyond a gaussian denoiser: Residual learning of deep CNN for image denoising", IEEE Transactions on Image Processing, 26.7 (2017): 3142-3155). Fig. 9 An image denoised by DnCNN, an image denoised by an example of the present system, and a zoomed-in and cropped portion of the input image (from left to right) are shown. The circled area in the leftmost image is the artifact area generated by the DnCNN method.

[0061] When the captured image is being denoised, the system is preferably applied before operations such as de-stitching and dynamic range compression.

[0062] Figure 5 and Figure 7 The entire network can be trained end-to-end using noisy images / clean ground truth image pairs. This can improve the accuracy of the resulting model.

[0063] Applicants hereby disclose separately each individual feature described herein and any combination of two or more such features. With the common knowledge of those skilled in the art, such features or combinations can be implemented as a whole based on this specification, without considering whether such features or combinations of features can solve any problem disclosed herein; and without causing any impact on the scope of the claims. This application shows that various aspects of the present invention can be composed of any such individual features or combinations of features. In view of the foregoing description, it will be obvious to those skilled in the art that various modifications can be made within the scope of the present invention.

Claims

1. An image processing device, characterized in that: A method for estimating noise in an image, wherein the image is represented by a set of pixels and each pixel has a value associated with the pixel in each of one or more channels, by: processing data derived from the image through a first trained model for detecting random noise to form a first noise estimate; processing data derived from the image to form a second noise estimate by a second trained model for detecting pixel extrema; wherein the first trained model and the second trained model include a processing architecture for: (a) processing the data derived from the image to gradually reduce the resolution through a first series of stages to form intermediate data; and (b) processing the intermediate data to gradually increase the resolution through a second series of stages to form a corresponding noise estimate, wherein there are jump connections between corresponding stages of the first series and the second series to provide feedthrough; The first noise estimate and the second noise estimate are combined to form an aggregated noise estimate.

2. The image processing device according to claim 1, characterized in that The first training model is used to detect Poisson noise and / or Gaussian noise.

3. The image processing device according to claim 2, characterized in that The first training model has higher accuracy in detecting Poisson noise and / or Gaussian noise than the second training model.

4. The image processing device according to any one of the preceding claims, characterized in that The second training model has higher accuracy in detecting bad pixel noise than the first training model.

5. The image processing device according to claim 1, characterized in that The apparatus is for subtracting the aggregated noise estimate from the image to form a denoised image.

6. The image processing device according to claim 1, characterized in that The device is used to process the image and the aggregated noise point estimate through a third training model to form a denoised image.

7. The image processing device according to claim 1, characterized in that The first series of stages of the second training model include: (a) a first stage for processing data derived from the image to reduce the resolution and increase the data depth to form second intermediate data, and (b) a second stage for processing the second intermediate data to reduce the resolution without increasing the data depth to form third intermediate data.

8. The image processing device according to claim 7, characterized in that The first level is the space-depth level.

9. The image processing device according to claim 7 or 8, characterized in that: The second stage is the max pooling stage.

10. The image processing device according to claim 7, characterized in that None of the first series of stages in the first training model increases data depth.

11. A method for training an image processing model, characterized in that: include: (a) receiving a plurality of pairs of images, wherein each pair of images represents a common scene, and a first image in each pair of images includes more noise than a second image in each pair of images; (b) For each pair of images: (i) processing data derived from the first image in each pair of images through a first model for estimating random noise in the images to form a first noise estimate; (ii) processing data derived from the first image in each pair of images through a second model for detecting pixel extrema to form a second noise estimate; (iii) combining the first noise estimate and the second noise estimate to form an aggregated noise estimate; (iv) estimating a difference between (A) the second image in each pair of images and (B) the first image in each pair of images denoised according to the aggregated noise estimate; (v) adapting the first model and the second model according to the estimated difference; The first model and the second model include a processing architecture for: (a) processing data derived from the first image in each pair of images to gradually reduce the resolution through a first series of levels to form intermediate data; and (b) processing the intermediate data to gradually increase the resolution through a second series of levels to form corresponding noise estimates, wherein there are jump connections between corresponding levels of the first series and the second series to provide feedthrough.

12. The method according to claim 11, characterized in that The first series of stages of the second model include: (a) a first stage for processing data derived from the first image in each pair of images to reduce the resolution and increase the data depth to form second intermediate data; and (b) a second stage for processing the second intermediate data to reduce the resolution without increasing the data depth to form third intermediate data.

13. An image processing device, characterized in that: Adapted from the method of claim 11 or 12.

14. An image processing device, characterized in that: The method comprises a processor and a memory, wherein the memory stores instructions executed by the processor in a non-transitory form to implement an image processing model adapted according to the method of claim 11 or 12.