Method for restoring blurred image containing noise based on network fusion
By integrating convolutional neural network, transformer network and diffusion model, multi-scale CNN and transformer are used to generate initial defuzzy results and residual details, and iterative reverse denoising under the diffusion framework, it solves the problem of difficulty in dealing with noise and dynamic blurred images in the prior art, and achieves efficient image restoration effect.
Patent Information
- Application Number
- CN202510405554.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively process images containing both noise and dynamic blur. Traditional methods rely on overly idealized assumptions and are difficult to obtain the necessary prior knowledge. Convolutional neural networks and transformer networks have their own limitations, and the diffusion model has shortcomings in generation speed and controllability.
A network fusion-based method is adopted to integrate convolutional neural network, transformer network and diffusion model, generate initial defuzzy results through multi-scale CNN, and use transformer to generate residual details, and iterative reverse denoising is performed under the diffusion framework, combining full-size Fourier and diffusion joint loss function for training, further enhancing robustness.
It realizes efficient restoration of images containing both noise and dynamic blur, improves the performance and robustness of image debum, and proves the feasibility of network fusion ideas in the field of image restoration.
Smart Images

Figure CN119991501A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image restoration, and in particular relates to a restoration method for noisy fuzzy images based on network fusion. Background Art
[0002] Image deblurring is a fundamental yet challenging task in computer vision, involving resolving various types of blur, such as motion blur, out-of-focus blur, Gaussian blur, and mixed blur. Motion blur, which can be further divided into global motion blur and local motion blur, is usually caused by camera shake or rapid object motion. The pathological nature of recovering a sharp image from a blurry image further complicates the problem. Traditional methods attempt to simplify this problem by making assumptions that deviate from realistic scenarios - such as uniform blur across the image - and rely heavily on prior knowledge during the deblurring process. In contrast, the advent of deep learning, especially convolutional neural networks, has marked a significant advance in this field. In order to improve the complexity and representation power of convolutional neural networks, researchers have proposed various structural improvement strategies, including multi-scale, multi-patch, and multi-temporal architectures.
[0003] Meanwhile, transformers, originally designed for natural language processing and other sequence-to-sequence tasks, have become a powerful tool in computer vision due to their core self-attention mechanism. Beyond this, generative diffusion models have gained prominence due to their impressive performance in tasks such as image generation, painting, and video generation. Unlike convolutional neural networks and transformers, which primarily employ regression methods, diffusion models leverage their powerful data generation capabilities, making them particularly effective in complex visual tasks. Together, these advances reflect the growing sophistication and diversity of techniques used to address the image deblurring challenge.
[0004] Traditional methods are often based on overly idealized assumptions, and it is difficult to obtain the necessary prior knowledge, making these methods unsuitable for handling real-world motion blurred images. In addition, the inherent limitations of convolutional neural networks - such as restricted kernel size, shallow network architecture, and limited receptive field - restrict their ability to effectively capture long-range pixel dependencies. In contrast, transformer networks, while capable of addressing the challenges of long-range pixel interactions, impose considerable computational overhead, resulting in significant demands on hardware resources and processing time. Similarly, diffusion models, while promising for generation tasks, suffer from slow generation speed and limited controllability, which can lead to undesirable artifacts, including irrelevant details or ghosting effects in the generated output.
[0005] In summary, different network strategies have their inherent problems, which are difficult to overcome by a single network strategy, which is one of the dilemmas that cannot be ignored in the current deep learning deblurring sub-network. In addition, limiting the model task to simple deblurring tasks without considering more complex deblurring situations also limits the development and application of deblurring sub-networks. In order to solve these two difficulties, we designed our model. Summary of the invention
[0006] In view of the above-mentioned technical problems, the purpose of the embodiments of the present application is to propose a method for restoring noisy blurred images based on network fusion to solve the technical problems mentioned in the above background technology part.
[0007] The integration of convolutional neural networks and transformers with diffusion models represents an innovative deblurring approach that aims to leverage the strengths of multiple network architectures while addressing their inherent limitations. The framework uses a multi-scale CNN to generate an initial deblurring result, and then uses an intra-strip and inter-strip attention transformation network to generate the residual between the initial output and the true image, further improving the final deblurring result. The two subnetworks are unified under the diffusion framework. That is, the powerful image generation ability of the diffusion model is utilized to diffuse the residual details from pure noise through continuous iterations. To alleviate the computational complexity associated with the transformer structure during the inverse denoising process, the model strategically balances the contributions of convolutional neural networks and transformers, optimizing efficiency without sacrificing performance. In addition, a novel full-scale Fourier and diffusion (FFD) joint loss function is introduced to coordinate the training of these components. To further enhance robustness, especially when dealing with noisy and severely blurred images, the system combines an external module denoising subnetwork, enabling the model to achieve the restoration task of images containing both noise and blur.
[0008] The present invention provides a method for restoring a noisy fuzzy image phenomenon based on network fusion, comprising the following steps:
[0009] Obtain an image to be restored that contains both noise and motion blur.
[0010] The restoration network for noisy and blurred images based on network fusion mainly includes three sub-networks, which are connected together by a series structure. The input image will pass through the three sub-networks in sequence to obtain the final restored image; therefore, the first thing to do is to build a denoising sub-network based on the transformer network framework, and input the image containing both noise and motion blur into it, so as to achieve image denoising and retain the blur information to the greatest extent and output the deblurred result, that is, an image containing only dynamic blur.
[0011] The next step is to build a preliminary deblurring subnetwork based on the convolutional neural network framework. The input is an image that only contains dynamic blur. The preliminary deblurring subnetwork will first downsample the image twice to obtain input images of three different scales. The network extracts features from input images of different scales, and encodes and decodes these features to obtain a preliminary deblurred image.
[0012] The last step in building the network is to construct a residual detail generation network based on the transformer and diffusion models. The image that has been initially deblurred is input. In the pre-training stage, the network will first obtain the residual details between the initial deblurred image and the real clear image, and then gradually add noise to transform the residual detail image into a pure noise image. Then, through the reverse denoising process of the diffusion model, a residual detail image is generated from the pure noise image. In the testing stage, the residual detail generation network takes a random pure noise image and a preliminary deblurred image as input, and obtains the residual details through reverse denoising of the diffusion model. Finally, the preliminary deblurred image and the residual detail image are combined element by element, and the network realizes the restoration task of images that contain both noise and dynamic blur. The emergence of this method solves the restoration task of images that contain both noise and blur, and at the same time proves the feasibility of network fusion ideas in the field of image restoration.
[0013] Based on Markov chain theory, the forward process of the diffusion model can be formally defined as:
[0014]
[0015] Among them, q(x t |x t-1 ) represents the forward process of the diffusion model, that is, the process of reasoning from the previous time step t-1 to the current time step t, t~{0,1,…,T} represents the steps from the clear image x0 to Gaussian noise, x t represents the noise image at time t, α t and β t are two predefined parameters, β t =1-α t , N represents Gaussian distribution. There are no additional parameters that need to be estimated in the forward process. From the above formula, we can deduce x t-1 to x t The process is:
[0016]
[0017] Where ε~N(0,1). Using similar ideas, the reverse process can be defined as follows:
[0018] p θ (x t-1 |xt )=N(x t-1 ;μ θ (x t ,t),∑ θ (x t ,t))
[0019] where p θ (x t-1 |x t ) represents the reverse process of the diffusion model, that is, the process of reverse deduction from the current time step to the previous time step, μ θ and∑ θ Respectively represent x t and t determine the mean and variance. Both are also key tasks for net estimation. Fortunately, the first clear image x0 is also a known prerequisite. Using x0, the above formula can be rewritten as:
[0020]
[0021] By combining the above formula with the Bayesian formula, the two key parameters of the reverse process can be expressed as:
[0022]
[0023] in and represent variance and absolute value respectively.
[0024] In the face image restoration method based on diffusion generation prior, the restoration method for noisy blurred images adopts the idea of network fusion. The so-called network fusion is an important idea in deep learning, which aims to improve the overall performance by combining the advantages of multiple neural network models; it mainly includes model integration, feature fusion, model distillation, dynamic network fusion and other ideas.
[0025] By adopting the method of model integration and dynamic network fusion, a restoration network for images containing both noise and dynamic blur is constructed, which consists of three sub-networks. The three sub-networks are a denoising sub-network for removing image noise and preserving the original blur information, a preliminary deblurring sub-network for preliminary deblurring, and a residual detail generation network for generating residual details.
[0026] In fact, the denoising subnetwork that removes image noise and preserves the original blur information is the same as the residual detail generation network at the network structure level. It is just because of the different functional designs and pre-training that the two subnetworks exhibit different functions. The backbone of the denoising subnetwork and the residual detail generation network includes an encoder, an intermediate attention module and a decoder. The encoder contains four encoding modules connected in sequence (each of which includes a residual module and a multi-layer perceptron) and a downsampling module, in which the minimum-scale encoding module and the next-small-scale encoding module each contain an encoding attention module; the decoder contains four decoding modules connected in sequence (each of which includes a residual module and two multi-layer perceptrons) and an upsampling module, in which the minimum-scale decoding module and the next-small-scale decoding module each contain an encoding attention module.
[0027] In the denoising sub-network and the residual detail generation network, the structure of the transformer network is used to obtain better results. Specifically: by adding the attention module, based on the transformer network in the attention module, we build the denoising sub-network into a complete transformer network; similarly, as the residual detail generation network of the diffusion model's reverse denoising process, the addition of the attention module makes it actually a fusion network of the diffusion model and the transformer network, so as to achieve better residual detail generation through deeper feature extraction.
[0028] During the training phase, the input of the denoising sub-network is the image to be restored that contains both noise and motion blur. The encoder extracts feature vector information through the encoding module and the encoding attention module, and builds a three-layer network structure through the downsampling module to obtain a deeper network depth; the intermediate attention module further extracts feature information and passes the information to the decoder; the decoder also decodes the information through the decoding module and the decoding attention module, and restores the information scale using the upsampling module; the denoising sub-network adjusts the parameters through the multi-scale loss function and the multi-scale frequency loss function to obtain a denoised image containing only blur; no additional network behavior is included during testing, and the loss function is not activated.
[0029] Different from the denoising sub-network, during the training phase, the input of the residual detail generation network is a pure Gaussian noise image and a preliminary deblurred image. After passing through the same network structure as the denoising sub-network, the result is the detail residual generated from the noise. The network adjusts the parameters through the diffusion loss function. Similarly, no additional network behavior is included during testing, and the loss function is not activated.
[0030] The preliminary deblurring subnetwork used for preliminary deblurring mainly adopts the Unet network structure in the convolutional neural network. The encoder contains three convolution modules, three residual modules and two feature fusion modules; the decoder contains three residual modules and two convolution modules; the feature transfer between the encoder and the decoder is realized through the full-scale channel feature fusion module.
[0031] The so-called full-scale channel feature fusion module is a module we designed to transfer feature vectors between the encoder and the decoder. It first receives the feature information of three different scales output by the encoder, and uses the scale conversion layer to unify it to the current scale, and then concatenates it. It then fuses these feature vectors through a 3*3 convolutional layer and a 1*1 convolutional layer, and inputs the final fusion result into the encoder of the corresponding scale.
[0032] During the training phase of the preliminary deblurring subnetwork, the input image will first be downsampled twice to obtain inputs of three different scales. After the image is converted to a feature vector, they are input into the encoders of the corresponding scales. It should be noted that after the large-scale feature vector is extracted, it will also be downsampled once to be input into the small-scale encoder. Subsequently, after receiving the encoded feature vector, decoders of different scales will decode the information and generate an image of the current scale as output, which corresponds one to one with the input. It should be noted that the small-scale decoded information will also be used as the input of the previous decoder after upsampling. After completing a deblurring, the network will evaluate the current deblurring effect through a multi-scale loss function and a multi-scale frequency loss function, and adjust the network parameters. The network behavior in the test phase is the same as that during training, and the loss function is not activated.
[0033] The residual detail generation network and the preliminary deblurring sub-network adopt the idea of dynamic network fusion and adopt the joint training method to control the weights of the two networks through the joint loss function. The specific loss function can be described as:
[0034] In the multi-scale loss function of the preliminary deblurring sub-network, the characteristics of multi-scale input and output of the network are fully utilized:
[0035]
[0036] Where K is the maximum scale of the network structure. is the deblurred output of scale k, S k is a clear image of the corresponding scale.
[0037] For the multi-scale frequency loss function of the preliminary deblurring subnetwork, since the purpose of deblurring is to recover the lost high-frequency components, it is also effective to minimize the difference operation in the free frequency space, which leads to the design method of additional auxiliary loss terms. Therefore, the multi-scale frequency loss is designed as an auxiliary loss term as follows:
[0038]
[0039] where F(·) represents the Fast Fourier Transform (FFT) that transfers the image signal to the frequency domain.
[0040] Combining the above two parts of the loss function, the composition of the complete loss function required for the preliminary deblurring sub-network is as follows:
[0041] L FSF =L cont +λL freq
[0042] where λ is the polynomial coefficient.
[0043] During the training phase, the following loss function is designed for the residual detail generation network:
[0044] As the second part of the two-network combination, the residual detail generation network is trained while taking into account the previous preliminary deblurring subnetwork. In the following formula, CP(·) represents the preliminary deblurring subnetwork operation, and Denoiser(·) represents the residual detail generation network operation:
[0045] x CP =CP(input blur )
[0046] x D =Denoiser(x CP )
[0047] The input blur represents the input image containing only motion blur, x CP represents the initial deblurred image, x D Represents the generated residual details.
[0048] L pixel =Ε||x CP -x gt ||2
[0049] L DM =Ε||x0-x D ||2
[0050] Where x0 represents the ideal residual detail between the initial deblurred image and the true sharp image.
[0051] L low =Ε||φ L ×(x CP -x gt )||2
[0052] L high =Ε||φ H ×(x0-x D )||2
[0053] where φ L and φ H They represent low-frequency filtering and high-frequency filtering respectively, and finally the above results are combined to obtain .
[0054]
[0055] where β0 and β1 represent the polynomial coefficients.
[0056] Combining all the loss function sub-formulas, we get the following complete loss function.
[0057] L total =L FSF +αL Diffusion
[0058] Where α represents the polynomial coefficient.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] (1) This paper proposes an image restoration model based on the diffusion model, which successfully realizes image restoration under complex degradation conditions. The success of this model also proves the inclusiveness of the diffusion model. Many different network schemes can be applied under the framework of the diffusion model to achieve a unified ultimate goal.
[0061] (2) The present invention designs a full-scale channel feature fusion block for the convolutional neural network to realize cross-scale multi-channel feature fusion; designs an encoding self-attention block and an intermediate self-attention block for the diffusion model transformer structure to realize the feature information collection in different directions in the image; and designs a full-scale Fourier and diffusion joint loss function (FFD) to ensure the smooth progress of network training.
[0062] (3) In order to improve the robustness of the entire model, the present invention designs an additional network specifically for removing noise interference in noisy blurred images and retaining simple blur information as much as possible. In this way, when facing a more complex deblurring task, it is only necessary to send the image to the denoising subnetwork for denoising and then input it into the original deblurring subnetwork to obtain the required clear image to the greatest extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0064] Figure 1 is a method flow chart of an embodiment of the present invention;
[0065] Figure 2 is a complete structural diagram of an embodiment of the present invention;
[0066] Figure 3 is a schematic diagram of the structure of a preliminary defuzzification subnetwork in an embodiment of the present invention;
[0067] Figure 4 Schematic diagram of the structure of a detail residual generation network and a preliminary restoration network in an embodiment of the present invention;
[0068] Figure 5 is a structural diagram of an image feature extraction module in an embodiment of the present invention;
[0069] Figure 6 is a structural diagram of a feature fusion module in an embodiment of the present invention;
[0070] Figure 7 is a structural diagram of a full-scale channel feature fusion module in an embodiment of the present invention;
[0071] Figure 8 is a structural diagram of an attention module within a strip in an embodiment of the present invention;
[0072] Fig. 9 It is a structural diagram of the inter-strip attention module in an embodiment of the present invention. DETAILED DESCRIPTION
[0073] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0074] Figure 1 It is shown that a method for restoring a noisy fuzzy image phenomenon based on network fusion provided by an embodiment of the present application includes the following steps:
[0075] S1 obtains the image to be restored that contains noise and motion blur.
[0076] S2 denoises images containing both noise and motion blur through a denoising sub-network, and preserves the original blur information to the greatest extent possible.
[0077] Specifically, the denoising sub-network is identical to the residual detail generation network at the network structure level, but due to the differences in functional design and pre-training, the two sub-networks exhibit different functions; its network backbone consists of three parts: encoder, intermediate attention module and decoder; the encoder contains four encoding modules (including one residual module and one multi-layer perceptron) and one encoding attention module; the decoder contains four encoding modules (including one residual module and two multi-layer perceptrons) and one decoding attention module.
[0078] S3 uses a preliminary deblurring subnetwork to perform preliminary deblurring on images containing only motion blur.
[0079] Specifically, the preliminary deblurring subnetwork mainly adopts the Unet network structure in the convolutional neural network. The encoder contains three convolution modules, three residual modules and two feature fusion modules; the decoder contains three residual modules and two convolution modules; the feature transfer between the encoder and the decoder is realized through the full-scale channel feature fusion module.
[0080] The S4 residual detail generation network uses the input preliminary deblurred image to generate residual details from pure noise.
[0081] Specifically, since the residual detail generation network and the denoising sub-network for generating residual details have the same structure, they will not be described in detail; the residual detail generation network and the preliminary deblurring sub-network adopt the idea of dynamic network fusion and the method of joint training. The weights of the two networks are controlled by the loss function, so that the two networks can cooperate with each other and jointly achieve image deblurring.
[0082] S5 finally combines the results of the two deblurring sub-networks to obtain the final deblurring result.
[0083] S6 outputs the obtained result, and the image at this time no longer contains noise and motion blur.
[0084] Figure 2It is a schematic diagram of the complete structure of the entire model. Specifically, a denoising sub-network based on the transformer network framework is constructed. By inputting images containing both noise and motion blur into it, it is possible to denoise the image and retain the blur information to the greatest extent and output the deblurred result. A preliminary deblurring sub-network based on the convolutional neural network framework is constructed. By inputting images containing only dynamic blur, the preliminary deblurring sub-network will first downsample the image twice to obtain input images of three different scales. The network extracts features from input images of different scales, and encodes and decodes these features to obtain preliminary deblurring results. A residual detail generation network based on both the transformer and diffusion models is constructed. The input image is Through the preliminary deblurred image, in the pre-training stage, the network will first obtain the residual details between the preliminary deblurred image and the real clear image, and then transform the residual detail image into a pure noise image through gradual denoising, and then generate a residual detail image from the pure noise image through the reverse denoising process of the diffusion model; in the testing stage, the residual detail generation network takes the random pure noise image and the preliminary deblurred image as input, and obtains the residual details through reverse denoising of the diffusion model; finally, the preliminary deblurred image and the residual detail image are combined together, and the network realizes the restoration task of the image containing both noise and dynamic blur; the emergence of this method solves the restoration task of images containing both noise and blur, and at the same time proves the feasibility of network fusion ideas in the field of image restoration.
[0085] Figure 3 The detailed structure of the preliminary deblurring subnetwork is shown; specifically, the preliminary deblurring subnetwork first scales the input image to 1 / 2 and 1 / 4 of the original scale by interpolation to achieve feature extraction of images from different scales. The image feature extraction module extracts spatial and local texture information through convolution operations. The convolution block obtains the feature information of the input image by stacking 1×1 and 3×3 convolution layers. In order to make full use of the feature vector space information at different scales, the feature fusion module not only extracts the current scale information between different scales, but also receives the feature information of the previous scale, integrates the feature information of the two scales, and outputs the feature information of the current scale. The feature information of the three scales is fused by the full-scale channel feature fusion module at different scales to generate the fused feature information of the corresponding scale and input it into the subsequent decoder. The decoder integrates the output of the full-scale channel feature fusion module of the current scale and the output information of the previous scale to generate the initial deblurring result of the current scale. In other words, the decoder will eventually generate three different scales of deblurring results, corresponding to the three scales of the input. Finally, the preliminary deblurring subnetwork selects the deblurring result with the same scale as the original input as the output result of the network.
[0086] The full-scale channel feature fusion module is mainly used to solve the problem of lack of information flow communication in traditional convolutional neural networks. It fuses the input information of all scales and outputs the fusion information of the current scale. The full-scale channel feature fusion module first receives input information of different scales, unifies the information scale by downsampling, and then passes through a 1×1 convolution layer, a relu activation layer, a 3×3 convolution layer, and finally outputs it. If the operation of the full-scale channel feature fusion module is expressed in formula form, it can be expressed as:
[0087]
[0088] in refers to the output of the ACFF module at scale i, Refers to the output of the encoder at scale i, ↑ represents upsampling, and ↓ represents downsampling.
[0089] Figure 4 The detailed structure of the residual detail generation network is shown. Compared with the general diffusion model, the residual detail generation network uses a transformer network structure based on Unet instead of a traditional linear network structure, and chooses to calculate the image residual instead of the entire sharp image to simplify the complexity of the entire network as much as possible and improve the network's running speed. In order to achieve the reverse denoising process, the residual detail generation network uses intra-strip and inter-strip self-attention blocks. The structure is superimposed by the interlocking intra-strip self-attention module and the inter-strip self-attention module. Figure 8 and Fig. 9 As shown in Figure 1, a coded self-attention block (CSAB) and a middle self-attention block (MSAB) are designed. The main difference is that CSAB contains only one set of superposition modules, while MSAB contains six sets of superposition modules, which aims to restore more realistic image details at different positions in the denoising network. For training, the detail residual generation network will be trained end-to-end together with the preliminary deblurring subnetwork, rather than separately, to ensure that the two modules complement each other and achieve better deblurring effects. The specific structures of the image feature extraction module, feature fusion module and full-scale channel feature fusion module are shown in Figure 1. Figure 5 , Figure 6 and Figure 7 shown.
[0090] All attention modules consist of two parts: intra-strip self-attention block and inter-strip self-attention block. Figure 8 and Fig. 9The specific structures of the core modules in the encoding attention module and the intermediate attention module: the intra-strip attention module and the inter-strip attention module are shown; for these two modules, the input feature tensor is divided into two parts according to the channel dimension, named horizontal features and vertical features, and the query, key, and value are obtained through convolution and cutting. The intra-strip attention module performs attention calculations on the query, key, and value in the vertical and horizontal directions respectively, while the inter-strip attention module performs feature fusion to obtain the fused query, fused key, and fused value after retrieving the query, key, and value, and then performs attention calculations on the fused query, fused key, and fused value.
[0091] Two widely used indicators are used to quantitatively evaluate the entire dataset: Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM). In terms of pure deblurring ability, the detailed comparison results of GoPro, RealBlur and REDS datasets are shown in Tables 1 and 2. The two sets of data show that ICTD has achieved good performance in the field of image deblurring.
[0092] Table 1
[0093]
[0094] Noise is an inherent feature of real images, and when noise is introduced, it destroys the blur kernel and motion information in the original blurred image. Although current deep learning-based deblurring subnetworks do not rely on explicit blur kernel estimation or motion trajectories, they still have difficulty producing effective deblurring results in the presence of noise.
[0095] The comparison results shown in Table 2 are all on the self-made dataset GoPro-Noisy. The comparison results show that all models, including ICTD without the denoising sub-network module, cannot effectively restore the deblurred GoPro-Noisy dataset. However, this problem is significantly alleviated when the denoising sub-network is incorporated into the ICTD network. The first row of the table highlights the difference between low-quality images and clear images in the original GoPro-Noisy dataset. Many models cannot address the challenges posed by noisy blurry images, and some even exacerbate the problem. In contrast, ICTD with a denoising sub-network can successfully restore images. Specifically, compared with the original GoPro-Noisy dataset, the combination of ICTD and the denoising sub-network improves PSNR by 5.75dB, and the performance of both PSNR and SSIM indicators exceeds that of other models.
[0096] Table 2
[0097]
[0098]
[0099] The above contents are descriptions of the preferred embodiments of the present application and the related technical principles. Those skilled in the art should understand that the scope of the invention covered by the present application is not limited to the technical solutions constituted by the specific combination of the above technical features, but should also include other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above invention concept. For example, the above features can be interchanged with the technical features with similar functions disclosed in this application (but not limited to) to form a technical solution.
Claims
1. A method for restoring noisy blurred images based on network fusion, characterized in that: The following steps are involved: Step 1, obtaining an image to be restored that contains both noise and motion blur; Step 2: Build a denoising sub-network based on the transformer network framework, input the image to be restored, and obtain an image containing only dynamic blur; Step 3: construct a preliminary deblurring subnetwork based on the convolutional neural network framework, input an image containing only dynamic blur, and obtain a preliminary deblurred image; Step 4: Construct a residual detail generation network based on transformer and diffusion model, input the image that has been initially deblurred, and restore the image that contains both noise and dynamic blur.
2. The method for restoring a noisy blurred image based on network fusion according to claim 1, characterized in that: The preliminary deblurring subnetwork is specifically implemented as follows: an image containing only dynamic blur is inputted, the preliminary deblurring subnetwork first downsamples the image twice to obtain input images of three different scales, the network extracts features from input images of different scales, and encodes and decodes these features to obtain preliminary deblurring results.
3. The method for restoring a noisy blurred image based on network fusion according to claim 2, characterized in that: The residual detail generation network is specifically implemented as follows: an image that has been initially deblurred is input, and in the pre-training stage, the residual details between the initial deblurred image and the real clear image are first obtained, and then the residual detail image is converted into a pure noise image through gradual denoising, and then a residual detail image is generated from the pure noise image through a reverse denoising process of a diffusion model; in the testing stage, the residual detail generation network takes a random pure noise image and a preliminary deblurred image as input, and obtains residual details through reverse denoising of a diffusion model; finally, the preliminary deblurred image and the residual detail image are added element by element and combined together to achieve the restoration task of an image containing both noise and dynamic blur.
4. The method for restoring a noisy blurred image based on network fusion according to claim 3 is characterized in that: The denoising sub-network and the residual detail generation network have the same structure, both of which include an encoder, an intermediate attention module and a decoder; the encoder includes four encoding modules and a downsampling module connected in sequence, wherein the minimum-scale encoding module and the next-small-scale encoding module each include an encoding attention module; the decoder includes four decoding modules and an upsampling module connected in sequence, wherein the minimum-scale decoding module and the next-small-scale decoding module each include an encoding attention module; Each of the encoding modules includes a residual module and a multi-layer perceptron; each of the decoding modules includes a residual module and two multi-layer perceptrons.
5. The method for restoring a noisy blurred image based on network fusion according to claim 4 is characterized in that: During the training phase, the input of the denoising subnetwork is the image to be restored that contains both noise and dynamic blur. The encoder extracts feature vector information through the encoding module and the encoding attention module, and constructs a three-layer network structure through the downsampling module; the intermediate attention module extracts feature information and passes the information to the decoder; the decoder also decodes the information through the decoding module and the decoding attention module, and restores the information scale using the upsampling module; the denoising subnetwork adjusts the parameters through the multi-scale loss function and the multi-scale frequency loss function to obtain a denoised image containing only blur.
6. The method for restoring a noisy blurred image based on network fusion according to claim 5, characterized in that: During the training phase, the input of the residual detail generation network is a pure Gaussian noise image and a preliminary deblurred image. After passing through the same network structure as the denoising sub-network, the result is the detail residual generated from the noise. The network adjusts parameters through the diffusion loss function.
7. The method for restoring a noisy blurred image based on network fusion according to claim 6, characterized in that: The preliminary deblurring subnetwork adopts the Unet network structure in the convolutional neural network, and the encoder includes three convolution modules, three residual modules and two feature fusion modules; the decoder includes three residual modules and two convolution modules; the feature transfer between the encoder and the decoder is realized by the full-scale channel feature fusion module; The full-scale channel feature fusion module is a module that transfers feature vectors between the encoder and the decoder. It first receives feature information of three different scales output by the encoder, unifies them to the current scale using a scale conversion layer, and then concatenates them. It then fuses these feature vectors through a convolutional layer and inputs the final fusion result into the encoder of the corresponding scale.
8. The method for restoring a noisy blurred image based on network fusion according to claim 7, characterized in that: All attention modules consist of two parts: intra-strip self-attention block and inter-strip self-attention block. For these two modules, the input feature tensor is divided into two parts according to the channel dimension, named horizontal feature and vertical feature respectively, and the query, key and value are obtained through convolution and cutting; the intra-strip attention module performs attention calculation on the query, key and value in the vertical and horizontal directions respectively, while the inter-strip attention module performs feature fusion to obtain fused query, fused key and fused value after retrieving the query, key and value, and then performs attention calculation on the fused query, fused key and fused value.
9. The method for restoring a noisy blurred image based on network fusion according to claim 8, characterized in that: The residual detail generation network and the preliminary deblurring sub-network adopt the idea of dynamic network fusion and adopt a joint training method to control the weights of the two networks through a joint loss function. The specific loss function is described as: For the multi-scale frequency loss function of the preliminary deblurring sub-network, the multi-scale frequency loss L freq As an auxiliary loss term design, the composition of the complete loss function of the preliminary deblurring subnetwork is as follows: THE FSF =L cont +λL freq Where λ is the polynomial coefficient, L cont is the multi-scale loss function of the preliminary deblurring subnetwork; During the training phase, the following loss function is designed for the residual detail generation network: CP(·) represents the preliminary deblurring sub-network operation, and Denoiser(·) represents the residual detail generation network operation: x CP =CP(input blur ) x D =Denoiser(x CP ) The input blur represents the input image containing only motion blur, x CP represents the initial deblurred image, x D Represents the generated residual details; L pixel =E||x CP -x gt ||2 L DM =Ε||x0-x D ||2 Where x0 represents the ideal residual detail between the initial deblurred image and the true clear image; L low =E||φ L ×(x CP -x gt )||2 L high =E||φ H ×(x0-x D )||2 where φ L and φ H Represent low-frequency filtering and high-frequency filtering respectively, and the above results are combined to obtain; Where β0 and β1 represent the polynomial coefficients; Combining all the loss function sub-formulas, we get the following complete loss function: L total =L FSF +αL Diffusion Where α represents the polynomial coefficient.
Citation Information
Patent Citations
Transform and multi-scale feature-based low-dose CT image denoising method
CN114998154A
Cited By
Image restoration method and system based on spiking neural network
CN121190334A
Multi-scale image deblurring method based on potential space condition diffusion model
CN121414625A