Image deblurring method based on U-Net network

By adding a multi-dense frequency domain attention module and a dual feature adaptive fusion module to the U-Net network, combined with a wavelet transformation and a convolution prediction module, the problem of poor image defuzzing effect in the prior art is solved, and an image defuzzing effect with higher quality and clarity is achieved.

CN120070251AActive Publication Date: 2025-05-30CHANGCHUN UNIV OF SCI & TECH

Patent Information

Application Number
CN202510139752.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

The existing image defuzzing method based on improved U-Net model has problems with low contrast and unnatural color in the defuzzing effect, and there is risk of parameter redundancy and overfitting, and lacks attention mechanism and detail extraction capabilities.

Method used

The image defuzzing method based on U-Net network is adopted, and a network model including U-Net network, a reference-free evaluation model, a stage fusion module, a multi-density frequency domain attention module, a coarse frequency domain feature evaluation module, a convolution prediction module and a dual feature adaptive fusion module are constructed. The feature extraction and recovery are combined with discrete wavelet transform and inverse wavelet transform, and the feature extraction and defuzzing effect are improved through the multi-density frequency domain attention module and a dual feature adaptive fusion module.

Benefits of technology

The deblurred image is achieved to be more in line with the visual characteristics of the human eye, improve the accuracy and image quality of the deblurred image, and make the output image clearer and more detailed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070251A_ABST
    Figure CN120070251A_ABST
Patent Text Reader

Abstract

The invention discloses an image deblurring method based on a U-Net network, and relates to the technical field of image processing. Comprising the following steps: preparing a data set; constructing a network model; training a network model; optimizing the model and evaluating the performance; according to the invention, a multi-dense frequency domain self-attention network and a double feature adaptive fusion network are added into a U-Net network. The multi-dense frequency domain self-attention network extracts features of different frequencies through a plurality of convolutional layers and captures global long-range dependency information in combination with a self-attention mechanism, so that higher-quality image recovery and reconstruction results can be generated in a decoding stage; and the double feature adaptive fusion network fuses the feature map obtained by the coarse and fine frequency domain feature evaluation module and the feature map obtained in the decoding stage, so that the decoder can better improve the deblurring accuracy and the quality of the generated image, and the output image is clearer and finer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing machines, and in particular to an image deblurring method based on a U-Net network. Background Art

[0002] Image deblurring is a key technology in digital image processing, aiming to eliminate blurring and improve clarity and quality. Blurring is usually caused by device movement, lens problems, or environmental factors, affecting information volume and readability. Deblurring technology improves image quality and is applicable to computer vision tasks such as object detection and image recognition.

[0003] This technology is widely used in fields such as daily photography, medical imaging, satellite analysis, and security monitoring. In the medical field, clear images contribute to accurate diagnosis; in security monitoring, it provides reliable recognition functions. The development of deblurring technology is of great significance for improving image quality and analysis accuracy.

[0004] The Chinese patent publication number is "CN114549361B", and its patent name is "An Image Motion Deblurring Method Based on an Improved U-Net Model". The improvements include using a 3×3 convolutional kernel and a Leaky ReLU activation function. The encoder extracts features in four stages, combining depthwise separable convolution, residual convolution, and Haar wavelet transform. The decoder processes information through four stages, using skip connections and inverse wavelet transform. The image obtained by this method has a low contrast and unnatural colors, which does not conform to the human visual effect. At the same time, multiple depthwise separable convolutions and residual convolutions may lead to parameter redundancy, increasing the risk of overfitting, and the lack of an attention mechanism fails to highlight key regions and lacks the ability to enhance the detail extraction of the model. Summary of the Invention

[0005] The technical solution of the present invention to solve the above technical problems is to provide an image deblurring method based on a U-Net network, including the following steps:

[0006] Step 1, prepare the data sets: Prepare data sets one, two, and three for network training. Among them, data set one is the GoPro Dataset data set, and methods such as random scaling, inversion, and translation are used to expand the data set one by inversion and translation, while providing multi-scale information; data set two is the RealBlur Dataset data set for model testing; data set three is the Dataset data set for model fine-tuning; the size of each image is 512×512;

[0007] Step 2, construct a network model: Construct a network model including a U-Net network, a no-reference evaluation model, a stage fusion module, a multi-dense frequency domain attention module, a coarse and fine frequency domain feature evaluation module, a convolutional prediction module, and a dual feature adaptive fusion module;

[0008] The U-Net network includes an encoder, a decoder, a multi-feature extraction network, a stage feature fusion module, a multi-dense frequency domain attention module, a coarse and fine frequency domain feature evaluation module, a dual feature adaptive fusion module, and a convolutional prediction module; among them, the encoder and decoder are composed of multiple convolutional blocks; the multi-feature extraction network contains convolutional layers with different dilation rates for capturing different scale information; the stage feature fusion module is used for flexible interaction of features at different levels; the multi-dense frequency domain attention module is used to enhance the feature extraction ability of key regions; the coarse and fine frequency domain feature evaluation module is divided into a low-frequency information extraction branch and a high-frequency information extraction branch to extract low-frequency and high-frequency features respectively; the dual feature adaptive fusion module is used to fuse different features and enhance important features; the convolutional prediction module is used to predict the image features after deblurring;

[0009] Step 3, train the network model: Use the prepared dataset to train the network model. During the training process, discrete wavelet transform and inverse wavelet transform are used for feature extraction and restoration, and the number of channels of the feature map is adjusted by 1×1 convolution to limit the model complexity;

[0010] Step 4, optimize the model and evaluate the performance: Optimize the model using three loss functions: deblurring loss, encoder feature loss, and prediction loss, and use peak signal-to-noise ratio and structural similarity as evaluation metrics to evaluate the model performance.

[0011] Further, in Step 2, both the encoder and decoder of the U-Net network are composed of multiple convolutional blocks. Some convolutional blocks use 3×3 depthwise separable residual convolution, and the remaining convolutional blocks are composed of 3×3 convolution, Leaky ReLU activation function, spatial attention, channel attention, and 1×1 convolution.

[0012] Further, in Step 2, the multi-feature extraction network includes 2 1×1 convolutional layers, 2 3×3 convolutional layers, 1 convolutional layer with a dilation rate of 1, 1 convolutional layer with a dilation rate of 3, and an activation function. The network is divided into two branches. The 3×3 convolution with different dilation rates is used to increase the receptive field. The feature map after the fusion of the two branches passes through 1×1 convolution and is fused with the input image through a residual connection.

[0013] Further, in Step 2, the stage feature fusion module includes multiple 1×1 convolutional layers, max pooling, average pooling, BN layer, linear interpolation, and CBAM module, and is used for feature fusion and enhancement of the feature maps obtained in three stages of the encoder.

[0014] Further, in step 2, the multi-dense frequency domain attention module includes a multi-scale frequency feature extraction module and a self-attention module, which are used to adaptively select important frequency information and enhance the feature representation ability.

[0015] Further, in step 2, the coarse and fine frequency domain feature evaluation module is divided into a low-frequency information extraction branch and a high-frequency information extraction branch. The low-frequency information extraction branch fuses the feature map after upsampling with the low-frequency - low-frequency feature map included after band decomposition. The high-frequency information extraction branch fuses the feature maps containing high-frequency - low-frequency, low-frequency - high-frequency, and high-frequency - high-frequency information after band decomposition. Finally, the feature maps containing high-frequency and low-frequency information are fused and then subjected to inverse wavelet transform, and then fused with the input feature map through a channel attention network to obtain the final output.

[0016] Further, in step 2, the dual feature adaptive fusion module includes a convolutional layer, a 1×1 convolution, an activation function layer, a softmax layer, a linear transformation layer, and a normalization layer, which are used to fuse the feature maps obtained by the coarse and fine frequency domain feature evaluation module and the feature maps obtained by the decoder, and enhance important features.

[0017] Further, in step 2, the convolutional prediction module includes a resize layer, a 1×1 convolutional layer, an activation function, and a 3×3 convolutional layer, which are used to predict the deblurred image features and calculate the loss with the quality score feature map of the no-reference evaluation model, so that the feature map obtained by the encoder approximates the quality score feature map of the no-reference evaluation model.

[0018] Further, the image deblurring method based on the U-Net network further includes the following steps:

[0019] Step 5: Fine-tune the model: Input the dataset into the network model for training, compare the obtained results with the image evaluation metrics, and then fine-tune the model parameters;

[0020] Step 6, save the model: Solidify the parameters of the finally optimized U-Net network model, save the network model, and deploy it to the hardware system.

[0021] Compared with the prior art, the present invention has the following beneficial effects:

[0022] (1) The present invention embeds the quality knowledge of the image into the encoder and decoder, dynamically adjusts the weights of the feature layers according to the image quality to adapt to different degrees of blur. Experiments prove that the deblurred images of the network proposed in this paper are more in line with the human visual characteristics.

[0023] (2) In the present invention, a multi-dense frequency domain attention module and a dual feature adaptive fusion module are added to the U-Net network. The multi-dense frequency domain attention module extracts features of different frequencies through multiple convolutional layers and captures global long-range dependence information by combining the self-attention mechanism, which helps to generate higher-quality image restoration and reconstruction results in the decoding stage. The dual feature adaptive fusion module fuses the feature maps obtained by the coarse and fine frequency domain feature evaluation modules with the feature maps obtained in the decoding stage, enabling the decoder to better improve the accuracy of deblurring and the quality of the generated image, making the output image clearer and more detailed.

[0024] (3) The present invention designs a coarse and fine frequency domain feature evaluation module, which extracts features of the feature map in the first stage of the decoder according to high-frequency information and low-frequency information, respectively obtaining the local details and overall structure of the image. The two are fused through convolution and then restored by inverse wavelet transform (IDWT), and finally enhanced image features are output and sent to the decoder and the convolutional prediction network, improving the deblurring effect of the decoder while better predicting the quality information of the feature map. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0026] Figure 1 It is a flowchart of the steps of the image deblurring method based on the U-Net network of the present invention;

[0027] Figure 2 It is a schematic diagram of the framework of the network model of the present invention;

[0028] Figure 3 It is a schematic diagram of the structure of the U-net network of the present invention;

[0029] Figure 4 It is a schematic diagram of the structure of the stage feature fusion module of the present invention;

[0030] Figure 5 It is a schematic diagram of the structure of the multiple feature extraction network of the present invention;

[0031] Figure 6 It is a schematic diagram of the structure of the multi-dense frequency domain attention module of the present invention;

[0032] Figure 7 It is a schematic diagram of the structure of the coarse and fine frequency domain feature evaluation module of the present invention;

[0033] Figure 8Schematic diagram of the dual - feature adaptive fusion module of the present invention;

[0034] Figure 9 Schematic diagram of the convolutional prediction module of the present invention;

[0035] Figure 10 Schematic diagram of the convolutional discrete wavelet transform of the present invention;

[0036] Figure 11 Schematic diagram of the inverse wavelet transform convolution of the present invention;

[0037] Figure 12 Schematic diagram of the convolutional block in the U - Net network of the present invention;

[0038] Figure 13 Hardware schematic diagram of an image de - blurring system based on improved U - Net and quality perception. Detailed implementation manners

[0039] The present invention proposes an image de - blurring method based on the U - Net network, aiming to improve the accuracy of de - blurring and the quality of the generated image, making the output image clearer and more detailed.

[0040] The following will illustrate the image de - blurring method based on the U - Net network proposed by the present invention in specific embodiments:

[0041] In the technical solution of this embodiment, as Figure 1 shown, an image de - blurring method based on the U - Net network includes the following steps:

[0042] Step 1, Prepare the dataset: Prepare dataset one, dataset two, and dataset three for network training. Among them, dataset one is the GoPro Dataset dataset. Use methods such as random scaling, inversion, and translation to expand dataset one by inversion and translation, providing multi - scale information so that the model can extract effective features at different resolutions, improving the processing ability for different - scale blurred regions; dataset two is the RealBlur Dataset dataset, which is used for model fine - tuning to further optimize the performance of the model; dataset three is the Dataset dataset; By testing the model on this dataset, verify its generalization ability in various blurred situations. Each image in each dataset has a size of 512x512. Ensure the consistency of different datasets during processing in the network model to avoid instability in the training process caused by image size differences.

[0043] Step 2, construct a network model: Construct a network model including a U-Net network, a no-reference evaluation model, a stage fusion module, a multi-dense frequency domain attention module, a coarse and fine frequency domain feature evaluation module, a convolutional prediction module, and a dual feature adaptive fusion module;

[0044] The U-Net network includes an encoder, a decoder, a multi-feature extraction network, a stage feature fusion module, a multi-dense frequency domain attention module, a coarse and fine frequency domain feature evaluation module, a dual feature adaptive fusion module, and a convolutional prediction module; among them, the encoder and decoder are composed of multiple convolutional blocks; the multi-feature extraction network contains convolutional layers with different dilation rates for capturing different scale information; the stage feature fusion module is used for flexible interaction of features at different levels; the multi-dense frequency domain attention module is used to enhance the feature extraction ability of key regions; the coarse and fine frequency domain feature evaluation module is divided into a low-frequency information extraction branch and a high-frequency information extraction branch to extract low-frequency and high-frequency features respectively; the dual feature adaptive fusion module is used to fuse different features and enhance important features; the convolutional prediction module is used to predict the image features after deblurring;

[0045] Step 3, train the network model: Use the dataset prepared in step (1) to train the network model. During the training process, discrete wavelet transform and inverse wavelet transform are used for feature extraction and restoration, and the number of channels of the feature map is adjusted by 1×1 convolution to limit the model complexity;

[0046] Step 4, optimize the model and evaluate the performance: Optimize the model using three loss functions: deblurring loss, encoder feature loss, and prediction loss, and use peak signal-to-noise ratio and structural similarity as evaluation metrics to evaluate the model performance.

[0047] Step 5: Fine-tune the model: Input the dataset into the network model for training, compare the obtained results with image evaluation metrics, and then fine-tune the model parameters;

[0048] Step 6, save the model: Solidify the parameters of the finally optimized U-Net network model, save the network model, and deploy it to the hardware system.

[0049] Furthermore, in step 2, both the encoder and decoder of the U-Net network are composed of multiple convolutional blocks. Some convolutional blocks adopt 3×3 depthwise separable residual convolution, and the remaining convolutional blocks are composed of 3×3 convolution, Leaky ReLU activation function, spatial attention, channel attention, and 1×1 convolution.

[0050] Further, in step 2, the multi-feature extraction network includes 2 1×1 convolutional layers, 2 3×3 convolutional layers, 1 convolutional layer with a dilation rate of 1, 1 convolutional layer with a dilation rate of 3, and an activation function. The network is divided into two branches. The 3×3 convolutional layers with different dilation rates are used to increase the receptive field. The feature map after the fusion of the two branches passes through a 1×1 convolution and is fused with the input image through a residual connection.

[0051] Further, in step 2, the stage feature fusion module includes multiple 1×1 convolutional layers, max pooling, average pooling, BN layers, bilinear interpolation, and a CBAM module, which are used for feature fusion and enhancement of the feature maps obtained in the three stages of the encoder.

[0052] Further, in step 2, the multi-dense frequency domain attention module includes a multi-scale frequency feature extraction module and a self-attention module, which are used to adaptively select important frequency information and improve the feature representation ability.

[0053] Further, in step 2, the coarse and fine frequency domain feature evaluation module is divided into a low-frequency information extraction branch and a high-frequency information extraction branch. The low-frequency information extraction branch fuses the upsampled feature map with the low-frequency-low-frequency feature map included after band decomposition. The high-frequency information extraction branch fuses the feature maps containing high-frequency-low-frequency, low-frequency-high-frequency, and high-frequency-high-frequency information after band decomposition. Finally, the feature map containing high-frequency and low-frequency information is fused and then subjected to inverse wavelet transform, and then fused with the input feature map through a channel attention network to obtain the final output.

[0054] Further, in step 2, the dual feature adaptive fusion module includes a convolutional layer, 1×1 convolution, an activation function layer, a softmax layer, a linear transformation layer, and a normalization layer, which are used to fuse the feature map obtained by the coarse and fine frequency domain feature evaluation module and the feature map obtained by the decoder, and enhance important features.

[0055] Further, in step 2, the convolutional prediction module includes a resizing layer, 1×1 convolutional layer, an activation function, and a 3×3 convolutional layer, which are used to predict the deblurred image features and calculate the loss with the quality score feature map of the no-reference evaluation model, so that the feature map obtained by the encoder approximates the quality score feature map of the no-reference evaluation model.

[0056] Furthermore, in step 2, the quality-aware feature map is extracted through a no-reference image quality assessment model, and the quality features are obtained by aligning the feature maps. The prediction loss is calculated using the quality features and the features predicted by the predictor to constrain the learning of the encoder. And the image quality knowledge is integrated into the encoder; by reusing the encoder, the generated deblurred image is input into the encoder again, its features are extracted and compared with the quality features of the clear image, and the quality knowledge of the image is embedded into the decoder to help the model pay more attention to quality when generating clear images, ensuring that the output image is both clear and has a good visual effect.

[0057] Example 2:

[0058] An image deblurring method based on the U-Net network, comprising the following steps:

[0059] Step 1, prepare the dataset:

[0060] Prepare dataset one for network training to train the entire network model. Dataset one is the GoProDataset dataset, and methods such as random scaling, inversion, and translation are used to expand the dataset by inversion and translation, while providing multi-scale information; dataset two is the RealBlur Dataset dataset for model fine-tuning; dataset three is the Dataset dataset for model testing. Each image has a size of 512×512.

[0061] Step 2, construct the network model:

[0062] The structure of the network model is as Figure 2 shown, including a U-Net network, a no-reference evaluation model, a stage fusion module, a multi-dense frequency domain attention module, a coarse and fine frequency domain feature evaluation module, a convolutional prediction module, and a dual feature adaptive fusion module.

[0063] Among them, the U-Net network is as Figure 3 shown. The U-Net network encoder and decoder are composed of 19 convolutional blocks. Convolutional blocks two, four, six, eight, nine, twelve, fifteen, and eighteen are 3×3 depthwise separable residual convolutions. The remaining convolutional blocks are composed of 1 3×3 convolution, a Leaky ReLU activation function, spatial attention, channel attention, and 3 1×1 convolutions. Its structure is as Figure 12As shown in the figure. In the encoding stage, features are extracted through Convolution Block 1, Convolution Block 2, Convolution Block 3, Convolution Block 4, Convolution Block 5, Convolution Block 6, Convolution Block 7, and Convolution Block 8, and downsampling is performed through wavelet transform. In the decoding stage, through Convolution Block 10, Convolution Block 11, Convolution Block 12, Convolution Block 13, Convolution Block 14, Convolution Block 15, Convolution Block 16, Convolution Block 17, Convolution Block 18, and Convolution Block 19, and upsampling is performed through inverse wavelet transform to reduce information loss during the reconstruction process and improve the deblurring effect. During the connection between the encoder and the decoder, a multi-dense frequency domain attention module. In the skip connections between the encoding stage and the decoding stage, three-stage feature fusion modules are added. The stage feature fusion module is used for flexible interaction of features at different levels to capture more details. The multi-dense frequency domain attention module is added in the fourth layer of the U-Net to improve the feature extraction ability of key regions and reduce redundant information.

[0064] The U-Net network includes a multiple feature extraction network, a decoder, a stage feature fusion network, a multi-dense frequency domain attention module, and a decoder. As Figure 5 shown, the multiple feature extraction network includes 2 1×1 convolutional layers (Convolution Layer 1 and Convolution Layer 2), 2 3×3 convolutional layers (Convolution Layer 3 and Convolution Layer 4), 1 convolutional layer with a dilation rate of 1 (Convolution Layer 5), 1 convolutional layer with a dilation rate of 3 (Convolution Layer 6), and an activation function. The role of this convolutional layer is to extract features from blurred images of different scales of the input. In order to better capture information at different scales and enhance the feature representation ability, the network is divided into two branches. The 3×3 convolutional layers with different dilation rates are used to increase the receptive field. The feature map after the fusion of the two branches passes through a 1×1 convolution and is fused with the input image through a residual connection. The feature extraction formula of the multiple feature extraction network is as follows

[0065] F 1 =Conv 3×3 (Conv 1×1 (Input))

[0066] F 2 =Conv 3×3 (Conv 1×1 (Input));

[0067] F = RrLU(Conv 1×1 (Cat(Conv 3×3,rate=1 (F 1 ),Conv 3×3,rate=3 (F 2 )))+Input);

[0068] In the formula, F 1 and F 2 represent the feature maps after 1×1 and 3×3 convolutions, and F represents the finally obtained feature map.

[0069] As shown Figure 4 in the figure, the stage feature fusion module consists of 7 1×1 convolutional layers, 4 max pooling layers, 3 average pooling layers, 3 BN layers, 3 bilinear interpolations (resize), and 3 CBAM modules. The feature maps obtained from the three stages of the encoder are fed into the network through convolutional layer 1, convolutional layer 2, and convolutional layer 3. The sizes of the feature maps are unified through bilinear interpolation operations. Each input feature map extracts features at different scales through max pooling and average pooling. The pooled feature maps pass through the CRAM module to further enhance the important features. The feature maps of different branches are fused through an addition operation to form a comprehensive feature map. The fused feature map is processed through multiple convolutional layers (convolutional layer 4, 5, 6, 7), and each convolutional layer is followed by a BN layer to standardize the features. Finally, the feature maps of different paths are fused and output.

[0070] P max,i = MaxPool(F i );

[0071] P avg,i = AvgPool(F i );

[0072] A i = CBAM(P max,i , P avg,i );

[0073] F fused = A 1 + A 2 + A 3 ;

[0074] F 4 = BN(Conv 1×1 (F fused ));

[0075] G i = BN(Conv 1×1 (F 4 )) × F i ;

[0076] G = G 1 + G 2 + G 3 ;

[0077] Where P max,i and P avg,i represent the max pooling and average pooling of each branch, and G is the final output.

[0078] As Figure 6As shown in the figure, the multi-dense frequency domain attention module consists of a multi-scale frequency feature extraction module and a self-attention module. The multi-scale frequency feature extraction module consists of 3 1×1 convolutional layers (convolutional layer one, convolutional layer five, and convolutional layer six), 3 3×3 convolutional layers (the dilation rates of convolutional layer two, convolutional layer three, and convolutional layer four are 1, 2, and 3 respectively), a Fourier transform, and a quantization matrix. In the multi-dense frequency domain attention module, the role of the quantization matrix is to adaptively select important frequency information, distinguish which low-frequency and high-frequency information should be retained, and improve the feature representation ability. The obtained feature map is fused with the feature map obtained through the self-attention network to obtain a feature map containing different scale frequencies. The self-attention module consists of 1 normalization layer, 3 linear transformation layers, 4 dimension reshaping layers (Reshape), 1 softmax layer, and 3 1×1 convolutional layers (convolutional layer seven and convolutional layer eight). The self-attention mechanism is used to capture long-range dependencies to dynamically adjust the attention area. By fusing with the multi-scale frequency feature extraction module, a feature map of different scale frequencies is obtained, and the output channels and dimensions are adjusted through 1 1×1 convolution (convolutional layer nine). The expression formula of the multi-dense frequency domain attention module is as follows:

[0079] Multi-scale frequency feature extraction module:

[0080] (1) Multi-scale extraction part of the convolutional layer:

[0081] F 1 = FFT(F multi ) × quantization matrix;

[0082] F 2 = IFFT(F 6 );

[0083] F 3 = Conv 1×1 (F 2 );

[0084] (2) Self-attention part:

[0085]

[0086] Final output:

[0087] F = Conv 1×1 (F 3 + Conv 1×1 (F attention ));

[0088] Such as Figure 7As shown in the figure, the thick and thin frequency domain feature evaluation module consists of 7 1×1 convolutional layers (convolutional layers one, two, three, four, seven, ten, twelve), 5 3×3 convolutional layers (convolutional layers five, six, eight, nine, eleven), 1 upsampling layer, 1 pooling layer, 1 wavelet transform, and a frequency band decomposition layer. It is divided into a low-frequency information extraction branch and a high-frequency information extraction branch. The low-frequency information extraction branch fuses the feature map after upsampling with the low-frequency - low-frequency feature map included after frequency band decomposition, and further extracts low-frequency features through a residual network. The high-frequency information extraction branch fuses the feature maps containing high-frequency - low-frequency, low-frequency - high-frequency, and high-frequency - high-frequency information after frequency band decomposition and further extracts high-frequency features through a residual network. The feature maps containing high-frequency and low-frequency information are fused and then inverse wavelet transform is performed. After passing through the channel attention network again, the feature maps emphasize both frequency information and channel information, and finally are fused with the input feature map to obtain the final output. The network expression formula is as follows:

[0089] Low-frequency information extraction branch:

[0090] F low = Upsample(Enhance(F t ));

[0091] F low_processed = GELU(Conv 3×3 (GELU(Conv 3×3 (Conv 1×1 (F low )))));

[0092]

[0093] In the formula, F t is the feature map after the convolution layer three after the fusion of the two input feature maps, is the final output feature map of the low-frequency information obtained after the residual connection.

[0094] High-frequency information extraction branch:

[0095] F high = F h-h + F l-h + F h-l ;

[0096] F high_processed =

[0097] Conv 1×1 (GELU(Conv 3×3 (GELU(Conv 3×3 (Conv 1×1 (F high ))))));

[0098] F high_out = F high_processed + F high ;

[0099] In the formula, F h-h , F l-h and F h-l respectively represent the feature maps containing high-frequency - high-frequency, low-frequency - high-frequency, and high-frequency - low-frequency information after wavelet transform. F high_out is the final output feature map of high-frequency information obtained after residual connection.

[0100] Final output:

[0101] G output = Conv 1×1 (IDWT(F high_out + F low_out )) + F t ;

[0102] As Figure 8 shown, the dual-feature adaptive fusion module consists of one 3×3 convolutional layer (convolutional layer three), four 1×1 convolutions (convolutional layer one, convolutional layer two, convolutional layer four, and convolutional layer five), two activation function layers, two softmax layers, two linear transformation layers, and one normalization layer. The feature map obtained by the coarse and fine frequency domain feature evaluation module enters convolutional layer two after upsampling, and after a segmentation operation, it is element-wise multiplied with the feature map obtained by the decoder to apply weights to different features, thereby enhancing important features and improving the expression ability and performance of the decoder. The segmented feature map passes through another parallel branch consisting of convolutional layer four, normalization layer, activation function, and convolutional layer five, and is added to another branch to extract features from different perspectives or levels and enrich feature information. Through the combination of linear transformation, activation function, and softmax, efficient integration of cross-modal information is achieved, and finally, the fused feature map is obtained.

[0103] Feature Figure 1 Feature extraction:

[0104] F 1 = Softmax(Conv 3×3 (Conv 1×1 (Input 1 )))

[0105] Feature Figure 2 Feature extraction:

[0106] F 2 = Split(Conv 1×1 (Input 2 ))

[0107] Feature map fusion multiplication:

[0108] F 3 = F 1 × F 2 ;

[0109] Residual branch:

[0110] F 4 = Conv 1×1 (ReLU(BN(Conv 1×1 (F 2 ))));

[0111] Feature map fusion addition:

[0112] F 5 = F 3 × F 4 ;

[0113] Final output:

[0114] F output = Softmax(Linear(ReLU(Linear(F 5 ))));

[0115] As Figure 9 shown, the convolutional prediction module consists of 3 Reshape layers, 4 1×1 convolutional layers (Convolutional Layer 1, Convolutional Layer 2, Convolutional Layer 3, and Convolutional Layer 4), 1 activation function, and 1 3×3 convolutional layer (Convolutional Layer 5). The input to the prediction network is three feature maps obtained through the coarse and fine frequency domain feature evaluation module and two upsamplings. After reshaping and convolutional layers, they are added and fused. The fused feature map passes through 1×1 convolution, activation function, and 3×3 convolution to obtain the predicted feature map. The feature map obtained by the prediction network and the quality score feature map of the no-reference evaluation model are used to calculate the loss, so that the feature map obtained by the encoder approximates the quality score feature map of the no-reference evaluation model, and the quality knowledge of the image is embedded into the encoder. The formula of the convolutional prediction network is as follows:

[0116] F 1 = Conv 1×1 (Reshape(Input 1 ));

[0117] F 2 = Conv 1×1 (Reshape(Input 2 ));

[0118] F 3 = Conv 1×1 (Reshape(Input 3 ));

[0119] F 4 = F 1 + F 2 + F 3 ;

[0120] F output = Conv 3×3 (ReLU(Conv 1×1 (F 4 )));

[0121] Wherein F 1 , F 2 and F 3 are respectively the outputs of three inputs passing through the resizing layer and the convolutional layer, and F output is the predicted feature map obtained by the prediction network.

[0122] As Figure 12 shown, in the U-Net network, convolutional block one, convolutional block three, convolutional block five, convolutional block seven, convolutional block eleven, convolutional block twelve, convolutional block fourteen, convolutional block fifteen, and convolutional block sixteen are composed of 1 3×3 convolutional layer (convolutional layer one), 3 1×1 convolutional layers (convolutional layer one, convolutional layer two, and convolutional layer three), 1 spatial attention module, 1 channel attention module, and 1 activation function layer. The input feature map is fused through the spatial attention module and the channel attention module, and then fused with the feature map passing through the 3×3 convolution. Finally, the output is obtained through the Leaky ReLU activation function. By adding channel attention and spatial attention in the residual connection of the convolutional block, the encoder and decoder can respectively focus on different channels and spatial positions of the feature map, improving the model's ability to capture important features, enhancing the feature expression ability of the U-Net network, and improving the model performance. The expression formula of the convolutional block is as follows:

[0123] CA output = CA(Conv 1×1 (Input));

[0124] SA output = SA(Conv 1×1 (Input));

[0125] F CA = CA output × Input;

[0126] F SA = SA output × Input;

[0127] F 1 = Conv 1×1 (Cat(F CA , F SA));

[0128] F output = LeakyReLU(F 1 + Conv 3×3 (Input));

[0129] In the formula, CA output and SA output represent the outputs after spatial attention and channel attention, F CA and F SA represent the outputs after fusion with the input image respectively after attention, F 1 represents the fusion of the spatial attention and channel attention feature maps, F output represents the final output of the convolutional block.

[0130] Step 3, train the network model:

[0131] Use the dataset to train the network model. During the training process, discrete wavelet transform and inverse wavelet transform are used for feature extraction and restoration, and 1×1 convolution is used to adjust the number of channels of the feature map to limit the model complexity.

[0132] The discrete wavelet transform is as Figure 10 shown, and the inverse wavelet transform is as Figure 11 shown. Since DWT will make the number of channels of the output feature map become four times that of the input feature map, the increase in the number of feature channels will lead to a significant increase in the number of parameters in the subsequent network layers, which may increase the model complexity, resulting in overfitting and low generalization ability, and will also lead to a significant increase in resource occupancy. Therefore, in order to be consistent with the mainstream image restoration, the present invention first uses 1×1 convolution to halve the number of channels of the input feature map, and then passes through DWT to limit the increase in the number of channels of the output feature map. There are two reasons for adding 1×1 convolution after IWT. One is to match the downsampling CDWT, and match the number of channels in the downsampling process, which helps the information correspondence process between CDWT and IWTC; the other is to make the size of the feature map in each layer of the feature restoration stage match that in each layer of the feature extraction stage. Since skip connections are added between each layer of the feature extraction stage and the feature restoration stage in this algorithm, by adding the feature map obtained in each layer of the feature extraction stage to the upsampled result of the previous layer in the corresponding feature restoration stage and then inputting it into the network block of this layer for restoration, the size of the feature map in each layer of the feature restoration stage needs to match that in each layer of the feature extraction stage.

[0133] CDWT formula:

[0134] F CDWT = DWT(Conv 1×1 (F input ( ));

[0135] IWTC formula:

[0136] F IWTC = Conv 1×1 (IWT(F input ));

[0137] Step 4, select a suitable loss function and optimal evaluation metrics;

[0138] To ensure the robustness of the network, retain more image information, and fully extract image features, the present invention selects three losses to optimize the model, calculating the deblurring loss between the deblurred output image and the target clear image, the encoder loss between the features output by the encoder and the quality features extracted by the reference-free evaluation model, and the prediction loss between the deblurred image and the true clear image. All three losses are calculated using the mean square error.

[0139] Deblurring loss:

[0140] l = MSE(I out , I gt );

[0141] Encoder feature loss:

[0142] l e = MSE(f p , f k );

[0143] Prediction loss:

[0144] l d = MSE(f out , f gt );

[0145] Total loss:

[0146] l t = λ 1 l + λ 2 l e + λ 3 l d ;

[0147] Where λ 1 , λ 2 and λ 3 are the weight hyperparameters of each loss function, used to adjust the influence of different loss terms.

[0148] Select the peak signal-to-noise ratio and structural similarity as suitable evaluation metrics. The peak signal-to-noise ratio is an evaluation metric based on the error between pixel points, used to sensitively evaluate image quality. The structural similarity index measures the similarity between images from three aspects: brightness, contrast, and structure, and is an index used to evaluate the similarity degree of digital images. The peak signal-to-noise ratio and structural similarity are defined as follows.

[0149]

[0150] MAX is the maximum possible pixel value of the image. I and K are the pixel values of the original image and the distorted image respectively, and M and N are the dimensions of the image.

[0151]

[0152] u x and u y are the average values of the images x and y, and are the variances of the images x and y, σ xy is the covariance of the images x and y, C 1 and C 2 are constants for stabilization to avoid a zero denominator.

[0153] Set the number of training times to 200. The learning rate for the first 100 training processes is set to 0.001, and the learning rate for the last 100 training processes gradually decreases from 0.001 to 0; the network parameter optimizer selects the Adam optimizer.

[0154] Step 5: Fine-tune the model: Input the dataset into the network model for training, compare the obtained results with the image evaluation metrics, and then fine-tune the model parameters;

[0155] Step 6, Save the model: Solidify the parameters of the finally optimized U-Net network model, save the network model, and deploy it to the hardware system.

[0156] An image deblurring system based on improved U-Net and quality perception, as Figure 13 shown. The system is based on the Xilinx Zynq UltraScale+ MPSoC ZCU104 development board and consists of an FPGA hardware system, an HDMI monitor, and a computer. The FPGA system is divided into two parts: PS (Processing System) and PL (Programmable Logic). The PS side loads the stored image data, U-Net model weights, and related parameters through an SD card, and performs preprocessing operations such as normalizing, resizing, and format conversion on the image data. The PL side contains a CNN IP core, which is responsible for performing efficient image deblurring inference tasks.

[0157] The PS side transfers the preprocessed image data, model weights, and parameters from the DDR memory to the PL side through AXI-DMA. The CNN IP core in the PL side consists of multiple functional modules, including a multi-dense frequency-domain self-attention module and a dual-feature adaptive fusion module, which are used to improve the detail recovery effect of blurred images. The CNN IP core is hardware-optimized, adopting a pipeline design and fixed-point quantization, and can efficiently complete the inference task.

[0158] After the inference is completed, the PL side transfers the results back to the PS side through AXI-FIFO. The PS side performs post-processing (denoising and format conversion) on the output image and displays the deblurring effect in real time through the HDMI monitor. The computer, as an auxiliary device, is responsible for writing the trained model and data to the SD card and supports remote monitoring and debugging.

[0159] The present invention constructs an image deblurring method based on an improved U-Net and quality perception, and realizes image deblurring in an end-to-end manner. It does not require the establishment of a complex degradation model and avoids the research on different blur types of images. Moreover, by using hardware acceleration and data transmission technologies, the real-time performance and processing efficiency of the deblurring process are ensured. In the present invention, the PS side is responsible for control and data management, while the PL side processes high-load CNN inference tasks, and the overall constitutes an efficient image processing system. Under the same conditions, by calculating the relevant metrics of the images obtained by the existing methods, the feasibility and superiority of this method are further verified. The comparison of the relevant metrics of the prior art (CN114549361B in the background art) and the technical solution of Invention Embodiment 2 is shown in Table 1. Two relevant metrics, namely the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM), commonly used in the deblurring network, are listed in the table.

[0160] Table 1 Comparison of relevant metrics between the prior art and the method proposed in the invention

[0161]

[0162] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An image deblurring method based on U-Net network, characterized in that: The following steps are involved: Step 1, prepare the data set: prepare data set 1, data set 2 and data set 3, where data set 1 is the GoProDataset data set; Dataset 2 is the RealBlur Dataset; Dataset 3 is Dataset Step 2, building a network model: building a network model including a U-Net network, a reference-free evaluation model, a stage fusion module, a multi-dense frequency domain attention module, a coarse and fine frequency domain feature evaluation module, a convolution prediction module, and a dual feature adaptive fusion module; The U-Net network includes an encoder, a decoder, a multiple feature extraction network, a stage feature fusion module, a multi-dense frequency domain attention module, a coarse and fine frequency domain feature evaluation module, a dual feature adaptive fusion module and a convolution prediction module; wherein the encoder and the decoder are composed of multiple convolution blocks; the multiple feature extraction network includes convolution layers with different expansion rates for capturing information of different scales; the stage feature fusion module is used for flexible interaction of features at different levels; the multi-dense frequency domain attention module is used to improve the feature extraction capability of key areas; the coarse and fine frequency domain feature evaluation module is divided into a low-frequency information extraction branch and a high-frequency information extraction branch to extract low-frequency and high-frequency features respectively; the dual feature adaptive fusion module is used to fuse different features and enhance important features; the convolution prediction module is used to predict the image features after deblurring; Step 3, training the network model: Use the prepared data set to train the network model. During the training process, discrete wavelet transform and inverse wavelet transform are used for feature extraction and restoration, and the number of feature map channels is adjusted through 1×1 convolution to limit the model complexity. Step 4: Optimize the model and evaluate the performance: Use three loss functions, namely deblurring loss, encoder feature loss, and prediction loss, to optimize the model, and use peak signal-to-noise ratio and structural similarity as evaluation indicators to evaluate the model performance.

2. The image deblurring method based on U-Net network according to claim 1, characterized in that: In step 2, the encoder and decoder of the U-Net network are composed of multiple convolution blocks, some of which use 3×3 depth-separable residual convolution, and the remaining convolution blocks are composed of 3×3 convolution, Leaky ReLU activation function, spatial attention, channel attention and 1×1 convolution.

3. The image deblurring method based on U-Net network according to claim 1, characterized in that: In step 2, the multiple feature extraction network includes 2 1×1 convolutional layers, 2 3×3 convolutional layers, 1 convolutional layer with a dilation rate of 1, 1 convolutional layer with a dilation rate of 3 and an activation function. The network is divided into two branches. 3×3 convolutions with different dilation rates are used to increase the receptive field. The feature maps fused from the two branches are fused with the input image through a residual connection after a 1×1 convolution.

4. The image deblurring method based on U-Net network according to claim 1, characterized in that: In step 2, the stage feature fusion module includes multiple 1×1 convolutional layers, maximum pooling, average pooling, BN layers, linear interpolation and CBAM modules, which are used for feature fusion and enhancement of feature maps obtained in the three stages of the encoder.

5. The image deblurring method based on U-Net network according to claim 1, characterized in that: In step 2, the multi-dense frequency domain attention module includes a multi-scale frequency feature extraction module and a self-attention module, which are used to adaptively select important frequency information and improve feature representation capabilities.

6. The image deblurring method based on U-Net network according to claim 1, characterized in that: In step 2, the coarse and fine frequency domain feature evaluation module is divided into a low-frequency information extraction branch and a high-frequency information extraction branch. The low-frequency information extraction branch fuses the feature map after upsampling with the low-frequency-low-frequency feature map contained after frequency band decomposition. The high-frequency information extraction branch fuses the feature map containing high-frequency-low-frequency, low-frequency-high-frequency and high-frequency-high-frequency information after frequency band decomposition. Finally, the feature map containing high-frequency and low-frequency information is fused and then inverse wavelet transformed, and then fused with the input feature map through the channel attention network to obtain the final output.

7. The image deblurring method based on U-Net network according to claim 1, characterized in that: In step 2, the dual feature adaptive fusion module includes a convolution layer, a 1×1 convolution, an activation function layer, a soft maximization layer, a linear transformation layer and a normalization layer, which are used to fuse the feature map obtained by the coarse and fine frequency domain feature evaluation module and the feature map obtained by the decoder, and enhance important features.

8. The image deblurring method based on U-Net network according to claim 1, characterized in that: In step 2, the convolution prediction module includes a resizing layer, a 1×1 convolution layer, an activation function and a 3×3 convolution layer, which are used to predict the deblurred image features and make a loss with the quality score feature map of the no-reference evaluation model so that the feature map obtained by the encoder is close to the quality score feature map of the no-reference evaluation model.

9. The image deblurring method based on U-Net network according to claim 1, characterized in that: The following steps are also included: Step 5: Fine-tune the model: Input the data set into the network model for training, compare the obtained results with image evaluation indicators, and then fine-tune the model parameters; Step 6, save the model: solidify the parameters of the final optimized U-Net network model, save the network model, and deploy it to the hardware system.

Citation Information

Patent Citations

  • Image deblurring algorithm based on block multi-scale convolutional neural network

    CN112801901A

  • Image deblurring algorithm based on knowledge distillation and deep neural network

    CN114677304A

  • Depth deblurring method based on adaptive fuzzy kernel estimation

    CN114841897A

  • Image denoising method and apparatus based on wavelet high-frequency channel synthesis

    US20240161251A1

Cited By

  • Lightweight wavelet convolution guide wire segmentation network model and double guide wire generation method

    CN120339271A

  • Image deblurring method based on improved U-Net

    CN120451003A

  • An image deblurring method based on improved U-Net

    CN120451003B

  • Multi-source remote sensing image radiation normalization method and system based on frequency domain decomposition

    CN120451592A

  • Blurred image deblurring method and system

    CN120746892A