Methods and systems for deblurring blurred images
By leveraging feature enhancement to expand the neighborhood attention Transformer block and multi-frequency feature fusion, combined with self-attention mechanism and Charbonnier loss, the problem of low detail and efficiency in existing deblurring models for non-uniformly blurred images is solved, thereby improving the quality and efficiency of image deblurring.
Patent Information
- Application Number
- CN202511241984.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-02
AI Technical Summary
Existing deblurring models struggle to effectively recover small but important details when dealing with non-uniformly blurred images, and they are computationally inefficient. Traditional methods have limited ability to capture global dependencies on high-resolution feature maps and are computationally complex.
We employ a feature enhancement-enhanced neighborhood attention Transformer block, combining multi-frequency feature fusion and self-attention mechanisms. High and low frequency components are separated by discrete cosine transform. The activation function is replaced by a feature enhancement module, and the model is optimized by combining content loss, frequency loss, and Charbonnier loss.
It improves the quality of image deblurring, better captures long-range dependencies and local details of images, balances global and local information, and improves computational efficiency and robustness.
Smart Images

Figure CN120746892B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electromechanical equipment, and in particular to a method and system for deblurring blurred images. Background Technology
[0002] Currently, research in the field of image deblurring can be roughly divided into two directions: global deblurring methods and local deblurring methods. Although Transformer performs well in capturing long-range dependencies, traditional methods such as Uformer use locally enhanced window Transformer blocks in the network to reduce computational complexity. This limits its ability to capture more global dependencies on high-resolution feature maps. While it is good at learning long-term dependencies from input, it has weaknesses in modeling subtle local details within and between local patches. This means that traditional Transformer may not be able to effectively recover small but important details in the image.
[0003] Existing deblurring models typically assume uniform blurring, which doesn't reflect the dynamic nature of real-world scenes. Non-uniform blurring is often random, making it difficult to model using specific priors. Furthermore, previous methods involved solving non-convex optimization problems, resulting in very long computation times. In some high-performance deblurring methods, some deblurring quality may be sacrificed to improve speed or reduce computational cost. For example, while Sharpformer improves speed, it may not capture certain details as well as other more time-consuming methods.
[0004] In summary, a method and system for deblurring blurred images are needed to address the shortcomings of existing technologies. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a method and system for deblurring blurred images, aiming to solve the aforementioned problems.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for deblurring a blurred image, comprising the following steps:
[0007] Step S1: Input the original blurred image. Input the original blurred image that needs to be deblurred into the processing network and split it into two paths.
[0008] Step S2: Downsampling and Discrete Cosine Transform: Downsample the original blurred image to generate a small-scale blurred image, and use Discrete Cosine Transform to divide the small-scale original image into high-frequency components and low-frequency components.
[0009] Step S3: Feature extraction. The separated high-frequency and low-frequency components are sent to the codec for processing to obtain high-frequency component features and low-frequency component features.
[0010] Step S4: Multi-frequency feature fusion. Upsample the high-frequency and low-frequency component features processed by the encoder and decoder to restore them to the original image size. Integrate the restored high-frequency and low-frequency features with the original scale-blurred image processed by the encoder to achieve feature depth fusion.
[0011] Step S5: Model optimization. Calculate the total loss of the model using the loss function, optimize the model, and balance global and local information, as well as high-frequency and low-frequency information.
[0012] Step S6: Output the deblurred image, input the fused features into the decoder for learning, and output the deblurred image;
[0013] Step S7: Evaluate the model performance using peak signal-to-noise ratio and structural similarity index, and adjust the network parameters based on the experimental results.
[0014] Optionally, in step S3, the encoder and decoder include an encoder and a decoder. The encoder includes convolutional blocks and residual blocks, and the decoder includes two feature-enhanced extended neighborhood attention modules.
[0015] Optionally, the feature enhancement and expansion neighborhood attention module includes a self-attention mechanism and a feedforward neural network, wherein the feature enhancement module replaces the activation function in the feedforward neural network.
[0016] Optionally, the feature enhancement module enhances the feature in the following ways:
[0017] The features are first divided into three parts according to the three channels through a convolutional layer, and then the two parts are added pixel by pixel. The results are fed into a convolutional layer for fusion. The fused results are multiplied pixel by pixel, and finally passed through another convolution.
[0018] Optionally, the self-attention mechanism performs short-range and long-range learning by adjusting the expansion factor, in the following ways:
[0019] Set the parameters for input to the self-attention mechanism, given an expansion factor and neighborhood size, calculate the attention weights of the tokens, define the attention weights as a matrix, and finally calculate the output features.
[0020] Optionally, the lower bound of the expansion factor is 1, and the upper bound of the expansion factor is the largest integer less than or equal to the ratio of feature size to kernel size.
[0021] Optionally, the multi-frequency feature fusion in step S4 is performed in the following way:
[0022] Step A1: Divide the original image features, high-frequency component features and low-frequency component features into three parts according to three channels, keep one part, and exchange the remaining two parts with the other two features so that the exchanged feature groups all contain three-channel features.
[0023] Step A2: Each of the swapped feature groups is then combined with cross-frequency contextual information through a 1×1 convolutional layer to obtain fused features;
[0024] Step A3: Select one fusion feature and multiply it pairwise with the other two fusion features to obtain two intermediate result features;
[0025] Step A4: Aggregate the two intermediate feature results locally using a 3×3 convolutional layer to obtain aggregated features;
[0026] Step A5: Normalize the aggregated features and divide them into two complementary features. Through two parallel branches, obtain linear and nonlinear features to complete the deep feature fusion.
[0027] Optionally, in step S5, the total loss of the model is calculated using a loss function in the following way:
[0028] The loss function is divided into content loss, frequency loss, and Charbonnier loss. The content loss, frequency loss, and Charbonnier loss are calculated separately. The three losses are multiplied by their corresponding weights and then summed to obtain the total loss of the model.
[0029] Optionally, the calculation of content loss is performed in the following ways:
[0030] Calculate the norms of the deblurred image and the corresponding sharp image for high-frequency features and the norms of the deblurred image and the corresponding sharp image for low-frequency features, divide each norm by the corresponding normalization factor, and then add the two together.
[0031] Frequency loss is calculated in the following way:
[0032] First, the deblurred image and the clear image are fed into the discrete cosine transform to separate the high-frequency coefficients and low-frequency coefficients. The norms of the high-frequency coefficients and low-frequency coefficients of the deblurred image and the clear image are calculated separately, divided by the corresponding normalization factor, and then the two are added together.
[0033] The Charbonnier loss is calculated as follows:
[0034] Calculate the norm of the deblurred image corresponding to the high-frequency features, add an empirical constant, square the result, and take the square root; calculate the norm of the sharp image corresponding to the high-frequency features, add an empirical constant, square the result, and finally add the two results together.
[0035] A blurred image deblurring system, employing the blurred image deblurring method, includes an input processing module, a downsampling and discrete cosine transform module, a feature extraction module, a multi-frequency feature fusion module, a model optimization module, an output processing module, and an evaluation and adjustment module;
[0036] The input processing module is used to input the original blurred image and split it into two paths;
[0037] The downsampling and discrete cosine transform module is used to downsample the original blurred image to generate a small-scale blurred image, and to divide the image into high-frequency components and low-frequency components through discrete cosine transform.
[0038] The feature extraction module is used by the encoder to extract features using convolutional blocks and residual blocks, while the decoder includes two feature enhancement, expansion neighborhood attention modules to further process high-frequency and low-frequency components, respectively obtaining their respective feature representations;
[0039] The multi-frequency feature fusion module is used to perform upsampling to restore high-frequency and low-frequency features to their original size and integrate them with the original scale-blurred image to achieve deep feature fusion.
[0040] The model optimization module is used to calculate the total loss of the model and adjust the model parameters accordingly, balancing global and local information as well as high- and low-frequency information, to optimize model performance.
[0041] The output processing module is used to input the fused features into the decoder for learning and finally output the deblurred image.
[0042] The evaluation and tuning module is used to evaluate model performance using metrics such as peak signal-to-noise ratio and structural similarity index, and to adjust network parameters based on experimental results.
[0043] The beneficial effects of this invention are:
[0044] 1. In this invention, by using feature enhancement and Transformers with different expansion factors, this module enables the model to learn more fully the global and local information in blurred images. This helps to capture long-range dependencies and subtle local details in the image. It abandons the commonly used activation function and replaces it with the feature enhancement module, thereby enhancing the network's flexibility and strong generalization ability.
[0045] 2. In this invention, high-frequency information and low-frequency information are processed separately and then fused into the original blurred image, which provides guidance for deblurring and helps to improve the quality of the final output image. By finely processing and fusing information of different frequencies, this module helps the model to better balance the relationship between global and local information, and between high-frequency and low-frequency information.
[0046] 3. In this invention, the combination of content loss, frequency loss and Charbonnier loss is used to guide the model optimization process, enabling the model to maintain robustness while being computationally efficient. Charbonnier loss is insensitive to outliers, which makes it perform well on image data with noise and outliers. At the same time, its gradient continuity also helps to avoid gradient explosion or vanishing problems in the optimization process. Attached Figure Description
[0047] Figure 1 This is a schematic diagram of a method flow of the present invention.
[0048] Figure 2 This is a schematic diagram of a system network according to the present invention.
[0049] Figure 3 This is a schematic diagram of a feature-enhanced attention network according to the present invention.
[0050] Figure 4 This is a schematic diagram of a feature enhancement feedforward network according to the present invention.
[0051] Figure 5 This is a schematic diagram of a multi-frequency feature fusion network according to the present invention.
[0052] Figure 6 This is a visualization of the comparative experiment of the present invention on the GoPro dataset and the HIDE dataset.
[0053] Figure 7 This is a visualization of the comparative experiment of the present invention on the RealBlur dataset. Detailed Implementation
[0054] To more clearly illustrate the technical solutions in the embodiments of the invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0055] like Figures 1 to 7 As shown, a method for deblurring a blurred image includes the following specific components:
[0056] like Figure 1 and Figure 2As shown, the original blurred image is input into the network and split into two paths. The upper part first downsamples to generate a smaller-scale blurred image, then uses Discrete Cosine Transform (DCT) to divide the smaller-scale original image into high-frequency and low-frequency components, which are then fed into the encoder and decoder respectively. The encoder consists of convolutional blocks and residual blocks, and the decoder includes two feature enhancement, expansion, and attention Transformers. After processing by the encoder and decoder, the high-frequency and low-frequency components are obtained as FHF and FLF respectively. Both are upsampled and fed into the multi-frequency feature fusion module, where they are fused with the original scale blurred image processed by the encoder, and then fed into the decoder together. The high-frequency image, low-frequency image, and deblurred image generated in the intermediate process all participate in the calculation of the loss function to guide model optimization.
[0057] The core structure of the Transformer model includes a self-attention mechanism and a feedforward neural network. The self-attention mechanism is a classic attention mechanism that allows the model to consider the relationships between all pixels in an image while processing image data, without being restricted by their positions in the image. The model automatically associates the information of any two pixels in a sequence, thereby capturing long-range dependencies in the image.
[0058] like Figure 3 As shown, the Feature Enhancement Expanded Neighborhood Attention Transformer is a flexible and powerful sparse global attention module. In its feedforward network, it abandons the commonly used activation function and replaces it with a feature enhancement module.
[0059] Feature enhancement modules such as Figure 4 As shown, the features are first divided into three parts by channel through a convolutional layer, then added pixel by pixel in pairs, and the results are fed into a convolutional layer for fusion. The fused results are multiplied pixel by pixel, and finally passed through another convolution.
[0060] In the self-attention module, expanded neighborhood attention is used. DiNA, as a flexible self-attention mechanism, performs short-range and long-range learning by adjusting the expansion factor without introducing additional complex modules. Let the feature map input to DiNA be... Let S be the number of tags, d be the dimension, and let Q and K be the linear projections of the query and key of S. Let B(i, j) be the relative positional deviation between two tags i and j. Given an expansion factor... Given a neighborhood size k, the attention weight of the i-th token is defined as an equation.
[0061] ,
[0062] The corresponding matrix Its elements are linear projections of the k neighboring values of the i-th label, expressed as an equation.
[0063] ,
[0064] The output feature mapping of the i-th token is represented by the equation.
[0065] ,
[0066] In DiNA, the lower bound of the dilation factor is 1, and the upper bound is calculated as the largest integer less than or equal to m / h, where m is the feature size and h is the kernel size. Increasing the value of the dilation factor expands the size of the self-attention window, meaning the network can learn dependencies over longer distances. To capture local and global blurring patterns, a cascaded structure is designed, where the dilation factor... Take the lower bound value and the upper bound value in turn.
[0067] Multi-frequency feature fusion:
[0068] The high-frequency and low-frequency coefficients of the decoder, which consists of two layers of feature enhancement, expansion, and attention Transformers, will undergo an inverse discrete cosine transform again, converting them from the frequency domain back to the spatial domain. At this point, the original blurred image, after passing through the encoder, arrives at the feature fusion module, ready to be combined with the high-frequency and low-frequency features to enter the multi-frequency feature fusion module.
[0069] like Figure 5 As shown, the original image features, high-frequency features, and low-frequency features are each divided into three parts according to channels. One part is retained, and the remaining two parts are swapped with the other two features. After the swap, each part is combined with cross-frequency context information through a 1x1 convolutional layer. The resulting fused features are multiplied pixel-wise to obtain two intermediate results. These two intermediate results are fed into a 3x3 convolutional layer to achieve local content aggregation. After normalization, the features are further divided into two complementary features. Through two parallel branches, linear features and non-linear features in the GELU-gated path are obtained.
[0070] CFC Loss Function
[0071] The loss function guides model optimization, and its quality largely determines the model's performance. For the model proposed in this chapter, the loss function consists of three parts: content loss, frequency loss, and Charbonnier loss. Content loss measures the difference in content between the generated and original images and is widely used in image restoration tasks such as image deblurring. Content loss guides the network to learn how to restore image content by comparing the differences in high-level features between blurred and sharp images; its formula is as follows:
[0072] ,
[0073] in, This represents the image after deblurring. This indicates the corresponding clear image. This represents the deblurred image after fusing high-frequency and low-frequency features. Indicates the sum after downsampling A clear image of the same size. and It is a normalization factor, set as N1 = W x H x 3, N2 = W / 2 x H / 2 x 3. This represents the L1 norm.
[0074] Frequency loss primarily involves calculating the loss of high-frequency and low-frequency features in the model. In this chapter, the original image is divided into high-frequency and low-frequency components for separate processing, and the frequency loss is calculated for each component. The frequency loss function is defined as follows:
[0075]
[0076] In the formula, and N represents the high-frequency and low-frequency coefficients separated by the discrete cosine transform, respectively. HF =N LF = W / 2 x H / 2 x 3. After separating high-frequency and low-frequency information, the model will not only focus on high-frequency information, but low-frequency information will also receive explicit attention.
[0077] To improve the performance of deblurring algorithms by finding a more efficient loss function, we attempted to introduce the Charbonnier loss, which balances robustness and computational simplicity. The basic principle of the Charbonnier loss is to replace the traditional squared loss with a continuous and smooth function, thereby reducing the impact of outliers on the loss function. In image deblurring, this means that even with noise or extreme values, the loss function remains stable, and the model is not overly sensitive to these values. In our model, the Charbonnier loss is defined as shown in the equation.
[0078]
[0079] in, Based on past experience, it is set to 10. -3 The Charbonnier loss is insensitive to outliers, making it perform well on image data containing noise and outliers. Furthermore, the gradient of the Charbonnier loss is continuous, which helps avoid gradient explosion or vanishing problems during optimization, resulting in a more stable optimization process. Most importantly, the Charbonnier loss is relatively simple to compute, requiring no complex operations, thus offering high computational efficiency in practical applications.
[0080] By combining the three losses, we can obtain the total loss of the model.
[0081] ,in, , .
[0082] The characteristics of the feature-enhanced extended neighborhood attention block were observed in the feature-enhanced feedforward network and multi-frequency feature fusion ablation experiment. The feature-enhanced feedforward network EFFN and multi-frequency feature fusion MFFF were added to the original network in turn for observation. The experimental results are shown in Table 1.
[0083] Table 1. Experimental results of FEDiNA network ablation
[0084]
[0085] The original network did not include a feature enhancement feedforward network and multi-frequency feature fusion. The feature enhancement feedforward network was replaced with a traditional ReLU activation function, and the multi-frequency feature fusion was replaced with an asymmetric feature fusion model. The resulting PSNR was only 31.74, and the SSIM was only 0.921. When a feature enhancement feedforward network was added, the global and local information of the blurred image was fully learned. The features enhanced by the feature enhancement feedforward network were fed into the FEDiNAT module with different dilation factors to achieve long-distance feature fusion. The PSNR improved by 1.69, and the SSIM improved by 0.016. To reintegrate the high-frequency and low-frequency features obtained by inverse discrete cosine transform into the blurred image, multi-frequency feature fusion was added to the feature enhancement feedforward network. The PSNR and SSIM were further improved. Compared with the original model, the PSNR improved by 8.29%, and the SSIM improved by 5.21%. Therefore, the proposed feature enhancement feedforward network and multi-frequency feature fusion are effective.
[0086] Ablation experiments were conducted on the loss functions. Different loss functions provide different guidance for the model, and the experiments are shown in Table 2.
[0087] Table 2 Ablation Experiment Results of Loss Function
[0088]
[0089] If only content loss is used, the model tends to capture the structure and content of the image. Deep features of the network typically encode most of the low-frequency information of the image, while shallow features typically encode most of the high-frequency information; content loss is usually calculated on deep features. When frequency loss is added, the model's PSNR and SSIM increase by 2.28 dB and 0.032 dB, respectively. Frequency loss calculates high-frequency and low-frequency information "openly," thus explicitly pursuing a balance between them. Adding Charbonnier loss further improves the model's performance. Charbonnier loss not only retains the advantage of L2 loss's sensitivity to small error values, allowing for fine-tuning of weights during training, but it is also smooth within its domain, meaning it can provide continuous and differentiable gradients during gradient descent optimization. Combining these three loss functions allows the model to learn both global and local information, as well as high-frequency and low-frequency information.
[0090] The choice of dilation factor in the FEDiNA block also affects the deblurring effect. In the model, the dilation factor used in the stages containing high-frequency and low-frequency images is set to 0.5. and The dilation factor used in the stage where the original image is located is 1 and In the parameter sensitivity analysis experiment, the expansion factor was set as shown in Table 3.
[0091] Table 3. Results of parameter sensitivity analysis of FEDiNA expansion factor
[0092]
[0093] When the FEDiNA blocks of the decoders throughout the entire network use 1 and 2 respectively When the expansion factor is used as the network's scaling factor, the model achieves the highest PSNR and SSIM, at 34.63 dB and 0.972 dB respectively, but the processing time is also the longest, requiring 1.848 seconds. If the scaling factor of the first FEDiNA block in each decoder is changed to 0.25... If this happens, the entire network will lose some local features, which will manifest as a decrease in PSNR and SSIM; when the expansion factor of the first FEDiNA block in each decoder is increased to 0.5... Therefore, whether it is high-frequency information, low-frequency information, or the original image, some information will be lost again, and PSNR and SSIM will decrease again. Although both increases in the dilation factor reduce time, the loss of PSNR and SSIM is significant. It is not worthwhile to exchange the reduction in PSNR and SSIM for the reduction in time.
[0094] While the final determined dilation factor in the model does not produce the optimal PSNR and SSIM, they are only 0.02 and 0.003 different from the optimal results, respectively, and are also suboptimal in terms of time. It is assumed that high-frequency and low-frequency information undergo downsampling before entering the encoder-decoder. At this point, a larger dilation factor can simulate a larger receptive field, allowing for the learning of more long-distance dependencies. When the fused high-frequency and low-frequency information is fed into the original image for further deblurring, a smaller dilation factor allows for more careful learning of image details, compensating for the initial loss of high-frequency and low-frequency information. Therefore, the chosen dilation factor configuration is most suitable for this model.
[0095] Finally, the parameter sensitivity analysis experiment on the number of FEDiNA blocks is shown in Table 4.
[0096] Table 4. Results of parameter sensitivity analysis on the number of FEDiNA blocks
[0097]
[0098] If four FEDiNA blocks are used alternately to process high-frequency, low-frequency, and original image information respectively, the model's PSNR and SSIM will improve slightly, but the computation time will increase significantly. Although such a model can fit complex functions and noise in the training data, it carries the risk of overfitting, resulting in a model with lower generalization ability. In the other two sets of experiments, the number of FEDiNA blocks in the high-frequency and low-frequency information processing stages was kept consistent. When the high-frequency and low-frequency information processing stages have one more alternating FEDiNA block than the original image processing stage, the model can learn more information of different frequencies, but the global information only contains one alternation, which is insufficient for supplementing the upsampled high-frequency and low-frequency information. When the original image processing stage has one more alternating FEDiNA block than the high-frequency and low-frequency information processing stages, the model can learn sufficient global information, but the details learned after downsampling high-frequency and low-frequency learning are insufficient. The PSNR and SSIM values of the model using two FEDiNA blocks per stage are comparable to those of the model using four FEDiNA blocks per stage, and the time is about 43.3% less than that of the model using four FEDiNA blocks per stage.
[0099] Looking at the data alone cannot accurately convey the model's deblurring effect. In this section, we will compare images before and after deblurring side by side, such as... Figure 6The upper part shows the model's results on the GoPro dataset. As can be seen from the image, after processing by the model, the boundaries of each number on the billboard are very clear. The same numbers (such as two consecutive 5s) appear to stick together after processing by Stripformer and DeblurDiNAT-S, while the model can clearly separate the consecutive numbers, resulting in a visually clearer image.
[0100] Figure 6 The lower half shows the deblurring results of the model on the HIDE dataset. The HIDE dataset contains many images of people in motion, such as a pedestrian carrying a backpack with letters printed on the shoulder strap. These letters are motion blurred due to the pedestrian's movement. The model restores the original content of the letters to the greatest extent possible, while MPRNet, SRN, and DeblurDiNAT-S restore almost no letters. Stripformer and DeblurDiNAT-L can analyze the general outline of the letters, but the specific content is still indistinguishable.
[0101] Figure 7 The upper and lower sections respectively present the results of comparative experiments on the RealBlur-R and RealBlur-J datasets. The results on the RealBlur-R dataset are mostly satisfactory, with some numbers and object outlines processed clearly. However, the night scene lights processed by MSSNet are relatively blurry; DeblurGAN-v2 loses some color information of the lights, altering part of the image content. Similar problems exist with MSSNet and DeblurGAN-v2 on the RealBlur-J dataset, but the results from the [previous model] are much clearer.
[0102] A blurred image deblurring system, employing the blurred image deblurring method, includes an input processing module, a downsampling and discrete cosine transform module, a feature extraction module, a multi-frequency feature fusion module, a model optimization module, an output processing module, and an evaluation and adjustment module;
[0103] The input processing module is used to input the original blurred image and split it into two paths;
[0104] The downsampling and discrete cosine transform module is used to downsample the original blurred image to generate a small-scale blurred image, and to divide the image into high-frequency components and low-frequency components through discrete cosine transform.
[0105] The feature extraction module is used by the encoder to extract features using convolutional blocks and residual blocks, while the decoder includes two feature enhancement, expansion neighborhood attention modules to further process high-frequency and low-frequency components, respectively obtaining their respective feature representations;
[0106] The multi-frequency feature fusion module is used to perform upsampling to restore high-frequency and low-frequency features to their original size and integrate them with the original scale-blurred image to achieve deep feature fusion.
[0107] The model optimization module is used to calculate the total loss of the model and adjust the model parameters accordingly, balancing global and local information as well as high- and low-frequency information, to optimize model performance.
[0108] The output processing module is used to input the fused features into the decoder for learning and finally output the deblurred image.
[0109] The evaluation and tuning module is used to evaluate model performance using metrics such as peak signal-to-noise ratio and structural similarity index, and to adjust network parameters based on experimental results.
[0110] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions or improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for deblurring a blurred image, characterized in that, Includes the following steps: Step S1: Input the original blurred image. Input the original blurred image that needs to be deblurred into the processing network and split it into two paths. Step S2: Downsampling and Discrete Cosine Transform: Downsample the original blurred image to generate a small-scale blurred image, and use Discrete Cosine Transform to divide the small-scale original image into high-frequency components and low-frequency components. Step S3: Feature extraction. The separated high-frequency and low-frequency components are sent to the codec for processing to obtain high-frequency component features and low-frequency component features. Step S4: Multi-frequency feature fusion. Upsample the high-frequency and low-frequency component features processed by the encoder and decoder to restore them to the original image size. Integrate the restored high-frequency and low-frequency features with the original scale-blurred image processed by the encoder to achieve feature depth fusion. Step S5: Model optimization. Calculate the total loss of the model using the loss function, optimize the model, and balance global and local information, as well as high-frequency and low-frequency information. The total loss of the model is calculated using the loss function in the following way: The loss function is divided into content loss, frequency loss and Charbonnier loss. The content loss, frequency loss and Charbonnier loss are calculated separately. The three losses are multiplied by their corresponding weights and then added together to obtain the total loss of the model. Step S6: Output the deblurred image, input the fused features into the decoder for learning, and output the deblurred image; Step S7: Evaluate the model performance using peak signal-to-noise ratio and structural similarity index, and adjust the network parameters based on the experimental results.
2. The method for deblurring a blurred image according to claim 1, characterized in that, In step S3, the encoder and decoder include an encoder and a decoder. The encoder includes convolutional blocks and residual blocks, and the decoder includes two feature enhancement, expansion, and neighborhood attention modules.
3. The method for deblurring blurred images according to claim 2, characterized in that, The feature enhancement extended neighborhood attention module includes a self-attention mechanism and a feedforward neural network, in which the feature enhancement module replaces the activation function.
4. The method for deblurring blurred images according to claim 3, characterized in that, The feature enhancement module is enhanced in the following ways: The features are first divided into three parts according to the three channels through a convolutional layer, and then the two parts are added pixel by pixel. The results are fed into a convolutional layer for fusion. The fused results are multiplied pixel by pixel, and finally passed through another convolution.
5. The method for deblurring blurred images according to claim 3, characterized in that, The self-attention mechanism performs short-range and long-range learning by adjusting the expansion factor, in the following ways: Set the parameters for input to the self-attention mechanism, given an expansion factor and neighborhood size, calculate the attention weights of the tokens, define the attention weights as a matrix, and finally calculate the output features.
6. The method for deblurring a blurred image according to claim 5, characterized in that, The lower bound of the expansion factor is 1, and the upper bound of the expansion factor is the largest integer less than or equal to the ratio of feature size to kernel size.
7. The method for deblurring a blurred image according to claim 1, characterized in that, The multi-frequency feature fusion in step S4 is performed in the following way: Step A1: Divide the original image features, high-frequency component features and low-frequency component features into three parts according to three channels, keep one part, and exchange the remaining two parts with the other two features so that the exchanged feature groups all contain three-channel features. Step A2: Each of the swapped feature groups is then combined with cross-frequency contextual information through a 1×1 convolutional layer to obtain fused features; Step A3: Select one fusion feature and multiply it pairwise with the other two fusion features to obtain two intermediate result features; Step A4: Aggregate the two intermediate feature results locally using a 3×3 convolutional layer to obtain aggregated features; Step A5: Normalize the aggregated features and divide them into two complementary features. Through two parallel branches, obtain linear and nonlinear features to complete the deep feature fusion.
8. The method for deblurring a blurred image according to claim 1, characterized in that, The calculated content loss is achieved through the following methods: Calculate the norms of the deblurred image and the corresponding sharp image for high-frequency features and the norms of the deblurred image and the corresponding sharp image for low-frequency features, divide each norm by the corresponding normalization factor, and then add the two together. Frequency loss is calculated in the following way: First, the deblurred image and the clear image are fed into the discrete cosine transform to separate the high-frequency coefficients and low-frequency coefficients. The norms of the high-frequency coefficients and low-frequency coefficients of the deblurred image and the clear image are calculated separately, divided by the corresponding normalization factor, and then the two are added together. The Charbonnier loss is calculated as follows: Calculate the norm of the deblurred image corresponding to the high-frequency features, add an empirical constant, square the result, and take the square root; calculate the norm of the sharp image corresponding to the high-frequency features, add an empirical constant, square the result, and finally add the two results together.
9. A system for deblurring blurred images, employing the method for deblurring blurred images as described in any one of claims 1-8, characterized in that, It includes an input processing module, a downsampling and discrete cosine transform module, a feature extraction module, a multi-frequency feature fusion module, a model optimization module, an output processing module, and an evaluation and adjustment module; The input processing module is used to input the original blurred image and split it into two paths; The downsampling and discrete cosine transform module is used to downsample the original blurred image to generate a small-scale blurred image, and to divide the image into high-frequency components and low-frequency components through discrete cosine transform. The feature extraction module is used by the encoder to extract features using convolutional blocks and residual blocks, while the decoder includes two feature enhancement, expansion neighborhood attention modules to further process high-frequency and low-frequency components, respectively obtaining their respective feature representations; The multi-frequency feature fusion module is used to perform upsampling to restore high-frequency and low-frequency features to their original size and integrate them with the original scale-blurred image to achieve deep feature fusion. The model optimization module is used to calculate the total loss of the model and adjust the model parameters accordingly, balancing global and local information as well as high- and low-frequency information, to optimize model performance. The output processing module is used to input the fused features into the decoder for learning and finally output the deblurred image. The evaluation and tuning module is used to evaluate model performance using metrics such as peak signal-to-noise ratio and structural similarity index, and to adjust network parameters based on experimental results.
Citation Information
Patent Citations
Image deblurring method based on multi-scale frequency separation network
CN115272113A