Double-branch image deblurring method and device based on frequency domain perception
By introducing a dual-branch structure network into the defuzzy network, combining spatial domain and frequency domain information, the problem of poor defuzzy effect in the prior art is solved, and better image defuzzy effect and computing efficiency are achieved.
Patent Information
- Application Number
- CN202510072669.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The existing defuzzing network structure is complex, and the defuzzing effect is poor, making it difficult to take into account both spatial domain and frequency domain information.
The dual-branch image defuzzing method based on frequency domain perception is adopted. The dual-branch structure network combines spatial domain and frequency domain information, and the D-FM module is used to efficiently integrate information to achieve defuzzing of images.
It improves the image debuffering effect, maintains the characteristics of simplicity and efficiency, while taking into account the details recovery and computing efficiency, improving the robustness and generalization ability of image debuffering.
Smart Images

Figure CN119991498A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a dual-branch image deblurring method and device based on frequency domain perception. Background Art
[0002] In the field of computer vision, image deblurring is one of the important image restoration tasks in computer vision, which aims to restore blurry images to clear images. Image blur is usually caused by the device imaging system or external factors, mainly including motion blur, defocus blur and Gaussian blur. In particular, image motion blur caused by camera or object motion is the most concerned and challenging problem.
[0003] In the problem of deblurring a single moving image, traditional deblurring methods are mostly based on the Richardson-Lucy method and the Wiener deconvolution method, such as the texture-preserving two-step deblurring method (see Chen F, Huang X, Chen W. Texture-preserving image deblurring [J]. IEEE Signal processing letters, 2010, 17 (12): 1018-1021.) and the blind image deblurring method based on the local maximum gradient prior (see Chen L, Fang F, Wang T, et al. Blind image deblurring with local maximum gradient prior [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2019: 1742-1750.). The improved versions of these traditional methods usually rely on the estimation of the blur kernel and the prior knowledge of the image to restore the image through deconvolution or regularization. However, these traditional methods are highly dependent on accurate estimation of blur kernels. In practical applications, due to the existence of complex blur kernels, it is often difficult to make accurate estimates and difficult to cope with complex situations in real scenes. With the continuous advancement of deep learning technology, more and more methods based on deep neural networks (DNNs) are applied to the task of deblurring single motion images. Unlike traditional methods, DNNs can automatically learn the features of blurred images through a large amount of training data, without relying on accurate blur kernel estimation like traditional methods. The powerful nonlinear characteristics of DNNs enable them to handle complex blur patterns, including non-uniform or spatially varying blur phenomena, greatly enhancing the robustness and generalization ability of deblurring algorithms.
[0004] Although DNN has made significant progress in single motion image deblurring, there are still some limitations in existing research. At present, most DNN-based single motion image deblurring algorithms either focus on processing spatial domain information or frequency domain information, and it is difficult to take both into account at the same time. In the spatial domain processing method, the network extracts image features through the convolution layer and uses the deconvolution or upsampling layer to restore the clear image, focusing on processing pixel-level features, effectively enhancing the local details and edge features of the image (see Wang X, Girshick R, Gupta A, et al. Non-local neural networks [C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2018: 7794-7803. Woo S, Park J, Lee JY, et al. Cbam: Convolutional block attention module [C] / / Proceedings of the European conference on computer vision (ECCV). 2018: 3-19.). However, this method ignores the frequency information of the image and is unable to fully recover high-frequency details.
[0005] In the frequency domain processing method, the model converts the image into frequency space to extract and enhance features, which is particularly helpful in recovering high-frequency information and removing blur components. However, this method usually relies on complex operations such as Fourier transform or wavelet transform. Over-reliance on frequency domain processing not only greatly increases the complexity of the model, but also easily ignores the spatial context information of the image. The MLWNet method proposed by Gao et al. introduces a learnable wavelet transform module and a multi-scale loss function, and uses the multi-scale architecture of SIMO to improve the image deblurring effect and detail recovery ability while better adapting to real scenes. However, since this method relies too much on frequency domain processing, although its performance is good, due to the high model complexity, the computing resource requirements are also relatively large (see Gao X, Qiu T, Zhang X, et al. Efficient multi-scale network with learnable discrete wavelet transform for blind motion deblurring [C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2024: 2733-2742.). Summary of the invention
[0006] In view of this, the purpose of the present invention is to provide a dual-branch image deblurring method and device based on frequency domain perception, which is used to solve the technical problems of the complex structure of the existing deblurring network and the poor deblurring effect.
[0007] The present invention provides a dual-branch image deblurring method based on frequency domain perception, comprising the steps of:
[0008] S1: Obtain a sample image set, perform image preprocessing on the sample image set, and obtain a training image set;
[0009] S2: construct a deblurring network, train the deblurring network through a training image set, and obtain a trained deblurring network;
[0010] S3: Deblur the image using the trained deblurring network.
[0011] Preferably, step S1 specifically comprises:
[0012] S11: Load a sample image, and perform cropping, flipping, rotating, and format conversion operations on the sample image in sequence to obtain a training image;
[0013] S12: Repeat step S11 until all sample images are traversed to obtain a training image set.
[0014] Preferred:
[0015] The defuzzification network includes: a first FEM module, an encoder, a dual-branch structure network, a decoder, and a second FEM module;
[0016] The first FEM module is connected to the encoder, the encoder is connected to the decoder, and the decoder is connected to the second FEM module;
[0017] The encoder, dual-branch structure network and decoder are connected in sequence.
[0018] Preferred:
[0019] The dual-branch structure network includes a first dual-branch module and a second dual-branch module in parallel;
[0020] The first double-branch module and the second double-branch module have the same structure;
[0021] The first dual-branch module includes: a NAF module, an LFA module and a D-FM module;
[0022] The encoder is connected to the NAF module and the LFA module;
[0023] The NAF module and the LFA module are connected to the D-FM module;
[0024] The D-FM module is connected to the decoder;
[0025] The LFA module includes: wavelet decomposition module, convolutional network, high-frequency feature enhancement module, wavelet fusion module and feature distillation module;
[0026] The encoder is connected to the wavelet decomposition module;
[0027] The wavelet decomposition module is connected with the convolutional network and the high-frequency feature enhancement module;
[0028] The convolutional network and high-frequency feature enhancement module are connected with the wavelet fusion module;
[0029] The wavelet fusion module, feature distillation module and D-FM module are connected in sequence.
[0030] Preferred:
[0031] The encoder includes a first ENAF module, a second ENAF module, a third ENAF module and a fourth ENAF module connected in sequence;
[0032] The decoder comprises a first DNAF module, a second DNAF module, a third DNAF module and a fourth DNAF module connected in sequence;
[0033] The first ENAF module is connected to the first DNAF module;
[0034] The second ENAF module is connected to the second DNAF module;
[0035] The third ENAF module, the first dual-branch module and the third DNAF module are connected in sequence;
[0036] The fourth ENAF module, the second dual-branch module and the fourth DNAF module are connected in sequence;
[0037] The fourth ENAF module includes 28 NAF modules, and the first ENAF module, the second ENAF module, the third ENAF module, the first DNAF module, the second DNAF module, the third DNAF module and the fourth DNAF module each include one NAF module.
[0038] Preferably, step S2 specifically comprises:
[0039] S21: inputting the training image into the first FEM module to obtain an initial feature image;
[0040] S22: inputting the initial feature image into the first ENAF module to obtain the first feature image; inputting the first feature image into the second ENAF module after downsampling operation to obtain the second feature image; inputting the second feature image into the third ENAF module after downsampling operation to obtain the third feature image; inputting the third feature image into the fourth ENAF module after downsampling operation to obtain the fourth feature image;
[0041] S23: inputting the fourth feature image into the second dual-branch module to perform feature extraction in the spatial domain and the frequency domain to obtain a second fused feature image, and inputting the second fused feature image into the fourth DNAF module to obtain a fourth output image;
[0042] S24: inputting the third feature image into the first dual-branch module to perform feature extraction in the spatial domain and the frequency domain to obtain a first fused feature image, and inputting the first fused feature image and the fourth output image after the upsampling operation into the third DNAF module to obtain a third output image;
[0043] S25: inputting the second feature image and the third output image after the upsampling operation into the second DNAF module to obtain the second output image; inputting the first feature image and the second output image after the upsampling operation into the first DNAF module to obtain the first output image;
[0044] S26: inputting the first output image into the second FEM module to obtain an initial deblurred image; superimposing the training image with the initial deblurred image to obtain a final deblurred image;
[0045] S27: Calculate the loss value through the final deblurred image and adjust the parameters of the deblurring network;
[0046] S28: Repeat steps S21-S27 until the loss value is less than a preset value to obtain a trained deblurring network.
[0047] Preferred:
[0048] The process of obtaining the first fused feature image and the second fused feature image is the same;
[0049] The specific process of obtaining the first fusion feature image is as follows:
[0050] The spatial domain feature image of the third feature image is extracted by the NAF module, the frequency domain feature image of the third feature image is extracted by the LFA module, and the spatial domain feature image and the frequency domain feature image are fused by the D-FM module to obtain a first fused feature image.
[0051] Preferred:
[0052] The process of obtaining the frequency domain feature image of the third feature image is specifically as follows:
[0053] The third feature image is transformed into a two-dimensional discrete wavelet through the wavelet decomposition module to obtain the low-frequency subband F LL , high frequency sub-band F LH , high frequency sub-band F HL and high frequency subband F HH ;
[0054] The low frequency subband F LL Input the convolutional network to obtain the enhanced low-frequency subband The high frequency subband F LH , high frequency sub-band F HL and high frequency subband F HH Input the high-frequency feature enhancement module to obtain the enhanced high-frequency sub-bands Enhanced high frequency subband and the enhanced high frequency subband
[0055] The enhanced low-frequency subband Enhanced high frequency subband Enhanced high frequency subband and the enhanced high frequency subband Input the wavelet fusion module to perform two-dimensional discrete wavelet inverse transform to obtain the inverse transform feature image F′ out ;
[0056] The inverse transformed feature image F′ out Input the feature distillation module to obtain the distilled feature image, and superimpose the distilled feature image with the third feature image to obtain the frequency domain feature image F of the third feature image. out .
[0057] A storage medium stores instructions and data for implementing the dual-branch image deblurring method based on frequency domain perception.
[0058] A dual-branch image deblurring device based on frequency domain perception comprises: a processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement the dual-branch image deblurring method based on frequency domain perception.
[0059] The present invention has the following beneficial effects:
[0060] A dual-branch structure network is set up in the deblurring network. The dual-branch structure network includes two parallel dual-branch modules. The dual-branch module extracts spatial domain features through the NAF module and frequency domain features through the LFA module, effectively combining the spatial domain and frequency domain, and using the D-FM module to efficiently integrate the information in the spatial domain and frequency domain. The dual-branch structure network enables the deblurring network to not only maintain the characteristics of simplicity and efficiency when processing frequency domain and spatial domain information, but also take into account detail recovery and computational efficiency. The dual branches do not interfere with each other, effectively improving the deblurring effect of the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 This is a flow chart of a method according to an embodiment of the present invention;
[0062] Figure 2 This is the structural diagram of the deblurring network;
[0063] Figure 3 It is the structural diagram of the double-branch structure network;
[0064] Figure 4 It is the structural diagram of the LFA module;
[0065] Figure 5 It is the structural diagram of D-FM module;
[0066] Figure 6 This is the structural diagram of the multi-scale channel attention mechanism module;
[0067] Figure 7 It is the data flow diagram of the encoder;
[0068] Figure 8 This is a structural diagram of the device according to an embodiment of the present invention;
[0069] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0070] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0071] Reference Figure 1 The present invention provides a dual-branch image deblurring method based on frequency domain perception, comprising the steps of:
[0072] S1: Obtain a sample image set, perform image preprocessing on the sample image set, and obtain a training image set;
[0073] As an example,
[0074] Step S1 is specifically as follows:
[0075] S11: Load a sample image, and perform cropping, flipping, rotating, and format conversion operations on the sample image in sequence to obtain a training image;
[0076] Specific:
[0077] First, a cropping strategy was used for sample images. The original images have a large resolution. In order to reduce the amount of calculation and improve training efficiency, all images used for training were cropped into multiple 512×512 pixel sub-images. This strategy can not only effectively reduce the size of the image and reduce processing time, but also increase the diversity of the data, because each image will be divided into multiple independent image blocks, introducing more image variants during the training process, and further improving the generalization ability of the model.
[0078] Secondly, in order to avoid overfitting of the model during training and improve the model's adaptability to different image deformations, the present invention adopts an enhancement method based on geometric transformation. Each image is randomly flipped and rotated with a probability of 0.5. Random flipping enables the model to learn the symmetry of the image in the horizontal and vertical directions, while random rotation helps the model adapt to the changes of the image at different angles. This data enhancement strategy not only increases the diversity of training data, but also enhances the robustness of the model, enabling it to maintain good performance when facing images of different perspectives and directions.
[0079] Finally, in order to improve the data reading speed and reduce memory usage during training, all training image sub-blocks are converted to lmdb format. lmdb (Lightning Memory-Mapped Database) format is an efficient data storage method that can quickly load and access data. By storing data in lmdb format, the data loading time can be reduced during large-scale data training, avoiding memory overflow problems caused by large amounts of data. Compared with traditional image storage methods, lmdb format can significantly improve data loading efficiency through memory mapping and compression storage, and provide better support for large-scale training.
[0080] S12: Repeat step S11 until all sample images are traversed to obtain a training image set.
[0081] S2: construct a deblurring network, train the deblurring network through a training image set, and obtain a trained deblurring network;
[0082] As an example,
[0083] The defuzzification network includes: a first FEM module, an encoder, a dual-branch structure network, a decoder, and a second FEM module;
[0084] The first FEM module is connected to the encoder, the encoder is connected to the decoder, and the decoder is connected to the second FEM module;
[0085] The encoder, dual-branch structure network and decoder are connected in sequence.
[0086] Specifically, the structure of the deblurring network (DFNet) is as follows Figure 2 As shown in the figure, the deblurring network is based on the classic U network framework. The encoder and decoder of DFNet are mainly composed of a nonlinear activation free feature extraction module (NAF). The dual-branch spatial domain feature processing is the NAF module, the frequency domain feature processing is the learnable frequency perception module (Learnable Frequency-domian Awareness, LFA) designed by the present invention, and the module for information fusion of the dual branches is the dual-branch-driven fusion module (D-FM).
[0087] As an example,
[0088] The dual-branch structure network includes a first dual-branch module and a second dual-branch module in parallel;
[0089] The first double-branch module and the second double-branch module have the same structure;
[0090] The first dual-branch module includes: a NAF module, an LFA module and a D-FM module;
[0091] The encoder is connected to the NAF module and the LFA module;
[0092] The NAF module and the LFA module are connected to the D-FM module;
[0093] The D-FM module is connected to the decoder;
[0094] The LFA module includes: wavelet decomposition module, convolutional network, high-frequency feature enhancement module, wavelet fusion module and feature distillation module;
[0095] The encoder is connected to the wavelet decomposition module;
[0096] The wavelet decomposition module is connected with the convolutional network and the high-frequency feature enhancement module;
[0097] The convolutional network and high-frequency feature enhancement module are connected with the wavelet fusion module;
[0098] The wavelet fusion module, feature distillation module and D-FM module are connected in sequence.
[0099] Specifically, the structure of the dual-branch network is as follows Figure 3 As shown in the figure, the dual-branch structure network is designed in the bottleneck layer and the second high layer of the deblurring network, which is designed with consideration of the deblurring performance and model complexity of the model. The dual branches can make full use of the spatial domain and frequency domain information, and the two branches do not interfere with each other. The spatial domain branch adopts a non-linear non-activated feature extraction module. This module extracts features through layer-by-layer convolution to capture local details and texture information of the image. This branch effectively reduces feature redundancy, maintains the lightweight of the model, and ensures that the local structure is maximized without introducing non-linear activation. The frequency domain branch performs multi-scale frequency domain decomposition of the image through two-dimensional discrete wavelet transform (2D-DWT) to extract feature information at different frequencies.
[0100] The structure of the learnable frequency domain perception module (LFA) is as follows Figure 4 As shown, most methods currently combine two-dimensional discrete wavelet transform with convolutional network. The steps of this design are: first, the features output by the encoder are decomposed by wavelet to obtain four sub-bands; then, the four bands are subjected to equal feature enhancement and receptive field expansion operations through the convolutional network; finally, the output of the convolutional network is subjected to inverse wavelet transform to obtain the output of the frequency domain branch. Generally, low-frequency components can express the smooth appearance of the global area, while high-frequency components can capture rich details of the local area, such as edges, textures and other complex features. However, performing equal operations on the four sub-bands lacks flexibility in processing different types of information. Therefore, the present invention optimizes the initial frequency domain branch to obtain a learnable frequency domain perception module (LFA), by combining the low-frequency sub-band with the convolutional network to increase the receptive field to extract global features, and combining the high-frequency sub-band with the feature enhancement module to extract local features, so that the frequency domain information can be more fully utilized and complex blur types can be better dealt with.
[0101] The structure of the branch-driven fusion module (D-FM) is as follows Figure 5 As shown in , the fusion module weights the features from the spatial domain branch and the frequency domain branch. The core of this module is the multi-scale channel attention mechanism module, as shown in Figure 6As shown in the figure, the attention mechanism has two paths, one is the local channel attention mechanism and the other is the global channel attention mechanism. The local path consists of two 1*1 convolutional layers, a normalization layer and a Relu activation function. In the local path, the feature map first needs to undergo a 1*1 convolution to reduce the number of channels and reduce the computational cost. Then batch normalization is performed to ensure the stability of the model, nonlinearity is introduced through the Relu activation function, and finally the initial number of channels is restored through a 1*1 convolution. The output is L(x), which has the same shape as the input feature map, so it can better retain and highlight the subtle details in the low-level features. The output of the local channel attention L(X) can be expressed as:
[0102]
[0103] Where X represents the input feature map, PWConv(·) represents the point-by-point convolution operation, and δ represents the ReLU activation function.
[0104] In the global path, the feature map first reduces the input features to C*1*1 through an adaptive average pooling operation, and then extracts global features through a series of convolutional layers, batch normalization layers, and ReLU activation functions. The output G(X) of the global channel attention can be expressed as:
[0105]
[0106] As an example,
[0107] The encoder includes a first ENAF module, a second ENAF module, a third ENAF module and a fourth ENAF module connected in sequence;
[0108] The decoder comprises a first DNAF module, a second DNAF module, a third DNAF module and a fourth DNAF module connected in sequence;
[0109] The first ENAF module is connected to the first DNAF module;
[0110] The second ENAF module is connected to the second DNAF module;
[0111] The third ENAF module, the first dual-branch module and the third DNAF module are connected in sequence;
[0112] The fourth ENAF module, the second dual-branch module and the fourth DNAF module are connected in sequence;
[0113] The fourth ENAF module includes 28 NAF modules, and the first ENAF module, the second ENAF module, the third ENAF module, the first DNAF module, the second DNAF module, the third DNAF module and the fourth DNAF module each include one NAF module.
[0114] As an example, the GoPro dataset is used as the training set, and the HIDE dataset and the GoPro dataset are used as the validation sets. The GoPro dataset contains 3214 pairs of blurry and clear images, of which 2103 pairs are used for training and 1111 pairs are used for testing. The model trained on the GoPro dataset will also be tested on the HIDE dataset.
[0115] Step S2 is specifically as follows:
[0116] S21: inputting the training image into the first FEM module to obtain an initial feature image;
[0117] Specifically, the images of the GoPro dataset obtained in the training phase after image preprocessing are first converted into feature embedding maps through a shallow feature extraction module (FEM). The core of the shallow feature extraction module is a convolution with a convolution kernel of 3×3. The convolution operation is used to extract preliminary local features of the input image, and the number of output channels of the feature map is adjusted. The adjusted feature map is used as the input of the encoder. This step ensures that the input image can be effectively processed by subsequent modules.
[0118] S22: inputting the initial feature image into the first ENAF module to obtain the first feature image; inputting the first feature image into the second ENAF module after downsampling operation to obtain the second feature image; inputting the second feature image into the third ENAF module after downsampling operation to obtain the third feature image; inputting the third feature image into the fourth ENAF module after downsampling operation to obtain the fourth feature image;
[0119] Specifically, the feature extraction of the encoder is divided into four stages, each of which contains a different number of NAF modules and a downsampling operation. The NAF module has the ability to capture local details. The downsampling operation is implemented through a convolutional layer to reduce the spatial size of the feature map and increase the number of channels. The module first loads the feature embedding map processed in step 1, initializes the number of channels to 64, and then traverses the number of NAF modules in each stage. Except for the bottleneck layer, which has 28 NAF modules, there is only one NAF module in other stages. Then, a downsampling layer is added to reduce the spatial size of the feature map and increase the number of channels, so as to gradually extract higher-level features. Finally, the number of channels is updated to prepare for the encoding layer of the next stage. The first three stages of the encoder only contain one nonlinear feature extraction module, in order to use fewer NAF modules on larger feature maps, so as to quickly perform preliminary feature extraction and downsampling, while reducing the amount of calculation and memory consumption. In the last stage (bottleneck layer), a larger number of NAF modules are used, which can perform deeper feature extraction on a smaller spatial size, thereby capturing more complex and abstract features. The entire encoder data flow is as follows: Figure 7 As shown;
[0120] S23: inputting the fourth feature image into the second dual-branch module to perform feature extraction in the spatial domain and the frequency domain to obtain a second fused feature image, and inputting the second fused feature image into the fourth DNAF module to obtain a fourth output image;
[0121] S24: inputting the third feature image into the first dual-branch module to perform feature extraction in the spatial domain and the frequency domain to obtain a first fused feature image, and inputting the first fused feature image and the fourth output image after the upsampling operation into the third DNAF module to obtain a third output image;
[0122] Specifically, the dual branches extract features in the spatial domain and frequency domain respectively, and the parallel dual branch structure is located in the bottleneck layer and the second highest layer of the network. The dual branch structure has a spatial domain branch and a frequency domain branch, which process the feature maps from the encoder respectively. The wavelet transform of the frequency domain branch can decompose and reconstruct these frequency components, helping the network to better capture and restore details. The optimized frequency domain perception module is more adaptable to the characteristics of motion blur. The second highest layer is the layer close to the bottleneck layer in the network. The spatial size of the feature map is slightly larger, but the number of channels is small. The dual branch module of this layer can extract and fuse features on a slightly larger spatial size to supplement the feature representation of the bottleneck layer. The bottleneck layer is the deepest layer in the network. The spatial size of the feature map is the smallest, but the number of channels is the largest. The dual branch module of this layer can extract and fuse deep features on the smallest spatial size. By introducing dual branch modules in the bottleneck layer and the second highest layer, the network can extract and fuse features at different scales, enhance the richness and diversity of feature representation, and help improve the expressiveness and generalization ability of the model.
[0123] The process of obtaining the first fused feature image and the second fused feature image is the same;
[0124] The specific process of obtaining the first fusion feature image is as follows:
[0125] The spatial domain feature image of the third feature image is extracted by the NAF module, the frequency domain feature image of the third feature image is extracted by the LFA module, and the spatial domain feature image and the frequency domain feature image are fused by the D-FM module to obtain a first fused feature image.
[0126] Further:
[0127] The process of obtaining the frequency domain feature image of the third feature image is specifically as follows:
[0128] The third feature image is transformed into a two-dimensional discrete wavelet through the wavelet decomposition module to obtain the low-frequency subband F LL , high frequency sub-band F LH , high frequency sub-band F HLand high frequency subband F HH ;
[0129] Specifically, these frequency sub-bands can be expressed as:
[0130] {F LL ,F LH ,F HL ,F HH = DWT(F 输入 ) (3)
[0131] Where DWT(·) represents two-dimensional discrete wavelet transform, F LL ,F LH ,F HL ,F HH They represent the characteristics of four different frequency sub-bands. In this way, the network can process high-frequency and low-frequency information separately.
[0132] The low frequency subband F LL Input the convolutional network to obtain the enhanced low-frequency subband The high frequency subband F LH , high frequency sub-band F HL and high frequency subband F HH Input the high-frequency feature enhancement module to obtain the enhanced high-frequency sub-bands Enhanced high frequency subband and the enhanced high frequency subband
[0133] Specifically, in order to effectively integrate and enhance these frequency features, the low-frequency sub-band F LL It is embedded in a convolutional network with 1x1 and 7x7 kernels to extract global features. At the same time, the remaining three high-frequency bands are enhanced by the high-frequency feature enhancement module (FMB) to extract local details.
[0134]
[0135] Among them, H conv×1×7 (·) represents a convolutional network with convolution kernels of 1×1 and 7×7, H FMB represents the feature enhancement module, and Represents the high-frequency component after enhancement.
[0136] The enhanced low-frequency subband Enhanced high frequency subband Enhanced high frequency subband and the enhanced high frequency subband Input the wavelet fusion module to perform two-dimensional discrete wavelet inverse transform to obtain the inverse transform feature image F′ out ;
[0137] The inverse transformed feature image F′ out Input the feature distillation module to obtain the distilled feature image, superimpose the distilled feature image with the third feature image to obtain the frequency domain feature image F of the third feature image out .
[0138] Specifically, the four frequency bands of the output are subjected to inverse wavelet transform to obtain the preliminarily processed feature map. Finally, the feature map is further enhanced by the efficient separable distillation module (ESDB) and added to the branch input to obtain the final frequency domain branch output, which can be expressed as follows:
[0139]
[0140] F out =H ESDB (F′ out )+F 输入 (9)
[0141] Where IDWT(·) represents the inverse two-dimensional discrete wavelet transform, H ESDB represents the feature distillation module, F′ out It represents the result of inverse transformation of the four components, F out Represents the output of the final frequency domain branch.
[0142] Further:
[0143] The branch driven fusion mechanism (D-FM) performs weighted fusion on different branch information. In order to effectively fuse the processed dual branch information, the present invention avoids the operation of directly adding the dual branch information, but instead uses a mechanism that can compare and weighted fuse the two path information.
[0144] S25: inputting the second feature image and the third output image after the upsampling operation into the second DNAF module to obtain the second output image; inputting the first feature image and the second output image after the upsampling operation into the first DNAF module to obtain the first output image;
[0145] Specifically, the decoder mainly acts on multi-scale feature maps, upsampling and deblurring them. Similar to the encoder, the decoder also has four stages, namely the lowest layer feature map processing, the middle layer feature map processing, the second highest layer feature map processing and the highest layer feature map processing. Each stage contains a NAF module and an upsampling operation. The highest layer (bottleneck layer) and the second highest layer process the low-resolution feature maps output by the dual branches. The middle layer stage directly fuses the output corresponding to the encoding layer, and then further extracts features. The lowest layer stage upsamples and increases the resolution of the output feature map of the second highest layer stage. The NAF module in the decoder is used to gradually restore the detail information of the feature map. The upsampling layer (convolution layer and pixel rearrangement) is to gradually restore the spatial size of the feature map while reducing the number of channels. The entire decoder restores the spatial size and detail information of the feature map layer by layer through multiple decoding blocks and upsampling layers, and finally generates an output of the same size as the input image.
[0146] S26: inputting the first output image into the second FEM module to obtain an initial deblurred image; superimposing the training image with the initial deblurred image to obtain a final deblurred image;
[0147] Specifically, the original feature embedding map is fused with the deblurred inference image using skip connections to obtain the final deblurred effect map of the model. Here, the role of the skip connection is to retain the highest resolution information and alleviate the gradient disappearance.
[0148] S27: Calculate the loss value through the final deblurred image and adjust the parameters of the deblurring network;
[0149] S28: Repeat steps S21-S27 until the loss value is less than a preset value to obtain a trained deblurring network.
[0150] S3: Deblur the image using the trained deblurring network.
[0151] As an example:
[0152] The present invention adopts peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) as evaluation indicators. In this embodiment, 16 methods are selected for comparison with the proposed method, and the selected methods are DeblurGAN-v2 (see paper: Deblurgan-v2: Deblurring (orders-of-magnitude) faster and better), MRDNet (see paper: Image deblurring method based on self-attention and residual wavelet transform), DBGAN (see paper: Deblurring by realistic blurring), MIMO-UNet (see paper: Rethinking coarse-to-fine approach in single image deblurring), BANet (see paper: Banet: a blur-aware attention network for dynamic scene deblurring), SPAIR (see paper: Spatially-adaptive image restoration using distortion-guided networks), SRN (see paper: Scale-recurrent network for deep image deblurring), SDWNet (see paper: Sdwnet: A straight dilated network with wavelet transformation for image deblurring), and SimpleNet (see paper: Perceptual variousness motion deblurring with light global context refinement), S2SVR (see paper: Unsupervised flow-aligned sequence-to-sequence learning for video restoration), the method proposed by Suin et al. (see paper: Spatially-attentive patch-hierarchical network for adaptive motion deblurring), MSSNet-small (see paper: Mssnet: Multi-scale-stage networkforsingle image deblurring), MPRNet (see paper: Multi-stage progressive image restoration), DMPHN (see paper: Deep stacked hierarchical multi-patch network forimage deblurring), the method proposed by Jiang et al. (see paper: Image blind motion deblurring method with longitudinal channel and wavelet dynamic convolution), MLWNet-width32 (see paper: Efficient multi-scale network with learnable discrete wavelet transform for blind motion deblurring). The test results on the GoPro and HIDE datasets are shown in Table 1. Ours represents the method of the present invention.
[0153] Table 1 Comparison of GoPro and HIDE dataset indicators
[0154]
[0155]
[0156] As can be seen from Table 1, on the GoPro and HIDE datasets, the comprehensive comparison of the two indicators of PSNR and SSIM shows that the method of the present invention can significantly improve the PSNR while effectively maintaining a high SSIM value. Compared with the advanced deblurring method MIMO-Unet, the PSNR of this method on the GOPro dataset is improved by 1.11, and the SSIM is improved by 0.035, and the PSNR on the HIDE dataset is improved by 1.08, and the SSIM is improved by 0.011; compared with the latest method MLWNet-width32, the PSNR of this method on the GoPro dataset is improved by 0.13, and the SSIM is improved by 0.026. The experimental results show that the method of the present invention can effectively improve the image deblurring effect.
[0157] See also Figure 8 , Figure 8 4 is a schematic diagram of the working of the hardware device of an embodiment of the present invention, wherein the hardware device specifically comprises: a dual-branch image deblurring device 401 based on frequency domain perception, a processor 402 and a storage medium 403.
[0158] A dual-branch image deblurring device 401 based on frequency domain perception: The dual-branch image deblurring device 401 based on frequency domain perception implements the dual-branch image deblurring method based on frequency domain perception.
[0159] Processor 402: The processor 402 loads and executes instructions and data in the storage medium 403 to implement the dual-branch image deblurring method based on frequency domain perception.
[0160] Storage medium 403: The storage medium 403 stores instructions and data; the storage medium 403 is used to implement the dual-branch image deblurring method based on frequency domain perception.
[0161] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0162] The serial numbers of the embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In a unit claim that lists several means, several of these means may be embodied by the same hardware item. The use of the words first, second, and third, etc. does not indicate any order and these words may be interpreted as identifiers.
[0163] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A dual-branch image deblurring method based on frequency domain perception, characterized in that: Includes steps: S1: Obtain a sample image set, perform image preprocessing on the sample image set, and obtain a training image set; S2: construct a deblurring network, train the deblurring network through a training image set, and obtain a trained deblurring network; S3: Deblur the image using the trained deblurring network.
2. The dual-branch image deblurring method based on frequency domain perception according to claim 1, characterized in that: Step S1 is specifically as follows: S11: Load a sample image, and perform cropping, flipping, rotating, and format conversion operations on the sample image in sequence to obtain a training image; S12: Repeat step S11 until all sample images are traversed to obtain a training image set.
3. The dual-branch image deblurring method based on frequency domain perception according to claim 1, characterized in that: The defuzzification network includes: a first FEM module, an encoder, a dual-branch structure network, a decoder, and a second FEM module; The first FEM module is connected to the encoder, the encoder is connected to the decoder, and the decoder is connected to the second FEM module; The encoder, dual-branch structure network and decoder are connected in sequence.
4. The dual-branch image deblurring method based on frequency domain perception according to claim 3 is characterized in that: The dual-branch structure network includes a first dual-branch module and a second dual-branch module in parallel; The first double-branch module and the second double-branch module have the same structure; The first dual-branch module includes: a NAF module, an LFA module and a D-FM module; The encoder is connected to the NAF module and the LFA module; The NAF module and the LFA module are connected to the D-FM module; The D-FM module is connected to the decoder; The LFA module includes: wavelet decomposition module, convolutional network, high-frequency feature enhancement module, wavelet fusion module and feature distillation module; The encoder is connected to the wavelet decomposition module; The wavelet decomposition module is connected with the convolutional network and the high-frequency feature enhancement module; The convolutional network and high-frequency feature enhancement module are connected with the wavelet fusion module; The wavelet fusion module, feature distillation module and D-FM module are connected in sequence.
5. The dual-branch image deblurring method based on frequency domain perception according to claim 4, characterized in that: The encoder includes a first ENAF module, a second ENAF module, a third ENAF module and a fourth ENAF module connected in sequence; The decoder comprises a first DNAF module, a second DNAF module, a third DNAF module and a fourth DNAF module connected in sequence; The first ENAF module is connected to the first DNAF module; The second ENAF module is connected to the second DNAF module; The third ENAF module, the first dual-branch module and the third DNAF module are connected in sequence; The fourth ENAF module, the second dual-branch module and the fourth DNAF module are connected in sequence; The fourth ENAF module includes 28 NAF modules, and the first ENAF module, the second ENAF module, the third ENAF module, the first DNAF module, the second DNAF module, the third DNAF module and the fourth DNAF module each include one NAF module.
6. The dual-branch image deblurring method based on frequency domain perception according to claim 5, characterized in that: Step S2 is specifically as follows: S21: inputting the training image into the first FEM module to obtain an initial feature image; S22: inputting the initial feature image into the first ENAF module to obtain the first feature image; inputting the first feature image into the second ENAF module after downsampling operation to obtain the second feature image; inputting the second feature image into the third ENAF module after downsampling operation to obtain the third feature image; inputting the third feature image into the fourth ENAF module after downsampling operation to obtain the fourth feature image; S23: inputting the fourth feature image into the second dual-branch module to perform feature extraction in the spatial domain and the frequency domain to obtain a second fused feature image, and inputting the second fused feature image into the fourth DNAF module to obtain a fourth output image; S24: inputting the third feature image into the first dual-branch module to perform feature extraction in the spatial domain and the frequency domain to obtain a first fused feature image, and inputting the first fused feature image and the fourth output image after the upsampling operation into the third DNAF module to obtain a third output image; S25: inputting the second feature image and the third output image after the upsampling operation into the second DNAF module to obtain the second output image; inputting the first feature image and the second output image after the upsampling operation into the first DNAF module to obtain the first output image; S26: inputting the first output image into a second FEM module to obtain an initial deblurred image; Superimpose the training image with the initial deblurred image to obtain the final deblurred image; S27: Calculate the loss value through the final deblurred image and adjust the parameters of the deblurring network; S28: Repeat steps S21-S27 until the loss value is less than a preset value to obtain a trained deblurring network.
7. The dual-branch image deblurring method based on frequency domain perception according to claim 6, characterized in that: The process of obtaining the first fused feature image and the second fused feature image is the same; The specific process of obtaining the first fusion feature image is as follows: The spatial domain feature image of the third feature image is extracted by the NAF module, the frequency domain feature image of the third feature image is extracted by the LFA module, and the spatial domain feature image and the frequency domain feature image are fused by the D-FM module to obtain a first fused feature image.
8. The dual-branch image deblurring method based on frequency domain perception according to claim 7, characterized in that: The process of obtaining the frequency domain feature image of the third feature image is specifically as follows: The third feature image is transformed into a two-dimensional discrete wavelet through the wavelet decomposition module to obtain the low-frequency subband F LL , high frequency sub-band F LH , high frequency sub-band F HL and high frequency subband F HH ; The low frequency subband F LL Input the convolutional network to obtain the enhanced low-frequency subband The high frequency subband F LH , high frequency sub-band F HL and high frequency subband F HH Input the high-frequency feature enhancement module to obtain the enhanced high-frequency sub-bands Enhanced high frequency subband and the enhanced high frequency subband The enhanced low-frequency subband Enhanced high frequency subband Enhanced high frequency subband and the enhanced high frequency subband Input the wavelet fusion module to perform two-dimensional discrete wavelet inverse transform to obtain the inverse transform feature image F′ out ; The inverse transformed feature image F′ out Input the feature distillation module to obtain the distilled feature image, and superimpose the distilled feature image with the third feature image to obtain the frequency domain feature image F of the third feature image. out .
9. A storage medium, characterized in that: The storage medium stores instructions and data for implementing the dual-branch image deblurring method based on frequency domain perception as described in any one of claims 1 to 8.
10. A dual-branch image deblurring device based on frequency domain perception, characterized in that: include: A processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement the dual-branch image deblurring method based on frequency domain perception as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Image deblurring method and system based on cavity double-residual multi-scale deep network
CN114723630A
Target detection method based on multi-characterization feature extraction method
CN116206123A
Food image segmentation method based on discrete wavelet attention network
CN116630964A
Method for realizing super-resolution for real-world text image through double-branch network capable of sensing multiple features
CN116703725A
Low-illumination image enhancement method based on multi-semantic feature fusion network
CN117408924A
Cited By
Freehand sketch generation system and method based on wavelet transform
CN121505070A