An image deblurring method based on dilated convolution with gradual change rate
By adopting a progressive rate of change hollow convolution and deformable convolution structure in the image defuzzy network, the problem of reconstructing characters in the prior art when restoring strong structure information content in traffic scenes is solved, and a more efficient image defuzzy effect is achieved.
Patent Information
- Application Number
- CN202111532682.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-12-15
AI Technical Summary
When existing deep learning methods restore content with strong geometric structure information such as license plates and street signs that are common in traffic scenes, they are likely to cause errors in reconstructing characters, affecting the effectiveness of the image defuzzing method.
The image defuzzy network is adopted based on asymmetry rate of void convolution, and the asymmetry rate of void convolution pyramid and deformable convolution structure are used to replace the repeated up-down sampling process to avoid feature misalignment, build joint loss function and adaptive dimensional clipping function, and optimize network training.
It effectively alleviates the problem of reconstructing character errors, improves the effectiveness of image defuzzing methods in traffic scenes, can accurately restore strong structural information content in blurred images, and improves image quality.
Smart Images

Figure CN114170111B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to an image deblurring method based on gradual change rate atrous convolution. Background Art
[0002] In recent years, with the price of image acquisition equipment gradually decreasing, various types of equipment have been widely used in various fields of actual production and life. In traffic scenes, whether it is traffic scene monitoring based on fixed probes, airborne cameras based on drones, or various scenes based on vehicle-mounted cameras, the following problems are inevitable: the image captured by the image acquisition device is blurred due to the shutter speed being too low in the shooting conditions, the relative movement speed between the object and the camera being too fast, etc. This type of image blur problem brings huge challenges to both human eye recognition tasks in real scenes and other computer vision tasks (such as license plate recognition and abnormal event monitoring).
[0003] The traditional solution is to simplify the blur model in the real scene and then build a model. These methods mainly include using different natural image prior knowledge to constrain the solution space to achieve the modeling of uniform blur, non-uniform blur, and blur that takes depth into consideration. These methods all involve a large number of manually designed parameters and costly calculations. In addition, the formation process of image blur in real scenes is more complicated than the modeling situation, and the prior reconstruction quality designed by traditional methods is extremely poor.
[0004] In recent years, with the widespread use of deep learning methods in low-level visual tasks, deep learning image deblurring methods trained with large-scale data sets have greatly improved the quality of reconstructed images and significantly improved the visual quality of reconstructed images. In addition, the image deblurring method designed based on convolutional neural networks realizes end-to-end image reconstruction and can be quickly ported to various deep learning embedded development boards to achieve real-time reconstruction and enhancement of images.
[0005] With the development of deep learning, there are a large number of image deblurring methods for real motion scenes. However, the network structures designed by these methods are often very complex, and a large number of repeated up- and down-sampling operations and the idea of generative adversarial networks are introduced in the network design. These methods will cause misalignment of image features, and then cause errors in reconstructing characters when restoring content with strong geometric structure information such as license plates and road signs commonly seen in traffic scenes, which seriously affects the effective use of image deblurring methods in traffic scenes. Summary of the invention
[0006] The technical problem to be solved by the present invention is to provide an image deblurring method based on gradually changing rate dilated convolution in response to the shortcomings of the above-mentioned prior art. The method replaces the repeated up and down sampling process with gradually changing dilated convolution to avoid the feature misalignment caused by the process, and is used to accurately restore the strong structural information content in the blurred images of real scenes.
[0007] In order to achieve the above technical objectives, the technical solution adopted by the present invention is:
[0008] An image deblurring method based on a gradual change rate dilated convolution, comprising:
[0009] Step 1: Construct an image deblurring network based on dilated convolution with a gradual change rate;
[0010] Step 2: construct the joint loss function of the image deblurring network in step 1;
[0011] Step 3: construct an adaptive size cropping function of the image deblurring network in step 1;
[0012] Step 4: Collect a number of clear and blurred image pairs to construct an image deblurring sample set;
[0013] Step 5: training the image deblurring network using the image deblurring sample set, the joint loss function and the adaptive size reduction function;
[0014] Step 6: Input the image to be enhanced into the trained image deblurring network to obtain a reconstructed image and achieve image deblurring.
[0015] To optimize the above technical solutions, the specific measures taken also include:
[0016] The image deblurring network based on the gradual change rate dilated convolution constructed in the above step 1 includes a feature extraction module, a gradual change rate dilated convolution pyramid module, a feature fusion module, and a feature distillation module;
[0017] Among them, the feature extraction module performs preliminary feature processing on the image and reduces the image resolution scale;
[0018] The progressive rate dilated convolutional pyramid module further processes the image features output by the feature extraction module and maps them to a high-dimensional space;
[0019] The feature fusion module upsamples the features output by the progressive change rate atrous convolution pyramid module to the original size and connects them with the shallow features of the network;
[0020] The feature distillation module further improves the output of the feature fusion module through deformable convolution to obtain the final output image.
[0021] The above feature extraction module consists of two convolutional layers, where the convolution kernel of the first convolutional layer is a 5*5 convolution; and the convolution kernel of the second convolutional layer is a 3*3 convolution.
[0022] The above-mentioned progressively changing rate atrous convolution pyramid module includes a plurality of parallel atrous convolution modules of pyramid structures;
[0023] The dilated convolution module is divided into odd-numbered dilated convolution module and even-numbered dilated convolution module;
[0024] Among them, the even-numbered atrous convolution module copies the input feature map into four copies, which enter the following four atrous convolution branches respectively:
[0025] The first branch has a dilated convolution rate of 1 and a convolution kernel of 3;
[0026] The second branch has a dilated convolution rate of 2 and a convolution kernel of 3;
[0027] The third branch has a dilated convolution rate of 4 and a convolution kernel of 3;
[0028] The fourth branch has a dilated convolution rate of 8 and a convolution kernel of 3.
[0029] After passing through four dilated convolution branches, the output feature map is spliced along the channel dimension. After passing through a 3*3 convolution layer and a channel shuffle layer, the input feature map is added to the output feature map after the channel shuffle layer to obtain the final output feature map.
[0030] In the odd-numbered atrous convolution module, the input feature map is copied into four copies and enters the following four atrous convolution branches respectively:
[0031] The first branch has a dilated convolution rate of 1 and a convolution kernel of 3;
[0032] The second branch has a dilated convolution rate of 3 and a convolution kernel of 3;
[0033] The third branch has a dilated convolution rate of 5 and a convolution kernel of 3;
[0034] The fourth branch has a dilated convolution rate of 7 and a convolution kernel of 3.
[0035] After passing through four dilated convolution branches, the output feature map is spliced along the channel dimension. After passing through a 3*3 convolution layer and a channel shuffle layer, the input feature map is added to the output feature map after the channel shuffle layer to obtain the final output feature map.
[0036] The above feature fusion module consists of a 3*3 convolution kernel and a 2x upsampling layer.
[0037] The above-mentioned feature distillation module includes two deformable convolutional layers. The first deformable convolutional layer has a convolution kernel of 3, 64 input channels, and 128 output channels; the second deformable convolutional layer has a convolution kernel of 3, 128 input channels, and 3 output channels.
[0038] In step 2 above, a joint loss function is constructed, including:
[0039] 1) Construct image similarity contrast loss L ssim :
[0040]
[0041] Wherein, SSIM is SSIM(x, y)=[l(x, y)] α [c(x,y)] β [s(x,y)] γ ;
[0042] X i , Y i , N are the i-th contrast element in the reconstructed image, the i-th contrast element in the real clear image, and the total number of contrast elements, respectively;
[0043] l(x, y) is the brightness component:
[0044]
[0045] Among them, x represents the reconstructed image, y represents the real clear image; μ x ,μ y Respectively represent the mean pixel grayscale value of the reconstructed image and the mean pixel grayscale value of the real clear image;
[0046] c(x, y) is the color component:
[0047]
[0048] Among them, σ x ,σ y They represent the normalized variance of the mean pixel value of the reconstructed image and the normalized variance of the mean pixel value of the real clear image respectively;
[0049] s(x, y) is the structural component:
[0050] s(x, y) = σ xy +C 3 / (σ x σ y +C 3 );
[0051] Take α=β=γ=1,C 1 =2.55, C2 =7.65, C 3 =3.55;
[0052] Among them, C 3 represents the constant term when the mean is close to 0;
[0053] 2) Construct image Charbonier loss L char :
[0054]
[0055] Among them, ∈ is the compensation factor;
[0056] 3) According to 1) and 2, we get the joint loss function L total :
[0057] L total =L char +μ*L ssim
[0058] Where μ is the loss adjustment coefficient.
[0059] In step 3 above, the adaptive size cropping function is:
[0060]
[0061] Epochs represents the number of training rounds of the network, and the cropping function indicates that a cropping size of 256*256 is used from the 1st to the 19th rounds, and a cropping size of 448*448 is used in the 20th round and thereafter.
[0062] The above-mentioned pictures to be enhanced are pictures of various viewing angles including traffic scene elements;
[0063] The various perspectives include perspectives taken by drone aerial photography, fixed vehicle-mounted cameras, and fixed surveillance probes;
[0064] The traffic scene elements include license plates, signs and road surface damage.
[0065] The present invention has the following beneficial effects:
[0066] The present invention uses an image deblurring network based on a progressively changing rate atrous convolution, a combined loss function and an adaptive size reduction function to alleviate the problem that the image deblurring network of the existing deep learning method will cause errors in reconstructing characters when restoring content with strong geometric structure information such as license plates and road signs commonly seen in traffic scenes, thereby promoting the effective use of image deblurring methods in traffic scenes.
[0067] The present invention constructs an image deblurring algorithm through a progressively changing rate atrous convolution pyramid and a deformable convolution structure to deblur blurred images, and image deblurring can be achieved without a large number of repeated up and down sampling.
[0068] The algorithm formed by the present invention has strong generalization ability, and can effectively alleviate image blur caused by reasons such as too low shutter speed in shooting conditions and too fast relative motion speed between the object and the camera, thereby improving the quality of images collected in the real world, especially for images with strong structured features such as traffic scene elements (such as license plates, signs, road surface diseases) in scenes from various perspectives (including but not limited to drone aerial pictures, fixed vehicle-mounted cameras, and fixed monitoring probes). It has a better restoration effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 is a flow chart of the method of the present invention;
[0070] Figure 2 It is a schematic diagram of the overall structure of the network of the present invention;
[0071] Figure 3 This is a schematic diagram of the pyramid structure of the odd-numbered hole convolution module;
[0072] Figure 4 Schematic diagram of the pyramid structure of the even-numbered dilated convolution module;
[0073] Figure 5 Comparison of the images reconstructed by the network of the present invention with other methods. DETAILED DESCRIPTION
[0074] The embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings.
[0075] See also Figure 1 , an image deblurring method based on gradual change rate dilated convolution, comprising:
[0076] Step 1: Construct an image deblurring network based on dilated convolution with a gradual change rate;
[0077] Step 2: construct the joint loss function of the image deblurring network in step 1;
[0078] Step 3: construct an adaptive size cropping function of the image deblurring network in step 1;
[0079] Step 4: Collect a number of clear and blurred image pairs to construct an image deblurring sample set;
[0080] Step 5: training the image deblurring network using the image deblurring sample set, the joint loss function and the adaptive size reduction function;
[0081] Step 6: Input the image to be enhanced into the trained image deblurring network to obtain a reconstructed image and achieve image deblurring.
[0082] See also Figure 2 In the embodiment, the image deblurring network based on the gradual change rate dilated convolution constructed in step 1 includes a feature extraction module, a gradual change rate dilated convolution pyramid module, a feature fusion module, and a feature distillation module;
[0083] Among them, the feature extraction module performs preliminary feature processing on the image and reduces the image resolution scale;
[0084] The progressive rate dilated convolutional pyramid module further processes the image features output by the feature extraction module and maps them to a high-dimensional space;
[0085] The feature fusion module upsamples the features output by the progressive change rate atrous convolution pyramid module to the original size and connects them with the shallow features of the network;
[0086] The feature distillation module further improves the output of the feature fusion module through deformable convolution to obtain the final output image.
[0087] In the embodiment, the feature extraction module is composed of two convolutional layers, wherein the convolution kernel of the first convolutional layer is a 5*5 convolution; and the convolution kernel of the second convolutional layer is a 3*3 convolution.
[0088] In an embodiment, the progressive change rate atrous convolution pyramid module includes a plurality of parallel atrous convolution modules of pyramid structure;
[0089] The dilated convolution module is divided into odd-numbered dilated convolution module and even-numbered dilated convolution module;
[0090] Among them, see Figure 3 In the odd-numbered atrous convolution module, the input feature map is copied into four copies and enters the following four atrous convolution branches respectively:
[0091] The first branch has a dilated convolution rate of 1 and a convolution kernel of 3;
[0092] The second branch has a dilated convolution rate of 3 and a convolution kernel of 3;
[0093] The third branch has a dilated convolution rate of 5 and a convolution kernel of 3;
[0094] The fourth branch has a dilated convolution rate of 7 and a convolution kernel of 3.
[0095] After passing through four dilated convolution branches, the output feature map is spliced along the channel dimension. After passing through a 3*3 convolution layer and a channel shuffle layer, the input feature map is added to the output feature map after the channel shuffle layer to obtain the final output feature map.
[0096] See also Figure 4 , the input feature map is copied into 4 copies in the even-numbered dilated convolution module, and enters the following 4 dilated convolution branches respectively:
[0097] The first branch has a dilated convolution rate of 1 and a convolution kernel of 3;
[0098] The second branch has a dilated convolution rate of 2 and a convolution kernel of 3;
[0099] The third branch has a dilated convolution rate of 4 and a convolution kernel of 3;
[0100] The fourth branch has a dilated convolution rate of 8 and a convolution kernel of 3.
[0101] After passing through four dilated convolution branches, the output feature map is spliced along the channel dimension. After passing through a 3*3 convolution layer and a channel shuffle layer, the input feature map is added to the output feature map after the channel shuffle layer to obtain the final output feature map.
[0102] In the embodiment, the feature fusion module is composed of a 3*3 convolution kernel and a 2-fold upsampling layer.
[0103] In an embodiment, the feature distillation module includes two deformable convolutional layers, the first deformable convolutional layer has a convolution kernel of 3, an input channel number of 64, and an output channel number of 128; the second deformable convolutional layer has a convolution kernel of 3, an input channel number of 128, and an output channel number of 3.
[0104] In the embodiment, in step 2, constructing a joint loss function includes:
[0105] 1) Construct image similarity contrast loss L ssim :
[0106]
[0107] Wherein, SSIM is SSIM(x, y)=[l(x, y)] α [c(x,y)] β [s(x,y)] γ ;
[0108] X i , Y i , N are the i-th contrast element in the reconstructed image, the i-th contrast element in the real clear image, and the total number of contrast elements, respectively;
[0109] l(x,y) is the brightness component:
[0110]
[0111] Among them, x represents the reconstructed image, y represents the real clear image; μ x ,μ y Respectively represent the mean pixel grayscale value of the reconstructed image and the mean pixel grayscale value of the real clear image, C 1 represents the constant term when the mean is close to 0;
[0112] c(x,y) is the color component:
[0113]
[0114] Among them, σ x ,σ y They represent the normalized variance of the pixel mean of the reconstructed image and the normalized variance of the pixel mean of the real clear image, respectively. 2 represents the constant term when the mean is close to 0;
[0115] s(x,y) is the structural component:
[0116] x(x,y)=σ xy +C 3 * / (σ x σ y +C 3 );
[0117] Take α=β=γ=1,C 1 =2.55,C 2 =7.65,C 3 =3.55;
[0118] Among them, C 3 represents the constant term when the mean is close to 0;
[0119] 2) Construct image Charbonier loss L char :
[0120]
[0121] Wherein, ∈ is a compensation factor, which is used to prevent the difference between the absolute values of two pixels from being 0. In the embodiment, ∈=1e-8;
[0122] 3) According to 1) and 2, we get the joint loss function L total :
[0123] L total =I char +μ*L ssim
[0124] Wherein, μ is the loss adjustment coefficient, and μ=1 in the embodiment.
[0125] In the embodiment, in step 3, the adaptive size cropping function is:
[0126]
[0127] Epochs represents the number of training rounds of the network, and the cropping function indicates that a cropping size of 256*256 is used from the 1st to the 19th rounds, and a cropping size of 448*448 is used in the 20th round and thereafter.
[0128] In the embodiment, the picture to be enhanced is a picture of various viewing angles including traffic scene elements;
[0129] The various perspectives include perspectives taken by drone aerial photography, fixed vehicle-mounted cameras, and fixed surveillance probes;
[0130] The traffic scene elements include license plates, signs and road surface damage.
[0131] See also Figure 5 The present invention uses an image deblurring network based on a progressively changing rate atrous convolution, a combined loss function and an adaptive size reduction function to alleviate the problem that the image deblurring network of the existing deep learning method will cause errors in reconstructing characters when restoring content with strong geometric structure information such as license plates and road signs commonly seen in traffic scenes, and can promote the effective use of image deblurring methods in traffic scenes. Figure 5 Reference methods 1-5 are:
[0132] Reference method 1. SRN: Scale-recurrent network for deep image deblurring. In 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 8174–8182, 2018.
[0133] Reference method 2, DeblurGANV2: Kupyn, T.Martyniuk, J.Wu, and Z.Wang.Deblurganv2: Deblurring (orders-of-magnitude) faster and better. In 2019 IEEE / CVFInternational Conference on Computer Vision (ICCV), pages 8877–8886, 2019
[0134] Reference method 3, Gao et al. method: H.Gao, X.Tao, X.Shen, and J.Jia.Dynamic scene deblurring with parameter selective sharing and nested skip connections. In2019IEEE / CVF Conference on Computer Vision and Pattern Recognition(CVPR), pages 3843–3851, 2019.
[0135] Reference method 4. MTRNN: Dongwon Park, Dong Un Kang, Jisoo Kim, and Se YoungChun. Multi-temporal recurrent neural networks for pro progressive non-uniformsingle image deblurring with incre mental temporal training. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, ComputerVision–ECCV 2020. Springer International Publishing, 2020
[0136] Reference method 5, DBGAN: Kaihao Zhang, Wenhan Luo, Yiran Zhong, Lin Ma, BjornStenger, Wei Liu, and Hongdong Li.Deblurring by realis tic blurring.InProceedings of the IEEE / CVF Conference on Computer Vision and PatternRecognition, pages 2737–2746, 2020.
[0137] The above are only preferred embodiments of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should be regarded as the protection scope of the present invention.
Claims
1. An image deblurring method based on dilated convolution with a gradual change rate, It is characterized in that include: Step 1: Construct an image deblurring network based on dilated convolution with a gradual change rate; Step 2: construct the joint loss function of the image deblurring network in step 1; Step 3: construct an adaptive size cropping function of the image deblurring network in step 1; Step 4: Collect a number of clear and blurred image pairs to construct an image deblurring sample set; Step 5: training the image deblurring network using the image deblurring sample set, the joint loss function and the adaptive size cropping function; Step 6: Input the image to be enhanced into the trained image deblurring network to obtain a reconstructed image and realize image deblurring; In step 2, a joint loss function is constructed, including: 1) Construct image similarity contrast loss L ssim : Wherein, SSIM is SSIM(x, y)=[l(x, y)] α [c(x,y)] β [s(x,y)] γ ; X i , Y i , N are the i-th contrast element in the reconstructed image, the i-th contrast element in the real clear image, and the total number of contrast elements, respectively; Among them, l(x, y) is the brightness component: Among them, x represents the reconstructed image, y represents the real clear image; μ x , μ y Respectively represent the mean pixel grayscale value of the reconstructed image and the mean pixel grayscale value of the real clear image; c(x, y) is the color component: Among them, σ x , σ y They represent the normalized variance of the mean pixel value of the reconstructed image and the normalized variance of the mean pixel value of the real clear image respectively; s(x, y) is the structural component: s(x, y) = σ xy + C 3 / (σ x σ y + C 3 ); Take α=β=γ=1, C 1 =2.55, C 2 =7.65, C 3 =3.55; 2) Construct image Charbonier loss L char : Among them, ∈ is the compensation factor; 3) According to 1) and 2, we get the joint loss function L total :L total =L char +μ*L ssim Where μ is the loss adjustment coefficient.
2. The image deblurring method based on gradual rate dilated convolution according to claim 1, It is characterized in that The image deblurring network based on the gradual change rate atrous convolution constructed in step 1 includes a feature extraction module, a gradual change rate atrous convolution pyramid module, a feature fusion module, and a feature distillation module; Among them, the feature extraction module performs preliminary feature processing on the image and reduces the image resolution scale; The progressive rate dilated convolutional pyramid module further processes the image features output by the feature extraction module and maps them to a high-dimensional space; The feature fusion module upsamples the features output by the progressive change rate atrous convolution pyramid module to the original size and connects them with the shallow features of the network; The feature distillation module further improves the output of the feature fusion module through deformable convolution to obtain the final output image.
3. The image deblurring method based on gradual change rate dilated convolution according to claim 2, It is characterized in that The feature extraction module consists of two convolutional layers, wherein the convolution kernel of the first convolutional layer is a 5*5 convolution; and the convolution kernel of the second convolutional layer is a 3*3 convolution.
4. The image deblurring method based on gradual rate dilated convolution according to claim 3, It is characterized in that The progressive rate dilated convolution pyramid module includes a plurality of parallel dilated convolution modules of pyramid structures; The dilated convolution module is divided into odd-numbered dilated convolution module and even-numbered dilated convolution module; Among them, the even-numbered atrous convolution module copies the input feature map into four copies, which enter the following four atrous convolution branches respectively: The first branch has a dilated convolution rate of 1 and a convolution kernel of 3; The second branch has a dilated convolution rate of 2 and a convolution kernel of 3; The third branch has a dilated convolution rate of 4 and a convolution kernel of 3; The fourth branch has a dilated convolution rate of 8 and a convolution kernel of 3. After passing through four dilated convolution branches, the output feature map is spliced along the channel dimension. After passing through a 3*3 convolution layer and a channel shuffle layer, the input feature map is added to the output feature map after the channel shuffle layer to obtain the final output feature map. In the odd-numbered atrous convolution module, the input feature map is copied into four copies and enters the following four atrous convolution branches respectively: The first branch has a dilated convolution rate of 1 and a convolution kernel of 3; The second branch has a dilated convolution rate of 3 and a convolution kernel of 3; The third branch has a dilated convolution rate of 5 and a convolution kernel of 3; The fourth branch has a dilated convolution rate of 7 and a convolution kernel of 3. After passing through 4 dilated convolution branches, the output feature map is spliced along the channel dimension. After passing through a 3*3 convolution layer and a channel shuffle layer, the input feature map is added to the output feature map after the channel shuffle layer to obtain the final output feature map.
5. The image deblurring method based on gradual change rate dilated convolution according to claim 4, It is characterized in that The feature fusion module consists of a 3*3 convolution kernel and a 2x upsampling layer.
6. The image deblurring method based on gradual rate dilated convolution according to claim 5, It is characterized in that The feature distillation module includes two deformable convolutional layers, the first deformable convolutional layer has a convolution kernel of 3, an input channel number of 64, and an output channel number of 128; the second deformable convolutional layer has a convolution kernel of 3, an input channel number of 128, and an output channel number of 3.
7. The image deblurring method based on gradual rate dilated convolution according to claim 1, It is characterized in that In step 3, the adaptive size cropping function is: Epochs represents the number of training rounds of the network, and the cropping function indicates that a cropping size of 256*256 is used from the 1st to the 19th rounds, and a cropping size of 448*448 is used in the 20th round and thereafter.
8. An image deblurring method based on gradual change rate dilated convolution according to any one of claims 1 to 7, It is characterized in that The pictures to be enhanced are pictures of various viewing angles including traffic scene elements; The various perspectives include perspectives taken by drone aerial photography, fixed vehicle-mounted cameras, and fixed surveillance probes; The traffic scene elements include license plates, signs and road surface damage.