Lightweight two-dimensional code image restoration method based on deep learning
Through the lightweight QR code image restoration method of deep learning, combined with the hybrid convolution module and attention mechanism, the various degradation problems of QR code images in complex scenes are solved, efficient and real-time image restoration effect is achieved, and the decoding rate and recognition rate of QR codes are improved.
Patent Information
- Application Number
- CN202510452722.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-22
AI Technical Summary
When processing QR code images, the prior art has problems such as insufficient adaptability to various degradation types, limited recovery effects, limited real-time and resource limitations, excessive dependence on training data and manual parameter adjustment, and lack of support for real data sets, making it difficult to achieve efficient and real-time image restoration in complex scenarios.
The lightweight QR code image restoration method based on deep learning is adopted, combined with the hybrid convolution module to fuse frequency domain information on spatial feature extraction, pay attention to global and local information through the attention mechanism, and use the QR-DN1.0 data set and Real-ESRGAN to generate degraded images. Combined with the multi-scale pyramid module, multi-branch color enhancement module and linear space module, we improve image details and contrast and design a lightweight network structure.
It significantly improves the quality of QR code images, improves PSNR and SSIM indicators, enhances decoding rate and recognition rate, adapts to a variety of degradation types, meets the needs of real-time and computing resources, and reduces the computational complexity.
Smart Images

Figure CN120355580A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image restoration, and particularly relates to a lightweight QR code image restoration method based on deep learning. Background Art
[0002] With the rapid development of Internet of Things technology and information industry, the two-dimensional code (QR, Quick Response) as an efficient information storage and transmission tool has been widely used in fields such as logistics management, mobile payment, electronic bills, healthcare, and intelligent transportation [1]. Compared with traditional one-dimensional barcodes, two-dimensional codes have advantages such as large data storage capacity, diverse encoded data types, and fast detection speed. However, in actual application scenarios, two-dimensional codes are often affected by the environment or human factors, resulting in damaged image quality, such as stains, scratches, light reflections, blurring, etc. These problems will significantly reduce the recognizability of two-dimensional codes and even lead to scanning failures.
[0003] By manually designing the restoration process, traditional methods have made some progress in solving the problem of damaged two-dimensional code image quality. For example, Van Gennip et al. [3] proposed a regularization method for solving the blurring problem of QR code images, which performs well under specific blurring conditions. Pan et al. [4] proposed a deblurring algorithm based on dark channel prior, which can effectively handle blurring problems in natural scenes. In addition, Rioux et al. proposed and evaluated a blind deblurring and denoising method based entirely on Kullback-Leibler divergence [5]. However, these methods usually take a long time and have limited restoration effects on severely motion-blurred or partially occluded QR codes, making it difficult to meet the robust recognition requirements of two-dimensional codes in complex scenarios.
[0004] In recent years, deep learning-based methods have made remarkable progress in the field of QR code deblurring. Tiwari et al. [6] combined the Ridgelet Transform (RT) and the Radial Basis Function (RBF) neural network to handle the barcode deblurring problem, but the algorithm complexity is relatively high and it is difficult to promote. Pu et al. [7] used a double Convolutional Neural Network (CNN) for deblurring, but the blur kernel needs to be manually set during the training process, which deviates from the complexity of blur in real scenarios. Li et al. [8] proposed a feature extraction architecture based on the encoder-decoder framework to directly achieve QR code deblurring in an end-to-end manner, and showed excellent performance in handling motion blur tasks in dynamic scenarios. Although deep learning methods are superior to traditional methods in terms of restoration ability, there are still some problems. For example, artifacts may appear in the restored image, the edge contours are not clear enough, the computational complexity of the network model is difficult to meet the real-time or performance requirements of embedded devices, and there may be compound degradation situations for two-dimensional codes in reality.
[0005] In addition, there is currently a lack of QR code datasets in real dynamic scenarios specifically for training neural networks, which limits the further development and application of deep learning models. To improve the adaptability of two-dimensional code decoding algorithms to various complex degradation scenarios, the diversity and authenticity of the dataset play a crucial role. However, when dealing with damaged two-dimensional codes, existing methods usually rely on synthetic datasets or simple degradation simulations, which are difficult to truly reflect the problems in complex real scenarios, such as various types of degradation like stains, scratches, light reflections, and partial missing. Therefore, the lack of support from real and rich datasets has become one of the important limitations in two-dimensional code decoding research, and there is an urgent need for an efficient solution that can adapt to complex scenarios and comprehensively handle the quality problems of two-dimensional code images, which can not only improve the restoration ability of damaged two-dimensional codes but also take into account real-time performance and optimization of computing resources.
[0006] For degradation simulation, in recent years, some super-resolution and image degradation techniques have provided new possibilities for generating high-quality degraded simulation datasets. For example, Real-ESRGAN is a general super-resolution generative adversarial network [2], which can generate realistic degradation effects on images, including blur, noise, low resolution, etc., and these characteristics can be used to synthesize damaged two-dimensional code image datasets in real scenarios.
[0007] Real-ESRGAN generates blurred images through high-order degradation modeling to simulate the complex degradation process in real scenarios: it uses randomly mixed blur kernels (such as Gaussian, motion, anisotropic blur, etc.) to dynamically generate blur kernels of different sizes, directions, and intensities, couples and superimposes blurring with operations such as downsampling, noise, and compression through a multi-stage degradation chain, and introduces random parameters to enhance diversity, thereby covering the complex blur forms caused by optical limitations, motion jitter, or transmission compression in real shooting, and providing more realistic data for the model. The degraded QR code dataset generated in this way can more realistically reflect the problems that may occur in actual application scenarios, providing more effective support for the research and optimization of decoding algorithms.
[0008] When dealing with damaged QR code images, existing technologies mainly adopt image preprocessing methods. Common denoising techniques include mean filtering, Gaussian filtering, etc., which reduce noise interference by smoothing the image and improve the clarity of the image. At the same time, contrast enhancement methods (such as histogram equalization and gamma correction) can effectively enhance the image contrast, making the edges of the QR code more clearly visible. To alleviate the blur problem, Laplacian filtering or other non-linear sharpening techniques are usually also applied to highlight image details. These traditional methods have improved the recognizability of QR codes to a certain extent, but most of them are designed for single problems and are difficult to handle QR code images with multiple complex degradation types. Especially when the image has serious missing or occlusion, the restoration effect is very limited.
[0009] In recent years, deep learning technologies have been introduced into the field of QR code image restoration, showing strong potential. Through models such as convolutional neural networks (CNNs) and generative adversarial networks (GANs), more complex image feature extraction and reconstruction can be achieved, and the processing effects for problems such as blur and dirt are better than traditional methods. However, existing deep learning methods still have certain limitations. For example, they have a large number of model parameters and high computational complexity, making it difficult to apply in embedded devices or real-time scenarios. In addition, these methods have a high dependence on training data and may have poor generalization ability when the data diversity is insufficient. Especially in the field of QR code image processing, these problems are particularly prominent. When dealing with damaged QR code images, existing technologies mainly have the following problems:
[0010] 1. Insufficient adaptability to multiple degradation types
[0011] Current methods are usually optimized for specific types of image degradation (such as blur or noise), but it is difficult to handle the situation where multiple degradations (such as dirt, scratches, blur, and partial missing) coexist in complex scenarios.
[0012] 2. Limited restoration effect
[0013] For the loss or severe occlusion of local information in the QR code image, existing methods often struggle to fully reconstruct the QR code content, resulting in insufficient decoding success rate and recognition accuracy.
[0014] 3. Real-time performance and resource constraints
[0015] Although deep learning methods have strong restoration capabilities, they usually have high computational complexity, require high hardware performance, and are difficult to meet the needs of embedded devices or real-time application scenarios.
[0016] 4. Over-reliance on training data and manual parameter tuning
[0017] Deep learning methods have high requirements for the quality and diversity of training data, making it difficult to ensure robustness in actual scenarios. In addition, some traditional methods require manual parameter adjustment, further reducing the degree of automation of processing.
[0018] 5. Lack of support for real datasets
[0019] In the development and evaluation process of existing methods, they often rely on synthetic datasets or QR code images in simple scenarios, and cannot truly reflect the degradation characteristics of QR codes in complex actual environments (such as scratches, stains, light reflection, etc.). The lack of support for real datasets may lead to insufficient generalization ability of the model or algorithm, making it difficult to adapt to the diverse and complex scenarios in actual applications. Summary of the Invention
[0020] The purpose of the present invention is to solve the problems existing in the prior art and provide a lightweight QR code image restoration method based on deep learning to meet the dual requirements of restoration effect and real-time performance in actual applications.
[0021] To achieve the above purpose, the technical solution of the present invention is: a lightweight QR code image restoration method based on deep learning, which combines a hybrid convolution module, fuses frequency domain information on the basis of spatial feature extraction, enhances the details and contrast of the QR code image, refines the image color, and pays attention to global and local information through an attention mechanism to refine the image details, greatly improving the image quality while ensuring computational efficiency, and finally completing the restoration of the QR code image.
[0022] Furthermore, for the establishment of the QR code dataset, the QR-DN1.0 dataset and the method of generating additional QR code images are adopted to construct the training set, test set and validation set. At the same time, degraded images are generated based on the QR-DN1.0 dataset and the generated additional QR code images. When using the QR-DN1.0 dataset and the generated additional QR code images to generate degraded images, Real-ESRGAN is used to generate blurred images; then, by generating lines and overlaying them on the images, scratches are simulated; blurred patterns are overlaid on the images to simulate dirt; the values of the RGB channels are modified to simulate color offsets caused by the influence of light sources; according to the attenuation mask, the color values are adjusted pixel by pixel to simulate local aging, and motion blur is added to the test set alone to test the generalization performance.
[0023] Furthermore, the method includes three stages:
[0024] In the rough restoration stage, hybrid convolution is used to extract initial features from the input image, that is, a 3-channel RGB image, and the number of channels is expanded to 32; then, based on the features extracted by the hybrid convolution, the local details of the image are improved, the color of the QR code is corrected, and the contrast is enhanced through the multi-scale pyramid module and the multi-branch color enhancement module respectively, and the QR code is roughly restored.
[0025] In the color refinement stage, the quality of the QR code image is improved by operating on the frequency components through the spatial-frequency domain enhancement module; specifically, first, the outputs of the multi-scale pyramid module and the multi-branch color enhancement module are fused. In the initial stage of fusion, the output feature maps are added element by element to fully integrate the information of different enhancement modules; then, the SFBlock module is designed to enhance the global and local detail features through the joint operation of the spatial domain and the frequency domain; in the spatial domain, the SFBlock module first extracts global information through global pooling, and then fuses the local and global information to enhance the model's perception ability of information at different scales; finally, all these spatial domain features are fused and prepared to be passed to the frequency domain processing stage; in the frequency domain, the amplitude and phase information of the image is decomposed using the fast Fourier transform FFT, the amplitude component is convolved and enhanced and reconstructed; subsequently, the frequency domain features are restored to the spatial domain through the inverse fast Fourier transform IFFT and fused with the spatial domain features; after passing through the spatial-frequency domain enhancement module, the feature expression ability of the image features is further enhanced through hybrid convolution; then, the multi-branch color enhancement module is used to further perform color correction and balance optimization on the QR code image, and the color of the QR code is refined.
[0026] In the detailed refinement stage, through the linear space module LSA, it not only focuses on global features but also can be used to focus on local detailed features, taking into account both computational efficiency and image quality. Specifically, the LSA is used to enhance the image restoration ability. First, the input features are spatially enhanced through a 3x3 depthwise separable convolution to extract local texture and edge information, and then normalized through a normalization layer. Next, a global dependency relationship between features is captured through a Mamba-Like Linear Attention (MLLA) mechanism to obtain a feature representation that fuses global context. Subsequently, the information flow is enhanced through a residual connection to further retain the correlation between low-order and high-order features. Next, a Multilayer Perceptron (MLP) further learns the globally enhanced features to improve the non-linear expression ability. At the same time, to generate a spatial attention map, the feature map is reduced to 1 / 8 of its original size through a 1x1 convolution. Then, the feature expression ability is enhanced through ReLU activation, and a single-channel spatial attention map is generated again through a 1x1 convolution to represent the importance weights of each spatial position. Finally, the spatial attention map is normalized through Sigmoid activation to ensure that the weights are in the range of [0, 1]. After the processing of each module, the hybrid convolution Conv1 is used for initial feature extraction, the hybrid convolution Conv2 further fuses features after the SFBlock module, and the hybrid convolution Conv3 is used to transform the features back into a 3-channel RGB image, and finally the restored image is output. At the same time, the hybrid convolution Conv4 is used as an intermediate output, located after the SFBlock module, to provide stagewise feature extraction results for loss calculation feedback to ensure the stability of the network during training.
[0027] Furthermore, the hybrid convolution is specifically implemented as follows:
[0028] The input features of the hybrid convolution are respectively passed through three convolution operations, namely: pointwise convolution, depthwise separable convolution with hierarchical learning, and wavelet transform convolution. The output of each convolution will be L2-normalized;
[0029] The depthwise separable convolution with hierarchical learning adopts a layer-by-layer learning strategy. Specifically, first, a 1x1 depth convolution is used to independently weight each channel, and the channel characteristics are adjusted separately to enable the channel to obtain adaptive learning ability. Then, pointwise convolution is used to fuse the adjusted features to learn complex inter-channel relationships. The formula for the overall process is as follows:
[0030] X i =w j *X j (j = 1, 2, 3…)
[0031]
[0032] Among them, X j Represented as the input feature of the mixed convolution, X i is the output after each channel is loaded with weights, giving each channel adaptive capabilities, w i and w j is a learnable weight parameter; X d It is the output of the deep separable convolution of hierarchical learning, and Norm represents L2 normalization;
[0033] For some complex features, one-time mapping may still be required. In this case, point-by-point convolution can directly perform weighted learning on all channels, and its weights and biases are learned at one time, so as to more comprehensively integrate features and learn complex inter-channel relationships, further optimize feature learning, and improve feature expression capabilities.
[0034]
[0035] Among them, X p Represents the output of point-by-point convolution;
[0036] Wavelet transform convolution uses wavelet transform to perform convolution operations in wavelet space, obtains a larger receptive field in the spatial domain, can extract richer local features, and avoids a significant increase in the number of parameters, while maintaining a lightweight design while enhancing local features;
[0037] X w =WTConv(X j )
[0038]
[0039] Among them, WTConv represents wavelet convolution, w w is a learnable parameter, X w is the final output of the wavelet transform convolution;
[0040] By dynamically adjusting the weights of these three convolutions, hybrid convolution can effectively adapt to various QR code damage situations and significantly improve the image restoration effect;
[0041] X′=a*X d +b*X p +c*X w
[0042] Among them, X' is the final output, a, b, c are learnable parameters for weighted fusion.
[0043] Furthermore, the spatial-frequency domain enhancement module is specifically implemented as follows:
[0044] In the spatial domain of the spatio-frequency domain enhancement module, the features obtained in the previous stage are applied with residual connection and adaptive average pooling to maintain gradient transmission and obtain global information. After that, the fast Fourier transform is applied to convert the image to the frequency domain to extract frequency features. After frequency domain enhancement, the image is converted back to the spatial domain through the inverse Fourier transform. The formula of the spatio-frequency domain enhancement module is as follows:
[0045] P s = Pooling(X ce + X de )
[0046]
[0047]
[0048]
[0049] A, P = FFT(X s )
[0050] A' = Conv 1×1 (LeakyReLU(Conv 1×1 (A)))
[0051]
[0052] X o = α * X s + β * X f
[0053] Among them, X de represents the output of the multi-scale pyramid module, X ce represents the output of the multi-branch color enhancement module, P s is the global information extracted by the adaptive average pooling Pooling, is the global feature after pooling upsampled to the size of the x ce feature map, C s is the output after the features are fused through a 1x1 convolution Conv 1×1 , X s is the output of performing a residual connection on C s to ensure stable gradients; A and P represent amplitude and phase respectively, FFT is the Fourier transform, A' is the output after the amplitude is strengthened by convolution, and LeakyReLU is used as a non-linear activation function to ensure that the network can learn complex features without completely ignoring negative value information. X f is the output transferred to the spatial domain through the inverse Fourier transform IFFT, X o is the weighted fusion between spatio-frequency domain features, and α and β are weight parameters.
[0054] Furthermore, α and β are 0.9 and 0.1 respectively.
[0055] Furthermore, the linear space module LSA is specifically implemented by the following formula:
[0056] X c = Conv 3×3 (X in ) + X in
[0057] G m = MLLA(Norm(X c ))
[0058] X m = G m + X c
[0059] G n = MLP(Norm(X m ))
[0060] X n = X m + G n
[0061] X r = Reshape(X n )
[0062] X lsa = σ(Conv 1×1 (ReLU(Conv 1×1 (X r )))) * X r
[0063] Among them, X c is the result of initially enhancing the input X in through convolution and using residual connections to retain the original input features. G m is the global context information extracted by MLLA. Norm represents normalization. Conv 3×3 represents a 3x3 convolution. X m is the residual connection between G m and X c . G n is the output that enhances the non-linear expression ability through MLP. X n is the residual connection between X m and G n . X ris the output after tensor reshaping (Reshape). ReLU introduces a non-linear operation for the next convolution, enhancing the feature expression ability and filtering negative value noise, making the subsequent attention weights (Sigmoid) more focused on effective features. σ represents the Sigmoid activation function, and Conv 1×1 represents a 1x1 convolution, and X lsa is the result obtained by strengthening the spatial attention of the feature map.
[0064] Furthermore, the multi-scale pyramid module is specifically implemented as follows:
[0065] F1 = ReLU(Conv 3×3 (X))
[0066] X 101 = AvgPool 128 (X)
[0067] X 102 = AvgPool 64 (X)
[0068] X 103 = AvgPool 32 (X)
[0069] X 1010 = Upsample(ReLU(Conv 1×1 (X 101 )), H, W)
[0070] X 1020 = Upsample(ReLU(Conv 1×1 (X 102 )), H, W)
[0071] X 1030 = Upsample(ReLU(Conv 1×1 (X 103 )), H, W)
[0072] F2 = Concat(X 1010 , X 1020 , X 1030 , F1)
[0073] X de = Tanh(Conv 3×3 (F2))
[0074] where X ∈ R C×H×W is the input feature map, C is the number of channels, H is the height of the input feature map, W is the width of the input feature map, and AvgPool kDenotes the average pooling operation with a kernel size of k. Upsample is enlarged using nearest neighbor interpolation to match the original feature size, i.e., (H,W). Concat represents the concatenation operation in the channel dimension. The role of ReLU is to introduce a non-linear operation. Finally, Tanh is used to control the output range. Conv 3×3 Denotes a 3x3 convolution, Conv 1×1 Denotes a 1x1 convolution, F1, X 101 、X 102 、X 103 、X 1010 、X 1020 、X 1030 、F2 represent the intermediate calculation results respectively, X de Denotes the output of the multi-scale pyramid module.
[0075] Furthermore, the multi-branch color enhancement module is specifically implemented as follows:
[0076] X1,X2,X3,X4 = Chunk(X)
[0077] X′1 = Conv 1×1 (X1)
[0078] X′2 = Conv 1×1 (X2)
[0079] X′3 = Conv 1×1 (X3)
[0080] X″1 = IN(X′1)
[0081] X″2 = IN(X′2)
[0082] X″3 = IN(X′3)
[0083] X″′1 = Conv 1×1 (X″1)
[0084] X″′2 = Conv 1×1 (X″2)
[0085] X″′3 = Conv 1×1 (X″3)
[0086] M = X″′1 + X″′2 + X″′3 + X4
[0087] X ce = Concat(X″′1,X″′2,X″′3,M)
[0088] Where, X ∈ R C×H×Wis the input feature map, which is evenly divided into 4 groups X1, X2, X3, X4 along the channels. X1, X2, and X3 are each upsampled to C / 2 channels through a 1×1 convolution Conv1×1 to obtain X'1, X'2, and X'3. IN represents Instance Normalization. X″1, X″2, and X″3 are the intermediate calculation results. Then, a 1×1 convolution Conv1×1 is used to change the number of channels back to C / 4 channels to obtain X″′1, X″′2, and X″′3. X4 is connected to X″′1, X″′2, and X″′3 with a residual connection to obtain M. Concat performs channel dimension concatenation, and finally the number of channels is restored to the same as the input, C, X ce represents the output of the multi-branch color enhancement module.
[0089] Furthermore, the training objective of the method is divided into a perceptual loss L vgg , a structural similarity loss SSIM loss, and a Charbonnier Loss, which are combined in a weighted manner:
[0090] L total =λ1L vgg1 +λ2L ssim +λ3L c +λ4L vgg2
[0091] where λ1, λ2, λ3, and λ4 are all weight parameters; L vgg1 is the perceptual loss between the network prediction value and the true value, and L vgg2 is the perceptual loss between the intermediate graph output by the hybrid convolution Conv4 and the true value;
[0092] The formula for the structural similarity loss SSIM loss is as follows
[0093]
[0094] L SSIM =1 - SSIM(x, y)
[0095] where μ I and μ K are the means of image I and image K, σ I and σ K are the variances of image I and image K, σ IK is the covariance of image I and image K, and C1 and C2 are constants;
[0096] The formula for the Charbonnier Loss is as follows:
[0097]
[0098] Among them, ∈ is a small constant used to prevent numerical instability caused by the sum of squares approaching zero, N is the total number of samples in a batch, and X i is the predicted value, and Y i is the true value.
[0099] Compared with the prior art, the present invention has the following beneficial effects: The method of the present invention combines a hybrid convolution module, fuses frequency domain information on the basis of traditional spatial feature extraction, enhances the details and contrast of the QR code image, refines the image color, and pays attention to global and local information through an attention mechanism to refine the image details, greatly improving the image quality while ensuring the computational efficiency. The experimental results on the real QR code image dataset show that the PSNR and SSIM metrics corresponding to the present invention are significantly better than several existing mainstream image restoration methods, and the decoding rate and recognition rate metrics are sub-optimal among all comparison methods, and the method of the present invention has the characteristics of being lightweight. Description of the Drawings
[0100] Figure 1 is a degraded QR code. From left to right, and from top to bottom are blur, scratch, dirt, color shift, old, and motion blur in sequence.
[0101] Figure 2 is the network model architecture of the method of the present invention.
[0102] Figure 3 is the QR code dataset establishment process.
[0103] Figure 4 are QR code images generated by different algorithms. Detailed Implementation Manner
[0104] The technical solution of the present invention will be specifically described below with reference to the drawings.
[0105] The present invention provides a lightweight QR code image restoration method based on deep learning, which combines a hybrid convolution module, fuses frequency domain information on the basis of spatial feature extraction, enhances the details and contrast of the QR code image, refines the image color, and pays attention to global and local information through an attention mechanism to refine the image details, greatly improving the image quality while ensuring the computational efficiency, and finally completes the QR code image restoration.
[0106] The following is the specific implementation process of the present invention.
[0107] A lightweight QR code image restoration method based on deep learning in the present invention aims to propose a lightweight QR code recognition algorithm for various QR code degradation situations. First, a degraded dataset is generated in ways such as Real-ESRGAN. After that, the primary goal of the method is to achieve clearer image restoration through end-to-end learning processing, obtain superior visual effects, and minimize the number of computational parameters as much as possible. The overall framework of the method of the present invention is as Figure 2 shown.
[0108] 1. Establish a dataset
[0109] QR code images may be affected by various degradation factors in actual applications, such as blur, dirt, scratches, background particles, and uneven color, etc. To address these issues, in this study, the present invention adopts the QR-DN1.0 dataset
[16] as the target image because this dataset is real QR codes, containing 1,500 training QR code images and 750 test images. To evaluate the performance of the model, the present invention divides the 750 test images into a validation set of 450 images and a test set of 300 images. To further enhance the robustness of the model, the present invention supplements the training set to 3,000 images, the validation set to 800 images, and the test set is adjusted to 1,000 images by generating additional QR code images. When using the QR-DN1.0 dataset and the generated high-quality QR code images to generate degraded images, the present invention uses Real-ESRGAN to simulate the blur effect of real image degradation. Although it is mainly a super-resolution model, in this study, the present invention explores its application in generating blurred images. In addition, the present invention adopts various methods to simulate common degradation situations of QR code images, such as dirt, scratches, uneven color, background particles, and motion blur, etc. These methods include image processing techniques such as adding noise, simulating scratch marks, and adjusting color balance to ensure that the model can perform well in various real scenarios. Specifically, by using Real-ESRGAN technology, the present invention can generate blurred images from high-quality original QR code images; and additionally simulate scratches by generating lines and superimposing them on the images; simulate the dirt effect by superimposing blurred patterns on the images; modify the values of the RGB channels to simulate color offset caused by the influence of light sources; simulate local aging effects by adjusting color values pixel by pixel according to the attenuation mask, and add motion blur to the test set alone to test the generalization performance. Figure 1 These are the degraded images of the 6 kinds of damaged situations. As Figure 3 shown, the process of the present invention for making QR code data. First, the present invention collects a QR code dataset, preprocesses the dataset, then divides it into a training set, a validation set, and a test set. After that, the data is processed by adding noise, and then the rationality of the established dataset is tested using evaluation metrics such as peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), decoding rate, and recognition rate.
[0110] 2 Spatial-Frequency Domain Enhancement Modules
[0111] Since the spatial domain cannot contain all the details and color information. To obtain more feature information, the method of the present invention hopes that the network can perform frequency domain enhancement, which refers to processing the image in both the spatial domain and the frequency domain to improve the overall effect of the image. As Figure 2 shown, first, in the spatial domain of this module, the features obtained in the previous stage are applied with residual connection and adaptive average pooling to maintain gradient transmission and obtain global information. Then, the fast Fourier transform is applied to convert the image to the frequency domain, extract frequency features, and in order to reduce the computational amount, the present invention enhances the amplitude that contains more information and reconstructs the phase. These frequency domain features can help restore the overall contrast of the image and reduce the color deviation and blurring caused by the environment. After frequency domain enhancement, the image is converted back to the spatial domain through inverse Fourier transform for further processing and adjustment. The formula of this module is as follows:
[0112] P s = Pooling(X ce + X de ) (1)
[0113]
[0114]
[0115] X s = X ce + X de + C s (4)
[0116] Wherein, P s is the global information extracted by adaptive average pooling, is to upsample the pooled global features to the size of the x ce feature map, C s is the output after fusing the features through 1x1 convolution, X s is to perform residual connection on C s to ensure the stable output of the gradient. Next, the frequency domain formula is as follows:
[0117] A, P = FFT(X s ) (5)
[0118] A' = Conv 1×1 (LeakyReLU(Conv 1×1 (A))) (6)
[0119] X f = IFFT(A', P) (7)
[0120] X o = α * X s + β * X f (8)
[0121] Wherein, A and P respectively represent amplitude and phase, A' is the output after the amplitude is enhanced by convolution, and LeakyReLU is used as a non-linear activation function to ensure that the network can learn complex features without completely ignoring negative value information. X f is the output transferred to the spatial domain through IFFT, X o is the weighted fusion between the spatial-frequency domain features, and α and β are weight parameters, which are 0.9 and 0.1 respectively.
[0122] 3. Hybrid Convolution
[0123] After the existing lightweight image enhancement algorithm finishes processing the module, it usually uses a single type of convolution operation for fusion, such as pointwise convolution or depthwise separable convolution. However, this method has certain computational limitations when dealing with complex scenarios. To address the above problems, the present invention proposes a hybrid convolution layer, such as Figure 2 shown, which combines convolution operations of multiple kernel types to improve the computational efficiency of the network and simultaneously retain a strong feature learning ability. The present invention combines pointwise convolution, depthwise separable convolution with hierarchical learning, and wavelet convolution
[14] , and utilizes their complementarity in channel fusion, enhanced feature modeling ability, and local feature extraction, thereby significantly improving the feature expression ability. At the same time, compared with directly using large kernel convolution, it reduces the computational cost while ensuring the accuracy, achieving the goal of lightweight. Through this hybrid convolution, the model can reduce the dependence on hardware resources while retaining the ability to learn feature details.
[0124] Specifically, the input of the hybrid convolution module will pass through three convolution operations respectively: pointwise convolution, depthwise separable convolution with hierarchical learning, and wavelet transform convolution. The output of each convolution will be L2-normalized to ensure that their amplitudes are in the same range and prevent the output of a certain convolution from dominating.
[0125] The improved depthwise separable convolution adopts a layer-by-layer learning strategy, so the present invention calls it depthwise separable convolution with hierarchical learning. Specifically, first, a 1x1 depth convolution is used to independently weight each channel, adjusting the channel characteristics separately to enable the channel to obtain adaptive learning ability. Then, pointwise convolution is used to perform channel fusion on the adjusted features to learn complex inter-channel relationships, effectively reducing the burden of pointwise convolution in learning complex mapping relationships and enhancing the feature modeling ability. This hierarchical learning method reduces the learning difficulty while maintaining a low number of parameters. The formula for this process is as follows:
[0126] Xi =w j *X j (j=1,2,3…) (9)
[0127] X d =Norm(∑ i (w i *X i )) (10)
[0128] Among them, X j Represented as the input feature of the mixed convolution, X i is the output after each channel is loaded with weights, giving each channel adaptive capabilities, w i and w j is a learnable weight parameter. d is the output of the convolution. Secondly, for some complex features, one-time mapping may still be required. In this case, point-by-point convolution can directly perform weighted learning on all channels, and its weights and biases are learned at one time, so as to more comprehensively integrate features and learn complex inter-channel relationships, further optimize feature learning, and improve feature expression capabilities.
[0129] X p =Norm(∑ j (w j *X j )) (11)
[0130] Among them, X p represents the output of point-wise convolution. Finally, wavelet convolution
[14] By using wavelet transform and performing convolution operation in wavelet space, a larger receptive field is obtained in the spatial domain, which can extract richer local features, avoid a significant increase in the number of parameters, and maintain a lightweight design while enhancing local features. By dynamically adjusting the weights of these three convolutions, the hybrid convolution module of the present invention can effectively adapt to various QR code damage situations and significantly improve the image restoration effect.
[0131] X w =WTConv(X j ) (12)
[0132] X w =Norm(∑ w (w w *X w )) (13)
[0133] X'=a*X d +b*X p +c*X w (14)
[0134] Among them, WTConv represents wavelet convolution, w w is a learnable parameter, X w is the final output of the convolution. By dynamically adjusting the weights of these three convolutions, X' is the final output, a, b, c are learnable parameters for weighted fusion, and Norm represents L2 normalization.
[0135] 4. Linear Space Module
[0136] The self-attention mechanism has achieved remarkable success in the field of image processing. However, it brings huge computational complexity, which is unacceptable for many tasks.
[15] By introducing a linear transformation, the global context information is captured and an appropriate weight is assigned to each pixel. Compared with the traditional self-attention mechanism, this attention mechanism has lower computational complexity and saves memory and computing resources while maintaining performance. Spatial attention focuses on the perception of local details. This method proposes an LSA module, such as Figure 2 As shown in Figure 1, linear attention and spatial attention are combined to mine effective information at different levels. The LSA module not only focuses on global features, but also enhances the perception of local details, optimizes attention distribution, ensures that the model focuses on key areas, and suppresses background interference. The specific description will be introduced in detail in the backbone network section. The formula of this module is as follows:
[0137] X c =Conv 3×3 (X in )+X in (15)
[0138] G m =MLLA(Norm(X c )) (16)
[0139] X m =G m +X c (17)
[0140] Among them, X c is the input X in Perform preliminary convolution enhancement and use residual connections to retain the original input features, G m is the global context information extracted by MLLA, X m G m With X c The residual connection is made.
[0141] G n =MLP(Norm(X m )) (18)
[0142] Xn = X m + G n (19)
[0143] X r = Reshape(X n )(20)
[0144] X lsa = σ(Conv 1×1 (ReLU(Conv 1×1 (X r )))) * X r (21)
[0145] where G n is the output after enhancing the non - linear expression ability through MLP, X n is the residual connection of X m and G n , X r is the output after tensor reshaping. ReLU introduces non - linear operations for the next convolution, enhancing the feature expression ability and filtering negative - value noise, making the subsequent attention weights (Sigmoid) more focused on effective features. σ represents the Sigmoid activation function, and X lsa is the enhancement of spatial attention to the feature map.
[0146] 5. Other modules
[0147] Figure 2 The multi - scale pyramid module [6] (MPM) in
[0148]
[0149] F1 = ReLU(Conv 3×3 (X)), (22)X 101 = AvgPool 128 (X), (23)
[0150] X 102 = AvgPool 64 (X), (24)
[0151] X 103 = AvgPool 32 (X), (25)
[0152] X 1010 = Upsample(ReLU(Conv 1×1 (X 101 ))), H, W), (26)
[0153] X 1020 = Upsample(ReLU(Conv 1×1 (X 102 ))), H, W), (27)
[0154] X 1030 = Upsample(ReLU(Conv 1×1 (X 103 ))), H, W), (28)
[0155] F2 = Concat(X 1010 , X 1020 , X 1030 , F1), (29)
[0156] X de = Tanh(Conv 3×3 (F2)), (30)
[0157] where X ∈ R C×H×W is the input feature map, C is the number of channels, H is the height of the input feature map, and W is the width of the input feature map. AvgPool k represents the average pooling operation with a kernel size of k. Upsample uses nearest-neighbor interpolation for upscaling to match the original feature size, i.e., (H, W). Concat represents the channel dimension concatenation operation. The role of ReLU is to introduce non-linear operations, and finally Tanh is used to control the output range.
[0158] Figure 2 The multi-branch color enhancement module [6] (MCEM) in
[0159] X1, X2, X3, X4 = Chunk(X), (31)
[0160] X'1 = Conv 1×1 (X1), (32)
[0161] X'2 = Conv 1×1 (X2), (33)
[0162] X'3 = Conv 1×1 (X3), (34)
[0163] X″1 = IN(X′1), (35)
[0164] X″2 = IN(X′2), (36)
[0165] X″3 = IN(X′3), (37)
[0166] X″′1 = Conv 1×1 (X″1), (38)
[0167] X″′2 = Conv 1×1 (X″2), (39)
[0168] X″′3 = Conv 1×1 (X″3), (40)
[0169] M = X″′1 + X″′2 + X″′3 + X4, (41)
[0170] X ce = Concat(X″′1,X″′2,X″′3,M), (42)
[0171] where X ∈ R C×H×W is the input feature map, which is evenly divided into 4 groups along the channels. Among them, Conv1×1 is used to increase the dimension to C / 2 channels, IN represents Instance Normalization, and then Conv1×1 is used to change the number of channels back to C / 4. X4 directly skips the previous changes and plays the role of residual connection. Concat performs channel dimension splicing and finally restores to the same number of channels C as the input.
[0172] 6. Backbone Network
[0173] As Figure 2 shown, the method of the present invention is generally composed of three stages, namely rough restoration, color refinement, and detail refinement. Each stage is targeted at different image degradation problems for targeted enhancement to ensure the final output of high-quality QR code images. First, in the first stage, the present invention uses hybrid convolution to extract initial features from the input image (a 3-channel RGB image) and expands the number of channels to 32, aiming to provide a richer information basis for subsequent tasks. Hybrid convolution can improve the computational efficiency of the network and at the same time retain strong feature learning ability. Then, based on the features extracted by hybrid convolution, the local details of the image are improved, the color of the QR code is corrected, and the contrast is enhanced through the multi-scale pyramid module and the multi-branch color enhancement module. The QR code is roughly restored, laying a foundation for subsequent color refinement.
[0174] In the second stage, the quality of the QR code image is mainly improved by operating on the frequency components through the spatial-frequency domain enhancement module. Especially in the presence of noise, blur, low contrast, color imbalance, etc., by enhancing the details, clarity, and anti-interference ability of the image, the QR code can be more easily scanned and decoded under various conditions, improving the user experience and the efficiency of the system. Specifically, first, the outputs of the color enhancement module and the detail enhancement module are further fused to enhance the feature representation ability. In the initial stage of fusion, the output feature maps are added element-wise to fully integrate the information of different enhancement modules. Then, the SFBlock module designed in the present invention enhances the global and local detail features through the joint operation of the spatial domain and the frequency domain, and optimizes the color performance. In the spatial domain, the SFBlock module first extracts global information through global pooling, and then fuses the local and global information to enhance the model's perception ability of information at different scales. Finally, all these spatial domain features are fused and prepared to be transferred to the frequency domain processing stage. In the frequency domain, the amplitude and phase information of the image is decomposed using the fast Fourier transform (FFT). The amplitude component is convolved and enhanced and reconstructed because the amplitude component is more vulnerable to degradation in the process of degradation. It is necessary to enhance the brightness, contrast, and detail intensity of the restored image to optimize the color information, make the color more natural and balanced, and improve the visual quality and decoding robustness of the image. The phase component is relatively stable, so only reconstruction is performed to reduce the computational amount. Subsequently, the frequency domain features are restored to the spatial domain through the inverse fast Fourier transform (IFFT) and fused with the spatial domain features, making full use of the complementary information in the spatial-frequency domain, thereby enhancing the details and global performance of the QR code image. After passing through the spatial-frequency domain enhancement module, the network further enhances the feature expression ability of the image features through hybrid convolution. Then, the multi-branch color enhancement module is used to further perform color correction and balance optimization on the QR code image, not only improving the color restoration degree, but also effectively eliminating color deviation, ensuring that the image color is more natural and consistent.
[0175] In the third stage, since the QR code image usually has a highly structured pattern, containing fine local information (such as module boundaries, positioning patterns, etc.), and at the same time, global information (such as the overall layout, contrast, etc.) of the QR code needs to be considered to ensure effective decoding. Therefore, the present invention uses the LSA module to not only focus on global features but also be used to focus on local detail features, taking into account both computational efficiency and image quality, so as to better restore complex QR code images. Specifically, the LSA module is used to strengthen the image restoration ability. First, the input features are spatially enhanced through a 3x3 depthwise separable convolution to extract local texture and edge information, and then normalized through a normalization layer to stabilize the feature distribution and accelerate convergence. Then, through the MLLA module
[15] Capture the global dependencies between features to obtain a feature representation that incorporates global context. Subsequently, enhance the information flow through residual connections to further preserve the correlation between low-order and high-order features. Next, the MLP module further learns the globally enhanced features to improve the non-linear expression ability. Meanwhile, to generate the spatial attention map, the feature map is reduced to 1 / 8 of its original size through 1x1 convolution to reduce the computational overhead. Then, enhance the feature expression ability through ReLU activation and generate a single-channel spatial attention map through 1×1 convolution again, representing the importance weights of each spatial position. Finally, the spatial attention map is normalized through Sigmoid activation to ensure that the weights are within the range of [0,1]. The enhanced feature map generated by the MLP module is weighted according to the generated spatial attention map, and the model automatically focuses on the key regions while suppressing the interference of irrelevant regions, thereby significantly improving the detail restoration ability and spatial domain expression ability.
[0176] Meanwhile, after the processing of each module, the hybrid convolution Conv1 is used for initial feature extraction, the hybrid convolution Conv2 further fuses features after the SFBlock, and the hybrid convolution Conv3 is used to convert the features back to a 3-channel RGB image, and finally the restored image is output. At the same time, the hybrid convolution Conv4 is used as an intermediate output, located after the SFBlock, to provide the phased feature extraction results for loss calculation feedback to ensure the stability of the network during the training process.
[0177] 7. Training Objectives
[0178] The lightweight QR code image restoration method based on deep learning proposed by the method of the present invention is an end-to-end network that can be optimized by combining loss terms. The entire training objective can be divided into the perceptual loss L vgg , SSIM loss, Charbonnier Loss, which are combined in a weighted manner:
[0179] L total =λ1L vgg1 +λ2L ssim +λ3L c +λ d L vgg2 (43)
[0180] Among them, λ1, λ2, λ3, λ4 are all weight parameters. In the present invention, λ1 is set to 0.2, λ2 is set to 0.5, λ3 is set to 1, and λ4 is set to 0.2. L vgg1 is the perceptual loss between the network prediction value and the true value, and L vgg2 is the perceptual loss between the intermediate map output by the hybrid convolution Conv4 and the true value;
[0181] The structural similarity loss SSIM loss formula is as follows
[0182]
[0183] L SSIM = 1 - SSIM(x, y)
[0184] where μ I and μ K are the means of image I and image K, σ I and σ K are the variances of image I and image K, σ IK is the covariance of image I and image K, and C1 and C2 are constants;
[0185] The formula for Charbonnier Loss is as follows:
[0186]
[0187] where ∈ is a small constant used to prevent numerical instability when the sum of squares approaches zero, N is the total number of samples in a batch, X i is the predicted value, and Y i is the true value. This loss function can be regarded as a smoothed L1 loss, avoiding the limitations of L1 loss at points of discontinuous gradients.
[0188] 8. Experiments
[0189] To evaluate the effectiveness of the method of the present invention on real QR code data, the present invention conducted experiments using its own made dataset and compared with several recently published algorithms, including FiveA+
[12] , IRNeXt
[10] , LIR[9], DnCnn
[13] , LiteEnhanceNet
[11] . Among them, IRNeXt and LIR are general image restoration methods. Currently, there is no lightweight general image restoration method. Therefore, in order to compare lightweight methods with existing general image restoration methods (such as IRNeXt and LIR) under fair conditions, the present invention reduced the parameters of these two methods. Since the original model parameter quantities of these two methods are huge and have a large advantage in computing resources and inference speed compared with lightweight methods, the present invention took specific measures, such as reducing the number of channels and the number of residual blocks, to reduce the parameter quantities of these models to the same order of magnitude as the method of the present invention for a more fair comparison. Although the reduction of these parameters has a certain impact on the model performance, the purpose of doing so is to more realistically compare the advantages of lightweight methods in terms of performance. The method of the present invention uses PSNR and SSIM as the main evaluation indicators. In addition, the present invention also uses the decoding rate and recognition rate as additional indicators.
[0190] 6.1 Parameter Selection
[0191] During the training process, the number of epochs is set to 1000, and the batch sizes for both training and validation are 1. The Adam optimizer is used, and the initial learning rate is set to 4×10 -4 , and the default values of β1 and β2 are 0.5 and 0.999 respectively. To dynamically adjust the learning rate, the present invention uses a CyclicLR scheduler, and the initial momentum is set to 0.9 and 0.999. The data augmentation strategy includes horizontal flipping and random rotations of 90 degrees, 180 degrees, and 270 degrees.
[0192] 6.2 Qualitative Comparison
[0193] To compare the qualitative performance of the method of the present invention with several other image restoration algorithms, the present invention outputs the predicted QR code images on the test set respectively. Figure 4 The images predicted by the method of the present invention and other methods on the test set are shown. It can be seen from Figure 4 that the method proposed by the present invention can not only restore the clarity of the QR code image, but also significantly improve the decoding rate and recognition rate. Specifically, when dealing with various degradation types (such as blur, noise, dirt, etc.), this method can restore higher-quality QR code images and still has strong generalization ability in the case of degradation not involved in training. Further comparison reveals that the method of the present invention has obvious advantages in detail restoration. The method of the present invention is particularly outstanding in detail restoration, and can more accurately repair the edges and complex structures of the QR code, ensuring the decodability of the QR code and thus improving the overall recognition effect.
[0194] 6.3 Quantitative Comparison
[0195] Table 1 shows the evaluation results on the QR code test set. The present invention compares five algorithms: Five A+, IRNeXt, LIR, DnCnn, and LiteEnhanceNet. Since the general image restoration algorithms IRNeXt and LIR among them are not lightweight algorithms and are difficult to compare, the present invention performs operations such as reducing the number of channels on these two models. For the IRNet network, the basic number of channels is reduced from 32 to 8, the second downsampling operation is removed, and the number of residual blocks is changed from 14 to 1, but its frequency-domain loss function is not changed. For LIR, the number of left modules, right modules, and bottom modules is all changed to 1, and multiple modules are not accumulated, and they are converted into the same order of magnitude of the number of parameters for easy comparison.
[0196] Table 1 Quantitative Comparison on the QR Code Dataset
[0197] Method PSNR SSIM Decoding Rate Recognition Rate Params FiveA+ 18.36 0.7969 50.60% 49.20% 9k IRNeXt 20.83 0.8344 67.50% 66.50% 56.96k LIR 21.12 0.8296 65.60% 64.10% 34.38k DnCnn 21.06 0.8525 59.3% 58.3% 558.34k LiteEnhanceNet 16.18 0.7420 51.1% 49.3% 13.7k Ours (C-16) 22.16 0.8726 61.90% 60.90% 13.65k Ours (C-32) 24.37 0.8915 65.60% 64.40% 49.33K
[0198] In addition, the method proposed by the present invention uses two configurations with different numbers of channels, namely C-16 with 16 basic channels, which further reduces the number of parameters while ensuring a relatively high restoration quality and is suitable for resource-constrained devices. C-32 with 32 basic channels further improves the restoration quality compared to C-16, especially performing better in terms of PSNR and SSIM metrics, while still maintaining a relatively low number of parameters.
[0199] It can be seen from the quantitative comparison that the method proposed by the present invention performs excellently in terms of the quality of restored images, and at the same time has relatively high decoding and recognition rates. In addition, the C-16 version is more lightweight and suitable for deployment on embedded devices, while the C-32 version achieves a higher restoration quality while maintaining a relatively low number of parameters.
[0200] References:
[0201] [1] Dong H, Liu H, Li M, et al. An Algorithm for the Recognition of Motion-Blurred QR Codes Based on Generative Adversarial Networks and Attention Mechanisms[J]. International Journal of Computational Intelligence Systems, 2024, 17(1): 83.
[0202] [2] Wang X, Xie L, Dong C, et al. Real-esrgan: Training real-world blind super-resolution with pure synthetic data[C]. International Conference on Computer Vision Workshops(ICCVW). 2021: 1905 - 1914.
[0203] [3] Van Gennip Y, Athavale P, Gilles J, et al. A regularization approach to blind deblurring and denoising of QR barcodes[J]. IEEE Transactions on Image Processing, 2015, 24(9): 2864 - 2873.
[0204] [4] Pan J, Sun D, Pfister H, et al. Blind image deblurring using darkchannel prior[C]. IEEE conference on Computer Vision and Pattern Recognition. 2016:1628-1636.
[0205] [5] Rioux G, Scarvelis C, Choksi R, et al. Blind deblurring of barcodes via Kullback-Leibler divergence[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019, 43(1):77-88.
[0206] [6] Tiwari S. Blind restoration of motion blurred barcode images using ridgelet transform and radial basis function neural network[J]. Electronic Letters on Computer Vision and Image Analysis, 2014, 13(3).
[0207] [7] Pu H, Fan M, Yang J, et al. Quick response barcode deblurring via doubly convolutional neural network[J]. Multimedia Tools and Applications, 2019, 78:897-912.
[0208] [8] Li J, Zhang D, Zhou M C, et al. A motion blur QR code identification algorithm based on feature extracting and improved adaptive thresholding[J]. Neurocomputing, 2022, 493:351-361.
[0209] [9] Fan D, Yue T, Zhao X, et al. LIR: A Lightweight Baseline for Image Restoration[J]. CoRR, 2024.
[0210]
[10] Cui Y, Ren W, Yang S, et al. Irnext: Rethinking convolutional network design for image restoration[C]. International Conference on Machine Learning. 2023.
[0211]
[11] Zhang S, Zhao S, An D, et al. LiteEnhanceNet: A lightweight network for real-time single underwater image enhancement[J]. Expert Systems with Applications, 2024, 240: 122546.
[0212]
[12] Jiang J, Ye T, Bai J, et al. Five A+Network: You Only Need 9K Parameters for Underwater Image Enhancement[J]. arXiv preprint arXiv:2305.08824, 2023.
[0213]
[13] Zhang K, Zuo W, Chen Y, et al. Beyond a gaussian denoiser: Residual learning of deep cnn for image denoising[J]. IEEE Transactions on Image Processing, 2017, 26(7): 3142 - 3155.
[0214]
[14] Finder S E, Amoyal R, Treister E, et al. Wavelet Convolutions for Large Receptive Fields.arXiv 2024[J]. arXiv preprint arXiv:2407.05848.
[0215]
[15] Han D, Wang Z, Xia Z, et al. Demystify Mamba in Vision: A Linear Attention Perspective[J]. arXiv preprint arXiv:2405.16605, 2024.
[0216]
[16] Monfared M, Koochari A, Monshianmotlagh R. QR-DN1.0: A new distorted and noisy QRs dataset[J]. Data in Brief, 2021, 39: 107605.
[0217] The above are the preferred embodiments of the present invention. All changes made according to the technical solution of the present invention, when the functional effects produced do not exceed the scope of the technical solution of the present invention, shall fall within the protection scope of the present invention.
Claims
1. A lightweight QR code image restoration method based on deep learning, characterized in that, Combined with the hybrid convolution module, it fuses frequency-domain information on the basis of spatial feature extraction, enhances the details and contrast of the QR code image, refines the image color, and pays attention to global and local information through the attention mechanism to refine the image details. While ensuring the computational efficiency, it greatly improves the image quality and finally completes the restoration of the QR code image.
2. The lightweight QR code image restoration method based on deep learning according to claim 1, characterized in that For the establishment of the QR code dataset, the QR-DN1.0 dataset and the method of generating additional QR code images are adopted to construct the training set, test set and validation set. At the same time, degraded images are generated based on the QR-DN1.0 dataset and the additional generated QR code images. When using the QR-DN1.0 dataset and the additional generated QR code images to generate degraded images, Real-ESRGAN is used to generate blurred images; then, by generating lines and superimposing them on the image, scratches are simulated; blurred patterns are superimposed on the image to simulate dirt; the values of the RGB channels are modified to simulate the color shift caused by the influence of the light source. According to the attenuation mask, the color values are adjusted pixel by pixel to simulate local aging, and motion blur is added to the test set alone to test the generalization performance.
3. The lightweight QR code image restoration method based on deep learning according to claim 1, characterized in that The method includes three stages: In the rough restoration stage, the hybrid convolution is used to extract the initial features of the input image, that is, the 3-channel RGB image, and the number of channels is expanded to 32; then, based on the features extracted by the hybrid convolution, the local details of the image are improved, the color of the QR code is corrected, and the contrast is enhanced respectively through the multi-scale pyramid module and the multi-branch color enhancement module, and the QR code is roughly restored. In the color refinement stage, the quality of the QR code image is improved by operating on the frequency components through the spatial-frequency domain enhancement module; specifically, first, the outputs of the multi-scale pyramid module and the multi-branch color enhancement module are fused. In the initial stage of fusion, the output feature maps are added element by element to fully integrate the information of different enhancement modules; then, the SFBlock module is designed to enhance the global and local detail features through the joint operation of the spatial domain and the frequency domain; in the spatial domain, the SFBlock module first extracts the global information through global pooling, and then fuses the local and global information to enhance the model's perception ability of information at different scales; finally, all these spatial domain features are fused and prepared to be passed to the frequency domain processing stage; in the frequency domain, the fast Fourier transform FFT is used to decompose the amplitude and phase information of the image, and the amplitude component is convolved and enhanced and reconstructed; subsequently, the frequency domain features are restored to the spatial domain through the inverse fast Fourier transform IFFT and fused with the spatial domain features; after passing through the spatial-frequency domain enhancement module, the feature expression ability of the image features is further enhanced through the hybrid convolution; then, the multi-branch color enhancement module is used to further perform color correction and balance optimization on the QR code image, and the color of the QR code is refined. In the detailed refinement stage, through the linear space module LSA, it not only focuses on global features but also can be used to focus on local detailed features, taking into account both computational efficiency and image quality. Specifically, the LSA is used to enhance the image restoration ability. First, the input features are spatially enhanced through a 3x3 depthwise separable convolution to extract local texture and edge information, and then normalized through a normalization layer. Next, the global dependencies between features are captured through a linear attention mechanism MLLA similar to Mamba to obtain a feature representation that fuses global context. Subsequently, the information flow is enhanced through residual connections to further retain the correlation between low-order and high-order features. Next, the multi-layer perceptron MLP further learns the globally enhanced features to improve the non-linear expression ability. At the same time, to generate a spatial attention map, the feature map is reduced in dimension to 1 / 8 of the original through a 1x1 convolution. Then, the feature expression ability is enhanced through ReLU activation, and a single-channel spatial attention map is generated again through a 1x1 convolution, representing the importance weights of each spatial position. Finally, the spatial attention map is normalized through Sigmoid activation to ensure that the weights are in the range of [0,1]. After the processing of each module, the hybrid convolution Conv1 is used for initial feature extraction, the hybrid convolution Conv2 further fuses features after the SFBlock module, and the hybrid convolution Conv3 is used to convert the features back to a 3-channel RGB image, and finally the restored image is output. At the same time, the hybrid convolution Conv4 is used as an intermediate output, located after the SFBlock module, to provide the phased feature extraction results for loss calculation feedback to ensure the stability of the network during training.
4. The lightweight QR code image restoration method based on deep learning according to claim 3, characterized in that The hybrid convolution is specifically implemented as follows: The input features of the hybrid convolution are respectively passed through three convolution operations, namely: pointwise convolution, depthwise separable convolution with hierarchical learning, and wavelet transform convolution. The output of each convolution will be L2-normalized. The depthwise separable convolution with hierarchical learning adopts a layer-by-layer learning strategy. Specifically, first, a 1x1 depth convolution is used to independently weight each channel, individually adjusting the channel characteristics to enable the channel to obtain adaptive learning ability, and then the pointwise convolution is used to fuse the adjusted features to learn the complex inter-channel relationships. The formula for the overall process is as follows: X i = w j * X j (j = 1, 2, 3…) Among them, X j represents the input feature of the hybrid convolution, and X i is the output after loading weights for each channel, endowing each channel with adaptive ability, and w i and w j are learnable weight parameters; X d is the output of the depthwise separable convolution for hierarchical learning, and Norm represents L2 normalization; For some complex characteristics, a one-time mapping may still be required. In this case, the pointwise convolution can directly weight and learn all channels, and its weights and biases are learned at one time, so as to more comprehensively integrate features and learn complex inter-channel relationships, further optimizing feature learning and enhancing the feature expression ability. Among them, X p represents the output of pointwise convolution; The wavelet transform convolution uses wavelet transform to perform convolution operations in the wavelet space, obtaining a larger receptive field in the spatial domain, capable of extracting richer local features, while avoiding a significant increase in the number of parameters, maintaining a lightweight design while enhancing local features. X w = WTConv(X j ) Among them, WTConv represents wavelet convolution, and w w is a learnable parameter, and X w is the final output of the wavelet transform convolution; By dynamically adjusting the weights of these three convolutions, the hybrid convolution can effectively adapt to various damaged QR code situations and significantly improve the image restoration effect. X′ = a * X d + b * X p + c * X w Among them, X' is the final output, and a, b, c are learnable parameters for weighted fusion.
5. The lightweight QR code image restoration method based on deep learning according to claim 3, characterized in that The spatio-frequency domain enhancement module is specifically implemented as follows: In the spatial domain of the spatio-frequency domain enhancement module, the features obtained in the previous stage are applied with residual connection and adaptive average pooling to maintain gradient transmission and obtain global information. Then, the fast Fourier transform is applied to convert the image into the frequency domain to extract frequency features. After frequency domain enhancement, the image is converted back to the spatial domain through the inverse Fourier transform. The formula representation of the spatio-frequency domain enhancement module is as follows: P s = Pooling(X ce + X de ) X s = X ce + X de + C s A, P = FFT(X s ) A' = Conv 1×1 (LeakyReLU(Conv 1×1 (A))) X f = IFFT(A', P) X o = α * X s + β * X f Among them, X de represents the output of the multi-scale pyramid module, and X ce represents the output of the multi-branch color enhancement module. P s is the global information extracted by adaptive average pooling. is the upsampled global feature after pooling to the size of the x ce feature map. C s is the output after fusing the features through a 1x1 convolution Conv 1×1 . X s is the output of performing a residual connection on C s to ensure stable gradients. A and P represent amplitude and phase respectively. FFT is the Fourier transform. A' is the output after the amplitude is enhanced by convolution. LeakyReLU is used as the non-linear activation function. X f is the output transferred to the spatial domain through the inverse Fourier transform IFFT. X o is the weighted fusion between the spatial-frequency domain features. α and β are weight parameters.
6. The lightweight QR code image restoration method based on deep learning according to claim 5, characterized in that α and β are 0.9 and 0.1 respectively.
7. The lightweight QR code image restoration method based on deep learning according to claim 3, characterized in that The linear spatial module LSA, the specific implementation formula is expressed as follows: X c = Conv 3×3 (X in ) + X in G m = MLLA(Norm(X c )) X m = G m + X c G n = MLP(Norm(X m )) X n = X m + G n X r = Reshape(X n ) X lsa = σ(Conv 1×1 (ReLU(Conv 1×1 (X r )))) * X r Among them, X c performs preliminary convolutional enhancement on the input X in and uses residual connections to retain the original input features. G m is the global context information extracted by MLLA. Norm represents normalization, and Conv 3×3 represents a 3x3 convolution. X m is the residual connection between G m and X c . G n is the output after enhancing the non-linear expression ability through MLP. X n is the residual connection between X m and G n . X r is the output after tensor reshaping Reshape. ReLU introduces non-linear operations for the next convolution. σ represents the Sigmoid activation function. Conv 1×1 represents a 1x1 convolution. X lsa is the result obtained by strengthening the spatial attention of the feature map.
8. The lightweight QR code image restoration method based on deep learning according to claim 3, characterized in that The multi-scale pyramid module, the specific implementation formula is as follows: F1 = ReLU(Conv 3×3 (X)) X 101 = AvgPool 128 (X) X 102 = AvgPool 64 (X) X 103 = AvgPool 32 (X) X 1010 = Upsample(ReLU(Conv 1×1 (X 101 ))), H, W) X 1020 = Upsample(ReLU(Conv 1×1 (X 102 )), H, W) X 1030 = Upsample(ReLU(Conv 1×1 (X 103 )), H, W) F2 = Concat(X 1010 , X 1020 , X 1030 , F1) X de = Tanh(Conv 3×3 (F2)) where X ∈ R C×H×W is the input feature map, C is the number of channels, H is the height of the input feature map, W is the width of the input feature map, AvgPool k represents the average pooling operation with a kernel size of k, Upsample is enlarged using nearest neighbor interpolation to match the original feature size, i.e., (H, W), Concat represents the channel dimension concatenation operation, ReLU introduces a non-linear operation for the next convolution, and finally Tanh is used to control the output range, Conv 3×3 represents a 3x3 convolution, Conv 1×1 represents a 1x1 convolution, F1, X 101 、X 102 、X 103 、X 1010 、X 1020 、X 1030 、F2 represent the intermediate calculation results respectively, and X de represents the output of the multi-scale pyramid module.
9. The lightweight QR code image restoration method based on deep learning according to claim 3, characterized in that The multi-branch color enhancement module, the specific implementation formula is as follows: X1, X2, X3, X4 = Chunk(X) X′1 = Conv 1×1 (X1) X′2 = Conv 1×1 (X2) X′3 = Conv 1×1 (X3) X″1 = IN(X′1) X″2 = IN(X′2) X″3 = IN(X′3) X″′1 = Conv 1×1 (X″1) X″′2 = Conv 1×1 (X″2) X″′3 = Conv 1×1 (X″3) M = X″′1 + X″′2 + X″′3 + X4 X ce = Concat(X″′1, X″′2, X″′3, M) where X ∈ R C×H×W is the input feature map, which is evenly divided into 4 groups X1, X2, X3, X4 along the channels. X1, X2, and X3 are respectively upsampled to C / 2 channels through 1×1 convolution Conv1×1 to obtain X'1, X'2, and X'3. IN represents Instance Normalization. X″1, X″2, and X″3 are the intermediate calculation results. Then, 1×1 convolution Conv1×1 is used to change the number of channels back to C / 4 channels to obtain X″′1, X″′2, and X″′3. X4 is residual-connected to X″′1, X″′2, and X″′3 to obtain M. Concat performs channel dimension splicing, and finally restores to the same number of channels C as the input, X ce represents the output of the multi-branch color enhancement module.
10. The lightweight QR code image restoration method based on deep learning according to claim 3, wherein, The training objectives of the method are divided into a perceptual loss L vgg , a structural similarity loss SSIM loss, and a Charbonnier loss, and are combined in a weighted manner: L total = λ1L vgg1 + λ2L ssim + λ3L c + λ4L vgg2 where λ1, λ2, λ3, λ4 are all weight parameters; L vgg1 is the perceptual loss between the network prediction value and the true value, and L vgg2 is the perceptual loss between the intermediate graph output by the hybrid convolution Conv4 and the true value; The formula of the structural similarity loss SSIMloss is as follows L SSIM = 1 - SSIM(x, y) where μ I and μ K are the means of images I and K, σ I and σ K are the variances of images I and K, σ IK is the covariance of images I and K, and C1 and C2 are constants; The formula of the Charbonnier Loss is as follows: where ∈ is a small constant to prevent numerical instability when the sum of squares approaches zero, N is the total number of samples in a batch, X i is the predicted value, and Y i is the true value.
Citation Information
Cited By
Underwater image enhancement method and system based on double-domain collaboration
CN121053048A