Two-dimensional code image restoration method based on encoder-decoder architecture
The QR code image restoration method using an encoder-decoder architecture with FEB and VMFB modules addresses the limitations of existing methods by improving noise modeling and geometric structure recovery, enhancing the precision and decoding success rate of QR codes in complex damage scenarios.
Patent Information
- Application Number
- CN202510357756.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-15
AI Technical Summary
The existing QR code image repair methods have limited effects in severe damage, making it difficult to accurately restore the geometric structure and key information of the QR code. The existing deep learning methods lack targetedness in QR code image recovery, making it difficult to adapt to the specific noise mode and occlusion of the QR code.
Using an encoder-decoder architecture method, the feature extraction module FEB and the variance modulation fusion module VMFB are introduced, combining the simple gating mechanism SG and the dual-domain strip attention mechanism DSAM, and through multi-scale feature extraction and global variance information fusion, the image recovery accuracy and decoding success rate are improved.
It significantly improves the recovery accuracy and decoding success rate of QR code images, and can effectively handle complex QR code corruption scenarios, ensuring the accurate recovery of the geometric structure of the image and key information.
Smart Images

Figure CN120318121A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image restoration, and particularly relates to a method for repairing QR code images based on an encoder-decoder architecture. Background Art
[0002] The QR Code is a highly efficient information storage and transmission technology widely used in fields such as mobile payment, logistics tracking, identity authentication, industrial automation, and the intelligent Internet of Things (IoT). Due to its high storage density, fault tolerance, and easy scannability, QR Codes have become popular in daily life and industrial scenarios. However, in actual applications, QR Codes often fail to decode due to factors such as smudging, occlusion, blurring, noise interference, and uneven illumination, affecting the user experience and system stability.
[0003] Regarding the repair of QR code images, traditional image processing techniques such as noise removal, edge detection, and image interpolation have been widely studied and applied. These methods aim to improve the recognition rate of QR Codes by enhancing image contrast, removing noise, and restoring blurred areas. For example, median filtering and mean filtering are commonly used to remove noise in QR code images [1], while Gaussian filtering is used to smooth the image and reduce background interference [2]. However, the repair effects of these methods are limited, especially when the damage degree of the QR code image is relatively severe, and accurate QR code decoding still cannot be guaranteed [3].
[0004] With the rapid development of deep learning technology, the proposed image restoration networks have achieved remarkable results. Their powerful capabilities in image repair have shown broad application potential in QR code image repair tasks. DnCNN [4] adopts residual learning and a convolutional neural network architecture, focusing on image denoising. By removing noise and restoring details, it significantly improves the quality of the image. DRUNet [5] combines the encoding-decoding structure of U-Net with deep residual learning, enhancing the ability to retain details in image denoising and also being applicable to various image repair tasks. DSANet [6] adopts a dual-domain strip attention mechanism, combining spatial and frequency strip attention units to perform feature aggregation and modulation in the spatial and frequency domains respectively. In this way, DSANet can efficiently capture multi-scale context information and improve the accuracy and effect of image restoration. FocalNet [7] emphasizes the key regions in the image through multi-scale feature extraction and a focusing mechanism, reducing background interference and helping to improve the detail restoration of the image. SFNet [8] strengthens the learning of the geometric shape of the image through a shape attention mechanism, and can better repair geometric deformations or occlusions in QR code images.
[0005] Although existing image inpainting methods have made some progress in improving the quality and readability of QR code images, there are still multiple technical problems that limit their application in complex QR code damage scenarios.
[0006] Traditional image inpainting methods (such as filtering for denoising, edge enhancement, interpolation for completion, and frequency domain transformation) perform well in repairing slightly damaged QR code images, but are limited in severely damaged cases. These methods have weak recovery capabilities for local defects, blurring, or occlusion of QR codes. For example, median filtering and Gaussian filtering are mainly used for noise reduction but cannot effectively repair structural damage, resulting in the loss of key information. In addition, filtering and smoothing may cause the edges of QR codes to become blurred, affecting the decoding rate. Interpolation methods can fill in missing areas, but often rely on surrounding pixels for filling and cannot recover real data, leading to decoding errors.
[0007] Deep learning image inpainting methods (such as DnCNN, DRUNet, DSANet, FocalNet, SFNet) perform well in tasks such as denoising, deblurring, and dehazing. However, these methods are mainly designed for natural images and may have certain limitations in QR code image restoration. Existing methods pay more attention to texture information and content reconstruction, while QR code restoration requires precise geometric structures. The lack of specialized modeling easily leads to edge blurring or misalignment, affecting decoding. In addition, the noise patterns of QR codes are different from common noises and are affected by printing, scanning, occlusion, etc. Existing methods are difficult to adapt to these characteristics, resulting in artifacts or distortions. At the same time, due to the lack of training datasets specifically for the problems of QR code images, these methods are difficult to be optimized specifically, and the restoration effect is limited. Summary of the Invention
[0008] The purpose of the present invention is to provide a QR code image inpainting method based on an encoder-decoder architecture.
[0009] To achieve the above purpose, the technical solution of the present invention is: a QR code image inpainting method based on an encoder-decoder architecture, which introduces two core modules, namely a feature extraction module FEB and a variance modulation fusion module VMFB, to improve the restoration accuracy, decoding success rate, and recognition rate of QR code images.
[0010] In an embodiment of the present invention, FEB realizes the efficient extraction of multi-scale features and detailed information of QR code images by integrating a simple gating mechanism SG and a dual-domain strip attention mechanism DSAM, while enhancing the noise modeling ability to solve the problem of feature loss in complex degradation scenarios.
[0011] In an embodiment of the present invention, the specific implementation of FEB is as follows:
[0012] First, the FEB is designed based on residual blocks to enhance the ability to model noise and damaged features in QR code images;
[0013] Secondly, the dual-domain strip attention mechanism DSAM is introduced. By extracting context information in the row and column directions in the spatial domain and performing frequency separation and modulation in the frequency domain, the features in the QR code image can be expressed more comprehensively;
[0014] At the same time, the LayerNorm operation is added to the FEB to effectively smooth the optimization process and accelerate convergence;
[0015] Furthermore, a simple gating mechanism SG is adopted in the FEB to simplify the network structure and reduce the channel dimension; SG first divides the input tensor into two branches T1 and T2, and the two branches are multiplied element-wise to strengthen the flow of effective information. The calculation formula is as follows:
[0016] T1, T2 = Split(T)
[0017]
[0018] where Split represents the channel division operation, represents element-wise multiplication.
[0019] In an embodiment of the present invention, the VMFB enhances the ability to model non-local information and optimizes the image restoration effect by extracting global variance information and adaptively fusing it with the decoder features.
[0020] In an embodiment of the present invention, the VMFB is specifically implemented as follows:
[0021] For encoder features and decoder features at the same scale, the VMFB is processed through a dual-branch structure; in the encoder branch, first, E is downsampled and 1×1 convolution is performed to extract low-resolution non-local structural features
[0022] E n = Conv 1×1 (D(X))
[0023] where D(·) represents adaptive max pooling with a scaling factor of 4; next, the channel dimension variance statistical information of E is extracted and integrated with E n to generate the modulated feature E l :
[0024]
[0025] El = Conv 1×1 (E n + σ 2 (E))
[0026] Wherein, is the channel - dimension variance of E, M is the total number of pixels of the feature map, e k and m respectively represent the pixel value and the channel mean;
[0027] To fuse and modulate the feature with the original feature, after performing an up - sampling operation on E l and passing through a non - linear activation, modulation is achieved by element - wise multiplication with the original feature:
[0028] E e = E ⊙ U(GELU(E l ))
[0029] Wherein, GELU(·) represents the GELU activation function, U(·) represents nearest - neighbor up - sampling; finally, the decoder feature D is convolved by 1×1 and then fused with the modulated feature E e to generate the final aggregated feature C ED :
[0030] C ED = E e + Conv 1×1 (D).
[0031] In an embodiment of the present invention, the method uses the QR - DN1.0 dataset, and by simulating the degradation problems that the QR - code images may encounter in reality, a representative degraded dataset is constructed. The degraded dataset includes different types of QR - code image degradation situations, including blur, dirt, scratches, background particles, and color unevenness, ensuring that the model can handle the complex damage situations of QR - code images in practical applications.
[0032] In an embodiment of the present invention, the method adopts an encoder - decoder architecture, combines multiple feature extraction modules FEB and a variance modulation fusion module VMFB; during the training process, by artificially simulating the degradation effect of QR - code images, it is used as the network input First, 3×3 convolution is used to extract shallow features, and their dimensions are converted to C×H×W, where H×W represents the spatial dimension and C represents the number of channels; subsequently, the shallow features are processed through 4 layers of symmetric encoder-decoder structures and mapped to a higher dimension; in each layer of the encoder-decoder, FEB is used to extract the key features of image degradation; a dual-domain strip attention mechanism DSAM is introduced in FEB, and combined with a simple gating mechanism SG to extract the degradation features of the QR code image; subsequently, through VMFB, the features output by each layer of the encoder and decoder are fused to extract global variance information, enhancing the expression ability of features and the modeling ability of high-dimensional noise features; strided convolution and transposed convolution operations are adopted in the corresponding layers of the encoder and decoder to achieve downsampling and upsampling of features.
[0033] In an embodiment of the present invention, there is no bias in the overall network design.
[0034] In an embodiment of the present invention, the method designs the network to use the L1 loss function as the optimization objective during the training process. The L1 loss function calculates the per-pixel difference between the predicted image and the real image, and can effectively measure the accuracy of image restoration; specifically, for each pair of input images I input and its corresponding real image I target , the L1 loss function evaluates the difference between the restored result and the target image by calculating its absolute error:
[0035]
[0036] where N is the total number of pixels in the image, and I input (i) and I target (i) represent the values of the restored image and the real image at the i-th pixel respectively; the L1 loss function sums and averages the absolute values of the errors, enabling the network to optimize subtle pixel differences, thereby promoting the quality and readability of the restored image.
[0037] By minimizing the L1 loss, the network can gradually adjust its parameters to optimize the accuracy of QR code image restoration, ensuring that the structure and content of the QR code image can be accurately restored under different damage conditions.
[0038] The present invention also provides a computer-readable storage medium, on which computer program instructions that can be run by a processor are stored. When the processor runs the computer program instructions, the method steps as described in any of the above can be implemented.
[0039] Compared with the prior art, the present invention has the following beneficial effects: A method for restoring QR code images based on an encoder-decoder architecture, the innovation of which lies in introducing two core modules, namely a feature extraction module (FEB) and a variance modulation fusion module (VMFB), significantly improving the restoration accuracy, decoding success rate, and recognition rate of QR code images. Specifically, the FEB module integrates a simple gating mechanism (SG) and a dual-domain strip attention mechanism (DSAM) to achieve efficient extraction of multi-scale features and detailed information of QR code images, while enhancing the noise modeling ability and effectively solving the problem of feature loss in complex degradation scenarios. The VMFB module extracts global variance information and adaptively fuses it with the decoder features, significantly improving the modeling ability of non-local information and further optimizing the image restoration effect. The present invention has broad application prospects in the field of QR code image restoration, especially suitable for scenarios of restoring QR code images with high noise, low quality, or partial damage. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is the network model architecture of the method of the present invention.
[0041] Figure 2 It is the evaluation of the simulated QR code degraded image.
[0042] Figure 3 It is the evaluation of the real QR code degraded image. DETAILED IMPLEMENTATION MANNER
[0043] The technical solutions of the present invention will be specifically described below with reference to the accompanying drawings.
[0044] The present invention provides a method for restoring QR code images based on an encoder-decoder architecture, introducing two core modules, namely a feature extraction module FEB and a variance modulation fusion module VMFB, to improve the restoration accuracy, decoding success rate, and recognition rate of QR code images.
[0045] The following is the specific implementation process of the present invention.
[0046] 1 QR code image restoration dataset
[0047] To effectively train and evaluate the QR code image restoration model, the present invention uses the QR-DN1.0 [9] dataset and constructs a representative degraded dataset by simulating the degradation problems that QR code images may encounter in reality. This dataset includes different types of QR code image degradation situations, such as blur, dirt, scratches, background particles, and color unevenness, ensuring that the model can handle the complex damage situations of QR code images in actual applications. The dataset contains 3,000 training images, 800 validation images, and 1,000 test images, covering a variety of actual QR code styles, as the target images for training and testing.
[0048] 2 Overall architecture
[0049] As Figure 1 shown, this algorithm adopts an encoder-decoder architecture, combined with multiple feature extraction modules (FEB) and a variance modulation fusion module (VMFB). During the training process, the present invention artificially simulates the degradation effect of the two-dimensional code image and uses it as the network input First, 3×3 convolution is used to extract shallow features, and its size is converted to C×H×W, where H×W represents the spatial dimension and C represents the number of channels. Subsequently, the shallow features are processed through 4 layers of symmetric encoder-decoder structures and mapped to a higher dimension. In each layer of the encoder-decoder, the FEB is used to extract the key features of image degradation. The present invention introduces a dual-domain strip attention mechanism [6] in the FEB and combines it with a simple gating mechanism
[10] to more effectively extract the degradation features of the two-dimensional code image, thereby enhancing the network's modeling ability for these features. Subsequently, through the VMFB, the features output by each layer of the encoder and decoder are fused to extract global variance information and enhance the expression ability of the features and the modeling ability of high-dimensional noise features.
[0050] To better restore the spatial resolution of the two-dimensional code image and reduce information loss, the present invention adopts strided convolution and transposed convolution operations in the corresponding layers of the encoder and decoder to achieve downsampling and upsampling of features. Finally, to improve the adaptability of the model to images with different degrees of damage, the entire network is designed without bias to ensure robustness under different damage conditions.
[0051] 3 Feature extraction module
[0052] Inspired by the basic module of the efficient image restoration network, the present invention is based on the residual block in the design of the feature extraction module (FEB) to enhance the modeling ability for noise and damaged features in the two-dimensional code image. As Figure 1 shown, the FEB combines multiple components and improves the restoration effect of the two-dimensional code image under different damage conditions through multi-dimensional feature extraction.
[0053] Although traditional residual structures can effectively alleviate the problem of gradient disappearance in deep networks, in QR code image restoration, the simple stacking of convolutional layers in traditional residual blocks may not be able to fully capture the noise features with spatial and geometric correlations in QR code images. In addition, the noise patterns of QR code images often have unique distribution characteristics in the spatial and frequency domains. Relying solely on spatial domain modeling may lead to information loss. Therefore, the present invention introduces a dual-domain strip attention mechanism (DSAM). This mechanism extracts context information in the row and column directions in the spatial domain and performs frequency separation and modulation in the frequency domain, so as to more comprehensively represent the features in QR code images. To further improve the training stability and feature representation ability of the network, the present invention adds a layer normalization (LayerNorm) operation to the FEB, thereby effectively smoothing the optimization process and accelerating convergence.
[0054] Inspired by NAFNet
[10] , the present invention does not use traditional non-linear activation functions in the FEB, but adopts a simple gating mechanism (SG) to simplify the network structure and reduce the channel dimension. It can be regarded as a simplified version of the efficient non-linear activation function gated linear unit (GLU), but does not reduce the expression ability of the network. SG first divides the input tensor into two branches T1 and T2, and the two branches are multiplied element by element to strengthen the effective information flow. The calculation formula is as follows:
[0055] T1, T2 = Split(T)
[0056]
[0057] where Split represents the channel division operation, represents element-by-element multiplication. Based on the residual structure, combining the simple gating mechanism and the dual-domain strip attention mechanism significantly improves the modeling ability of QR code image noise and damaged features while retaining the advantages of residual connections.
[0058] 4 Variance Modulation Fusion Block (VMFB)
[0059] In image inpainting tasks, non-local information can capture the relationships between similar regions in the image, thus helping to more accurately restore damaged or missing image information. Variance information, as a global statistical feature, can effectively reflect the overall change law of image content, and in QR code images, it helps to capture the overall geometric structure and noise distribution characteristics of the image. Based on this, the present invention designs a variance modulation fusion block (VMFB), which modulates features by extracting the global variance information of encoder features and adaptively fuses them with decoder features, thereby enhancing the network's modeling ability for non-local information in QR code images.
[0060] Specifically, for encoder features at the same scale and decoder features VMFB is processed through a dual-branch structure. In the encoder branch, the present invention first downsamples E and performs 1×1 convolution to extract low-resolution non-local structural features
[0061] E n = Conv 1×1 (D(X))
[0062] where D(·) represents adaptive max pooling with a scaling factor of 4. Next, the present invention extracts the channel dimension variance statistical information of E and integrates it with E n to generate the modulated feature E l :
[0063]
[0064] E l = Conv 1×1 (E n + σ 2 (E))
[0065] where is the channel dimension variance of E, M is the total number of pixels in the feature map, e k and m represent the pixel value and the channel mean respectively.
[0066] To fuse the modulated feature and the original feature, the present invention performs an upsampling operation on E l and, after passing through a non-linear activation, multiplies it element-wise with the original feature to achieve modulation:
[0067] E e = E ☉ U(GELU(E l ))
[0068] where GELU(·) represents the GELU activation function and U(·) represents nearest neighbor upsampling. Finally, the decoder feature D is convolved with a 1×1 convolution and then fused with the modulated feature E e to generate the final aggregated feature C ED :
[0069] C ED = E e + Conv 1×1 (D)
[0070] 5 Training Objectives
[0071] The proposed algorithm is an end-to-end network. The present invention uses the L1 loss function as the optimization objective during the training process. The L1 loss function calculates the per-pixel difference between the predicted image and the ground truth image, and can effectively measure the accuracy of image restoration. Specifically, for each pair of input images I input and its corresponding ground truth image I target , the L1 loss function evaluates the difference between the restoration result and the target image by calculating its absolute error:
[0072]
[0073] where N is the total number of pixels in the image, and I input (i) and I target (i) represent the values of the restored image and the ground truth image at the i-th pixel, respectively. By summing and averaging the absolute values of the errors, the L1 loss function enables the network to optimize the subtle pixel differences, thereby promoting the quality and readability of the restored image.
[0074] By minimizing the L1 loss, the network can gradually adjust its parameters to optimize the accuracy of QR code image restoration, ensuring that the structure and content of the QR code image can be accurately restored under different damage conditions.
[0075] 6 Experiments
[0076] To evaluate the performance and adaptability of the algorithm of the present invention, motion artifacts not involved in the training set were additionally added to the test images, and they were compared with several representative image restoration or denoising algorithms, including DnCNN, DRUNet, DSANet, FocalNet, and SFNet. Since DRUNet requires a noise level map as input, in order to ensure its applicability to QR code image restoration, the present invention removed the noise level map. To quantify the restoration effects of each algorithm, the present invention used multiple evaluation metrics, including PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity Index), decoding rate, and recognition rate. PSNR and SSIM are used to measure the quality and structural similarity of the restored image, the decoding rate reflects the readability of the restored QR code image, and the recognition rate measures the recognition accuracy of the image in practical applications.
[0077] 6.1 Quantitative Comparison
[0078] Table 1 shows the evaluation results on the QR-DN1.0 simulated degraded image test set. DRUnet and DnCNN are mainly deep learning algorithms for image denoising, but their modeling capabilities are limited when dealing with diverse image degradation problems, so they perform worse than the algorithm of the present invention in various metrics. FocalNet and SFNet perform excellently in image detail and structure restoration, achieving relatively high PSNR and SSIM values. However, when emphasizing detail restoration, these two methods fail to fully reconstruct the precise geometric structure of the QR code, resulting in slight errors or information loss during the decoding process of the QR code image. In contrast, the method proposed by the present invention achieves the best performance in terms of decoding rate and recognition rate, while the performance of PSNR and SSIM is comparable to that of SFNet.
[0079] Table 1 Quantitative comparison of QR code restoration algorithms on the QR-DN10 dataset
[0080] Network architecture PSNR SSIM Decoding rate Recognition rate DnCnn 21.06 0.8525 59.3 58.3 DRUNet 25.96 0.9024 76.5 75 DSANet 28.21 0.9050 76.8 76.4 FocalNet 28.59 0.9110 76.1 75.5 SFNet 30.44 0.9165 76.9 76.0 Proposed 28.40 0.9140 77.0 76.5
[0081] 6.2 Qualitative comparison
[0082] 6.2.1 Simulated images
[0083] Figure 2 The comparison of the restoration effects of each algorithm on the simulated degraded QR code images is shown. From the visualization results, the method of the present invention shows significant advantages in QR code image restoration compared with FocalNet and SFNet. It can not only accurately reconstruct the geometric structure and edge features of the QR code, but also effectively suppress noise interference, making the restored image have a high decoding accuracy. However, DnCNN generates some artifacts during the restoration process of the second image, affecting the structural integrity and information readability of the QR code; DnCNN and DRUNet show excessive smoothing in the edge area when processing the last test image, resulting in blurred boundaries between the QR code positioning points and the data area, reducing the geometric structure fidelity and decoding success rate; DSANet shows overall blurring and detail loss in the restoration of the third image, affecting the clarity of the QR code and the accurate recognition of key information points, reducing the usability of the QR code.
[0084] 6.2.2 Real images
[0085] To verify the generalization performance of the algorithm in actual application scenarios, the present invention conducts restoration tests on multiple groups of real degraded QR code images. As Figure 3As shown, due to differences in imaging devices, environmental conditions, and degrees of damage, these test samples exhibit different degradation characteristics, including blur, noise, and scratches. The experimental results show that the method of the present invention can, to a certain extent, handle various degradation situations, restore the original structure of the QR code, and ensure its decodability. Specifically, in the first group of tests, different degrees of noise residue or blur occurred in other comparison algorithms, while this algorithm maintained good clarity; in the second and third groups of images, obvious noise problems occurred in DSANet and FocalNet, which may be due to the insufficient ability of their network structures to model specific types of noise; in the fourth group of samples with scratches, except for DRUNet and the algorithm of the present invention, other methods failed to effectively remove the scratch marks, affecting the recognition effect of the QR code.
[0086] The present invention also provides a computer-readable storage medium, on which computer program instructions that can be run by a processor are stored. When the processor runs the computer program instructions, the method steps as described in any of the above can be implemented.
[0087] References:
[0088] [1] Appiah O, Martey E M, Quayson E, et al. Quick Median Filtering Algorithm for Denoising QR Code Images[J]. Ghana Journal of Technology, 2024, 8(1): 11 - 20.
[0089] [2] Wu Q, He Y, Luo Y. Research on QR code image processing on the LED screen[C] / / International Conference on Optics and Image Processing (ICOIP 2021). SPIE, 2021, 11915: 154 - 160.
[0090] [3] Zheng J, Zhao R, Lin Z, et al. EHFP - GAN: Edge - enhanced hierarchical feature pyramid network for damaged QR code reconstruction[J]. Mathematics, 2023, 11(20): 4349.
[0091] [4] Zhang K, Zuo W, Chen Y, et al. Beyond a gaussian denoiser: Residual learning of deep cnn for image denoising[J]. IEEE transactions on image processing, 2017, 26(7): 3142 - 3155.
[0092] [5] Zhang K, Li Y, Zuo W, et al. Plug-and-play image restoration with deep denoiser prior[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, 44(10): 6360 - 6376.
[0093] [6] Cui Y, Knoll A. Dual-domain strip attention for image restoration[J]. Neural Networks, 2024, 171: 429 - 439.
[0094] [7] Cui Y, Tao Y, Bing Z, et al. Selective frequency network for image restoration[C] / / The Eleventh International Conference on Learning Representations. 2023.
[0095] [8] Cui Y, Ren W, Cao X, et al. Focal network for image restoration[C] / / Proceedings of the IEEE / CVF international conference on computer vision. 2023: 13001 - 13011.
[0096] [9] Monfared M, Koochari A, Monshianmotlagh R. QR-DN1.0: A new distorted and noisy QRs dataset[J]. Data in Brief, 2021, 39: 107605.
[0097]
[10] Chen L,Chu X,Zhang X,et al.Simple baselines for image restoration[C] / / European conference on computer vision.Cham:Springer Nature Switzerland,2022:17-33.。
[0098] The above are the preferred embodiments of the present invention. All changes made according to the technical solution of the present invention, when the functions and effects produced do not exceed the scope of the technical solution of the present invention, shall fall within the protection scope of the present invention.
Claims
1. A QR code image restoration method based on an encoder-decoder architecture, characterized in that, Two core modules, namely the Feature Extraction Block (FEB) and the Variance Modulation Fusion Block (VMFB), are introduced to improve the recovery accuracy, decoding success rate, and recognition rate of QR code images.
2. The method for repairing a two-dimensional code image based on an encoder-decoder architecture according to claim 1, wherein The FEB integrates the Simple Gating Mechanism (SG) and the Dual-Domain Strip Attention Mechanism (DSAM) to efficiently extract multi-scale features and detailed information of QR code images, while enhancing the noise modeling ability to solve the problem of feature loss in complex degradation scenarios.
3. The method for repairing a two-dimensional code image based on an encoder-decoder architecture according to claim 1 or 2, characterized in that The specific implementation of the FEB is as follows: First, the FEB is designed based on residual blocks to enhance the modeling ability of noise and damaged features in QR code images; Second, the Dual-Domain Strip Attention Mechanism (DSAM) is introduced. By extracting context information in the row and column directions in the spatial domain and performing frequency separation and modulation in the frequency domain, the features in QR code images can be more comprehensively expressed; At the same time, Layer Normalization (LayerNorm) operations are added to the FEB to effectively smooth the optimization process and accelerate convergence; Furthermore, a simple gating mechanism SG is adopted in FEB to simplify the network structure and reduce the channel dimension; SG first divides the input tensor into two branches T1 and T2, and the two branches are multiplied element by element to strengthen the effective information flow. The calculation formula is as follows: T1, T2 = Split(T) Among them, Split represents the channel division operation, represents element-wise multiplication.
4. The method for repairing a two-dimensional code image based on an encoder-decoder architecture according to claim 1, wherein The VMFB extracts global variance information and adaptively fuses it with decoder features to enhance the modeling ability of non-local information and optimize the image recovery effect.
5. The method for repairing a two-dimensional code image based on an encoder-decoder architecture according to claim 1 or 4, characterized in that, The specific implementation of the VMFB is as follows: For encoder features of the same scale and decoder features The VMFB is processed through a dual-branch structure; in the encoder branch, first, E is downsampled and 1×1 convolution is performed to extract low-resolution non-local structural features E n = Conv 1×1 (D(X)) Among them, D(·) represents adaptive max pooling with a scaling factor of 4; next, extract the variance statistical information of the channel dimension of E and integrate it with E n to generate the modulated feature E l : E l = Conv 1×1 (E n + σ 2 (E)) Among them, is the channel dimension variance of E, M is the total number of pixels of the feature map, e k and m represent the pixel value and the channel mean respectively; To fuse the modulation features and the original features, for E l perform an upsampling operation and, after passing through a non-linear activation, perform an element-wise multiplication with the original features to achieve modulation: E e = E⊙U(GELU(E l )) Among them, GELU(·) represents the GELU activation function, and U(·) represents nearest neighbor upsampling; finally, the decoder feature D is convolved with a 1×1 kernel and fused with the modulation feature E e to generate the final aggregated feature C ED : C ED = E e + Conv 1×1 (D).
6. The method for repairing a two-dimensional code image based on an encoder-decoder architecture according to claim 1, wherein This method uses the QR-DN1.0 dataset and constructs a representative degraded dataset by simulating the degradation problems that QR code images may encounter in reality. The degraded dataset includes different types of QR code image degradation situations, including blur, dirt, scratches, background particles, and color unevenness, to ensure that the model can handle the complex damage situations of QR code images in practical applications.
7. The method for repairing a two-dimensional code image based on an encoder-decoder architecture according to claim 1, wherein The method adopts an encoder-decoder architecture, combines multiple feature extraction modules FEB and variance modulation fusion module VMFB; during the training process, by artificially simulating the degradation effect of the QR code image, it is used as the network input. First, 3×3 convolution is used to extract shallow features and convert their size to C×H×W, where H×W represents the spatial dimension and C represents the number of channels; subsequently, the shallow features are processed through 4 layers of symmetric encoder-decoder structures and mapped to a higher dimension; in each layer of the encoder-decoder, FEB is used to extract the key features of image degradation; a dual-domain strip attention mechanism DSAM is introduced in FEB, and combined with a simple gating mechanism SG, the degradation features of the QR code image are extracted; subsequently, through VMFB, the features output by each layer of the encoder and decoder are fused to extract global variance information, enhancing the expression ability of features and the modeling ability of high-dimensional noise features; strided convolution and transposed convolution operations are adopted in the corresponding layers of the encoder and decoder to achieve downsampling and upsampling of features.
8. The method for repairing a two-dimensional code image based on an encoder-decoder architecture according to claim 7, wherein The entire network design has no bias.
9. The method for repairing a two-dimensional code image based on an encoder-decoder architecture according to claim 1, wherein The method designs the network to use the L1 loss function as the optimization objective during the training process. The L1 loss function calculates the per-pixel difference between the predicted image and the real image, and can effectively measure the accuracy of image restoration. Specifically, for each pair of input images I input and its corresponding real image I target , the L1 loss function evaluates the difference between the restoration result and the target image by calculating its absolute error: where N is the total number of pixels in the image, I input (i) and I target (i) represent the values of the restored image and the ground truth image at the i-th pixel, respectively; the L1 loss function sums and averages the absolute values of the errors, enabling the network to optimize for subtle pixel differences, thereby promoting the quality and readability of the restored image. By minimizing the L1 loss, the network can gradually adjust its parameters to optimize the accuracy of QR code image recovery, ensuring that the structure and content of QR code images can still be accurately recovered under different damage conditions.
10. A computer-readable storage medium storing computer program instructions executable by a processor. When the processor runs the computer program instructions, the method steps described in any one of claims 1-9 can be implemented.