Two-dimensional code image classification method based on double-domain contrast learning and sample adaptive enhancement

Through the QR code image classification network QRMatch, which is enhanced with dual-domain comparison learning and sample adaptive enhancement, the problem of QR code image classification in the existing technology is solved in complex backgrounds and limited label data, and it achieves higher classification accuracy and robustness, which is especially suitable for QR code image classification tasks with scarce label data.

CN120298800APending Publication Date: 2025-07-11MINJIANG UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510451405.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing QR code image classification methods perform poorly in scenarios with complex backgrounds and limited labeling data, and it is difficult to effectively utilize labelless data. In addition, traditional contrast learning and data enhancement strategies have limitations, and it is difficult to capture global mode and frequency domain features, negative sample selection is unstable, and simple sample enhancement effect is limited.

Method used

QRMatch, a QR code image classification network based on dual-domain contrast learning and sample adaptive enhancement, is adopted to dynamically adjust the enhancement strategy of simple samples through parallel feature domain and frequency domain comparison learning, combining momentum comparison learning and frequency comparison regularization, and improve the quality of feature representation and model robustness.

Benefits of technology

It significantly improves the accuracy and robustness of QR code image classification, especially suitable for scenarios where label data is scarce, and improves the recognition ability and anti-interference performance of the model in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298800A_ABST
    Figure CN120298800A_ABST
Patent Text Reader

Abstract

The invention relates to a two-dimensional code image classification method based on double-domain contrast learning and sample adaptive enhancement, and belongs to the field of semi-supervised image classification. According to the method, a double-domain contrast learning module (Freq-MOCO) is designed. The module carries out comparative learning in parallel in a feature domain and a frequency domain, the feature domain learning mainly focuses on spatial features of images, and the frequency domain learning focuses on spectrum features and global structure information. Besides, a sample adaptive enhancement module (SAA) is also designed, and the learning effect of the samples is further improved by identifying the simple samples and performing more diversified enhancement on the simple samples. The SAA module selects a simple sample by using historical loss information, and applies a richer enhancement strategy to the simple sample, so as to ensure that the sample which cannot effectively promote model learning originally can generate positive influence on model training through diversified enhancement modes. The algorithm provided by the invention has better performance than several recent algorithms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of semi-supervised image classification, and particularly relates to a QR code image classification method based on dual-domain contrast learning and sample adaptive enhancement. Background Art

[0002] Accurately classifying QR code images is crucial for multiple application scenarios such as mobile payment, logistics tracking, and information security. Differences in the shooting environment, printing quality, and damage degree of QR codes are the main factors affecting the recognition and classification effects. Existing QR code image classification methods mainly include traditional feature extraction and machine learning strategies, as well as deep learning methods that have emerged in recent years. Although traditional methods have high computational efficiency, they perform poorly in scenarios with complex backgrounds and limited labeled data. Deep learning methods such as the MobileNet[1] network, although significantly improving the classification accuracy, due to the use of fully supervised training methods, rely heavily on labeled data and it is difficult to achieve large-scale data coverage and rapid iteration in practical applications.

[0003] To solve these problems, semi-supervised learning strategies are gradually attracting attention. By combining a small amount of labeled data and a large amount of unlabeled data, they provide a more robust and scalable solution for QR code image classification. For example, Sohn[2] et al. proposed a semi-supervised learning method combining consistency regularization and pseudo-labels, using weak augmentation and strong augmentation and screening high-confidence predictions, which significantly improved the semi-supervised image classification performance. Li[3] et al. optimized for data distribution imbalance and class long-tail problems, enhancing the model's adaptability to imbalanced data through mutual information maximization, multi-view augmentation, and dynamic adjustment mechanisms. Cai[4] et al. applied Transformer to large-scale semi-supervised visual tasks, introducing consistency regularization and label smoothing in the self-attention module, and multi-stage training significantly reduced the dependence on a large amount of labeled data. Duan[5] et al. proposed a soft label dynamic adjustment method, expanding and shrinking high- and low-confidence samples respectively during training to capture subtle differences and suppress noise, thereby improving the accuracy of fine-grained classification.

[0004] In current QR code semi-supervised classification, although contrast learning can effectively utilize unlabeled data, there are two main limitations: First, it overly relies on spatial feature extraction and is difficult to capture global patterns and frequency domain features; second, randomly selected negative samples may lead to unstable learning processes and affect the model training effect. In addition, traditional data augmentation strategies treat all unlabeled samples equally, resulting in simple samples being easily correctly classified even after augmentation, with their loss approaching zero and limited contribution to model training.

[0005] To overcome the limitations of traditional contrastive learning and data augmentation strategies, researchers have proposed various improvement methods. In terms of contrastive learning, Momentum Contrastive Learning [6] (MOCO) generates "dynamic" negative samples through a momentum encoder, solving the problem of "difficulty in selecting negative samples" and improving training stability, but increasing computational resource consumption; Simsiam [7] adopts a Siamese network architecture to optimize the similarity between enhanced image representations, avoiding the construction of negative samples, but performing poorly in complex tasks and being sensitive to augmentation strategies. In terms of data augmentation, UDA [8] combines unsupervised augmentation and consistency regularization to improve the robustness of the model in scenarios with scarce labels, but excessive augmentation may introduce noise and affect stability; Augmentation Learning [9] automatically optimizes the strategy based on reinforcement learning to enhance data diversity, but is susceptible to augmentation strategies and has problems of overfitting or noise.

[0006] However, these contrastive learning techniques are difficult to effectively capture spatial and frequency domain feature information simultaneously, and the augmentation strategy is difficult to take into account simple samples that contribute less to network learning, limiting their performance in QR code image classification. Therefore, although these innovative methods have improved the model performance to a certain extent, there are still limitations when dealing with diverse and complex QR code images, and other technical means need to be combined to achieve higher classification accuracy and robustness. Summary of the Invention

[0007] The purpose of the present invention is to provide a QR code image classification method based on dual-domain contrastive learning and sample adaptive augmentation, and a new QR code classification network is proposed - a semi-supervised classification network QRMatch for QR codes based on dual-domain contrastive learning and sample adaptive augmentation. This network decomposes the QR code classification task into three stages: a supervised learning stage, an unsupervised learning stage integrating contrastive learning, and a sample adaptive augmentation stage, so as to enhance the model's recognition ability for QR code images and improve the accuracy and robustness of classification.

[0008] To achieve the above purpose, the technical solution of the present invention is: a QR code image classification method based on dual-domain contrastive learning and sample adaptive augmentation, including:

[0009] Design a dual-domain contrastive learning module, which consists of parallel feature domain contrastive learning and frequency domain contrastive learning. The dual-domain contrastive learning module performs contrastive learning in parallel in the feature domain and the frequency domain. Among them, feature domain learning mainly focuses on the spatial features of the image, while frequency domain learning focuses on spectral features and global structure information;

[0010] Design a sample adaptive augmentation module, which selects simple samples that have not effectively contributed during the training process based on historical loss monitoring and threshold determination, and applies a more abundant augmentation strategy to them.

[0011] Furthermore, the method adopts a typical semi-supervised learning framework, which includes three key stages: 1) The supervised learning stage, where the model is trained using labeled samples, the loss is calculated and the network parameters are updated through backpropagation to establish a preliminary feature representation; 2) The unsupervised learning stage combined with contrastive learning, where unlabeled data is utilized by generating high-confidence pseudo-labels, and at the same time, an improved momentum contrastive learning structure is introduced, integrating the feature domain contrastive learning and the newly added frequency domain contrastive learning branch to enhance the quality of the feature representation and the discriminative and generalization abilities of the model; 3) The sample adaptive enhancement stage, which accurately identifies the simple samples that have not effectively contributed during the training process and applies more diverse enhancement strategies to them to ensure that these samples can effectively promote model optimization.

[0012] Furthermore, the dual-domain contrastive learning module is based on the infrastructure of momentum contrastive learning and introduces frequency contrast regularization to form a dual-domain collaborative representation learning paradigm; among them,

[0013] The feature domain branch adopts the momentum contrastive learning framework, which includes a query encoder and a key encoder. The two encoders share the same network structure, but the parameter update methods are different: the query encoder is directly updated through gradient descent, while the key encoder updates its parameters through the momentum update mechanism:

[0014] θ k ←mθ k +(1 - m)θ q

[0015] Among them, θ k and θ q represent the parameters of the key encoder and the query encoder respectively, and m ∈ [0, 1) is the momentum coefficient, which controls the parameter update rate;

[0016] During the training process, the query image x q and the key image x k , x q and x k are different augmented versions of the same image. They respectively generate the normalized feature vectors q and k through their respective encoders. The system maintains a negative sample queue of size K, and the network is optimized through the following contrastive loss function:

[0017]

[0018] Among them, τ is the temperature coefficient, k i is the negative sample feature in the queue, q T represents the transpose of the feature vector q, and exp(·) represents the exponential function;

[0019] In the frequency domain branch, the Frequency Contrast Regularization (FCR) mechanism is introduced to construct a contrastive learning task in the Fourier frequency domain. The FCR mechanism consists of two core components: the frequency domain negative sample pool maintenance strategy and the frequency domain contrast calculation module. First, to ensure the effectiveness of frequency domain contrastive learning, a negative sample pool maintenance mechanism is designed. The model maintains a queue Q that contains the original images of the latest negative samples. This queue operates in parallel with the original feature queue in momentum contrastive learning. Then, the First-In-First-Out (FIFO) strategy is used to maintain the image queue, ensuring that the latest K sample images are always retained:

[0020] Q (t+1) = Tail K (Q (t) ∪ {x (t) )

[0021] where Q (t) represents the set of original images of negative samples stored in the queue at step t, x (t) represents the newly generated negative sample image at step t, ∪ means adding the new negative sample to the end of the queue, and Tail K (·) means only retaining the last K elements. When the queue length exceeds K, the earliest elements are discarded, so that the final size of the queue remains within K. During each forward pass, the model extracts the last K negative sample images from the queue for frequency domain contrastive learning:

[0022] n (t) = Tail K (Q (t) )

[0023] Here, n (t) represents the set of negative sample images used in the forward pass at step t. In this way, the last K elements are obtained from Q (t) to ensure that the frequency domain contrastive learning always uses the latest negative samples;

[0024] Then, frequency domain contrast calculation is performed. For the anchor image x q , the positive sample image x k (in contrastive learning, the query image can be regarded as the anchor image, and the key image can be regarded as the positive sample image), and the image x n selected from the negative sample queue, the two-dimensional fast Fourier transform (2D-FFT) is performed to obtain their representations in the frequency domain:

[0025]

[0026] where represents the Fourier transform, X q , X k , X nrepresent the corresponding output; then, in the frequency domain, calculate the L1 distance d between the anchor point and the positive sample ap , and the L1 distance d between the anchor point and multiple negative samples an :

[0027] d ap = ||X q - X k ||1, d an = |||X q - X n ||1

[0028] Finally, construct the frequency domain contrast loss function to make the distance between the anchor point and the positive sample in the frequency domain smaller than that with the negative sample:

[0029]

[0030] where N is the batch size, M is the number of negative samples selected for each anchor point, and ∈ is a numerical stability constant;

[0031] Ultimately, the model learns a robust representation of the QR code by jointly optimizing the spatial domain contrast loss and the frequency domain contrast loss:

[0032]

[0033] where λ is the weight coefficient for balancing the two loss terms.

[0034] Furthermore, the sample adaptive enhancement module optimizes simple samples through two modules, namely: the sample selection module and the sample enhancement module.

[0035] Furthermore, in the sample selection module, the sample adaptive enhancement module determines whether a sample is a simple sample by recording the historical consistency loss of each sample in each epoch; the historical loss H t (i) of each sample is updated according to the loss l t (i) of the sample in the corresponding epoch at each epoch, and the update formula is:

[0036] H t (i) = (1 - α)H t-1 (i) + α·l t (i)

[0037] where α is the decay factor, and l t (i) is the loss of sample u i at the t-th epoch; the sample adaptive enhancement module calculates the historical loss threshold τ of the sample using the OTSU method s , and then compares the loss l t(i) and τ s If the magnitude of t l s (i) ≤ τ

[0038] Furthermore, the OTSU method calculates the threshold τ s of the historical loss with the following formula:

[0039] τ s = OTSU(H1, H2,..., H N ).

[0040] Furthermore, in the sample enhancement module, the sample adaptive enhancement module applies a more complex enhancement strategy to simple samples; for non-simple samples, i.e., l t (i) ≥ τ s , the original enhancement strategy is continued; for simple samples, a new enhancement strategy A′(u i ) is adopted. By splicing two independent enhanced versions into a new image, the diversity of the image is increased. The formula is:

[0041] A(u i )1 = u i + N(0, σ 2 )

[0042] A(u i )2 = u i ⊙ m

[0043] A′(u i ) = Concat(A(u i )1, A(u i )2)

[0044] where A(u i )1 and A(u i )2 are two independent enhanced versions of the unlabeled sample u i . A(u i )1 represents the Gaussian noise operation, σ 2 is the variance of the Gaussian noise, A(u i )2 represents the random mask operation, m is a mask image, ⊙ represents the Hadamard product, and Concat represents the concatenation operation.

[0045] Furthermore, the method combines multiple loss functions with weights to form the final training objective for optimizing the network. The total loss function formula is as follows:

[0046]

[0047] where represents the loss of the supervised branch, where the cross-entropy loss Cross-Entropy Loss is used; represents the loss of the unsupervised branch, where the consistency loss Consistency Loss is used; represents the contrastive loss; λ1 and λ2 are hyperparameters used to control the weights of different loss terms;

[0048] For multi-class classification problems, the formula for the cross-entropy loss is expressed as:

[0049]

[0050] where C is the number of classes, y i is the value of the label on class i. If the sample belongs to class i, then y i = 1; otherwise, y i = 0, and p i is the predicted probability of the model for class i;

[0051] The consistency loss is expressed as:

[0052]

[0053] where E[·] represents taking the mean, is the data after weak augmentation, is the data after strong augmentation, f(x) is the prediction function of the model, is the loss function, and the cross-entropy loss is used to calculate the difference between the model output and the pseudo-label;

[0054] The contrastive loss is generated by the dual-domain contrastive learning module and is expressed as:

[0055]

[0056] where is the original momentum contrastive learning contrast loss, is the contrastive loss calculated in the frequency domain, aiming to maximize the similarity between the spectral representations of images, and λ is a hyperparameter used to control the relative weights of the two parts of the loss.

[0057] The present invention also provides a QR code image classification system based on dual-domain contrastive learning and sample adaptive augmentation, including a memory, a processor, and computer program instructions stored on the memory and executable by the processor. When the processor runs the computer program instructions, the method steps described above can be implemented.

[0058] The present invention also provides a computer-readable storage medium, on which computer program instructions capable of being run by a processor are stored. When the processor runs the computer program instructions, the method steps described in any of the above can be implemented.

[0059] Compared with the prior art, the present invention has the following beneficial effects: Description of the Drawings

[0060] Figure 1 It is the significant difference of different types of two-dimensional code images in the frequency domain.

[0061] Figure 2 It is the network model architecture of the method of the present invention.

[0062] Figure 3 It is a schematic diagram of the dual-domain contrast learning module. Detailed Embodiments

[0063] Next, in conjunction with the drawings, the technical solutions of the present invention will be specifically described.

[0064] The present invention provides a two-dimensional code image classification method based on dual-domain contrast learning and sample adaptive enhancement, including:

[0065] Design a dual-domain contrast learning module, which consists of parallel feature-domain contrast learning and frequency-domain contrast learning. The dual-domain contrast learning module performs contrast learning in parallel in the feature domain and the frequency domain. Among them, feature-domain learning mainly focuses on the spatial features of the image, while frequency-domain learning focuses on spectral features and global structure information;

[0066] Design a sample adaptive enhancement module. Based on historical loss monitoring and threshold determination, select the simple samples that have not effectively contributed during the training process, and apply richer enhancement strategies to them.

[0067] The following is the specific implementation process of the present invention.

[0068] Figure 1 It is the significant difference of different types of two-dimensional code images in the frequency domain.

[0069] A two-dimensional code image classification method based on dual-domain contrast learning and sample adaptive enhancement of the present invention proposes a new two-dimensional code classification network - a semi-supervised classification network QRMatch for two-dimensional codes based on dual-domain contrast learning and sample adaptive enhancement. This network decomposes the two-dimensional code classification task into three stages: a supervised learning stage, an unsupervised learning stage integrating contrast learning, and a sample adaptive enhancement stage, so as to enhance the model's recognition ability for two-dimensional code images and improve the accuracy and robustness of classification. As Figure 2As shown, the algorithm of the present invention proposes a dual-domain contrastive learning module and a sample adaptive enhancement module. The dual-domain contrastive learning module consists of parallel feature-domain contrastive learning and frequency-domain contrastive learning. This combination can comprehensively extract the differential features of the image in terms of spatial structure and spectral distribution, improving the model's ability to capture the details and overall texture of the QR code. At the same time, this module enhances the model's sensitivity to local subtle changes, optimizes the discrimination ability in different noise and interference environments, further improves the anti-interference performance and robustness, and thus realizes more accurate and efficient QR code classification. The sample adaptive enhancement module dynamically applies stronger data augmentation strategies to unlabeled samples that exhibit lower difficulty based on historical loss monitoring and threshold determination, thereby enhancing the model's adaptability and robustness to diverse samples. The combination of these two innovative modules gives full play to their respective advantages, realizes the efficient collaboration of contrastive learning and adaptive enhancement, and significantly improves the overall performance and application adaptability of QRMatch in the QR code classification task.

[0070] 1. Representation Learning and Data Augmentation

[0071] To enhance the discriminability and generalization ability of feature representations and ensure the full utilization of samples in training, the present invention proposes an innovative semi-supervised classification framework. By introducing advanced representation learning modules and data augmentation modules, this framework significantly improves the network's ability to understand and process QR code images. Specifically, the representation learning module adopts an improved momentum contrastive learning structure and adds a frequency-domain contrastive learning branch, improving the quality of feature representations and enhancing the discriminability and generalization ability of the model. To make more effective use of the easy-to-classify samples during training, the present invention designs a sample adaptive enhancement module that can identify the easy-to-classify samples with less contribution to the model during training, apply more diverse augmentation strategies to them, ensure their effective utilization in training, and dynamically adjust the data augmentation strategy according to the characteristics of the samples. The overall framework not only improves the accuracy and robustness of QR code image classification but is also particularly suitable for QR code classification tasks with scarce labeled data.

[0072] 2. Network Structure Design

[0073] As Figure 2As shown, the present invention adopts a typical semi-supervised learning framework, which includes three key stages: 1) The supervised learning stage, where the model is trained using labeled samples, the loss is calculated and the network parameters are updated by backpropagation to establish a preliminary feature representation; 2) The unsupervised learning stage combined with contrastive learning, where unlabeled data is utilized by generating high-confidence pseudo-labels, and at the same time, an improved momentum contrastive learning structure is introduced, integrating feature domain contrastive learning and a newly added frequency domain contrastive learning branch to improve the quality of feature representation and the discriminative and generalization abilities of the model; 3) The sample adaptive enhancement stage, which accurately identifies the "easy-to-classify samples" during the training process and applies more diverse enhancement strategies to them to ensure that these samples can effectively promote model optimization.

[0074] 3. Dual-domain contrastive learning module

[0075] The dual-domain contrastive learning module proposed by the present invention is a novel representation learning method. As Figure 3 shown, this method conducts contrastive learning in both the feature domain and the frequency domain to obtain a more robust and rich feature representation of the QR code image. This framework is based on the infrastructure of momentum contrastive learning and innovatively introduces frequency contrast regularization to form a dual-domain collaborative representation learning paradigm.

[0076] The feature domain branch adopts a momentum contrastive learning framework, which includes two networks: a query encoder and a key encoder. The two encoders share the same network structure, but the parameter update methods are different: the query encoder is directly updated by gradient descent, while the key encoder updates its parameters through a momentum update mechanism:

[0077] θ k ←mθ k +(1 - m)θ q

[0078] Among them, θ k and θ q represent the parameters of the key encoder and the query encoder respectively, and m ∈ [0, 1) is the momentum coefficient, which controls the parameter update rate.

[0079] During the training process, the query image x q and the key image x k (both are different enhanced versions of the same image) pass through their respective encoders to generate normalized feature vectors q and k. The system maintains a negative sample queue of size K, and optimizes the network through the following contrastive loss function:

[0080]

[0081] Among them, τ is the temperature coefficient, k i is the negative sample feature in the queue, and q Trepresents the transpose of the feature vector q, and exp(·) represents the exponential function.

[0082] In the frequency domain branch, the present invention introduces a Frequency Contrast Regularization (FCR) mechanism to construct a contrastive learning task in the Fourier frequency domain. This mechanism consists of two core components: a frequency domain negative sample pool maintenance strategy and a frequency domain contrast calculation module. First, to ensure the effectiveness of frequency domain contrastive learning, the present invention designs a dedicated negative sample pool maintenance mechanism: the model maintains a queue Q that contains the original images of the latest negative samples. This queue operates in parallel with the original feature queue of MOCO. Then, the present invention adopts a First-In-First-Out (FIFO) strategy to maintain the image queue, ensuring that the latest K sample images are always retained:

[0083] Q (t+1) = Tail K (Q (t) ∪{x (t)})

[0084] where Q (t) represents the set of original images of negative samples stored in the queue at the t-th step, x (t) represents the newly generated negative sample image at the t-th step, ∪ means adding the new negative sample to the end of the queue, and Tail K (·) means only retaining the latest K elements at the end (FIFIO strategy). When the queue length exceeds K, the earliest element is discarded, so that the final size of the queue remains within K. In each forward propagation, the model extracts the latest K negative sample images from the queue for frequency domain contrastive learning:

[0085] n (t) = Tail K (Q (t) )

[0086] Here, n (t) represents the set of negative sample images used in the forward propagation at the t-th step. In this way, the present invention obtains the latest K elements from Q (t) to ensure that frequency domain contrastive learning always uses the latest negative samples.

[0087] Then, the present invention performs frequency domain contrast calculation. The present invention performs a two-dimensional fast Fourier transform (2D-FFT) on the anchor image x q , the positive sample image x k , and the image x n selected from the negative sample queue to obtain their representations in the frequency domain:

[0088]

[0089] where, Denote the Fourier transform. Next, in the frequency domain, the present invention calculates the L1 distance d between the anchor point and the positive sample ap , and the L1 distance d between the anchor point and multiple negative samples an :

[0090] d ap = ||X q - X k ||1, d an = ||X q - X n ||1

[0091] Finally, the present invention constructs a frequency-domain contrast loss function to make the distance between the anchor point and the positive sample in the frequency domain smaller than the distance to the negative sample:

[0092]

[0093] where N is the batch size, M is the number of negative samples selected for each anchor point, and ∈ is a numerical stability constant.

[0094] Ultimately, the model learns a robust representation of the QR code by jointly optimizing the spatial-domain contrast loss and the frequency-domain contrast loss:

[0095]

[0096] where λ is the weight coefficient for balancing the two loss terms.

[0097] This dual-domain contrast learning framework effectively enhances the model's robustness to noise, blur, and geometric distortion by fusing spatial structure and frequency features. Its frequency-domain module effectively distinguishes QR code types by analyzing the unique spectral characteristics of different coding schemes. This multi-perspective feature understanding mechanism supports efficient semi-supervised learning with limited labeled data.

[0098] 4. Sample Adaptive Augmentation Module

[0099] The purpose of the Sample Adaptive Augmentation module (SAA) is to dynamically adjust the sample augmentation strategy to fully utilize the simple samples that did not contribute effectively during the training process. Since simple samples will still be correctly classified by the network with extremely high confidence after augmentation, resulting in their losses being close to zero and unable to promote further learning of the model.

[0100] SAA realizes the optimization of simple samples through two modules: the sample selection module and the sample augmentation module, as Figure 2 shown below. First, in the sample selection module, SAA determines whether a sample is a simple sample by recording the historical consistency loss of each sample in each epoch. The historical loss H of each sample t(i) At each epoch t, it is updated according to the loss l of the sample in this epoch t (i) The update formula is:

[0101] H t (i) = (1 - α)H t-1 (i) + α·l t (i)

[0102] where α is the decay factor, and l t (i) is the loss of sample u i at the t-th epoch. By this method, SAA can smooth the fluctuations of the sample loss and better reflect the long-term impact of the sample on model training. Then, SAA uses the OTSU

[10] method to perform threshold segmentation on the historical losses of the samples, automatically dividing the samples into simple samples and non-simple samples. The OTSU method calculates the threshold τ s :

[0103] τ s = OTSU(H1, H2,..., H N )

[0104] Then, compare the loss l t (i) of the current sample epoch with τ s . If l t (i) ≤ τ s , then the current sample is a simple sample. Next, in the sample enhancement module, SAA applies a more complex enhancement strategy to the simple sample. For non-simple samples (l t (i) ≥ τ s ), the original enhancement strategy is continued; while for simple samples, a new enhancement strategy A′(u i ) is adopted. This strategy increases the diversity of the image by splicing two independent enhanced versions into a new image. The formula is:

[0105] A(u i )1 = u i + N(0, σ 2 )

[0106] A(u i )2 = u i ☉m

[0107] A′(u i ) = Concat(A(u i )1, A(u i )2)

[0108] where A(u i )1 and A(ui )2 is the unlabeled sample u i and two independent enhanced versions, A(u i )1 represents the Gaussian noise operation, σ 2 is the variance of the Gaussian noise. A(u i )2 represents the random masking operation, m is a masked image, ⊙ represents the Hadamard product, and Concat represents the concatenation operation.

[0109] This method can greatly increase the diversity of the enhanced samples, thus prompting the model to better learn from these samples. Finally, SAA ensures the full utilization of simple samples by dynamically adjusting the augmentation strategy for each sample, avoiding their overly low contribution to training, and thus effectively improving the performance of the model.

[0110] 5. Training Objectives

[0111] The present invention combines multiple loss functions with weights to form the final training objective for optimizing the network:

[0112]

[0113] where represents the loss of the supervised branch, and here the cross-entropy loss

[11] (Cross-Entropy Loss) is used; represents the loss of the unsupervised branch, and here the consistency loss

[12] (Consistency Loss) is used; represents the contrastive loss [6]; λ1 and λ2 are hyperparameters used to control the weights of different loss terms.

[0114] The cross-entropy loss is a widely used loss function in deep learning, especially suitable for classification problems. It is used to measure the difference between two probability distributions, usually the difference between the probability distribution of the model output and the true label. For multi-class classification problems, the formula for the cross-entropy loss can be expressed as:

[0115]

[0116] where C is the number of classes, y i is the value of the label on class i. If the sample belongs to class i, then y i = 1, otherwise y i = 0, and p i is the predicted probability of the model for class i.

[0117] The consistency loss is used to measure the consistency of the model's prediction results when facing different transformations or data augmentations of the same input. Its core goal is to encourage the model to output similar prediction results when processing different representations of the same data. By maintaining this consistency, the model can better capture the underlying structure of the data, thereby improving the learning effect on unlabeled data.

[0118] In QRMatch, the calculation of the consistency loss is based on the model's prediction results after weak augmentation and strong augmentation

[13] . First, the model makes predictions on the data after weak augmentation, and when the prediction probability exceeds a set threshold, this prediction is used as a pseudo-label. Then, the model makes predictions on the strongly augmented version of the same data and calculates the prediction difference between the strongly augmented and weakly augmented data. The goal is to minimize this difference so that the model's predictions for the same image under different transformations are consistent. Specifically, the consistency loss in QRMatch can be expressed as:

[0119]

[0120] where, is the data after weak augmentation, is the data after strong augmentation, f(x) is the model's prediction function, is the loss function, and usually the cross-entropy loss is used to calculate the difference between the model output and the pseudo-label.

[0121] The contrastive loss is generated by the dual-domain contrastive learning module and can be expressed as:

[0122]

[0123] where, is the original MOCO contrastive loss, is the contrastive loss calculated in the frequency domain, aiming to maximize the similarity between the spectral representations of images, and λ is a hyperparameter used to control the relative weight of the two parts of the loss.

[0124] 6. Experiments

[0125] The model of the present invention was evaluated on the publicly available QR code image dataset. The QR Image V2

[14] dataset is an image set specifically used to evaluate the performance of QR code detection and decoding, which contains QR code images taken under various different conditions, aiming to test the robustness and performance of QR code detection algorithms. This dataset covers various types of QR code images, including blurred, brightness-changed, damaged QR codes, curved QR codes, perspective distortion, etc.

[0126] 6.1 Parameter Selection

[0127] The present invention resizes the images to 256×256 pixels and divides the dataset proportionally into 60% for training, 20% for validation, and 20% for testing. All images are processed with data augmentation, including random flipping, rotation, cropping, and brightness adjustment. A deep convolutional neural network based on the ResNet-50 architecture is adopted, with the optimizer selected as Adam, the initial learning rate set to 0.03, and the batch size of 16. The experiment is conducted on a machine equipped with a 10-core Intel Xeon CPU, 128GB of memory, and an NVIDIA RTX3090 GPU, using the PyTorch 2.3.1 framework. The loss function is cross-entropy loss, and semi-supervised training is carried out in combination with the pseudo-label generation strategy.

[0128] To evaluate the contribution of each module to the performance of the semi-supervised classification network, the present invention designs a series of ablation experiments. As shown in Table 1, after adding the dual-domain contrastive learning module (Freq-MOCO), the accuracy of the model is improved by 2.61%, and the F1 score is increased by 3.46%; after adding the sample adaptive augmentation module (SAA), the accuracy of the model is improved by 1.74%, and the F1 score is increased by 1.19%; when adding both modules simultaneously, the accuracy of the model is improved by 4.35%, and the F1 score is increased by 4.69%.

[0129] Table 1 Ablation Experiments

[0130]

[0131] 6.2 Quantitative Comparison

[0132] The present invention conducts comprehensive experiments on the publicly available QR code image dataset to evaluate the performance of the method of the present invention compared with other semi-supervised classification methods. The evaluation mainly focuses on the following key metrics: accuracy and F1 score. Table 2 shows the comparison results of the method proposed by the present invention with Mixmatch

[15] , Fixmatch [2], Flexmatch

[16] , Freematch

[17] , Softmatch

[18] , and Soc [5] on the QR Image V2 dataset. The experimental results show that the method of the present invention outperforms the existing methods in these two important metrics, demonstrating its excellent performance in the QR code image classification task.

[0133] Table 2 Results of Each Comparative Method on the QRImageV2 Dataset

[0134]

[0135] The experiments were conducted at three different labeled:unlabeled ratios: 1:9, 3:7, and 5:5. Generally speaking, as the ratio of labeled data increases, the performance of all algorithms improves, but the improvement amplitudes of different algorithms vary. Mixmatch performs mediocrely at the 1:9 ratio, but its performance improves as the label ratio increases. Nevertheless, its performance still lags behind other more advanced methods. Fixmatch performs stably at all ratios, especially at the 3:7 ratio. Although the performance improvement amplitude decreases after further increasing the labeled data. Flexmatch shows strong robustness at all ratios. Compared with Mixmatch and Fixmatch, its performance improvement is relatively stable. Freematch is slightly inferior to other advanced methods at the 1:9 ratio, but its performance improves as the label ratio increases and it performs most prominently at 3:7. However, at 5:5, the performance improvement is limited. Softmatch shows excellent overall classification results at the three label ratios, basically outperforming other comparison methods. Only when the label ratio is 3:7, its accuracy is slightly inferior to Freematch. However, the results at the three ratios still fail to completely surpass the method of the present invention. SoC performs prominently at the 3:7 and 5:5 ratios. Especially at these two ratios, it outperforms Mixmatch and Fixmatch, demonstrating strong adaptability. Finally, the method of the present invention surpasses all other comparison methods at all label ratios, verifying the effectiveness and advantages of the new method proposed by the present invention in semi-supervised learning.

[0136] Generally speaking, the experiments show that increasing the ratio of labeled data helps improve the performance of all algorithms. However, the method of the present invention performs excellently in all tests. Even at lower ratios of labeled data, it still demonstrates strong generalization ability and adaptability, setting a new benchmark in the field of semi-supervised QR code classification.

[0137] The present invention also provides a QR code image classification system based on dual-domain contrast learning and sample adaptive enhancement, including a memory, a processor, and computer program instructions stored on the memory and executable by the processor. When the processor runs the computer program instructions, it can implement the method steps as described in any of the above.

[0138] The present invention also provides a computer-readable storage medium, on which computer program instructions executable by the processor are stored. When the processor runs the computer program instructions, it can implement the method steps as described in any of the above.

[0139] References:

[0140] [1]Howard A G,Zhu M,Chen B,et al.Mobilenets:Efficient convolutional neural networks for mobile vision applications[J].arXiv preprint arXiv:1704.04861,2017.

[0141] [2]Sohn K,Berthelot D,Carlini N,et al.Fixmatch:Simplifying semi-supervised learning with consistency and confidence[J].Advances in Neural Information Processing Systems,2020,33:596-608.

[0142] [3]Li J,Xiong C,Hoi S C H.Comatch:Semi-supervised learning with contrastive graph regularization[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision.2021,9475-9484.

[0143] [4]Cai Z,Ravichandran A,Favaro P,et al.Semi-supervised vision transformers at scale[J].Advances in Neural Information Processing Systems,2022,35:25697-25710.

[0144] [5]Duan Y, Zhao Z, Qi L, et al. Roll with the Punches: Expansion and Shrinkage of Soft Label Selection for Semi-supervised Fine-Grained Learning[C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2024, 38(10): 11829-11837.

[0145] [6]He K, Fan H, Wu Y, et al. Momentum contrast for unsupervised visual representation learning[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2020, 9729-9738.

[0146] [7]Chen X, He K. Exploring simple siamese representation learning[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021, 15750-15758.

[0147] [8]Xie Q, Dai Z, Hovy E, et al. Unsupervised data augmentation for consistency training[J]. Advances in Neural Information Processing Systems, 2020, 33: 6256-6268.

[0148] [9] Frommknecht T, Zipf P A, Fan Q, et al. Augmentation Learning for Semi-Supervised Classification[C] / / DAGM German Conference on Pattern Recognition. Cham: Springer International Publishing, 2022, 85 - 98.

[0149]

[10] Otsu N. A threshold selection method from gray-level histograms[J]. Automatica, 1975, 11(285 - 296): 23 - 27.

[0150]

[11] Tieleman T. Lecture 6.5 - rmsprop: Divide the gradient by a running average of its recent magnitude[J]. COURSERA: Neural Networks for Machine Learning, 2012, 4(2): 26.

[0151]

[12] Sajjadi M, Javanmardi M, Tasdizen T. Regularization with stochastic transformations and perturbations for deep semi-supervised learning[J]. Advances in Neural Information Processing Systems, 2016, 29.

[0152]

[13] Tarvainen A, Valpola H. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results[J]. Advances in Neural Information Processing Systems, 2017, 30.

[0153]

[14] Abeles P. Study of QR Code Scanning Performance in Different Environments. V3[EB / OL]. https: / / boofcv.org / index.php?title=Performance:QrCode, 2019.

[0154]

[15] Berthelot D, Carlini N, Goodfellow I, et al. Mixmatch: A holistic approach to semi-supervised learning[J]. Advances in Neural Information Processing Systems, 2019, 32.

[0155]

[16] Zhang B, Wang Y, Hou W, et al. Flexmatch: Boosting semi-supervised learning with curriculum pseudo labeling[J]. Advances in Neural Information Processing Systems, 2021, 34:18408-18419.

[0156]

[17] Wang Y, Chen H, Heng Q, et al. Freematch: Self-adaptive thresholding for semi-supervised learning[J]. arXiv preprint arXiv:2205.07246, 2022.

[0157]

[18] Chen H, Tao R, Fan Y, et al. Softmatch: Addressing the quantity-quality trade-off in semi-supervised learning[J]. arXiv preprint arXiv:2301.10921, 2023.

[0158] The above are the preferred embodiments of the present invention. All changes made according to the technical solutions of the present invention, as long as the functions and effects produced do not exceed the scope of the technical solutions of the present invention, shall fall within the protection scope of the present invention.

Claims

1. A QR code image classification method based on dual-domain contrastive learning and sample adaptive enhancement, characterized in that, Including: Design a dual-domain contrastive learning module, which consists of parallel feature-domain contrastive learning and frequency-domain contrastive learning. The dual-domain contrastive learning module conducts contrastive learning in parallel in the feature domain and the frequency domain. Among them, feature-domain learning mainly focuses on the spatial features of images, while frequency-domain learning focuses on spectral features and global structural information. Design a sample adaptive augmentation module. Based on historical loss monitoring and threshold determination, select simple samples that have not effectively contributed during the training process and apply richer augmentation strategies to them.

2. The QR code image classification method based on dual-domain contrast learning and sample adaptive enhancement according to claim 1, wherein, The method adopts a typical semi-supervised learning framework, which includes three key stages: 1) The supervised learning stage, uses labeled samples to train the model, calculates the loss and backpropagates to update the network parameters, and establishes a preliminary feature representation. 2) The unsupervised learning stage combined with contrastive learning, utilizes unlabeled data by generating high-confidence pseudo-labels, and at the same time introduces an improved momentum contrastive learning structure. Fuse the feature-domain contrastive learning and the newly added frequency-domain contrastive learning branch. Improve the quality of feature representation and the discriminative and generalization abilities of the model. 3) The sample adaptive augmentation stage, accurately identifies simple samples that have not effectively contributed during the training process, applies richer augmentation strategies to them, and ensures that these samples can effectively promote model optimization.

3. The method for classifying QR code images based on dual-domain contrastive learning and sample adaptive enhancement according to claim 1, wherein The dual-domain contrastive learning module is based on the infrastructure of momentum contrastive learning and introduces frequency contrast regularization to form a dual-domain collaborative representation learning paradigm. Among them, The feature-domain branch adopts the momentum contrastive learning framework, which includes a query encoder and a key encoder. The two encoders share the same network structure, but the parameter update methods are different: the query encoder is directly updated by gradient descent, and the key encoder is updated by the momentum update mechanism: θ k ←mθ k +(1 - m)θ q where, θ k and θ q represent the parameters of the key encoder and the query encoder respectively, m ∈ [0, 1) is the momentum coefficient, which controls the parameter update rate; During the training process, the query image x q and the key image x k , x q and x k are different augmented versions of the same image. Respectively, through their own encoders, they generate the normalized feature vectors q and k. The system maintains a negative sample queue of size K and optimizes the network through the following contrastive loss function: where τ is the temperature coefficient, k i is the negative sample feature in the queue, q T represents the transpose of the feature vector q, and exp(·) represents the exponential function; In the frequency-domain branch, introduce the frequency contrast regularization FCR mechanism to construct a contrastive learning task in the Fourier frequency domain. The frequency contrast regularization FCR mechanism includes two core components: the frequency-domain negative sample pool maintenance strategy and the frequency-domain contrast calculation module. First, to ensure the effectiveness of frequency-domain contrastive learning, design a negative sample pool maintenance mechanism: the model maintains a queue Q containing the original images of the latest negative samples. This queue operates in parallel with the original feature queue of momentum contrastive learning, and then adopts the first-in-first-out (FIFO) strategy to maintain the image queue to ensure that the latest K sample images are always retained. Q (t+1) = Tail K (Q (t) ∪ { (t) ) Among them, Q (t) represents the set of original negative sample images stored in the queue at the t-th step, and x (t) represents the newly generated negative sample image at the t-th step. ∪ means adding the new negative sample to the end of the queue, and Tail K (·) means only keeping the last K elements. When the queue length exceeds K, the earliest elements are discarded so that the final size of the queue remains within K; during each forward propagation, the model extracts the last K negative sample images from the queue for frequency-domain contrastive learning: n (t) = Tail K (Q (t) ) Here, n (t) represents the set of negative sample images used in the t-th forward propagation step, so that the latest K elements are obtained from Q (t) to ensure that the frequency domain contrastive learning always uses the latest negative samples; Then, perform frequency-domain comparison calculations; for the query image, i.e., the anchor image x q , the key image, i.e., the positive sample image x k , and the image x n selected from the negative sample queue, perform the two-dimensional fast Fourier transform (2D-FFT) to obtain their representations in the frequency domain: Among them, represents the Fourier transform, and X q , X k , X n represent the corresponding outputs; then, in the frequency domain, calculate the L1 distance d ap between the anchor point and the positive sample, and the L1 distance d an between the anchor point and multiple negative samples: d ap = ||X q -X k ||1, d an = ||X q -X n ||1 Finally, construct a frequency-domain contrastive loss function to make the distance between the anchor and the positive sample in the frequency domain smaller than the distance to the negative sample: where N is the batch size, M is the number of negative samples selected for each anchor, and ∈ is a numerical stability constant; Finally, the model learns a robust representation of the QR code by jointly optimizing the spatial-domain contrastive loss and the frequency-domain contrastive loss: where λ is the weight coefficient for balancing the two loss terms.

4. The QR code image classification method based on dual-domain contrast learning and sample adaptive enhancement according to claim 1, characterized in that The sample adaptive augmentation module realizes the optimization of simple samples through two modules, namely: the sample selection module and the sample augmentation module.

5. The QR code image classification method based on dual-domain contrast learning and sample adaptive enhancement according to claim 4, characterized in that, In the sample selection module, the sample adaptive enhancement module determines whether a sample is an easy sample by recording the historical consistency loss of each sample in each epoch; the historical loss H of each sample t (i) In each epoch, according to the loss l of the sample in the corresponding epoch t (i) Update, and the update formula is: H t (i) = (1 - α)H t-1 (i) + α·l t (i) where α is the attenuation factor, l t (i) is the loss of sample u i at the t-th epoch; the sample adaptive enhancement module uses the OTSU method to calculate the historical loss threshold τ of the sample s , and then, compare the loss l t (i) of the current sample epoch with τ s . If l t (i) ≤ τ s , the current sample is a simple sample, thus automatically dividing the samples into simple samples and non-simple samples.

6. The method for classifying two-dimensional code images based on dual-domain contrast learning and sample adaptive enhancement according to claim 5, wherein, O Threshold τ for calculating historical losses by the TSU method s The formula is as follows: τ s = OTSU(H1, H2,..., H N ).

7. The method for classifying two-dimensional code images based on dual-domain contrastive learning and sample adaptive enhancement according to claim 5, wherein In the sample enhancement module, the sample adaptive enhancement module applies more complex enhancement strategies to simple samples; for non-simple samples, i.e., l t (i)≥τ s , the original enhancement strategy is continued; for simple samples, a new enhancement strategy A′(u i ) is adopted. By splicing two independent enhanced versions into a new image, the diversity of the image is increased. The formula is: A(u i )1 = u i + N(0, σ 2 ) A(u i )2 = u i ☉m A′(u i ) = Concat(A(u i )1, A(u i )2) Among them, A(u i )1 and A(u i )2 are two independent enhanced versions of the unlabeled sample u i . A(u i )1 represents Gaussian noise operation, and σ 2 is the variance of Gaussian noise. A(u i )2 represents random masking operation, m is a masked image, ⊙ represents Hadamard product, and Concat represents concatenation operation.

8. The method for classifying QR code images based on dual-domain contrastive learning and sample adaptive enhancement according to claim 1, wherein The method combines multiple loss functions with weights to form the final training objective to optimize the network. The formula of the total loss function is as follows: Among them, represents the loss of the supervised branch, and the Cross-Entropy Loss is used here; represents the loss of the unsupervised branch, and the Consistency Loss is used here; represents the contrastive loss; λ1 and λ2 are hyperparameters used to control the weights of different loss terms; For multi-classification problems, the formula of the cross-entropy loss is expressed as: Among them, C is the number of categories, and y i is the value of the label on category i. If the sample belongs to category i, then y i = 1; otherwise, y i = 0, and p i is the predicted probability of the model for category i. The consistency loss is expressed as: where E[·] represents taking the mean, is the data after weak augmentation, is the data after strong augmentation, and f(x) is the prediction function of the model, is the loss function, and cross-entropy loss is used to calculate the difference between the model output and the pseudo-label; The contrastive loss is generated by the dual-domain contrastive learning module and is expressed as: Among them, is the original momentum contrastive learning contrastive loss, is the contrastive loss calculated in the frequency domain, aiming to maximize the similarity between the spectral representations of images. λ is a hyperparameter used to control the relative weights of the two parts of the loss.

9. A QR code image classification system based on dual-domain contrast learning and sample adaptive enhancement, characterized in that, It includes a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, it can implement the method steps described in any one of claims 1-8.

10. A computer-readable storage medium, on which computer program instructions capable of being run by a processor are stored. When the processor runs the computer program instructions, it can implement the method steps described in any one of claims 1-8.

Citation Information

Cited By

  • Motor bearing fault detection system and method based on robust deep learning

    CN121210963A