Printing source identification method of two-dimensional code anti-counterfeit label based on unbalanced sample and related device
By combining a multi-head attention mechanism with a dynamic loss function, the problem of sample class imbalance in QR code anti-counterfeiting label printing source identification is solved, improving the identification accuracy and robustness. It is suitable for QR code anti-counterfeiting label printing source identification under smartphone acquisition conditions.
Patent Information
- Application Number
- CN202510958346.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-10-21
AI Technical Summary
Existing QR code anti-counterfeiting label printing source identification technology suffers from sample class imbalance, leading to decreased recognition accuracy and insufficient model robustness. This is especially true when smartphone image acquisition conditions increase recognition difficulty, and the small differences in QR code textures among similar printer models further limit recognition accuracy and robustness.
A multi-head attention mechanism is adopted to enhance the fusion of multimodal deep features of images and text. Combined with dynamic loss function adjustment, a CLIP-based fine-tuning model is designed. The multi-head attention mechanism captures local and global features in parallel, dynamically adjusts sample weights, alleviates the impact of class imbalance, and improves the recognition ability of minority classes.
It significantly improves the recognition accuracy and system stability for a few types of printers, adapts to complex data collection environments, enhances the accuracy and robustness of QR code anti-counterfeiting label printing source identification, and is compatible with smartphone data collection conditions, possessing strong practicality and promotional value.
Smart Images

Figure CN120823490A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition and anti-counterfeiting technology, and specifically to a multimodal optimization and dynamic loss fusion technology based on deep learning, a method and system for identifying the printing source of a QR code anti-counterfeiting label to solve the problem of imbalanced sample categories. The method and system are particularly suitable for collecting samples through smartphones and realizing high-precision printer model recognition in an environment with severely unbalanced category distribution. Background Art
[0002] With the rapid development of information technology and mobile smart devices, QR codes have become a crucial information carrier for product anti-counterfeiting and traceability. QR codes can store a wealth of product information, allowing consumers to quickly access product details by scanning them with their smartphones. This significantly improves the efficiency of product authenticity verification and effectively promotes market regulation and consumer protection.
[0003] However, QR codes lack physical anti-counterfeiting properties, making them easily copied and imitated by criminals. This has led to the proliferation of counterfeit and shoddy goods, seriously disrupting market order and threatening public safety. According to statistics, counterfeit and shoddy products in my country cause hundreds of billions of yuan in economic losses annually, impacting consumers, businesses, and even government tax revenue, with far-reaching consequences. Against this backdrop, technology that identifies the printed source of QR code anti-counterfeiting labels has become a crucial component in preventing and combating counterfeiting.
[0004] Currently, identifying the source of QR code prints faces numerous challenges, the most prominent of which is a severe imbalance in sample categories. In practice, due to significant differences in the market penetration and frequency of use of different printer models, training data contains abundant samples from the majority category, while samples from the minority category printers are extremely scarce. This imbalance causes traditional recognition models to favor the majority category during training and neglect feature learning for the minority category, severely impacting the model's ability to identify minority category printers and reducing the overall robustness and fairness of the system. Furthermore, minority category printers are often more closely associated with illegal copying activities, and insufficient recognition accuracy can easily lead to the entry of counterfeit products into the market, causing significant losses to consumers and businesses.
[0005] On the other hand, smartphones, as the primary devices for collecting QR code information, are subject to multiple influences, including lighting variations, shooting angle, jitter, and ambient noise. This leads to unstable QR code image quality and increases the difficulty of recognition. Traditional print source recognition technologies are often based on samples collected by high-resolution scanning devices, making them difficult to directly apply to smartphone environments. Furthermore, the subtle differences in texture features between QR codes printed by similar printer models are minimal, making existing technologies poorly suited for complex environments and for similar types of print sources, severely limiting recognition accuracy and robustness.
[0006] In summary, identifying the source of QR code anti-counterfeiting labels urgently requires a high-performance recognition technology solution that can effectively address sample category imbalance, integrate multimodal information, and adapt to smartphone image acquisition conditions. This solution should enhance the perception of subtle features through a multi-head attention mechanism, enrich feature representation through multimodal fusion of images and text, and incorporate dynamic loss function adjustment to mitigate training bias caused by category imbalance, significantly improving the recognition rate of minority categories and the stability of the overall system. This will not only help improve the accuracy and robustness of anti-counterfeiting label printer model identification, but also have important significance for maintaining fair market order, protecting consumer rights, and national economic security. Summary of the Invention
[0007] In response to the sample category imbalance problem in the existing QR code anti-counterfeiting label printing source identification technology and the resulting defects such as decreased recognition accuracy and insufficient model robustness, the present invention proposes an efficient, accurate and robust QR code anti-counterfeiting label printing source identification method and system based on multimodal feature fusion and dynamic loss adjustment, thereby effectively improving the recognition ability of a minority of categories of printers, enhancing the fairness and stability of the system, ensuring anti-counterfeiting effects, maintaining market order, and effectively protecting consumer rights.
[0008] In the first aspect, a method for identifying the printing source of a QR code anti-counterfeiting label based on sample imbalance includes the following steps:
[0009] Data preprocessing: We obtain a dataset of QR code images printed by nine different printers using a mobile phone. We resize the QR code images to 512×512 and split them into 64×64 image blocks. Each printer in the test set contains 1300 image blocks to ensure a balanced distribution of test set categories. The number of image blocks with the most categories in the training set is 9000, and the number of image blocks for the remaining categories of printers is N. i , where i is the category of printers. The training set is set as an unbalanced data set, and the exponential distribution method based on the imbalance parameter p (with values of 10, 20, and 50) is used to construct training samples. The calculation formula of the imbalance parameter p is: The image blocks N of the remaining categories of printers i The calculation formula is:
[0010] Network Model Construction: We designed and built an optimized CLIP fine-tuning model that integrates a multi-head attention mechanism. This model integrates multimodal deep features from images and text, enhancing the model's ability to perceive subtle textures in QR codes and semantic information in labels. The multi-head attention mechanism concurrently captures local and global features across different subspaces, particularly enhancing the model's sensitivity and discriminative power for sample-scarce categories, effectively mitigating the negative impact of data imbalance.
[0011] Dynamic Loss Fusion: An innovative design integrates label smoothing and contrastive learning to dynamically optimize the loss function. This adjusts the weights of minority and majority category samples during training, reducing the model's tendency to overfit to the majority category while improving the ability to distinguish minority category features, thereby enhancing overall recognition accuracy and model robustness.
[0012] Model training: Based on a stratified sampling strategy, a reasonable batch size is set, multiple rounds of iterative training are performed on the training and validation subsets, and the weight ratio of label smoothing and contrastive learning loss is dynamically adjusted to promote stable convergence of the model in an imbalanced sample environment, effectively improving the accurate identification of the source of minority category prints.
[0013] Print source identification: The trained multimodal optimization model is used to identify QR code anti-counterfeiting labels in the test set and output printer model and source results, significantly improving the accuracy of rare printer category identification and the overall performance of the system.
[0014] Secondly, the QR code anti-counterfeiting label printing source identification system based on sample imbalance includes:
[0015] Data processing unit: responsible for the collection, preprocessing and enhancement of QR code image data, using stratified sampling and data expansion strategies to balance the distribution of training data categories and provide high-quality data support for model training and evaluation;
[0016] Model training unit: Based on the CLIP fine-tuning model enhanced by multi-head attention and combined with a dynamic fusion loss function, it continuously optimizes feature expression and classification performance during training, effectively alleviating the problem of sample imbalance;
[0017] Recognition unit: Loads the trained model, identifies the print source of the actual input QR code image, accurately determines the printer model, outputs the results and evaluates the system performance in real time.
[0018] Application unit: Receives the QR code anti-counterfeiting label image from the user, calls the recognition unit to determine the print source, and feeds back the recognition results through the user interface, providing accurate and reliable technical support for practical applications such as anti-counterfeiting authentication;
[0019] The beneficial effects of the present invention are:
[0020] By introducing a multi-head attention mechanism and an optimized CLIP model to achieve multimodal feature fusion of images and text, the model's ability to capture the texture and semantic features of QR codes from minority categories of printers is enhanced. The loss function of dynamically fused label smoothing and contrastive learning is used to significantly alleviate training bias caused by sample imbalance, improving recognition fairness and robustness. This ensures that the model can adapt to real, complex collection environments and category distributions, accurately identifying rare printer models. The system is adaptable to smartphone collection conditions, does not require high-resolution hardware, and has strong practicality and promotional value. While improving the accuracy of identifying the printed source of QR code anti-counterfeiting labels, the overall solution effectively maintains market order, protects consumer rights, and provides strong technical support for the field of anti-counterfeiting and traceability. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a flow chart of the method for identifying the printing source of a QR code anti-counterfeiting label based on unbalanced samples according to the present invention;
[0022] Figure 2 This is the principle diagram of the head attention mechanism;
[0023] Figure 3 Schematic diagram of the optimized CLIP fine-tuning network model structure;
[0024] Figure 4 Detailed explanation of the source identification process for printing QR code anti-counterfeiting labels for imbalanced samples. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0026] Figure 1 A flowchart of a method for identifying the printing source of a QR code anti-counterfeiting label provided by an embodiment of the present invention specifically includes the following steps:
[0027] A QR code anti-counterfeiting label dataset is collected and preprocessed based on imbalanced samples. The present invention proposes a QR code image dataset preprocessing method for the problem of sample category imbalance to ensure the representativeness and diversity of the training data and enhance the model's recognition ability for different printer categories. First, QR code images printed by nine different models of printers are collected, all of which are collected by multiple smartphones, including Huawei Mate40 Pro, Huawei Nova5 Pro, Xiaomi Redmi K30, Apple iPhone X, Apple iPhone 13, etc., to simulate the real consumer usage environment and ensure that the collected images cover rich printing features and acquisition device differences. All QR code images are uniformly adjusted to 512×512 pixels in size to reduce the impact of acquisition resolution differences and improve the stability and accuracy of model training.
[0028] Each 512×512 pixel QR code image was then segmented into 64×64 pixel blocks, which served as the basic sample units for model training and testing. This helped the model discover more local features, reduced computational complexity, and improved training efficiency. Each printer in the test set had 1,300 image blocks, ensuring a balanced and representative distribution of test data categories, enabling accurate evaluation of the model's recognition performance on different printer categories.
[0029] To solve the problem of imbalanced sample categories in the training set, the number of image blocks of the printer with the largest category is set to 9000, and the number of image blocks of the remaining printer categories is N. i , where i represents the printer category number. The number of training samples is dynamically generated using an exponential distribution based on the imbalance parameter p (typical values are 10, 20, and 50). The calculation formula is:
[0030]
[0031] Among them, max i {|C i |} is the number of image blocks with the largest category of samples, min i {|C i |} is the minimum number of image patches in each category, C is the total number of categories, and i = 0, 2, ..., 8. This strategy precisely controls the sample size for each category, simulating the class imbalance found in real applications while avoiding the negative impact of extremely scarce samples on model training.
[0032] Finally, the training set was split into a training subset and a validation subset in a 4:1 ratio. The training subset was used for model parameter optimization, while the validation subset was used to monitor model performance during training and prevent overfitting. Through reasonable partitioning and targeted sampling, the model was able to fully learn the effective features of each category, improving the recognition accuracy of printers in a minority of categories. Furthermore, stable and accurate performance evaluation was achieved during the testing and validation phases, meeting the requirements for identifying the source of QR code anti-counterfeiting labels in complex and unbalanced data environments.
[0033] A pre-trained model based on CLIP was constructed. Based on the CLIP pre-trained model, a fine-tuning architecture incorporating a multi-head attention mechanism was designed. This architecture aims to leverage the rich semantic and visual features learned by CLIP during large-scale image-text alignment pre-training. Combined with the powerful expressive power of the multi-head attention structure, this architecture achieves more precise image and text feature fusion. By jointly encoding image and text, the CLIP model establishes a foundation for cross-modal semantic associations. This allows the model to not only understand the visual information of the QR code image but also enhance its understanding of the category semantics by integrating textual descriptions such as the printer model. By incorporating a multi-head attention mechanism into the CLIP visual encoder, the model can capture fine-grained texture features in the image in parallel across multiple subspaces. This is particularly critical for identifying minority class printers, as samples of these classes are scarce and texture variations are often subtle. The multi-head attention mechanism enables the model to simultaneously focus on both local and global features, effectively extracting more discriminative representations. Combined with textual features, after unified feature projection and normalization, the image and text information are fully integrated. During training, the model can better associate visual signals with semantic concepts, improving recognition and generalization for minority class samples. In summary, this fine-tuning architecture not only enhances the deep expression of features, but also compensates for the information loss caused by data imbalance, significantly improving the overall recognition performance and robustness of the printing source of QR code anti-counterfeiting labels.
[0034] The working principle of the multi-head attention mechanism module and its application in the identification of the printing source of QR code anti-counterfeiting labels based on imbalanced samples. The structure of the multi-head attention mechanism module is shown in the figure below. Figure 2 As shown. For the input feature sequence, it is first mapped into the query matrix Q, key matrix K and value matrix V through three learnable linear transformations:
[0035] Q=XW Q ,K=XW K ,V=XW V
[0036] Among them, the weight matrix W Q , W K , d k is a single-head attention dimension, and satisfies dk =d model / h, where h is the number of attention heads. Then, Q, K, and V are split into multiple groups of subspaces according to the h dimension, and the scaled dot product attention is calculated separately:
[0037]
[0038] The outputs of each head are then concatenated and linearly mapped to obtain the final multi-head attention output:
[0039] MultiHead(Q,K,V)=Concat(head1,…,head h )W O
[0040] in, is the linear mapping weight. This module also includes residual connections and layer normalization to improve the stability and expressiveness of model training:
[0041] Output=LayerNorm(X+MultiHead(Q,K,V))
[0042] The multi-head attention mechanism is introduced in the identification of the printing source of QR code security labels based on imbalanced samples, primarily due to its advantages in significantly improving the ability to capture fine-grained features, enhancing the model's discrimination of minority categories, and improving generalization and robustness. The texture differences between different printing sources in QR code security labels are subtle and complex. The multi-head attention mechanism, by computing multiple attention heads in parallel, independently captures local and global features in different subspaces, effectively mining the unique texture information of minority category printers and avoiding overlooking key details. Furthermore, faced with severe imbalance in training data categories, traditional single-head attention or convolutional methods tend to favor features from the majority category. The multi-head attention mechanism, however, facilitates information fusion from multiple perspectives, balancing and refining feature representations across categories, thereby enhancing the ability to distinguish minority category samples. Furthermore, the multi-head attention module's ability to dynamically adjust the focus of each attention head effectively filters out noise and irrelevant information, improving the model's adaptability and recognition performance for QR code images in diverse and complex acquisition environments (such as different mobile devices and changing lighting conditions), and comprehensively enhancing the stability and reliability of the recognition system.
[0043] The QR code anti-counterfeiting label printing source identification model proposed by the present invention is as follows Figure 3As shown in the figure, based on the CLIP (Contrastive Language-Image Pre-training) pre-training model proposed by OpenAI, combined with the Visual Transformer (ViT) architecture and multi-head attention mechanism, it achieves deep fusion and collaborative learning of image and text information. This structure effectively solves the problem of recognizing fine-grained features and scarce category samples that are difficult to capture with traditional single-modal recognition, significantly improving recognition accuracy and generalization capabilities.
[0044] The overall architecture consists of the following modules: Input Module: A text description input is generated using the template "a photo of xxxprinter" to generate a natural language category description. Image input is a pre-processed QR code anti-counterfeiting label image with a fixed size of 64×64 pixels. Encoder Module: A pre-trained CLIP text encoder converts the natural language input into a semantic vector. A visual encoder: A CLIP visual encoder based on the ViT architecture extracts deep image features using patch embedding and Transformer layers. A multi-head attention module: Multi-head self-attention is performed on the sequential image features output by the visual encoder to improve local and global texture understanding. A feature projection and normalization module: After multi-layer linear transformation, ReLU activation, and dropout processing, the text and image features are mapped to the same feature space and then normalized to ensure direct similarity calculation between the two modal vectors. A fusion and output module: This module calculates the similarity matrix (logits) between the image and text features, combines the predictions from the classification head, and uses a weighted fusion approach to output a comprehensive classification result. The joint loss function (label smoothing classification loss and contrastive learning loss) guides model training to achieve consistent optimization of visual and textual features.
[0045] Detailed explanation of the QR code anti-counterfeiting label printing source identification model proposed in this invention.
[0046] 1. Text input and text encoding
[0047] The printer category of a QR code anti-counterfeiting label is described in natural language, using a text input such as "a photo of {category name} printer." This text is then passed through the CLIP text encoder, which converts the natural language into a high-dimensional semantic vector. CLIP's text encoder uses a Transformer architecture to process token sequences and capture contextual dependencies. The resulting text feature vector has rich semantic expression capabilities and accurately reflects the semantic attributes of the printer category.
[0048] By inputting text descriptions, the model can leverage semantic information to aid in visual judgment, effectively compensating for fuzzy or incomplete image information. In situations of sample imbalance and scarcity, text features provide stable and highly discriminative category cues.
[0049] 2. Image Input and Visual Encoding
[0050] The input QR code image size is uniformly 64×64 pixels. This size has been experimentally verified to preserve the details of the QR code's anti-counterfeiting texture while ensuring efficient model input computation.
[0051] The visual encoder uses the ViT (Vision Transformer) model built into CLIP, replacing traditional convolutional neural networks with its powerful global modeling capabilities. ViT slices the input image into patches of fixed size (for example, 32×32) and linearly embeds each patch into a vector sequence, similar to a sequence of tokens in natural language. Subsequently, through multiple layers of Transformer encoders, the model adaptively learns local and global image features and establishes long-range dependencies between patches.
[0052] The ViT architecture has excellent feature abstraction capabilities, especially when processing images such as QR codes that are rich in texture details but have a regular overall structure, it can highlight key information and suppress interference.
[0053] 3. Multi-head attention module
[0054] In order to further enhance the capture of local texture and contextual feature interaction in the image, the present invention adds a multi-head self-attention mechanism based on the output of the visual encoder.
[0055] The multi-head attention mechanism splits the input features into multiple subspaces and computes attention in parallel, focusing on different regions and feature dimensions of the image. Each attention head independently captures local details and global information. The results from multiple heads are concatenated and merged, then linearly transformed. Finally, residual connections and layer normalization are combined to ensure learning stability.
[0056] This mechanism significantly improves the model's ability to capture fine-grained printing features and enhances the model's recognition ability, especially showing better recognition effects on the unique printing textures of a few categories of printers.
[0057] 4. Feature Projection and Normalization
[0058] The feature vectors generated by the visual and text encoders are mapped to the same dimensional space through a deep multi-layer perceptron (MLP) projection layer, which includes linear layers, ReLU activations, and Dropout layers.
[0059] The projected features are L2 normalized separately so that the image and text feature vectors have a uniform scale and distribution in the same feature space, which facilitates subsequent similarity calculation.
[0060] 5. Multimodal Fusion and Output
[0061] This structure designs a fusion strategy to perform weighted summation of the logits output by the classification network and the logits calculated based on CLIP image-text similarity:
[0062] CombinedLogits=0.8×ClassificationLogits+0.2×SimilarityLogits
[0063] Among them, classification logits are generated by the model classification head based on visual features, representing the discriminative ability obtained by model training; similarity logits are calculated by the dot product of image and text feature vectors and multiplied by the CLIP pre-trained scaling factor, reflecting the multimodal consistency between images and text.
[0064] By fusing the two parts of information, the model comprehensively utilizes the advantages of single-modal classification and multi-modal alignment to enhance the reliability and accuracy of the prediction.
[0065] This architecture is particularly well-suited for the task of identifying the source of printed QR code anti-counterfeiting labels based on imbalanced samples, primarily due to its integration of multimodal information. This allows the model to rely not only on visual features but also incorporate rich textual semantic information, effectively overcoming the recognition challenges associated with scarce sample categories and subtle texture variations. The ViT-based visual encoder simultaneously captures both local details and global structure in QR code images, while the multi-head attention mechanism enhances the model's sensitivity and ability to discern fine-grained printed textures by concurrently focusing on different image regions and feature dimensions. This design significantly improves the model's ability to discriminate minority categories and overall generalization performance. Furthermore, by mapping text descriptions and image features into a unified feature space and fusing them through similarity calculation, the complementary advantages of different modalities are enhanced, further improving recognition accuracy. Compared to traditional convolutional networks and unimodal models, this architecture leverages pre-trained CLIP model weights, offering enhanced cross-modal understanding and transferability. It is better suited to complex acquisition environments and real-world application scenarios with imbalanced samples, demonstrating superior stability and robustness. In summary, this multimodal fusion structure not only breaks through the limitations of traditional visual models and realizes efficient and accurate identification of the source of QR code anti-counterfeiting label printers, but also has good engineering feasibility and promotion potential, and is suitable for practical needs such as industrial-grade anti-counterfeiting traceability and quality control.
[0066] In the task of identifying the source of QR code anti-counterfeiting labels based on imbalanced samples, the traditional cross-entropy loss function (Cross-Entropy Loss) is the most commonly used loss in classification tasks and is defined as
[0067]
[0068] Where C is the total number of categories, y c is the one-hot encoding of the true label (the correct category is 1, the others are 0), p c is the model's predicted probability for class c (obtained through softmax). However, when the class distribution is uneven, the standard cross entropy does not adequately penalize the errors of minority class samples, causing the model to be biased towards the majority class, affecting generalization performance. To alleviate this problem, label smoothing is used, which converts hard labels from "1" and "0" into "soft labels" by redistributing the true label probabilities. Specifically,
[0069]
[0070] Where ε is a smoothing coefficient (such as 0.1). Label smoothing loss is usually implemented by calculating the Kullback-Leibler divergence between the predicted probability and the soft label:
[0071] L LS =KL(Softmax(p)||y LS )
[0072] Compared with traditional cross entropy, label smoothing effectively reduces the model's overconfidence in a single category and improves generalization and robustness in complex and noisy environments.
[0073] In addition, considering the multimodal characteristics (image and text) of the QR code printer source recognition task, a contrastive learning loss is designed to enhance the consistency of cross-modal features. i img With the text feature f j text , and calculate the similarity between the two:
[0074]
[0075] Where τ is the temperature coefficient, which adjusts the smoothness of the similarity distribution. The goal of contrastive learning loss is to maximize the similarity s of the diagonal elements corresponding to the correct image-text pairs. ii , while suppressing the similarity between negative samples, the loss function is defined as:
[0076]
[0077] Finally, the classification label smoothing loss and contrastive learning loss are combined through learnable weights α and β, and the sigmoid function is used to ensure that the weight range is (0, 1), forming a joint loss:
[0078] L=σ(β)·L LS +σ(α)·L contrast
[0079] Where σ(·) is the sigmoid activation function. This design enables the model to dynamically balance the contributions of the two losses and adapt to different needs during the training process. This loss function framework is suitable for the identification of the printed source of QR code anti-counterfeiting labels in an unbalanced sample environment. It is mainly reflected in the following aspects: label smoothing effectively alleviates the extreme confidence of the majority class, improves the learning effect of minority class samples, and enhances generalization ability; contrastive learning loss strengthens the cross-modal feature alignment of images and texts, and improves the model's discrimination power for minority classes; the dynamically adjusted weight mechanism ensures the stability and flexibility of the training process; joint optimization enables the model to take into account both unimodal classification accuracy and multimodal semantic consistency, significantly improving its robustness in complex environments and noise. In summary, from traditional cross entropy to label smoothing, and then integrating cross-modal contrastive learning, this joint loss provides a solid theoretical foundation and practical effect for the identification of the printed source of QR code anti-counterfeiting labels, significantly promoting the improvement of accuracy and stability.
[0080] Figure 4 A detailed explanation of the printing source identification process of QR code anti-counterfeiting labels based on imbalanced samples is shown. In the present invention, an optimized joint loss function is adopted, which combines the label smoothing classification loss and the multimodal contrastive learning loss as a key means to replace the traditional cross entropy loss function. Label smoothing converts hard labels into soft labels by adjusting the probability distribution of the true labels, effectively reducing the model's excessive dependence on majority category samples and alleviating the overfitting problem caused by sample imbalance. The contrastive learning loss uses the similarity constraint between image and text features to shorten the distance between multimodal features of the same category and increase the distance between features of different categories, thereby enhancing the alignment and discrimination of cross-modal information. During the training process, the joint loss enables the model to learn more generalized and discriminative multimodal feature representations, significantly improving the recognition ability of minority categories and unknown samples. The ViT visual encoder built into the CLIP model combines powerful pre-trained weights with the natural language text encoder to further enhance the capture of fine-grained category differences and cross-modal fusion effects through multimodal contrastive learning. At the same time, the present invention adopts The Xeon(R) Silver 4215 CPU, combined with CUDA 10.2, accelerates deep learning training, fully leveraging the advantages of GPU parallel computing and effectively improving training efficiency. In terms of software, the training process is built on the PyTorch 1.10.0 framework, in conjunction with Torchvision 0.11.0 and Torchaudio 0.10.0. An overall optimized loss function design achieves an adaptive balance between classification accuracy and cross-modal consistency. This enables the CLIP fine-tuning model to more stably and accurately extract fused image and text features when handling the complex and unbalanced data environment of QR code anti-counterfeiting label printing source identification, greatly enhancing the model's robustness and generalization capabilities.
[0081] The trained CLIP-based fine-tuning model is used to identify the printing source of the QR code anti-counterfeiting label images in the test set. During recognition, the QR code images in the test set are input into the model, and the printing source category of each image is accurately determined by relying on the multimodal discrimination ability of integrating visual features and text features during the training process. The CLIP model uses a visual encoder (such as ViT-B / 32) to extract fine-grained image semantic features, and combines it with the semantic vector generated by the text encoder to strengthen the cross-modal semantic alignment between different categories through a multimodal contrast learning mechanism. During the model training process, a joint label smoothing classification loss and cross-modal contrast loss are used to effectively alleviate the overfitting and category bias problems caused by the imbalance of sample categories, and significantly improve the recognition accuracy of minority categories and the overall generalization ability.
[0082] During recognition, image and text features are first projected and normalized, and then the similarity is calculated. The weighted fusion of the classification logits output by the classification head is then used to achieve final discrimination, enhancing the robustness and stability of the model. Based on the trained model, feature extraction and category prediction are performed on the test set QR code images, and their accuracy and performance indicators are evaluated. In actual application scenarios, the system can process newly collected QR code anti-counterfeiting label images in real time, calling the fine-tuned CLIP model to complete print source discrimination, providing accurate and reliable technical support for product anti-counterfeiting, helping to prevent counterfeit and shoddy goods and maintain market order.
[0083] The data processing module first collected a dataset of QR code images generated by nine different printers. All original QR code images were uniformly resized to 512×512 pixels, and each image was then divided into multiple 64×64 pixel blocks. In the test set, each printer corresponded to approximately 1,300 image blocks to ensure a balanced distribution of classes in the test set. The training set was designed to be unbalanced, with approximately 9,000 image blocks corresponding to the printer with the most classes.
[0084] During the training phase, a custom optimized combined loss function is employed, combining a label smoothing classification loss with a cross-modal contrastive learning loss. The label smoothing loss probabilistically smoothes the target label and calculates the difference between the predicted probability and the soft label using KL divergence, effectively alleviating the model's over-reliance on a single label in the training data and reducing the risk of overfitting. The cross-modal contrastive loss normalizes image and text features, calculates a similarity matrix between them, and employs a cross-entropy loss to maximize the similarity of correct image-text pairs and minimize the similarity between different categories, promoting a deep fusion of visual and textual information. The loss function adaptively coordinates the contributions of the label smoothing classification loss and the contrastive loss through two learnable weight parameters, achieving joint optimization and ensuring excellent recognition performance even under conditions of imbalanced sample categories. During training, the training and validation losses and accuracy are monitored in real time, and the training strategy and regularization parameters are dynamically adjusted to prevent overfitting.
[0085] After training, the recognition unit uses the model to perform multimodal feature fusion and discrimination on the input QR code image, outputs the print source category, and intuitively displays the anti-counterfeiting recognition results to the user through the application unit. The overall system fully utilizes the advantages of the CLIP model and the synergistic effect of label smoothing and contrast loss to significantly improve the accuracy and robustness of QR code anti-counterfeiting label print source recognition, providing solid technical support for anti-counterfeiting traceability.
[0086] Although the above embodiments have been described in detail, those skilled in the art may, based on their specific needs, make appropriate adjustments, improvements, or equivalent substitutions to the model structure, training strategy, and application process without violating the core concepts and technical principles of the present invention. Such modifications also fall within the scope of protection of the present invention.
Claims
1. A method for identifying the printing source of a QR code anti-counterfeiting label based on multimodal optimization and dynamic loss fusion, characterized in that: The following steps are involved: Data set preprocessing: We obtain a dataset of QR code images printed by nine different printers using a mobile phone. We resize the QR code images to 512×512 and split them into 64×64 image blocks. Each printer in the test set contains 1300 image blocks to ensure a balanced distribution of test set categories. The training set contains 9000 image blocks with the most categories, and N image blocks for the remaining categories of printers. i , where i is the category of printers. The training set is set as an unbalanced data set, and the exponential distribution method based on the imbalance parameter p (with values of 10, 20, and 50) is used to construct training samples. The calculation formula of the imbalance parameter p is: The image blocks N of the remaining categories of printers i The calculation formula is: Build a network model: Build an optimized CLIP fine-tuning model that integrates a multi-head attention mechanism to achieve deep fusion of multimodal features of QR code images and text information, enhancing the model's ability to perceive subtle features of minority categories; Dynamic loss fusion: Design an optimized loss function that dynamically integrates label smoothing and contrastive learning to alleviate model training bias caused by category imbalance, improve minority category recognition accuracy and system robustness; Model training: Using the network model and optimized loss function, setting a reasonable batch size and training rounds, and using stratified sampling to ensure class balance during training to improve the robustness of the model; Print source identification: Use the trained multimodal optimization model to identify the print source of the QR code anti-counterfeiting labels in the test set and output the final identification results.
2. The printing source identification method based on the QR code anti-counterfeiting label according to claim 1 is characterized in that: The multi-head attention mechanism module improves the model's ability to capture the subtle textures and features of QR codes by parallelly calculating multiple sets of attention weights, especially strengthening the feature representation of sample-scarce categories.
3. The method according to claim 1, characterized in that The optimized CLIP fine-tuning model includes a visual encoder and a text encoder, which achieve multimodal fusion by sharing a feature projection layer, thereby improving the performance of discriminating the printing source.
4. The method according to claim 1, wherein The dynamic fusion loss function combines label smoothing technology and contrastive learning to dynamically weight multi-category printer samples, alleviate the model's overfitting of the majority category and improve the ability to distinguish the minority category.
5. The method according to any one of claims 1 to 4, characterized in that Effective data processing and optimization strategies are adopted in the training process to further alleviate the impact of imbalance in training data categories on model performance.
6. A QR code anti-counterfeiting label printing source identification system based on a sample imbalance adaptive mechanism, characterized in that: include: Data processing unit: realizes the acquisition, preprocessing, category imbalance data enhancement and sampling strategy of QR code images; Model training unit: Based on the optimized CLIP fine-tuning model, combined with the multi-head attention mechanism and dynamic fusion loss function, the model is trained using a stratified sampling method to improve the ability to recognize minority categories; Identification unit: loads the trained multimodal fusion model, identifies the printing source of the input QR code anti-counterfeiting label, and outputs the printer category and confidence level; Application unit: Receives the QR code label input by the user, calls the recognition unit to complete the print source identification, and feeds back the results to the user interface to provide accurate anti-counterfeiting services.
7. The system according to claim 6, characterized in that The model training unit dynamically adjusts the weight ratio of label smoothing and contrastive learning loss during the training phase to adapt to the distribution of samples of different categories, ensuring that the model takes into account both accuracy and fairness.
8. The system according to any one of claims 6 to 7, characterized in that: The multi-head attention mechanism module is integrated into the key layer of the visual encoder, focusing on the fine-grained features of the input QR code image from multiple angles, and enhancing the local texture discrimination ability in an imbalanced sample environment.
Citation Information
Cited By
Open set product surface fine defect detection method and device based on machine touch sense
CN121074022A
Method and device for detecting subtle defects on open set product surfaces based on machine haptics
CN121074022B