A method, device, and electronic equipment for cigarette pack reflection recognition based on a contrastive learning model.
By using a contrastive learning model to distinguish between real and reflected images of cigarette packs, the problem of recognition error in existing technologies has been solved, enabling accurate cigarette pack quantity statistics and stable inventory management.
Patent Information
- Application Number
- CN202411473484.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-22
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-10-22
AI Technical Summary
Existing technology cannot effectively distinguish between the real image and the reflection image of a cigarette pack, leading to errors in cigarette pack quantity counting and affecting inventory management and merchant operational efficiency.
A contrastive learning model-based approach is adopted, which uses MSDA-YOLO, convolutional neural network and feature pyramid network to obtain the image to be identified from the whole image of the cigarette pack. The distance between the cigarette pack image and the reflection image is determined by the pre-trained contrastive learning model. The image type is distinguished according to the distance threshold and the reflection image is ignored.
Accurate identification of cigarette pack reflection images eliminates interference from quantity statistics, improves the accuracy of inventory management and system stability, reduces computing resource consumption, and increases processing speed and work efficiency.
Smart Images

Figure CN119964131B_ABST
Abstract
Description
[0001] This application relates to the field of image recognition technology, and in particular to a method, apparatus and electronic device for recognizing cigarette pack reflections based on a contrastive learning model. Background Technology
[0002] In smart retail and vending machine scenarios, customer interaction with products is becoming increasingly frequent, making precise merchandise management and inventory control crucial for improving customer satisfaction and operational efficiency. However, when items like cigarette packs are meticulously displayed in retail environments equipped with reflective surfaces (such as glass counters, stainless steel display racks, or display platforms made of any smooth material), technological challenges arise. System cameras, acting as the eyes of these smart retail systems, are responsible for capturing and identifying products on the shelves; however, under such specific conditions, they may indiscriminately record all images, including the actual cigarette packs and their mirror reflections.
[0003] Traditional cigarette pack recognition technology, while capable of identifying and counting items on shelves to some extent, often struggles to distinguish between real cigarette pack images and reflections created by light refraction or reflection. This technological bottleneck leads to a common problem: the system misidentifies reflections of cigarette packs as independent, real packs, resulting in discrepancies in subsequent inventory statistics and management.
[0004] Such errors in quantity statistics are extremely detrimental to merchants. Not only do they prevent merchants from accurately grasping the real-time inventory status of cigarette packs, but they can also trigger a series of chain reactions, such as over-replenishment or stockouts due to inaccurate inventory information, thereby affecting the customer shopping experience and brand trust. Furthermore, erroneous inventory data can mislead merchants' sales strategies and market forecasts, negatively impacting the company's operational efficiency and profitability in the long run. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this application provides a method, apparatus, and electronic device for recognizing cigarette pack reflections based on a contrastive learning model. The aim is to solve the technical problem in existing technologies where the inability to accurately distinguish between the reflection image and the real image of a cigarette pack leads to errors in cigarette pack quantity counting. Specifically:
[0006] In a first aspect, embodiments of this application provide a method for recognizing cigarette pack reflections based on a contrastive learning model. The method includes: acquiring a cigarette pack image to be identified from a whole image of a cigarette pack based on MSDA-YOLO, a convolutional neural network, and a feature pyramid network; determining the distance between the cigarette pack image to be identified and its reflection image based on a pre-trained contrastive learning model, and determining the cigarette pack type of the cigarette pack image to be identified based on the distance and a distance threshold; if the cigarette pack type of the cigarette pack image to be identified is the reflection image, then ignoring the cigarette pack image to be identified; if the cigarette pack type of the cigarette pack image to be identified is the cigarette pack body image, then identifying and acquiring cigarette pack information in the cigarette pack image to be identified, and performing cigarette pack quantity statistics based on the cigarette pack information.
[0007] Secondly, embodiments of this application provide a cigarette pack reflection recognition device based on a contrastive learning model. The device includes: an image acquisition module for acquiring a cigarette pack image to be identified from a whole cigarette pack image based on MSDA-YOLO, a convolutional neural network, and a feature pyramid network; a recognition module for determining the distance between the cigarette pack image to be identified and the cigarette pack reflection image based on a pre-trained contrastive learning model, and determining the cigarette pack type of the cigarette pack image to be identified based on the distance and a distance threshold; a first judgment module for ignoring the cigarette pack image if the cigarette pack type of the cigarette pack image to be identified is the cigarette pack reflection image; and a second judgment module for identifying and acquiring cigarette pack information in the cigarette pack image to be identified if the cigarette pack type of the cigarette pack image to be identified is the cigarette pack body image, and performing cigarette pack quantity statistics based on the cigarette pack information.
[0008] Thirdly, embodiments of this application provide an electronic device comprising: one or more processors, a memory, and one or more application programs. The one or more application programs are stored in the memory and configured to be executed by the one or more processors, and are configured to perform the method as described in the first aspect.
[0009] In the technical solution provided in this application, an image of a cigarette pack to be identified is obtained from the overall image of the cigarette pack based on MSDA-YOLO, a convolutional neural network, and a feature pyramid network. The distance between the image of the cigarette pack to be identified and its reflection image is determined based on a pre-trained contrastive learning model, and the type of cigarette pack in the image is determined according to the distance and a distance threshold. If the cigarette pack type of the image to be identified is a reflection image, the image is ignored. If the cigarette pack type of the image to be identified is the cigarette pack itself, the cigarette pack information in the image is identified, and the number of cigarette packs is counted based on this information. Therefore, by inputting the image of the cigarette pack to be identified into the pre-trained contrastive learning model, the reflection image of the cigarette pack can be accurately identified; simultaneously, when the reflection image is identified, it is automatically ignored, effectively eliminating reflection interference during cigarette pack counting. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments and drawings obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0011] Figure 1 The diagram shows a flowchart of a cigarette pack reflection recognition method based on a contrastive learning model proposed in an embodiment of this application.
[0012] Figure 2 The diagram shows a structural block diagram of a cigarette pack reflection recognition device based on a contrastive learning model proposed in an embodiment of this application.
[0013] Figure 3 A structural block diagram of an electronic device according to an embodiment of this application is shown. Detailed Implementation
[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0016] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0017] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0018] In smart retail and vending machine scenarios, customer interaction with products is becoming increasingly frequent, making precise merchandise management and inventory control crucial for improving customer satisfaction and operational efficiency. However, when items like cigarette packs are meticulously displayed in retail environments equipped with reflective surfaces (such as glass counters, stainless steel display racks, or display platforms made of any smooth material), technological challenges arise. System cameras, acting as the eyes of these smart retail systems, are responsible for capturing and identifying products on the shelves; however, under such specific conditions, they may indiscriminately record all images, including the actual cigarette packs and their mirror reflections.
[0019] Traditional cigarette pack recognition technology, while capable of identifying and counting items on shelves to some extent, often struggles to distinguish between real cigarette pack images and reflections created by light refraction or reflection. This technological bottleneck leads to a common problem: the system misidentifies reflections of cigarette packs as independent, real packs, resulting in discrepancies in subsequent inventory statistics and management.
[0020] Such errors in quantity statistics are extremely detrimental to merchants. Not only do they prevent merchants from accurately grasping the real-time inventory status of cigarette packs, but they can also trigger a series of chain reactions, such as over-replenishment or stockouts due to inaccurate inventory information, thereby affecting the customer shopping experience and brand trust. Furthermore, erroneous inventory data can mislead merchants' sales strategies and market forecasts, negatively impacting the company's operational efficiency and profitability in the long run.
[0021] Therefore, in order to solve the problems existing in the prior art, this application proposes a method, device and electronic device for recognizing cigarette pack reflections based on a contrastive learning model, which will be described below with reference to specific embodiments.
[0022] Please see Figure 1 , Figure 1This application illustrates an embodiment of a cigarette pack reflection recognition method based on a contrastive learning model, specifically:
[0023] Step S110: Obtain the image of the cigarette pack to be identified from the overall image of the cigarette pack based on MSDA-YOLO, convolutional neural network and feature pyramid network.
[0024] In this embodiment of the application, the overall image of the cigarette pack refers to a complete image that includes one or more cigarette packs and possible other background elements (such as shelves, reflective surfaces, other goods, etc.). The overall image of the cigarette pack can be a real-time image captured by a camera or other image acquisition device, or it can be an image obtained from other channels (such as an image manually uploaded by the merchant). The image of the cigarette pack to be identified refers to the image of a single cigarette pack separated from the overall image of the cigarette pack and prepared for further identification processing.
[0025] Specifically, after the camera acquires the overall image of the cigarette pack, it can segment and separate the cigarette pack in the overall image using MSDA-YOLO, convolutional neural network, and feature pyramid network to form individual cigarette pack images to be identified.
[0026] In some implementations, the process of obtaining a cigarette pack image to be identified from a whole image of a cigarette pack based on MSDA-YOLO, a convolutional neural network, and a feature pyramid network includes: acquiring a whole image of a cigarette pack; performing distortion correction processing on the whole image of the cigarette pack to obtain a distorted whole image of the cigarette pack; preprocessing the distorted whole image of the cigarette pack; detecting cigarette pack images in the preprocessed distorted whole image of the cigarette pack based on MSDA-YOLO, and determining the cigarette box region image based on the detection results; extracting image features from the cigarette box region image based on a convolutional neural network; inputting the image features into a feature pyramid network; determining the cutting region from the distorted whole image of the cigarette pack based on the multi-scale feature information and attention mechanism output by the feature pyramid network; and cutting the cutting region based on DeepLab to obtain a cut image of the cigarette pack; and straightening the obtained cut image of the cigarette pack to obtain the cigarette pack image to be identified.
[0027] Specifically, in processing cigarette pack images, we first acquire an overall image of the cigarette pack, containing at least the cigarette packs. Due to factors such as image acquisition equipment or environment, the image may be distorted, so distortion correction processing is needed to restore the original shape and proportions of the image, resulting in a distorted overall image of the cigarette pack. Next, to improve the effect of subsequent processing, the distorted overall image is preprocessed, and then MSDA-YOLO is used to scan the preprocessed distorted overall image of the cigarette pack to detect and identify the approximate location and range of the cigarette pack images. This allows us to crop out irrelevant scenes from the distorted overall image of the cigarette pack, obtaining the cigarette box region image. Preprocessing can include operations such as image normalization and image enhancement (e.g., contrast adjustment, noise suppression). Furthermore, to analyze the specific location of each cigarette pack in the cigarette box region image more deeply, a convolutional neural network is used to extract key features from the image. Inputting these image features into a feature pyramid network allows for a more accurate understanding of the cigarette pack information in the image. Then, based on the output multi-scale feature information and the attention map generated by the attention mechanism, the cutting region is accurately determined from the distorted overall image of the cigarette pack. Then, DeepLab is used to perform fine pixel-level segmentation of the cut area. This accurately separates individual cigarette pack cut images from the overall image after distortion correction. Finally, the cut images are corrected using a preset image straightening algorithm to maintain their vertical orientation, resulting in the final image of the cigarette pack to be identified. The preset image straightening algorithm can be, for example, affine transformation or perspective transformation.
[0028] Furthermore, in The cigarette pack cut image is obtained by cutting the cutting area using DeepLab. Subsequently, since cigarette packs may be obscured in practical applications, it is necessary to identify and count the obscured cigarette packs to ensure the accuracy of the cigarette pack count. Specifically, the occlusion images of cigarette packs are obtained from the preprocessed, distortion-free overall image of the cigarette packs using U-Net; the occlusion images are reconstructed using generative adversarial networks and context-aware mechanisms to obtain the reconstructed cigarette pack images; and the segmented cigarette pack images and the reconstructed cigarette pack images are then straightened to obtain the images of the cigarette packs to be identified.
[0029] Specifically, if there is occlusion of the cigarette pack, U-Net can be used to segment the occluded region from the preprocessed distorted overall image of the cigarette pack. Then, the occluded region is reconstructed based on the generative adversarial network (GAN) and context-aware mechanism to generate a reconstructed image of the cigarette pack. After that, the segmented image and the reconstructed image of the cigarette pack are straightened to obtain all the images of the cigarette pack to be identified from the overall image of the cigarette pack.
[0030] In some implementations, since the cigarette packs on the left and right sides are roughly the same shape during the sales process, the reliability of the reconstructed cigarette pack image can be determined based on the similarity between the reconstructed image and at least one of the cigarette pack images. If the similarity is lower than a preset threshold, the reconstructed image is considered to deviate significantly from the actual image, and the reconstruction fails. A secondary reconstruction can then be performed separately based on its position, or it can be discarded directly.
[0031] Step S120: Determine the distance between the cigarette pack image to be identified and the cigarette pack reflection image based on the pre-trained contrastive learning model, and determine the cigarette pack type of the cigarette pack image to be identified based on the distance and the distance threshold.
[0032] In this embodiment of the application, the cigarette pack type includes a cigarette pack body image and a cigarette pack reflection image.
[0033] Specifically, if the distance is greater than or equal to the distance threshold, the cigarette pack type of the image to be identified is considered to be a cigarette pack reflection image; if the distance is less than the distance threshold, the cigarette pack category of the image to be identified is considered to be a cigarette pack body image.
[0034] In this embodiment, the training process of the contrastive learning model can be as follows: Training samples are constructed by acquiring an image of the cigarette pack itself, an image of its reflection, and other images of the cigarette pack; wherein, the other images of the cigarette pack have the same orientation as the image of the cigarette pack but different scenes, and are not images of the cigarette pack's reflection; Triples are formed by randomly acquiring the image of the cigarette pack itself, other images of the cigarette pack, and images of its reflection from the training samples; Cigarette image features in the triples are extracted based on the convolutional neural network in the contrastive learning model, and the cigarette image features are converted into cigarette pack feature vectors; A first distance between the image of the cigarette pack itself and other images of the cigarette pack is determined based on the cigarette pack feature vectors and Euclidean distance, and a second distance between the image of the cigarette pack itself and its reflection is determined; The loss value corresponding to the triples is determined based on the first distance, the second distance, and Triplet Loss; The total loss value is determined by summing the loss values of all triples; The weights and biases of the contrastive learning model are updated based on the backpropagation algorithm and the mini-batch gradient descent algorithm according to the total loss value, iterating until the change in the loss value of the contrastive learning model meets the preset conditions to obtain the trained contrastive learning model.
[0035] In some implementations, images of the cigarette pack itself, its reflection, and other images of the cigarette pack can be collected from multiple sources to ensure the diversity and representativeness of the training samples, covering images of cigarette packs from different brands, models, and scenarios. Understandably, the initially collected images can undergo data cleaning to remove excessively blurry or repetitive images, and the cleaned images can be cropped, scaled, and normalized to ensure the consistency of the input data.
[0036] In some implementations, the specific architecture of the convolutional neural network can be, for example, ResNet, VGG, etc. Specifically, the input layer of the convolutional neural network receives a fixed-size image of the cigarette pack itself, an image of the cigarette pack's reflection, and other images of the cigarette pack. These images are then processed through a fully connected layer to obtain the features of the cigarette pack image, and finally, a fixed-length feature vector is output.
[0037] In this embodiment, Triplet Loss is the core loss function of the contrastive learning model. This function aims to minimize the distance between positive sample pairs (i.e., the image of the cigarette pack itself and other images of the cigarette pack) and maximize the distance between negative sample pairs (i.e., the image of the cigarette pack itself and its reflection), while maintaining a margin α to ensure that the distance between negative sample pairs is always greater than the distance between positive sample pairs plus this margin. Specifically, Triplet Loss is defined as follows:
[0038] ,
[0039] Where f(A) is the embedding representation of the cigarette pack body image sample, f(P) is the embedding representation of other cigarette pack image samples, f(N) is the embedding representation of the cigarette pack reflection image sample, and α is a hyperparameter used to control the minimum distance between the cigarette pack body image sample and the cigarette pack reflection image sample.
[0040] In some implementations, the preset conditions may be, for example, that the loss value no longer decreases significantly, that a certain number of iterations are reached, or that the performance on the validation set is optimal.
[0041] In some implementations, updating the weights and biases of the contrastive learning model based on the backpropagation algorithm and the mini-batch gradient descent algorithm according to the total loss value may include: determining the gradients of the weights and biases of the contrastive learning model based on the total loss value and the backpropagation algorithm; adjusting the first learning rate based on the gradients of the weights and biases and the Adam optimizer; and updating the weights and biases based on the gradients of the weights and biases, the first learning rate, and the mini-batch gradient descent algorithm.
[0042] The first learning rate is the initial learning rate. Specifically, during the training phase, the model receives a series of triplet data, calculates the loss value for each triplet through forward propagation, and these loss values are accumulated to form the total loss value. The total loss value reflects the difference between the model's current predictive ability for all training samples and the actual labels. To improve the model, the gradient of the total loss value with respect to the model's weights and biases can be calculated using the backpropagation algorithm. These gradients indicate how the weights and biases should be adjusted to reduce the total loss value. Specifically, the backpropagation algorithm starts from the output layer and calculates the gradient of each layer's parameters layer by layer until it reaches the input layer. However, directly updating the weights and biases based on these gradients may lead to unstable training or slow convergence. Therefore, this scheme uses the Adam optimizer to help adjust the learning rate to control the update step size of the weights and biases. It can automatically adjust the learning rate of each parameter and calculate the update value based on the first and second moment estimates of the gradient. In the Adam optimizer, an initial learning rate (i.e., the first learning rate) is set. The Adam optimizer calculates the exponential moving average of the gradient (i.e., the first moment estimate) and the exponential moving average of the squared gradient (i.e., the second moment estimate), and then adjusts the learning rate for each parameter based on these two estimates. Finally, the model's weights and biases are updated based on the gradients of the weights and biases, the learning rate adjusted by the Adam optimizer, and the mini-batch gradient descent algorithm. Understandably, based on the backpropagation algorithm, the mini-batch gradient descent algorithm, and the Adam optimizer, the weights and biases of the contrastive learning model can be effectively updated, thereby minimizing the total loss and improving model performance.
[0043] Furthermore, if the convergence effect is not ideal, it may be due to insufficient negative samples, thus requiring an increase in the number of negative samples. Specifically, the number of cigarette pack reflection images in the training samples is determined based on the changes in the loss value of the contrastive learning model during the iteration process; if insufficient, data augmentation is performed based on the cigarette pack reflection images to obtain pseudo cigarette pack reflection images, which are then added to the training samples.
[0044] Specifically, if the loss value decreases slowly or stagnates, it indicates that there are insufficient images of cigarette pack reflections. Therefore, data augmentation can be performed on the cigarette pack reflection data to obtain pseudo-cigarette pack reflection images, which are then used as negative samples to supplement the training samples. Data augmentation can involve operations such as rotating, scaling, flipping, cropping, and adding noise to the cigarette pack reflection images to generate pseudo-cigarette pack reflection images. Although these pseudo-images are not real reflections, they can simulate different reflection scenarios, thus helping the model learn more useful features. After supplementing the training samples with the generated pseudo-cigarette pack reflection images, the contrastive learning model can continue to be trained. As training progresses, the model will be exposed to more diverse negative samples, thereby enhancing its ability to distinguish between the cigarette pack itself and its reflection, improving the model's convergence performance, and ultimately increasing the model's recognition accuracy and generalization ability.
[0045] Furthermore, based on at least one training iteration, key training samples can be obtained from historical training results to focus training on samples that are difficult to distinguish. Specifically, a predetermined number of triplets with the highest loss values are obtained based on the number of triplets and the overall loss values; data augmentation is performed on the obtained triplets to obtain key training samples; the second learning rate is adjusted based on the number of key training triplets, the total loss value, and the change in the total loss value in the key training samples; and the weights and biases of the contrastive learning model are updated based on the total loss value, the backpropagation algorithm, the second learning rate, and the mini-batch gradient descent algorithm.
[0046] Specifically, the loss value of all triples during training is first calculated. Then, based on the magnitude of the loss value, a predetermined number of triples with the largest loss values (e.g., the top 10% or the top N) are selected as the key training targets. These triples typically represent samples that the model currently struggles to distinguish, so focusing on them for training is expected to significantly improve model performance. Next, data augmentation techniques are used to generate more variations of the images in the triples. Data augmentation can include various image transformation operations, such as rotation, scaling, cropping, flipping, adding noise, and possible color changes and brightness adjustments. Through data augmentation, multiple similar samples can be generated for each key triple, forming a rich set of key training samples. Then, the second learning rate is dynamically adjusted based on the number of key training triples (i.e., the richness of the key training samples), the total loss value (i.e., the overall performance of the current model), and the change in the total loss value (i.e., the rate of improvement in model performance). For example, if the total loss value remains high and decreases slowly, the second learning rate can be increased to accelerate the learning of key training samples; conversely, if the total loss value is already low and tends to stabilize, the second learning rate can be appropriately decreased to prevent overfitting. After acquiring the key training samples and adjusting the second learning rate, the loss value for the key training samples is first calculated and then weighted and summed with the total loss value. Next, the backpropagation algorithm is used to calculate the gradient of the total loss value with respect to the model weights and biases. Then, the step size of these gradients is adjusted according to the second learning rate. Finally, the mini-batch gradient descent algorithm is used to update the model's weights and biases. This approach maintains model stability while focusing training on samples that are difficult to distinguish, thereby improving the overall performance of the model.
[0047] The initial second learning rate is greater than the initial first learning rate. In some implementations, the initial second learning rate can also be determined based on the average of the loss values in a preset number of triples.
[0048] In some implementations, the cigarette pack feature vector can also be determined based on fused features. Specifically, cigarette pack image features are extracted from triples using a convolutional neural network in a contrastive learning model; the physical image features of the cigarette pack are determined based on the light source category and light source position corresponding to the cigarette pack body image, other cigarette pack images, and cigarette pack reflection image in the triples; and the cigarette pack feature vector is determined based on the convolutional neural network according to the cigarette pack image features, the first weight, the cigarette pack physical image features, and the second weight.
[0049] The light source type can be, for example, daylight, fluorescent lamp, LED, etc., and the light source position can be, for example, front light source, side light source, backlight, etc.
[0050] Specifically, by combining the previously extracted image features and physical image features, a comprehensive cigarette pack feature vector can be formed. This vector includes not only the image features of the cigarette pack itself but also the physical image features of the cigarette pack (lighting, angle, position, etc.). Based on actual sales conditions, since the physical features of the cigarette pack have considerable uncertainty, a first weight can be assigned to the image features, and a second weight to the physical image features. The second weight can be at most half the first weight. Understandably, since the fused image features and physical image features of the cigarette pack form a comprehensive feature, it is necessary to first unfold the fused feature into a one-dimensional vector before inputting it into the convolutional neural network to determine the cigarette pack feature vector.
[0051] Understandably, combining image features extracted by convolutional neural networks with physical image features of cigarette packs, and assigning appropriate weights to them to form the final cigarette pack feature vector, brings significant benefits in several aspects: 1. Improved richness and accuracy of feature representation: While relying solely on image features extracted by convolutional neural networks can capture high-level information of the image, it may overlook physical conditions directly related to image generation (such as lighting, angle, etc.). By introducing physical image features of the cigarette pack, such as light source type and location, this deficiency can be compensated for, making the feature vector more comprehensively reflect the essential characteristics of the cigarette pack. This combination method makes the feature vector contain both the detailed information of the image itself and the contextual information of image generation, thereby improving the richness and accuracy of feature representation. 2. Enhanced robustness and generalization ability of the model: In practical applications, cigarette pack images may be affected by various factors, such as changes in lighting, angle changes, occlusion, etc. By introducing physical image features, the model can better understand the impact of these factors on the image, thereby more accurately identifying the cigarette pack. This ability makes the model more robust when facing complex and ever-changing real-world scenarios. Meanwhile, because feature vectors contain more comprehensive information, the model can learn more useful feature combinations and patterns during training, thereby improving its generalization ability on new samples.
[0052] It should be noted that the loss value acquisition methods and feature extraction methods mentioned above can all be explained with reference to the ideas presented in the first explanation. The input data that needs to be adjusted can be adjusted accordingly.
[0053] Step S130: If the cigarette pack image to be identified is a cigarette pack reflection image, then the cigarette pack image to be identified is ignored.
[0054] In this embodiment, when the contrastive learning model identifies the cigarette pack image to be identified as a cigarette pack reflection image, the cigarette pack image to be identified is ignored, and no output recognition or count is performed. This mechanism ensures that the cigarette pack reflection image does not interfere with the final recognition result, avoids the problems of misidentification and duplicate counting, and improves the accuracy of recognition and the stability of the system.
[0055] Furthermore, ignoring reflection images means the model doesn't need to perform complex output recognition on these images. This reduces computational resource consumption, increases processing speed, and allows the model to process other, more valuable images faster. Ignoring reflection images also simplifies the processing flow. No additional labeling, classification, or filtering is required, reducing data processing workload and improving efficiency. Moreover, in practical applications, if the model incorrectly identifies a reflection image as the cigarette pack itself or another type, it may cause confusion and false alarms. This could lead to subsequent incorrect decisions or unnecessary interventions. By ignoring reflection images, such confusion and false alarms can be avoided, improving the system's stability and reliability.
[0056] Step S140: If the cigarette pack type of the cigarette pack image to be identified is a cigarette pack body image, then the cigarette pack information in the cigarette pack image to be identified is obtained, and the number of cigarette packs is counted based on the cigarette pack information.
[0057] In this embodiment, if the type of cigarette pack image to be identified is determined to be a cigarette pack body image, then the text and images in the cigarette pack body image are identified to determine the cigarette brand, specifications, etc., to which the cigarette pack image belongs. Then, the number of cigarette packs is counted based on the identified cigarette pack information. The cigarette pack body image is a real image of the cigarette pack.
[0058] In some implementations, after the identification and statistics are completed, feedback data of the cigarette pack image to be identified can be determined based on MSDA-YOLO; wherein, the feedback data includes the cutting accuracy and reconstruction effect; MSDA-YOLO, DeepLab and Generative Adversarial Network are adjusted according to the feedback data, and the adjusted MSDA-YOLO, DeepLab and Generative Adversarial Network are used as the new MSDA-YOLO, DeepLab and Generative Adversarial Network.
[0059] Specifically, after each recognition and statistical task is completed, MSDA-YOLO can be used to evaluate the performance of the task and generate feedback data. Based on the collected feedback data, MSDA-YOLO, DeepLab, and the generative adversarial network used in this application can be adjusted and optimized. The adjusted MSDA-YOLO, DeepLab, and generative adversarial network can then be used in the next execution of the cigarette pack reflection recognition method provided in this application. Adjustments may include fine-tuning model parameters, improving the network structure, and enhancing training data. Optimizing the cigarette pack reflection recognition method each time it is executed can gradually improve the accuracy and efficiency of subsequent recognitions.
[0060] In the technical solution provided in this application, an image of a cigarette pack to be identified is obtained from the overall image of the cigarette pack based on MSDA-YOLO, a convolutional neural network, and a feature pyramid network. The distance between the image of the cigarette pack to be identified and its reflection image is determined based on a pre-trained contrastive learning model, and the type of cigarette pack in the image is determined according to the distance and a distance threshold. If the cigarette pack type of the image to be identified is a reflection image, the image is ignored. If the cigarette pack type of the image to be identified is the cigarette pack itself, the cigarette pack information in the image is identified, and the number of cigarette packs is counted based on this information. Therefore, by inputting the image of the cigarette pack to be identified into the pre-trained contrastive learning model, the reflection image of the cigarette pack can be accurately identified; simultaneously, when the reflection image is identified, it is automatically ignored, effectively eliminating reflection interference during cigarette pack counting.
[0061] Please see Figure 2 , Figure 2 This application illustrates a cigarette pack reflection recognition device 100 based on a contrastive learning model, according to an embodiment of this application. The device includes an image acquisition module 110, a recognition module 120, a first judgment module 130, and a second judgment module 140. Specifically:
[0062] Image acquisition module 110 is used to acquire the image of the cigarette pack to be identified from the overall image of the cigarette pack based on MSDA-YOLO, convolutional neural network and feature pyramid network;
[0063] The recognition module 120 is used to determine the distance between the image of the cigarette pack to be recognized and the reflection image of the cigarette pack based on a pre-trained contrastive learning model, and to determine the type of cigarette pack in the image of the cigarette pack to be recognized based on the distance and a distance threshold.
[0064] The first judgment module 130 is used to ignore the cigarette pack image to be identified if the cigarette pack type of the image to be identified is a cigarette pack reflection image.
[0065] The second judgment module 140 is used to identify and obtain the cigarette pack information in the cigarette pack image if the cigarette pack type of the cigarette pack image to be identified is a cigarette pack body image, and to count the number of cigarette packs based on the cigarette pack information.
[0066] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For device embodiments, since they are basically similar to method embodiments, the descriptions are relatively simple; relevant parts can be referred to the descriptions in the method embodiments. Any processing method described in the method embodiments can be implemented in the device embodiments through corresponding processing modules, and will not be elaborated upon further in the device embodiments.
[0067] Please see Figure 3 , Figure 3 An electronic device 200 according to an embodiment of this application is shown. The electronic device 200 of this application may include one or more of the following components: a processor 210, a memory 220, and one or more application programs, wherein the one or more application programs may be stored in the memory 220 and configured to be executed by the one or more processors 210, and the one or more application programs are configured to perform the methods as described in the foregoing method embodiments.
[0068] Processor 210 may include one or more processing cores. Processor 210 connects to various parts within the electronic device 200 using various interfaces and lines, and performs various functions and processes data of the electronic device 200 by running or executing instructions, programs, code sets, or instruction sets stored in memory 220, and by calling data stored in memory 220. Optionally, processor 210 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 210 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 210, but may be implemented separately through a communication chip.
[0069] The memory 220 may include random access memory (RAM) or read-only memory (ROM). The memory 220 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 220 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a recognition function, a judgment function, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data created by the electronic device 200 during use, such as an image of a cigarette pack to be recognized, a reflection image of a cigarette pack, etc.
[0070] This application provides a method, apparatus, and electronic device for identifying cigarette pack reflections based on a contrastive learning model. The method acquires the image of the cigarette pack to be identified from the overall image of the cigarette pack using MSDA-YOLO, a convolutional neural network, and a feature pyramid network. It determines the distance between the image of the cigarette pack to be identified and its reflection image based on a pre-trained contrastive learning model, and determines the type of cigarette pack in the image of the cigarette pack to be identified based on the distance and a distance threshold. If the type of the cigarette pack to be identified is a reflection image, the image of the cigarette pack to be identified is ignored. If the type of the cigarette pack to be identified is the cigarette pack itself, the cigarette pack information in the image of the cigarette pack to be identified is acquired, and the number of cigarette packs is counted based on the cigarette pack information. Therefore, by inputting the image of the cigarette pack to be identified into the pre-trained contrastive learning model, the reflection image of the cigarette pack can be accurately identified; at the same time, the reflection image is automatically ignored when it is identified, effectively eliminating reflection interference during cigarette pack count.
[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for recognizing the reflection of cigarette packs based on a contrastive learning model, characterized in that, The method includes: A complete image of the cigarette pack is acquired, and distortion correction processing is performed on the complete image to obtain a distortion-corrected image of the cigarette pack. The distortion-corrected image is then preprocessed, and the cigarette pack image within the preprocessed image is detected using MSDA-YOLO. Based on the detection results, the cigarette box region image is determined. Image features are extracted from the cigarette box region image using a convolutional neural network. These image features are input into a feature pyramid network, and a cutting region is determined from the distortion-corrected image of the cigarette pack based on the multi-scale feature information and attention mechanism output by the feature pyramid network. The cutting region is then cut using DeepLab to obtain a cut image of the cigarette pack. The obtained cut image of the cigarette pack is then straightened to obtain the image of the cigarette pack to be identified. The distance between the image of the cigarette pack to be identified and its reflection image is determined based on a pre-trained contrastive learning model, and the type of cigarette pack in the image is determined according to the distance and a distance threshold. The training process of the pre-trained contrastive learning model includes: acquiring a cigarette pack body image, the cigarette pack reflection image, and other cigarette pack images to construct training samples; randomly selecting the cigarette pack body image, other cigarette pack images, and the cigarette pack reflection image from the training samples to form triples; extracting cigarette pack image features from the triples based on the convolutional neural network in the contrastive learning model, and converting the cigarette pack image features into cigarette pack feature vectors. The process involves: determining a first distance between the cigarette pack body image and other images of the cigarette pack based on the cigarette pack feature vector and Euclidean distance; determining a second distance between the cigarette pack body image and the cigarette pack reflection image; determining the loss value corresponding to the triplet based on the first distance, the second distance, and TripletLoss; summing the loss values of all triplets to determine the total loss value; updating the weights and biases of the contrastive learning model based on the backpropagation algorithm and mini-batch gradient descent algorithm according to the total loss value; iterating until the change in the loss value of the contrastive learning model meets the preset conditions to obtain the trained contrastive learning model. If the cigarette pack image to be identified is a cigarette pack reflection image, then the cigarette pack image to be identified is ignored. If the cigarette pack type of the image to be identified is the cigarette pack body image, then the cigarette pack information in the image to be identified is obtained, and the number of cigarette packs is counted based on the cigarette pack information.
2. The method according to claim 1, characterized in that, The method of updating the weights and biases of the contrastive learning model based on the backpropagation algorithm and mini-batch gradient descent algorithm according to the total loss value includes: The gradients of the weights and biases of the contrastive learning model are determined based on the total loss value and the backpropagation algorithm. The first learning rate is adjusted based on the gradient and Adam optimizer according to the weights and biases. The weights and biases are updated based on the gradient of the weights and biases, the first learning rate, and the mini-batch gradient descent algorithm. After updating the weights and biases based on the gradient of the weights and biases, the first learning rate, and the mini-batch gradient descent algorithm, the method further includes: Based on the change in the loss value of the contrastive learning model during the iteration process, determine whether the number of cigarette pack reflection images in the training samples is sufficient; If this is insufficient, data augmentation is performed based on the cigarette pack reflection image to obtain a pseudo cigarette pack reflection image, which is then added to the training samples.
3. The method according to claim 2, characterized in that, After updating the weights and biases of the contrastive learning model based on the total loss value using the backpropagation algorithm and mini-batch gradient descent algorithm, the method further includes: Based on the number of triplets and the overall loss value, obtain a preset number of triplets with the largest loss value; The obtained triples are augmented to obtain key training samples; The second learning rate is adjusted based on the number of key training triples, the total loss value, and the changes in the total loss value in the key training samples. The weights and biases of the contrastive learning model are updated based on the total loss value, the backpropagation algorithm, the second learning rate, and the mini-batch gradient descent algorithm.
4. The method according to claim 3, characterized in that, The method of extracting cigarette pack image features from the triples using a convolutional neural network in a contrastive learning model, and converting the cigarette pack image features into cigarette pack feature vectors, includes: The convolutional neural network in the contrastive learning model extracts the cigarette pack image features from the triples; Based on the light source category and light source position corresponding to the cigarette pack body image, other cigarette pack images, and cigarette pack reflection image in the triplet, the physical image features of the cigarette pack are determined. Based on the convolutional neural network, the cigarette pack feature vector is determined according to the cigarette pack image features, the first weight, the cigarette pack physical image features, and the second weight.
5. The method according to any one of claims 1-4, characterized in that, After obtaining the cigarette pack cut image by cutting the cutting region using DeepLab, the method further includes: The cigarette pack occlusion image in the preprocessed overall image of the cigarette pack after distortion removal is obtained based on U-Net; The occluded image of the cigarette pack is reconstructed based on generative adversarial networks and context-aware mechanisms to obtain a reconstructed image of the cigarette pack. The step of straightening the acquired cut image of the cigarette pack to obtain the image of the cigarette pack to be identified includes: The cut image and reconstructed image of the cigarette pack are straightened to obtain the image of the cigarette pack to be identified.
6. The method according to claim 5, characterized in that, After counting the number of cigarette packs based on the cigarette pack information, the method further includes: The feedback data for determining the image of the cigarette pack to be identified is based on the MSDA-YOLO; wherein, the feedback data includes cutting accuracy and reconstruction effect; Based on the feedback data, the MSDA-YOLO, DeepLab, and Generative Adversarial Network are adjusted, and the adjusted MSDA-YOLO, DeepLab, and Generative Adversarial Network are used as the new MSDA-YOLO, DeepLab, and Generative Adversarial Network.
7. A cigarette pack reflection recognition device based on a contrastive learning model, characterized in that, The device includes: An image acquisition module is used to acquire an overall image of the cigarette pack, perform distortion correction processing on the overall image of the cigarette pack to obtain a distortion-corrected overall image of the cigarette pack; preprocess the distortion-corrected overall image of the cigarette pack, detect the cigarette pack image in the preprocessed distortion-corrected overall image of the cigarette pack based on MSDA-YOLO, and determine the cigarette box region image based on the detection results; extract image features from the cigarette box region image based on a convolutional neural network; input the image features into a feature pyramid network, determine the cutting region from the distortion-corrected overall image of the cigarette pack based on the multi-scale feature information and attention mechanism output by the feature pyramid network, and cut the cutting region based on DeepLab to obtain a cigarette pack cutting image; and perform straightening processing on the acquired cigarette pack cutting image to obtain the cigarette pack image to be identified. The identification module is used to determine the distance between the image of the cigarette pack to be identified and its reflection image based on a pre-trained contrastive learning model, and to determine the cigarette pack type of the image of the cigarette pack to be identified based on the distance and a distance threshold. The training process of the pre-trained contrastive learning model includes: acquiring a cigarette pack body image, the cigarette pack reflection image, and other cigarette pack images to construct training samples; randomly selecting the cigarette pack body image, other cigarette pack images, and the cigarette pack reflection image from the training samples to form a triplet; extracting cigarette pack image features from the triplet based on a convolutional neural network in the contrastive learning model, and converting the cigarette pack image features into a cigarette pack feature vector; determining a first distance between the cigarette pack body image and other cigarette pack images based on the cigarette pack feature vector and Euclidean distance, and determining a second distance between the cigarette pack body image and the cigarette pack reflection image; and determining the cigarette pack type based on the first distance, the second distance, and the triplet distance. Loss determines the loss value corresponding to the triple; the loss values of all triples are added together to determine the total loss value; the weights and biases of the contrastive learning model are updated according to the total loss value based on the backpropagation algorithm and the mini-batch gradient descent algorithm; the iteration continues until the change of the loss value of the contrastive learning model meets the preset conditions to obtain the trained contrastive learning model. The first judgment module is used to ignore the cigarette pack image if the cigarette pack type of the image to be identified is the cigarette pack reflection image. The second judgment module is used to identify and obtain cigarette pack information in the cigarette pack image if the cigarette pack type of the cigarette pack image to be identified is the cigarette pack body image, and to count the number of cigarette packs based on the cigarette pack information.
8. An electronic device, characterized in that, include: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to perform the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Multi-target tracking method based on multi-scale deformable attention mechanism
CN116309725A
Terrain surveying and mapping method based on aerial surveying and mapping technology
CN118640878A