Cigarette identification method and system based on knowledge distillation, electronic equipment and medium
By applying knowledge distillation technology in the cigarette recognition system, pruning and constructing a lightweight model, the problems of high complexity and low recognition accuracy in the existing technology are solved, and efficient and real-time cigarette product identification is achieved.
Patent Information
- Application Number
- CN202510105698.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing cigarette product identification methods rely on deep learning models, resulting in high model complexity, high computing and storage resource requirements, high implementation cost, low real-time performance, and low recognition accuracy on edge devices or mobile devices.
Using a knowledge distillation method, the teacher model is pruned, the student model is constructed, and the knowledge and parameters of the teacher model are transferred to the student model through knowledge distillation, reducing the complexity of the model and improving identification efficiency and real-timeness.
It realizes efficient identification of cigarette products on edge devices or mobile devices, improves identification accuracy and real-time performance, and reduces computing and storage resource requirements.
Smart Images

Figure CN120070962A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a cigarette recognition method, system, electronic device and storage medium based on knowledge distillation. Background Art
[0002] In cigarette retail services, it is necessary to identify the cigarette specifications in the display cabinet to determine whether there are any violations; traditional cigarette specification identification usually adopts manual identification, with low identification efficiency and high work intensity; existing identification methods use deep learning models for cigarette specification identification, automatically identifying cigarette images uploaded by edge devices or mobile devices to reduce the work intensity; however, existing deep learning models have high complexity, high requirements for computing and storage resources, high implementation costs, and low real-time performance; while models that can be applied to edge devices or mobile devices are limited by computing resources and have low identification accuracy. Summary of the Invention
[0003] The main purpose of the embodiments of the present invention is to provide a cigarette recognition method, system, electronic device and storage medium based on knowledge distillation, which can improve the recognition efficiency and real-time performance and improve the recognition accuracy.
[0004] To achieve the above object, on the one hand, an embodiment of the present invention provides a cigarette recognition method based on knowledge distillation, the method comprising:
[0005] Obtain a display cigarette image to be recognized, input the display cigarette image to be recognized into a lightweight model for processing, and determine the cigarette specification information of the display cigarette image to be recognized; wherein, the cigarette specification information includes a cigarette specification code and a cigarette category, and the lightweight model is determined by the following method:
[0006] Process the obtained sample image dataset to determine a cigarette specification image dataset, and determine a training set and a test set according to a first preset ratio and the cigarette specification image dataset;
[0007] Construct a first model, train the first model according to a preset dataset to determine a teacher model, and prune the teacher model to determine a student model;
[0008] Train the student model according to the teacher model and the training set to determine a second model; test the second model according to the test set to determine the recognition accuracy;
[0009] Determine a lightweight model according to the recognition accuracy and a first preset threshold.
[0010] In some embodiments, the processing the obtained sample image dataset to determine a cigarette specification image dataset specifically includes:
[0011] Extract the sample image data in the sample image dataset to determine the cigarette brand dataset;
[0012] Match the cigarette brand dataset with a preset dictionary to determine the matching result; wherein, the matching result includes brand code information and cigarette category information;
[0013] Classify and label the sample image dataset according to the matching result to determine the cigarette brand image dataset.
[0014] In some embodiments, pruning the teacher model to determine the student model specifically includes:
[0015] Determine a model weight set according to the weights of the teacher model and a second preset ratio, and generate a weight mask according to the model weight set;
[0016] Prune the teacher model according to the weight mask to determine the pruned teacher model, and train the pruned teacher model according to the preset dataset to determine the updated pruned model;
[0017] Determine the student model according to the updated pruned model and a second preset threshold.
[0018] In some embodiments, determining the student model according to the updated pruned model and a second preset threshold specifically includes:
[0019] Determine the model pruning rate according to the updated pruned model, and compare the model pruning rate with the second preset threshold;
[0020] If the model pruning rate is greater than or equal to the second preset threshold, use the updated pruned model as the student model;
[0021] If the model pruning rate is less than the second preset threshold, use the updated pruned model as the teacher model, and return to execute determining the model weight set according to the weights of the teacher model and the second preset ratio, and generating a weight mask according to the model weight set; until the model pruning rate is greater than or equal to the second preset threshold, use the updated pruned model as the student model.
[0022] In some embodiments, training the student model according to the teacher model and the training set to determine a second model specifically includes:
[0023] Input the training set into the teacher model for processing to determine a first recognition result; input the training set into the student model for processing to determine a second recognition result;
[0024] Determine a target loss value according to the first recognition result and the second recognition result, and compare the target loss value with a third preset threshold;
[0025] If the target loss value is less than or equal to the third preset threshold, use the student model as the second model;
[0026] If the target loss value is greater than the third preset threshold, adjust the parameters of the student model, and return to execute the step of inputting the training set into the teacher model for processing to determine the first recognition result; input the training set into the student model for processing to determine the second recognition result; until the target loss value is less than or equal to the third preset threshold, use the student model as the second model.
[0027] In some embodiments, the determining the target loss value according to the first recognition result and the second recognition result specifically includes:
[0028] Determine a hard decision result and a soft decision result according to the second recognition result;
[0029] Calculate according to the soft decision result and the first recognition result to determine an evaporation loss value; calculate according to the hard decision result and the training set to determine a hard discrimination loss value;
[0030] Perform weighted summation according to the evaporation loss value and the hard discrimination loss value to determine the target loss value.
[0031] In some embodiments, the determining the lightweight model according to the recognition accuracy rate and the first preset threshold specifically includes:
[0032] Compare the recognition accuracy rate with the first preset threshold;
[0033] If the recognition accuracy rate is greater than or equal to the first preset threshold, use the second model as the lightweight model;
[0034] If the recognition accuracy rate is less than the first preset threshold, adjust the parameters of the second model, and return to execute the step of training the student model according to the teacher model and the training set to determine the second model; until the recognition accuracy rate is greater than the first preset threshold, use the second model as the lightweight model.
[0035] To achieve the above object, another aspect of the embodiments of the present invention provides a cigarette recognition system based on knowledge distillation, including:
[0036] A cigarette recognition module, which is used to obtain an image of the displayed cigarettes to be recognized, input the image of the displayed cigarettes to be recognized into a lightweight model for processing, and determine the cigarette product specification information of the image of the displayed cigarettes to be recognized; wherein, the cigarette product specification information includes a cigarette product specification code and a cigarette category, and the lightweight model is determined by the following method:
[0037] Process the obtained sample image dataset to determine a cigarette product specification image dataset, and determine a training set and a test set according to a first preset ratio and the cigarette product specification image dataset;
[0038] Construct a first model, train the first model according to a preset dataset to determine a teacher model, and prune the teacher model to determine a student model;
[0039] Train the student model according to the teacher model and the training set to determine a second model; test the second model according to the test set to determine the recognition accuracy rate;
[0040] Determine the lightweight model according to the recognition accuracy rate and a first preset threshold.
[0041] To achieve the above object, another aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the foregoing method is implemented.
[0042] To achieve the above object, another aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the foregoing method is implemented.
[0043] Implementing the embodiments of the present invention includes the following beneficial effects: This embodiment provides a cigarette recognition method, system, electronic device, and storage medium based on knowledge distillation. In this solution, the image of the displayed cigarettes to be recognized is input into a lightweight model for recognition, and the corresponding cigarette specification information is output. The lightweight model processes the acquired sample image dataset to obtain a cigarette specification image dataset, divides the cigarette specification image dataset according to a first preset ratio to obtain a training set and a test set; trains the constructed first model according to a preset dataset to obtain a teacher model; prunes the trained teacher model to construct a student model; trains the student model according to the divided training set and the teacher model to obtain a second model; tests the second model according to the test set to determine the recognition accuracy rate of the second model; determines the lightweight model by comparing the recognition accuracy rate with a first preset threshold, and uses the lightweight model to recognize the image of the displayed cigarettes. By pruning the trained teacher model to construct a student model, the model complexity is reduced, and the model recognition efficiency and real-time performance are improved; the constructed model is trained according to the trained teacher model and the training set, and the knowledge and parameters of the teacher model are transferred to the constructed student model through knowledge distillation, improving the recognition accuracy rate of the student model. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 FIG. is a schematic flowchart of the steps of a cigarette recognition method based on knowledge distillation provided by an embodiment of the present invention;
[0045] Figure 2 FIG. is a schematic flowchart of the steps of determining a cigarette specification image dataset in a cigarette recognition method based on knowledge distillation provided by an embodiment of the present invention;
[0046] Figure 3 FIG. is a schematic flowchart of the steps of pruning a teacher model in a cigarette recognition method based on knowledge distillation provided by an embodiment of the present invention;
[0047] Figure 4 FIG. is a schematic flowchart of the steps of determining a second model in a cigarette recognition method based on knowledge distillation provided by an embodiment of the present invention;
[0048] Figure 5 FIG. is a schematic flowchart of the steps of determining a third model in a cigarette recognition method based on knowledge distillation provided by an embodiment of the present invention;
[0049] Figure 6 FIG. is a schematic flowchart of the steps of calculating a target loss value in a cigarette recognition method based on knowledge distillation provided by an embodiment of the present invention;
[0050] Figure 7It is a schematic diagram of the step process for determining a student model in a cigarette recognition method based on knowledge distillation provided by an embodiment of the present invention;
[0051] Figure 8 It is a schematic diagram of the step process of a specific embodiment provided by an embodiment of the present invention;
[0052] Figure 9 It is a schematic diagram of an original sample diagram collected in a specific embodiment provided by an embodiment of the present invention;
[0053] Figure 10(a) - Figure 10(c) It is a schematic diagram of a cigarette product specification image data set in a specific embodiment provided by an embodiment of the present invention;
[0054] Figure 11(a) - Figure 11(b) They are respectively schematic diagrams of the structures of a teacher model and a student model in a specific embodiment provided by an embodiment of the present invention;
[0055] Figure 12 It is a schematic diagram of the structure of a residual block of a teacher model in a specific embodiment provided by an embodiment of the present invention;
[0056] Figure 13 It is a schematic diagram of the change curve of the recognition accuracy of the student model of group B for the training set and the test set in a specific embodiment provided by an embodiment of the present invention;
[0057] Figure 14(a) - Figure 14(d) It is a schematic diagram of the recognition result of the teacher model for the test set in a specific embodiment provided by an embodiment of the present invention;
[0058] Figure 14(e) - Figure 14(h) It is a schematic diagram of the recognition result of the student model for the test set in a specific embodiment provided by an embodiment of the present invention;
[0059] Figure 15 It is a schematic diagram of the change curve of the recognition accuracy of the teacher model for the training set and the test set in a specific embodiment provided by an embodiment of the present invention;
[0060] Figure 16 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0061] The following further elaborates the present invention in detail with reference to the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0062] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict.
[0063] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It is understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present invention described herein can be implemented in an order other than that illustrated or described herein.
[0064] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present invention have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the embodiments of the present invention are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.
[0065] An optional step flow of a cigarette recognition method based on knowledge distillation provided by an embodiment of the present invention includes:
[0066] Obtain an image of the displayed cigarettes to be recognized, input the image of the displayed cigarettes to be recognized into a lightweight model for processing, and determine the cigarette brand and specification information of the image of the displayed cigarettes to be recognized; wherein, the cigarette brand and specification information includes a cigarette brand and specification code and a cigarette category, and the lightweight model is determined by Figure 1 the method shown, Figure 1 is an optional flowchart for determining a lightweight model in a cigarette recognition method based on knowledge distillation provided by an embodiment of the present invention. The method may include, but is not limited to, steps S101 to S104.
[0067] Step S101: Process the obtained sample image dataset to determine a cigarette brand and specification image dataset, and determine a training set and a test set according to a first preset ratio and the cigarette brand and specification image dataset;
[0068] Step S102: Construct a first model, train the first model according to a preset dataset to determine a teacher model, and prune the teacher model to determine a student model;
[0069] Step S103: Train the student model according to the teacher model and the training set to determine a second model; test the second model according to the test set to determine the recognition accuracy rate;
[0070] Step S104: Determine a lightweight model according to the recognition accuracy rate and a first preset threshold, and recognize the image of the displayed cigarettes according to the lightweight model.
[0071] In the embodiments of the present invention, steps S101 to S104 shown in the schematic diagram preprocess the collected cigarette sample image dataset, extract the product specifications in the cigarette sample images, label and classify the extracted product specification images to obtain a cigarette product specification image dataset, and divide the cigarette product specification image model into a training set and a test set for subsequent model training; construct a deep learning network model, and train the constructed deep learning network according to the publicly available cigarette image dataset to learn the cigarette image features therein until the recognition accuracy of the deep learning network model meets certain requirements, and use the trained deep learning network model as the teacher model; prune the teacher model according to the preset pruning rate to obtain the student model; train the student model according to the teacher model and the training set, transfer the model parameters of the teacher model and the recognition ability of cigarettes to the student model through knowledge distillation, and use the test set to evaluate the performance of the lightweight model obtained by knowledge distillation. Use the student model that meets the performance requirements as the lightweight model, and apply the lightweight model to a mobile device or an edge device to recognize the cigarette images collected by the mobile device or the edge device, and output the product specification results corresponding to the cigarette images.
[0072] In step S101 of some embodiments, publicly available cigarette images can be collected through the network as the cigarette sample image dataset. It is also possible to collect images of cigarette display cabinets in cigarette retail stores offline to obtain the cigarette sample image dataset, and this is not limited thereto.
[0073] Please refer to Figure 2 , in some embodiments, step S101 may include but is not limited to steps S201 to S203:
[0074] Step S201, extract the sample image data in the sample image dataset to determine the cigarette product specification dataset;
[0075] Step S202, match the cigarette product specification dataset with a preset dictionary to determine the matching result; wherein, the matching result includes product specification code information and cigarette category information;
[0076] Step S203, classify and label the sample image dataset according to the matching result to determine the cigarette product specification image dataset.
[0077] In step S201 of some embodiments, the cigarette product specifications in the cigarette sample images in the sample image dataset are recognized, segmented and extracted through image recognition technology or manually, and the segmented and extracted cigarette product specification images are used as the cigarette product specification dataset for subsequent training of the student model; wherein, the cigarette product specifications include the brand and specifications of cigarettes.
[0078] In step S202 of some embodiments, the obtained cigarette specification data is matched with a preset cigarette specification code to determine the specification code corresponding to the segmented and extracted cigarette specification image and the corresponding specification information, and the cigarette images of different specifications are classified into different folders accordingly.
[0079] In step S203 of some embodiments, the cigarette specification images are labeled according to the specification codes and specification information matched from each cigarette specification image to obtain a cigarette specification image dataset; the training set and the test set are obtained by dividing the cigarette specification image dataset for subsequent training of the student model; in this embodiment, the cigarette specification image dataset is augmented by data augmentation techniques such as random scaling and cropping, and random flipping to provide more data for training the student model, thereby improving the performance of the student model.
[0080] Please refer to Figure 3 , in some embodiments, step S102 may include but is not limited to steps S301 to S303:
[0081] Step S301, determining a model weight set according to the weights of the teacher model and a second preset ratio, and generating a weight mask according to the model weight set;
[0082] Step S302, pruning the teacher model according to the weight mask to determine the pruned teacher model, and training the pruned teacher model according to a preset dataset to determine the updated pruned model;
[0083] Step S303, determining the student model according to the updated pruned model and a second preset threshold.
[0084] In step S301 of some embodiments, the constructed model is pre-trained according to a large-scale dataset publicly available on the network to obtain the teacher model; the weight parameters of the teacher model are obtained, sorted according to the magnitude of the weight values, and a certain proportion of the weights are selected from the sorting results as the pruning targets for model pruning; at the same time, a binary mask for the weights is generated according to the selected model weights for pruning the teacher model.
[0085] In step S302 of some embodiments, the corresponding model weights in the teacher model are set to zero according to the generated binary mask to implement pruning of the teacher model; then, the pruned model is re-trained according to a large-scale dataset publicly available on the network, and it is judged whether the re-trained pruned model meets the set requirements.
[0086] In step S303 of some embodiments, if the pruned model after retraining meets certain requirements, the current pruned model is used as the student model for subsequent lightweight model training; if the requirements are not met, the iterative magnitude pruning method is repeatedly executed on the teacher model until the pruned model meets certain requirements.
[0087] Please refer to Figure 4 , in some embodiments, step S303 may include but is not limited to steps S401 to S403:
[0088] Step S401, determining the model pruning rate according to the updated pruned model and comparing the model pruning rate with a second preset threshold;
[0089] Step S402, if the model pruning rate is greater than or equal to the second preset threshold, using the updated pruned model as the student model;
[0090] Step S403, if the model pruning rate is less than the second preset threshold, using the updated pruned model as the teacher model and returning to execute determining the model weight set according to the weights of the teacher model and the second preset ratio, and generating a weight mask according to the model weight set; until the model pruning rate is greater than or equal to the second preset threshold, using the updated pruned model as the student model.
[0091] In step S401 of some embodiments, determining the pruning rate for the pruned model after retraining, comparing the pruning rate with a preset threshold, and judging whether the number of model parameters and the model complexity of the pruned model meet the requirements for use on a mobile device or an edge device.
[0092] In step S402 of some embodiments, if the pruning rate of the current pruned model is greater than or equal to the preset threshold, it means that the computing resources of the mobile device or the edge device can meet the complexity and the number of parameters of the current pruned model. Using the current pruned model as the student model for subsequent lightweight model training can improve the recognition efficiency and real-time performance of the current pruned model when recognizing on a mobile device or an edge device.
[0093] In step S403 of some embodiments, if the pruning rate of the current pruned model is less than the preset threshold, the computing resources of the mobile device or the edge device still cannot meet the operation of the current pruned model; using the current pruned model as the teacher model and performing iterative magnitude pruning on the teacher model until the pruning rate of the pruned model meets the preset threshold, and using the last pruned model as the student model; in this embodiment, it can be judged according to the pruning rate of the pruned model, or according to the number of pruning times of the pruned model or the decrease in recognition accuracy of the pruned model within a certain range. The embodiments of the present invention are not limited thereto.
[0094] Please refer to Figure 5 In some embodiments, step S103 may include but is not limited to steps S501 to S504:
[0095] Step S501, input the training set into the teacher model for processing to determine the first recognition result; input the training set into the student model for processing to determine the second recognition result;
[0096] Step S502, determine the target loss value according to the first recognition result and the second recognition result, and compare the target loss value with a third preset threshold;
[0097] Step S503, if the target loss value is less than or equal to the third preset threshold, use the student model as the second model;
[0098] Step S504, if the target loss value is greater than the third preset threshold, adjust the parameters of the student model, and return to execute input the training set into the teacher model for processing to determine the first recognition result; input the training set into the student model for processing to determine the second recognition result; until the target loss value is less than or equal to the third preset threshold, use the student model as the second model.
[0099] In step S501 of some embodiments, the training set obtained by proportional division is respectively input into the obtained teacher model and student model. The teacher model processes the input training set and outputs the corresponding recognition result, which reflects the learning result of the teacher model on the cigarette brand characteristics; the student model processes the input training set and outputs the corresponding recognition result, calculates the loss value according to the recognition results of the two different models, and guides the training of the student model according to the loss value.
[0100] In step S502 of some embodiments, calculate the loss value according to the recognition results of the teacher model and the student model, compare the calculated loss value with the preset threshold, and train the student model according to the comparison result; in this embodiment, the loss value of the student model training includes the distillation loss value and the hard discrimination loss value; the distillation loss value represents the loss of transferring the model parameters and feature learning ability of the teacher model to the student model, and the hard discrimination loss value represents the loss between the recognition result of the student model for the input training set and the label or annotation of the input training set.
[0101] In step S503 of some embodiments, if the calculated loss value is less than or equal to the preset threshold, complete the training of the student model, use the test set to test the trained student model, and check whether the performance indicators of the student model meet the requirements. If the performance indicators do not meet the requirements, adjust the training parameters to retrain the second model.
[0102] In step S504 of some embodiments, if the calculated loss value is greater than a preset threshold, it is necessary to adjust the parameters of the student model, retrain according to the student model and the teacher model after parameter adjustment, process the training set to obtain their respective recognition results; calculate the current loss value according to the obtained recognition results, and compare the loss value with the preset threshold until the calculated loss value is less than or equal to the preset threshold. Take the currently trained student model as the second model and test the second model according to the test set.
[0103] Please refer to Figure 6 , in some embodiments, step S502 may include but is not limited to steps S601 to S603:
[0104] Step S601, determine the hard decision result and the soft decision result according to the second recognition result;
[0105] Step S602, calculate according to the soft decision result and the first recognition result to determine the evaporation loss value; calculate according to the hard decision result and the training set to determine the hard discrimination loss value;
[0106] Step S603, perform weighted summation according to the evaporation loss value and the hard discrimination loss value to determine the target loss value.
[0107] In step S601 of some embodiments, analyze and extract according to the recognition result of the student model on the training set to obtain the hard decision result and the soft decision result output by the student model; among them, the hard decision result is the cigarette category and the product specification category output by the student model, and the soft decision result is the probability distribution of the cigarette category and the product specification category output by the student model.
[0108] In step S602 of some embodiments, calculate the evaporation loss value according to the soft decision result of the student model and the recognition result of the teacher model. Take the recognition result output by the teacher model as the supervision information, judge the difference between the student model and the teacher model, and determine the result of the knowledge and feature learning transfer of the teacher model to the student model; in this embodiment, calculate the cross-entropy loss value of the soft decision result of the student model and the recognition result of the teacher model as the distillation loss value. The specific calculation formula is as follows:
[0109]
[0110] Among them, H(p,q) is the cross-entropy loss value, p is the target distribution output by the teacher model, q is the predicted matching distribution output by the student model, n is the number of sample images in the training set, p(x i ) is the target distribution of the teacher model for the i-th sample image, and q(x i ) is the predicted matching distribution of the student model for the i-th sample image.
[0111] Calculate the hard discrimination loss based on the hard decision results output by the student model and the labels annotated for the cigarette images in the training set, determine the difference between the output of the student model and the data in the training set, and determine the learning results of the student model for cigarette features.
[0112] In step S603 of some embodiments, perform weighted calculation based on the calculated distillation loss value and hard discrimination loss value to obtain the loss value for training the student model; adjust the training parameters of the student model according to the calculated loss value, such as the number of training iterations, learning rate, etc.
[0113] Please refer to Figure 7 , in some embodiments, step S104 may include but is not limited to steps S701 to S703:
[0114] Step S701, compare the recognition accuracy rate with a first preset threshold;
[0115] Step S702, if the recognition accuracy rate is greater than or equal to the first preset threshold, use the second model as the lightweight model;
[0116] Step S703, if the recognition accuracy rate is less than the first preset threshold, adjust the parameters of the second model, and return to perform training on the student model according to the teacher model and the training set to determine the second model; until the recognition accuracy rate is greater than the first preset threshold, use the second model as the lightweight model.
[0117] In step S701 of some embodiments, use the trained student model as the second model, input the test set into the second model for processing, evaluate the recognition accuracy rate of the second model, compare the recognition accuracy rate with the preset threshold, determine whether the second model meets the performance requirements, and use the student model that meets the performance requirements as the lightweight model to recognize the displayed cigarette images collected by the mobile device or edge device.
[0118] In step S702 of some embodiments, if the recognition accuracy rate of the second model meets the preset threshold, use the current second model as the lightweight model to perform real-time recognition on the collected displayed cigarette images according to the lightweight model.
[0119] In step S703 of some embodiments, if the recognition accuracy rate of the second model is less than the preset threshold, adjust the training parameters of the second model, and retrain the teacher model and the pruned student model until the recognition accuracy rate of the trained student model meets the preset threshold, and use the trained student model as the lightweight model for real-time recognition of the collected displayed cigarette images.
[0120] Next, in combination with specific application examples, the solutions of the embodiments of the present invention will be introduced and described in detail:
[0121] Please refer to Figure 8 , identify the displayed cigarettes in cigarette retail stores. Select 100 cigarette retail stores within the tobacco management jurisdiction, collect multiple images of the cigarette display cabinets in the selected cigarette retail stores, as shown in Figure 9 . And collect the cigarette product specification information sold within this management jurisdiction, including cigarette categories, the number of product specifications, the types of product specifications, and product specification codes. Generate a digital index dictionary based on the collected cigarette product specification information; retain the image with the highest clarity among the multiple images collected for each cigarette retail store to obtain the original sample atlas; use image recognition technology to segment and extract the original sample atlas, and classify and label it according to the digital index dictionary to collect the cigarette product specification images in the original sample atlas to obtain a cigarette product specification image dataset, as shown in Figures 10(a), 10(b), and 10(c). Among them, Figure 10(a) shows the correspondence between the cigarette product specification codes and cigarette types in the cigarette product specification image dataset, Figure 10(b) shows the classification folders classified according to cigarette types in the cigarette product specification image dataset, and Figure 10(c) shows the cigarette product specification images in the classification folders of the cigarette product specification image dataset; perform data augmentation on the obtained cigarette product specification image dataset using techniques such as random flipping and random scaling and cropping, and divide the augmented dataset according to a ratio of 8:2 to obtain a training set and a test set; based on the cigarette image features, construct a teacher model based on the network structure of Resnet18. The structure of the teacher model is shown in Figure 11(a), including several residual blocks as shown in Figure 12 . Each residual block consists of a convolutional layer, a batch normalization layer, an activation layer, and a residual connection, and pre-train the teacher model on the ImageNet large dataset; through multiple sets of comparative experiments, set different pruning strategies to construct corresponding student models; train and test the student models using the training set and the test set, and record the corresponding accuracy data as shown in the following table
[0122] Table 1
[0123] Group Test set accuracy A 94.92%±0.50% B 95.72%±0.64% C 93.03%±0.78%
[0124] Select Group B as the target pruning strategy according to the recorded data, and draw the corresponding change curve, as shown in Figure 13As shown in the figure, iterative magnitude pruning is performed on the pre-trained teacher model. Some weights in the teacher model are set to zero, and the channels of the intermediate layer of the teacher model are sequentially pruned to 16, 32, 64, and 128 to obtain the student model shown in Fig. 11(b). The divided training set is input into the teacher model and the student model respectively to obtain the target probability distribution output by the teacher model, as shown in Figs. 14(a), 14(b), 14(c), and 14(d), and the predicted probability distribution and recognition results output by the student model, as shown in Figs. 14(e), 14(f), 14(g), and 14(h). The target probability distribution output by the teacher model is used as supervision information to train the teacher model and the student model. Calculate the loss function of the student model according to the output of the teacher model, the output of the student model, and the labels of the training set, and test the teacher model and the student model according to the test set. Record the accuracy of the teacher model during the training process and the test process to obtain the figure shown in Figure 15 As shown, according to Figure 15 the content, as the teacher model is trained and tested, its accuracy increases synchronously. The teacher model approaches convergence and learns the discriminative inter-class features of cigarette images. Through the loss function, the features learned by the teacher model are transferred to the student model by means of knowledge distillation. Compare the parameters of the trained student model with the original teacher model to obtain the following table,
[0125] Table II
[0126] Category Teacher model Student model Accuracy 96.96%±0.38% 95.72%±0.64% Number of parameters 11241150 625166 Model size 42.88MB 2.40MB Computational complexity 1823.59M (FLOPs) 104.49M (FLOPs) Inference time 34.41ms / image 6.50ms / image
[0127] According to the content of Table II, the accuracy of the student model decreases slightly compared with that of the teacher model, but the number of parameters of the student model is reduced by 94.44%, the model size is reduced by 94.40%, the computational complexity is reduced by 94.27%, and the inference time is increased by 27.91 ms / image. Therefore, the student model obtained in the embodiment of the present invention can effectively reduce the model parameters, reduce the computational complexity, and save the inference time, while maintaining the same level of accuracy as the original model. At the same time, compare the comprehensive parameters of the student model obtained in the embodiment of the present invention with those of other models to obtain the following table information,
[0128] Table III
[0129] Model Accuracy F1 - Score VGG16 92.13%±0.29% 92.02%±0.37% VGG19 93.57%±0.98% 93.29%±0.19% InceptionV1 94.44%±0.27% 94.03%±0.11% InceptionV2 95.08%±0.34% 94.46%±0.85% Model of the embodiment of the present invention 95.72%±0.64% 94.82%±0.66%
[0130] According to the information in Table 3, compared with VGG16 and VGG19, the accuracy and F1-Score of the model in the embodiment of the present invention are respectively improved by 3.59%, 2.8% and 2.15%, 1.53%; compared with InceptionV1 and InceptionV2, the accuracy and F1-Score of the model in the embodiment of the present invention are respectively improved by 1.28%, 0.79% and 0.64%, 0.36%.
[0131] Apply the student model obtained in the embodiment of the present invention to the edge model to detect and identify the displayed cigarettes. Compare the student model of the embodiment of the present invention with other lightweight models to obtain the following table information,
[0132] Table 4
[0133] Model Accuracy Model size Inference time MobileNetv2 92.13%±0.29% 4.26MB 15.78ms / image MobileNetv3 93.88%±0.34% 5.43MB 17.49ms / image EfficientNetB0 94.47%±0.21% 5.32MB 16.27ms / image Model of the embodiment of the present invention 95.72%±0.64% 2.40MB 14.56ms / image
[0134] According to the information in Table 4, compared with MobileNetv2, the inference speed of the model in the embodiment of the present invention is increased by 1.22 ms / image, the model size is reduced by 43.66%, and the accuracy is improved by 3.59%; compared with MobileNetv3, the inference speed of the model in the embodiment of the present invention is increased by 2.93 ms / image, the model size is reduced by 55.80%, and the accuracy is improved by 1.84%; compared with EfficientNetB0, the inference speed of the model in the embodiment of the present invention is increased by 1.71 ms / image, the model size is reduced by 54.89%, and the accuracy is improved by 1.25%; Therefore, the model provided by the embodiment of the present invention has better performance than other lightweight models, with high inference efficiency and high accuracy.
[0135] Implementing the embodiments of the present invention includes the following beneficial effects: This embodiment provides a cigarette recognition method, system, electronic device, and storage medium based on knowledge distillation. This solution inputs the display cigarette image to be recognized into a lightweight model for recognition and outputs the corresponding cigarette product specification information. The lightweight model processes the obtained sample image dataset to obtain a cigarette product specification image dataset, divides the cigarette product specification image dataset according to a first preset ratio to obtain a training set and a test set; trains the constructed first model according to a preset dataset to obtain a teacher model; prunes the trained teacher model to construct a student model; trains the student model according to the divided training set and the teacher model to obtain a second model; tests the second model according to the test set to determine the recognition accuracy rate of the second model; determines the lightweight model by comparing the recognition accuracy rate with a first preset threshold, and uses the lightweight model to recognize the image of the display cigarette. By pruning the trained teacher model to construct a student model, the model complexity is reduced, and the model recognition efficiency and real-time performance are improved; the constructed model is trained according to the trained teacher model and the training set, and the knowledge and parameters of the teacher model are transferred to the constructed student model through knowledge distillation, improving the recognition accuracy rate of the student model.
[0136] The embodiments of the present invention also provide a cigarette recognition system based on knowledge distillation, which can implement the above-mentioned cigarette recognition method based on knowledge distillation. The system includes:
[0137] A cigarette recognition module, configured to obtain a display cigarette image to be recognized, input the display cigarette image to be recognized into a lightweight model for processing, and determine the cigarette product specification information of the display cigarette image to be recognized; wherein, the cigarette product specification information includes a cigarette product specification code and a cigarette category, and the lightweight model is determined by the following method:
[0138] Process the obtained sample image dataset to determine a cigarette product specification image dataset, and determine a training set and a test set according to a first preset ratio and the cigarette product specification image dataset;
[0139] Construct a first model, train the first model according to a preset dataset to determine a teacher model, and prune the teacher model to determine a student model;
[0140] Train the student model according to the teacher model and the training set to determine a second model; test the second model according to the test set to determine the recognition accuracy rate;
[0141] Determine the lightweight model according to the recognition accuracy rate and a first preset threshold.
[0142] It can be seen that the content in the above method embodiments is applicable to the system embodiments of the present application. The functions specifically implemented in the system embodiments of the present application are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0143] An embodiment of the present application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned cigarette recognition method based on knowledge distillation. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.
[0144] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented in the device embodiments of the present application are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0145] Please refer to Figure 16 , Figure 16 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0146] A processor 1601, which can be implemented by using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0147] A memory 1602, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1602 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of the present specification through software or firmware, the relevant program codes are stored in the memory 1602 and are called by the processor 1601 to execute a cigarette recognition method based on knowledge distillation in the embodiments of the present application;
[0148] An input / output interface 1603, which is used to implement information input and output;
[0149] A communication interface 1604, which is used to implement communication interaction between the device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.);
[0150] The bus 1605 transmits information among various components of the device (such as the processor 1601, the memory 1602, the input / output interface 1603, and the communication interface 1604).
[0151] Among them, the processor 1601, the memory 1602, the input / output interface 1603, and the communication interface 1604 are communicatively connected to each other inside the device through the bus 1605.
[0152] Among them, the memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. The memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a remote memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0153] In addition, an embodiment of the present application also discloses a computer program product or a computer program. The computer program product or the computer program is stored in a computer-readable storage medium. The processor of the computer device can read the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above method. Similarly, the content in the above method embodiments is applicable to this storage medium embodiment. The functions specifically implemented by this storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0154] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned cigarette recognition method based on knowledge distillation.
[0155] It can be understood that the content in the above method embodiments is applicable to this storage medium embodiment. The functions specifically implemented by this storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0156] It can be understood that all or some of the steps and systems disclosed in the above methods can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0157] The above is a specific description of the preferred embodiment of the present invention, but the present invention is not limited to the described embodiment. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A cigarette identification method based on knowledge distillation, characterized in that: The method comprises: Obtain a displayed cigarette image to be identified, input the displayed cigarette image to be identified into a lightweight model for processing, and determine the cigarette specification information of the displayed cigarette image to be identified; wherein the cigarette specification information includes a cigarette specification code and a cigarette category, and the lightweight model is determined by the following method: Processing the acquired sample image data set to determine a cigarette specification image data set, and determining a training set and a test set according to a first preset ratio and the cigarette specification image data set; Constructing a first model, training the first model according to a preset data set, determining a teacher model, and pruning the teacher model to determine a student model; The student model is trained according to the teacher model and the training set to determine a second model; the second model is tested according to the test set to determine the recognition accuracy; A lightweight model is determined according to the recognition accuracy and a first preset threshold.
2. The method according to claim 1, characterized in that The step of processing the acquired sample image data set to determine the cigarette specification image data set specifically includes: Extracting sample image data from the sample image data set to determine a cigarette specification data set; Matching the cigarette specification data set with a preset dictionary to determine a matching result; wherein the matching result includes specification code information and cigarette category information; The sample image dataset is classified and labeled according to the matching result to determine the cigarette specification image dataset.
3. The method according to claim 1, characterized in that The pruning of the teacher model to determine the student model specifically includes: Determining a model weight set according to the weight of the teacher model and a second preset ratio, and generating a weight mask according to the model weight set; Pruning the teacher model according to the weight mask to determine a pruned teacher model, and training the pruned teacher model according to the preset data set to determine an updated pruned model; The student model is determined according to the updated pruning model and a second preset threshold.
4. The method according to claim 3, characterized in that Determining the student model according to the updated pruning model and the second preset threshold specifically includes: Determine a model pruning rate according to the updated pruning model, and compare the model pruning rate with the second preset threshold; If the model pruning rate is greater than or equal to the second preset threshold, using the updated pruning model as the student model; If the model pruning rate is less than the second preset threshold, the updated pruned model is used as the teacher model, and the process returns to the step of determining the model weight set according to the weight of the teacher model and the second preset ratio, and generating a weight mask according to the model weight set; until the model pruning rate is greater than or equal to the second preset threshold, the updated pruned model is used as the student model.
5. The method according to claim 1, characterized in that The step of training the student model according to the teacher model and the training set to determine the second model specifically includes: Input the training set into the teacher model for processing to determine a first recognition result; input the training set into the student model for processing to determine a second recognition result; Determine a target loss value according to the first recognition result and the second recognition result, and compare the target loss value with a third preset threshold; If the target loss value is less than or equal to the third preset threshold, use the student model as the second model; If the target loss value is greater than the third preset threshold, the parameters of the student model are adjusted, and the process returns to the step of inputting the training set into the teacher model for processing to determine the first recognition result; inputting the training set into the student model for processing to determine the second recognition result; until the target loss value is less than or equal to the third preset threshold, the student model is used as the second model.
6. The method according to claim 5, characterized in that The determining of the target loss value according to the first recognition result and the second recognition result specifically includes: Determine a hard decision result and a soft decision result according to the second recognition result; Calculate according to the soft decision result and the first recognition result to determine the evaporation loss value; calculate according to the hard decision result and the training set to determine the hard discrimination loss value; The target loss value is determined by performing a weighted summation on the evaporation loss value and the hard discrimination loss value.
7. The method according to claim 1, characterized in that The step of determining the lightweight model according to the recognition accuracy and the first preset threshold specifically includes: Comparing the recognition accuracy with the first preset threshold; If the recognition accuracy is greater than or equal to the first preset threshold, using the second model as the lightweight model; If the recognition accuracy rate is less than the first preset threshold, the parameters of the second model are adjusted, and the process returns to execute the training of the student model according to the teacher model and the training set to determine the second model; until the recognition accuracy rate is greater than the first preset threshold, the second model is used as the lightweight model.
8. A cigarette identification system based on knowledge distillation, characterized in that: include: The cigarette recognition module is used to obtain a displayed cigarette image to be identified, input the displayed cigarette image to be identified into a lightweight model for processing, and determine the cigarette specification information of the displayed cigarette image to be identified; wherein the cigarette specification information includes a cigarette specification code and a cigarette category, and the lightweight model is determined by the following method: Processing the acquired sample image data set to determine a cigarette specification image data set, and determining a training set and a test set according to a first preset ratio and the cigarette specification image data set; Constructing a first model, training the first model according to a preset data set, determining a teacher model, and pruning the teacher model to determine a student model; The student model is trained according to the teacher model and the training set to determine a second model; the second model is tested according to the test set to determine the recognition accuracy; A lightweight model is determined according to the recognition accuracy and a first preset threshold.
9. An electronic device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to perform the method according to any one of claims 1 to 7 when executed by the processor.
Citation Information
Patent Citations
Model training method and device, equipment and storage medium
CN116976428A