Explanatable few-sample cut tobacco defect detection method
By constructing an interpretable method of detecting defects of tobacco defects, using ResNet-50 and channel attention mechanism, combining the metric learning module and abnormal score, the data dependence and environmental sensitivity problems in the detection of tobacco defects are solved, and high-precision, robustness and transparency are achieved.
Patent Information
- Application Number
- CN202510423028.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-08
AI Technical Summary
The existing tobacco defect detection technology has the problems of strong data dependence, high environmental sensitivity, poor generalization of small samples and lack of interpretability, especially in actual production, it is difficult to achieve efficient, accurate and transparent defect detection.
Using an interpretable method of detecting defects of tobacco wire with few samples, the similar/heterogeneous sample pair is constructed, and features are extracted using the ResNet-50 backbone network and channel attention submodule, combined with the metric learning module and the abnormal scoring module, a joint optimization loss function is constructed to achieve high-precision detection under the conditions of few samples.
It significantly reduces the labeling cost, improves the robustness and generalization capabilities of the model in complex lighting environments, provides a transparent decision-making process, and meets the needs of industrial quality inspection.
Smart Images

Figure CN120279358A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of industrial defect detection, and specifically relates to an interpretable few-shot tobacco cut filler defect detection method. Background Art
[0002] The existing tobacco cut filler defect detection technologies face many challenges in practical applications, mainly manifested in the following aspects:
[0003] Strong data dependence: Traditional deep learning methods require a large number of defect samples to be labeled for tobacco cut filler defect detection in order to train the model to identify various defect types. However, in actual production, defect samples are often scarce and difficult to obtain, resulting in extremely high labeling costs. Although some deep learning-based detection methods can theoretically achieve high-precision defect detection, due to the lack of sufficient labeled data, the training effect of the model is limited.
[0004] High environmental sensitivity: Tobacco cut filler defect detection is relatively sensitive to environmental conditions, especially changes in light intensity. Unstable lighting conditions will cause deviations in image feature extraction, thus affecting the stability of detection. For example, under different lighting conditions, the color and texture of tobacco cut filler may change, making it difficult for the model to accurately identify defects.
[0005] Poor few-shot generalization: In few-shot scenarios, existing metric learning methods are prone to overfitting. This means that the model performs well on the training data but poorly on new, unseen data, and it is difficult to accurately distinguish subtle differences in tobacco cut filler defects. This is a serious problem in actual production because the types and forms of tobacco cut filler defects may change with production conditions.
[0006] Lack of interpretability: Most existing deep learning models are black-box models and it is difficult to explain their decision-making processes. In industrial quality inspection, decision transparency is crucial because quality inspectors need to clarify the characteristics of tobacco cut filler defects for subsequent processing and improvement. However, black-box models cannot provide this transparency, resulting in low acceptability in actual applications. Summary of the Invention
[0007] The present invention is to solve the above-mentioned deficiencies of the existing technologies, and proposes an interpretable few-shot tobacco cut filler defect detection method, in order to achieve efficient detection of tobacco cut filler defects using few-shot learning, and improve the decision transparency of the model through an interpretable metric learning mechanism, so as to reduce the data labeling cost and improve the reliability of defect recognition in the industrial quality inspection scenario.
[0008] The present invention adopts the following technical solutions to achieve the above-mentioned invention purposes:
[0009] The characteristics of an interpretable few-shot tobacco cut filler defect detection method according to the present invention include the following steps:
[0010] Step 1: Obtain a set of tobacco cut filler image samples with labels , where: represents the i-th tobacco cut filler image sample, represents the i-th tobacco cut filler image, represents the true label of , and represents normal tobacco cut filler, represents defective tobacco cut filler; is the total number of samples;
[0011] Judge whether the true label of the i-th tobacco cut filler image is the same as the true label of the j-th tobacco cut filler image . If they are the same, then and are used as a pair of similar samples, and their common label and is set. Otherwise, and are used as a pair of dissimilar samples; and their common label is set; ;
[0012] Step 2: Construct an interpretable few-shot tobacco cut filler defect detection network, including: a feature extraction module, a feature reduction module, a metric learning module, and an anomaly scoring module;
[0013] Step 2.1: The feature extraction module is composed of a ResNet-50 backbone network and a channel attention sub-module, and processes and respectively, and correspondingly obtains the i-th enhanced feature that eliminates light intensity sensitivity and the j-th enhanced feature ;
[0014] Step 2.2: The feature reduction module is composed of a reduction network and a spatial average pooling module, and processes and respectively, and correspondingly obtains the i-th refined feature that retains channel characteristics and the j-th refined feature that retains channel characteristics;
[0015] Step 2.3: The metric learning module processes and , and obtains and the cosine similarity And the margin :
[0016] Step 2.4, the abnormal scoring module calculates The mean similarity with the normal feature library ;
[0017] Step 3, based on the cosine similarity and the margin as well as the true label and Construct a jointly optimized total loss function L;
[0018] Step 4, use the Adam optimizer to minimize the total loss function L, thereby updating the parameters of the interpretable few-shot tobacco cut filler defect detection network until the preset maximum number of iterations or the total objective loss L converges, so as to obtain the trained optimal interpretable few-shot tobacco cut filler defect detection model for defect detection and recognition of the input tobacco cut filler image dataset.
[0019] The characteristics of an interpretable few-shot tobacco cut filler defect detection method according to the present invention also lie in that step 2.1 includes the following steps:
[0020] Step 2.1.1, the ResNet-50 backbone network extracts and processes to output the i-th initial feature map , where H represents the length of the initial feature map, W represents the width of the initial feature map, and C represents the number of channels of the initial feature map;
[0021] Step 2.1.2, the channel attention sub-module processes to generate the i-th channel weight matrix , where , are the parameters to be trained, GAP represents the global average pooling operation, is the Sigmoid function;
[0022] Step 2.1.3, the channel attention sub-module performs channel-wise weighting on and to obtain the i-th enhanced feature , where , represents channel-wise multiplication;
[0023] Step 2.1.4, the channel attention sub-module performs L2 normalization on to obtain the i-th enhanced feature that eliminates the sensitivity to light intensity, and ;
[0024] Step 2.1.5. Process according to the process of Step 2.2.1 - Step 2.2.4 to generate the j-th enhanced feature that eliminates light intensity sensitivity .
[0025] Furthermore, Step 2.2 includes the following steps:
[0026] Step 2.2.1. The reduction network processes to obtain the i-th refined feature , where , are two weight matrices, , are two bias terms; is an activation function;
[0027] Step 2.2.2. The spatial average pooling module performs spatial average pooling on to obtain the i-th refined feature that retains channel characteristics ;
[0028] Step 2.2.3. Process according to the process of Step 2.3.1 - Step 2.3.2 to to obtain the j-th refined feature that retains channel characteristics .
[0029] Furthermore, Step 2.3 includes the following steps:
[0030] Step 2.3.1. Calculate the cosine similarity of and as , where represents transpose;
[0031] Step 2.3.2. If , then generate the margin between , , is the dynamic margin preset coefficient;
[0032] If , set to generate the margin between , , is the fixed margin value.
[0033] Furthermore, Step 2.4 includes the following steps:
[0034] Step 2.4.1. Set The refined features of the reserved channel characteristics corresponding to all the tobacco cuttings images with the true label of "0" in [the set] constitute the normal feature library ;
[0035] Step 2.4.2, calculate and mean similarity , where represents the refined feature of the reserved channel characteristics corresponding to the k-th tobacco cuttings image in [the set], and M represents the number of refined features in [the set].
[0036] Furthermore, the said Step 3 includes the following steps:
[0037] Step 3.1, construct the metric learning loss L1 using Equation (1):
[0038] (1)
[0039] In Equation (1), is the margin buffer threshold;
[0040] Step 3.2, construct the anomaly score loss L2 using Equation (2):
[0041] (2)
[0042] Step 3.3, construct the objective loss function L using Equation (3):
[0043] (3)
[0044] In Equation (3), is the hyperparameter for balancing the contributions of the two types of losses, .
[0045] An electronic device according to the present invention, comprising a memory and a processor, is characterized in that the memory is used to store a program for supporting the processor to execute the few-shot tobacco cuttings defect detection method according to any one of claims 1-6, and the processor is configured to execute the program stored in the memory.
[0046] A computer-readable storage medium according to the present invention, on which a computer program is stored, is characterized in that the computer program, when run by a processor, executes the steps of the few-shot tobacco cuttings defect detection method according to any one of claims 1-6.
[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0048] 1. Through the mechanism of dynamically generating homogeneous / heterogeneous tobacco cut samples, the present invention expands a single sample into multiple groups of training pairs, constructs similarity constraints using relationship labels between samples, breaks through the dependence on a large amount of labeled data in traditional methods, significantly reduces the labeling cost, and achieves high-precision tobacco cut defect detection under few-sample conditions.
[0049] 2. Through the channel attention mechanism and L2 normalization processing, the present invention strengthens the key feature channels by adaptively allocating channel weights, combines L2 normalization to eliminate the interference of illumination intensity differences on the extraction of tobacco cut features, improves the robustness of the model in complex illumination environments, and effectively reduces the false detection rate.
[0050] 3. Through the dynamic margin metric learning strategy, the present invention sets a dynamic margin for homogeneous sample pairs and a fixed margin for heterogeneous sample pairs, constrains the distribution of the feature space, enhances the generalization ability of the model in small-sample scenarios, and accurately distinguishes subtle differences in tobacco cut defects.
[0051] 4. Through an interpretable anomaly scoring module, the present invention constructs a normal feature library and calculates the mean similarity, quantifies the probability of tobacco cut defects, makes the model decision-making process transparent and traceable, and meets the strict requirements of industrial quality inspection for interpretability. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a flowchart of an interpretable few-shot tobacco cut defect detection method of the present invention;
[0053] Figure 2 It is a flowchart of the feature extraction module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] In this embodiment, as Figure 1 shown, an interpretable few-shot tobacco cut defect detection method is carried out according to the following steps:
[0055] Step 1: Obtain a set of tobacco cut image samples with labels , where: represents the i-th tobacco cut image sample, represents the i-th tobacco cut image, represents 's true label, and , represents normal tobacco cut, represents defective tobacco cut; is the total number of samples;
[0056] Judge the true label of the i-th tobacco cut image and the true label of the j-th tobacco cut image Whether they are the same. If so, then and are used as a pair of similar samples, and their common label is set. Otherwise, and are used as a pair of dissimilar samples; and their common label is set; by constructing pairs of similar / dissimilar samples, the problem of insufficient data in the few-shot scenario is solved; a single sample is extended to multiple pairs of samples (similar or dissimilar), enriching the training information without increasing the amount of annotation and improving the data utilization rate.
[0057] Step 2. Construct an interpretable few-shot cut tobacco defect detection network, including: a feature extraction module, a feature reduction module, a metric learning module, and an anomaly scoring module;
[0058] Step 2.1. The feature extraction module is composed of a ResNet-50 backbone network and a channel attention sub-module, as Figure 2 shown, and and are processed respectively, and the i-th enhanced feature and the j-th enhanced feature that eliminate the sensitivity to light intensity are obtained respectively;
[0059] Step 2.1.1. The ResNet-50 backbone network extracts and processes , and outputs the i-th initial feature map , where H represents the length of the initial feature map, W represents the width of the initial feature map, and C represents the number of channels of the initial feature map; the deep structure of ResNet-50 can capture details such as cut tobacco texture and edges, which is suitable for detecting tiny defects;
[0060] Step 2.1.2. The channel attention sub-module processes , and generates the i-th channel weight matrix , where , are parameters to be trained, GAP represents the global average pooling operation, is the Sigmoid function; by adaptively allocating weights the channels sensitive to defects (such as broken textures) are strengthened, and the irrelevant channels (such as reflective areas) are suppressed, improving the robustness of the model to light changes and avoiding false detections.
[0061] Step 2.1.3. The channel attention sub-module performs channel-wise weighting on and , and obtains the i-th enhanced feature , where , Indicates per-channel multiplication;
[0062] Step 2.1.4. The channel attention sub-module is subjected to L2 normalization to obtain the i-th enhanced feature that eliminates the sensitivity to light intensity , and ; the feature vector is constrained to unit length to eliminate the fluctuation of the feature modulus length caused by the difference in illumination intensity; ensure that the similarity calculation only depends on the feature direction, rather than the illumination intensity;
[0063] Step 2.1.5. Process according to the process of Step 2.2.1 - Step 2.2.4 to generate the j-th enhanced feature that eliminates the sensitivity to light intensity .
[0064] Step 2.2. The feature reduction module consists of a reduction network and a spatial average pooling module, and processes and respectively to obtain the i-th refined feature that retains the channel characteristics and the j-th refined feature that retains the channel characteristics ;
[0065] Step 2.2.1. The reduction network processes to obtain the i-th refined feature , where , are two weight matrices, , are two bias terms; is an activation function; compress high-dimensional features (such as 2048 dimensions) to low dimensions (such as 256 dimensions), reduce the subsequent calculation amount; reduce memory consumption, improve the inference speed, and at the same time retain the separability through non-linear activation (ReLU).
[0066] Step 2.2.2. The spatial average pooling module performs spatial average pooling on to obtain the i-th refined feature that retains the channel characteristics ; take the mean along the spatial dimension to eliminate the sensitivity to the feature position; enable the model to focus on global attributes (such as color, texture), rather than the local position of the defect;
[0067] Step 2.2.3. Process according to the process of Step 2.3.1 - Step 2.3.2 to obtain the j-th refined feature that retains the channel characteristics .
[0068] Step 2.3. The metric learning module processes and to obtain and Margin between :
[0069] Step 2.3.1, calculate and cosine similarity of , where represents transpose;
[0070] Step 2.3.2, if , then generate and margin between , is the preset coefficient of dynamic margin;
[0071] If , set to generate and margin between , is the fixed value of margin; During the feature learning process, for similar features, a loose distribution is allowed initially to ensure that the model can fully learn the diversity of features, and then the distribution range is gradually tightened to enhance the compactness and consistency of features and avoid overfitting of the model; For dissimilar features, especially the features of defects and normal samples, a strategy of forced separation is adopted to ensure that the difference between the two is obvious, thereby improving the defect recognition ability and discrimination of the model.
[0072] Step 2.4, processing of the anomaly scoring module:
[0073] Step 2.4.1, compose the refined features of the reserved channel characteristics corresponding to all cut tobacco images with true label "0" in into the normal feature library ; Aggregate the refined features of all normal samples as the benchmark for defect determination;
[0074] Step 2.4.2, calculate and mean similarity of , where represents the refined feature of the reserved channel characteristics corresponding to the k-th cut tobacco image in , M represents the number of refined features in The value also provides the top-3 most similar normal samples, and quality inspectors can intuitively understand the basis for the model's determination by comparing these similar feature samples. This design directly quantifies the probability of cut tobacco defects through numerical values. For example, when is determined to be a defect, the normal sample with the highest similarity is output at the same time to assist manual re-inspection, thereby realizing a transparent decision-making process.
[0075] Step 3. Construct the overall loss function L of joint optimization:
[0076] Step 3.1. Use Equation (1) to construct the metric learning loss L1:
[0077] (1)
[0078] In Equation (1), is the margin buffer threshold; through the margin buffer threshold , the feature similarity is allowed to fluctuate within a reasonable range; overfitting to samples near the boundary is avoided, and the generalization of the model is improved.
[0079] Step 3.2. Use Equation (2) to construct the anomaly score loss L2:
[0080] (2)
[0081] Step 3.3. Use Equation (3) to construct the objective loss function L:
[0082] (3)
[0083] In Equation (3), is a hyperparameter that balances the contributions of the two types of losses, .
[0084] Step 4. Use the Adam optimizer to minimize the overall loss function L, thereby updating the parameters of the interpretable few-shot cut tobacco defect detection network until the preset maximum number of iterations is reached or the overall objective loss L converges, so as to obtain the trained optimal interpretable few-shot cut tobacco defect detection model for defect detection and recognition of the input cut tobacco image dataset.
[0085] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0086] In this embodiment, a computer-readable storage medium stores a computer program on the computer-readable storage medium. When the computer program is run by a processor, it executes the steps of the above method.
Claims
1. An interpretable few-shot tobacco cut filler defect detection method, characterized in that, It includes the following steps: Step 1: Obtain a set of tobacco cuttings image samples with labels , where: represents the i-th tobacco cuttings image sample, represents the i-th tobacco cuttings image, represents the true label of, and , represents normal tobacco cuttings, represents defective tobacco cuttings; is the total number of samples; Judge the true label of the i-th cut tobacco image and the true label of the j-th cut tobacco image to see if they are the same. If they are the same, then take and as a pair of samples of the same category, and set their common label and . Otherwise, take and and as a pair of samples of different categories; and set their common label ; Step 2: Construct an interpretable few-shot tobacco cut filler defect detection network, including: a feature extraction module, a feature reduction module, a metric learning module, and an anomaly scoring module; Step 2.1: The feature extraction module consists of a ResNet-50 backbone network and a channel attention sub-module, and processes and respectively, and correspondingly obtains the i-th enhanced feature and the j-th enhanced feature ; Step 2.
2. The feature reduction module is composed of a reduction network and a spatial average pooling module, and processes and respectively, and correspondingly obtains the refined feature with the characteristics of the i-th retained channel and the refined feature with the characteristics of the j-th retained channel; Step 2.3, the metric learning module processes and to obtain the cosine similarity and and the margin : : Step 2.4, the anomaly scoring module calculates with the normal feature library mean similarity ; Step 3: Based on the cosine similarity and the margin as well as the true label and construct the jointly optimized total loss function L; Step 4: Use the Adam optimizer to minimize the total loss function L, thereby updating the parameters of the interpretable few-shot tobacco cut filler defect detection network until the preset maximum number of iterations is reached or the total target loss L converges, so as to obtain the trained optimal interpretable few-shot tobacco cut filler defect detection model for defect detection and recognition of the input tobacco cut filler image dataset.
2. The interpretable few-shot tobacco cut filler defect detection method according to claim 1, wherein Step 2.1 includes the following steps: Step 2.1.1, the ResNet-50 backbone network extracts and processes to output the i-th initial feature map , where H represents the length of the initial feature map, W represents the width of the initial feature map, and C represents the number of channels of the initial feature map; Step 2.1.
2. The channel attention sub-module processes to generate the i-th channel weight matrix , where , are the parameters to be trained, GAP represents the global average pooling operation, is the Sigmoid function; Step 2.1.3, the channel attention sub-module performs channel-wise weighting on and to obtain the i-th enhanced feature , where , represents channel-wise multiplication; Step 2.1.4, the channel attention sub-module performs L2 normalization processing to obtain the i-th enhanced feature that eliminates the sensitivity to light intensity , and ; Step 2.1.
5. Process according to the process of Step 2.2.1 - Step 2.2.4 to generate the j-th enhanced feature that eliminates light intensity sensitivity for . to generate the j-th enhanced feature that eliminates light intensity sensitivity .
3. The interpretable few-shot tobacco cut filler defect detection method according to claim 2, characterized in that, Step 2.2 includes the following steps: Step 2.2.1, the reduction network processes to obtain the i-th refined feature , where , are two weight matrices, , are two bias terms; is an activation function; Step 2.2.2, the spatial average pooling module performs spatial average pooling to obtain the refined feature of the i-th retained channel characteristic ; Step 2.2.3: Process according to the process of Step 2.3.1 - Step 2.3.2 to obtain the refined feature of the j-th retained channel characteristic . to obtain the refined feature of the j-th retained channel characteristic .
4. An interpretable few-shot tobacco cut filler defect detection method according to claim 3, characterized in that Step 2.3 includes the following steps: Step 2.3.1, calculate and cosine similarity , where represents transpose; Step 2.3.
2. If , then generate the margin between , where is the dynamic margin preset coefficient; If , set the margin and to a fixed margin value . 5. The interpretable few-shot tobacco cut filler defect detection method according to claim 4, wherein Said Step 2.4 includes the following steps: Step 2.4.
1. Compose a normal feature library from the refined features that retain channel characteristics corresponding to all tobacco leaf images with a true label of "0" in ; Step 2.4.2, calculate and mean similarity , where represents the refined feature of the retained channel characteristics corresponding to the k-th cut tobacco image in , and M represents the number of refined features in 6. The interpretable few-shot tobacco cut filler defect detection method according to claim 5, wherein Said Step 3 includes the following steps: Step 3.1: Construct the metric learning loss L1 using Equation (1): (1) In formula (1), is the margin buffer threshold; Step 3.2: Construct the anomaly scoring loss L2 using Equation (2): (2) Step 3.3: Construct the target loss function L using Equation (3): (3) In formula (3), is a hyperparameter that balances the contribution of the two types of losses, .
7. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor to execute the few-shot tobacco cut filler defect detection method described in any one of Claims 1-6, and the processor is configured to execute the program stored in the memory.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it executes the steps of the few-shot tobacco cut filler defect detection method described in any one of Claims 1-6.