A method for detecting stem tags based on X-ray vision
Through the X-ray vision-based mesh detection method, the generative adversarial network and mesh classification network are used to solve the problems of low efficiency and low accuracy of the existing cigarette products, and efficient and accurate mesh detection is achieved.
Patent Information
- Application Number
- CN202210179428.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-25
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-02-25
AI Technical Summary
The current cigarette products have low efficiency and low accuracy, resulting in high error detection rates and missed detection rates, affecting the combustion properties of the cigarette branches and the quality of the suction sensor.
Using the X-ray vision-based staple detection method, cigarette fluoroscopic images are collected through X-ray equipment, pseudo-labeled samples are generated using a generative adversarial network, and the final expanded samples are screened through screening indicators, which are used to train the staple classification network, and finally the cigarette samples are detected.
It improves the efficiency of meme detection, reduces the false detection rate and missed detection rate, and improves the accuracy and reliability of detection.
Smart Images

Figure CN114549485B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cigarette product detection, and in particular to a method for detecting stem pieces based on X-ray vision. Background Art
[0002] During the silk-making process, due to the influence of special physical properties and processing technology factors, stem pieces with larger widths or lengths are likely to appear. A stem piece refers to a tobacco stem in tobacco leaves that is similar in shape to a toothpick and has not expanded or has an expansion effect that does not meet the requirements for rolling. Stem pieces in cigarettes will increase the off-flavor and irritation, cause punctures and air leaks, and may cause phenomena such as the burning end bursting or extinguishing during combustion, which not only affects the combustibility of the cigarette, but also affects the sensory quality of smoking. At the same time, during the production and processing process, it will cause an increase in the weight deviation of the cigarette, affect the stability of the physical indicators of the cigarette, is not conducive to quality control and the normal operation of the equipment, and affects the equipment efficiency and various material consumption indicators. Currently, the detection method for stem pieces in cigarettes generally adopts the method of manual spot-checking, that is, using a blade to cut each cigarette one by one, stripping the tobacco leaves, and then visually inspecting the tobacco leaves. On the one hand, this method has low detection efficiency, and on the other hand, the accuracy is reduced due to the judgment of human factors. Therefore, how to automatically and accurately detect the presence of stem pieces in cigarettes to improve the detection efficiency of stem pieces and reduce the false detection rate or missed detection rate of tobacco stem detection is of great significance. Summary of the Invention
[0003] The present invention provides a method for detecting stem pieces based on X-ray vision, which solves the problems of low efficiency and low accuracy in the detection of stem pieces in existing cigarette products, and can improve the detection efficiency of stem pieces and reduce the false detection rate or missed detection rate of tobacco stem detection.
[0004] To achieve the above objectives, the present invention provides the following technical solutions:
[0005] A method for detecting stem pieces based on X-ray vision, comprising:
[0006] Randomly select cigarettes of different brands as detection objects, irradiate the detection objects with X-rays using an X-ray device, and obtain corresponding cigarette fluoroscopic images;
[0007] Use a generative adversarial network to generate multiple groups of pseudo-labeled samples from the cigarette fluoroscopic images, and screen the pseudo-labeled samples according to screening criteria to determine the final extended labeled samples;
[0008] Obtain the manually labeled samples of the detection objects, input the extended labeled samples into a preset stem piece classification network for pre-training, and use the manually labeled samples to adjust the training network;
[0009] Use the trained stem piece classification network to detect stem pieces in the tested cigarette samples.
[0010] Preferably, it further includes:
[0011] Taking the overall classification accuracy as the evaluation index of the trained meme tag classification network, the overall classification accuracy is calculated according to the formula where OA is the overall classification accuracy, N is the total number of samples, and Z is the number of correctly classified samples.
[0012] Preferably, it further includes:
[0013] Taking the screening index as the evaluation index of the trained meme tag classification network, the screening index is calculated according to the formula SDF n =αNFID n +βTR n , n ∈ [0, N], where SDF n is the evaluation score of the nth group of pseudo-labeled samples generated, NFID n ∈ [0, 1] is the normalized FID score, and TR n ∈ [0, 1] is the normalized training evaluation score, α is the weight coefficient of NFID n and β is the weight coefficient of TR n , and α + β = 1.
[0014] Preferably, a SinGAN model based on an improved loss function is used to generate multiple groups of the pseudo-labeled samples, and sample training is performed based on the SinGAN model. The loss function of the discriminator of the SinGAN model is:
[0015] The loss function of the generator of the SinGAN model is:
[0016] where D n is the nth discriminator, G n is the nth generator, χ is the joint sampling space of x n and , is the gradient penalty term, μ is the weight coefficient, is the pseudo-image generated by the (n + 1)th generator, is the pseudo-image generated by the nth generator, x n is the corresponding real image at each scale, z * is the randomly selected value before training, and L div is the ratio of the distance between the generated images to the distance between the noises.
[0017] Preferably, screening the pseudo-labeled samples according to the screening index includes:
[0018] After the training of the SinGAN model is completed, N + 1 groups of pseudo-images {PS N ,…,PS n ,…,PS 0} are generated. The SinGAN model generates the pseudo-labeled samples according to the pseudo-images.
[0019] According to the selected screening index SDF n evaluate the authenticity and diversity of the pseudo-labeled samples, and use the formula to evaluate the quality of the generated images, where F RS and represent the average of the feature vectors of the real image RS and the nth group of generated images PS n respectively, CF RS and represent the covariance matrices calculated from the feature vectors of RS and PS n respectively, and Tr(·) represents the trace of the matrix.
[0020] Preferably, the use of the trained stem label classification network to detect the stem labels of the tested cigarette samples includes:
[0021] Predict and classify the target samples through a Softmax classifier containing a focal loss function, and determine the loss function value according to the formula FL(p i )=-α i (1 - p i ) γ log(p i ), where FL(p i ) is the loss function value, p i represents the probability that the model predicts that the sample contains stem labels, γ represents the hyperparameter controlling "focus", and α i represents the contribution of positive and negative samples to the total loss.
[0022] Preferably, it further includes:
[0023] Select cigarettes of different brands as the detection objects and make the collected data into a dataset, and detect and verify the trained stem label classification network according to the dataset to determine whether the final detection rate of the stem labels reaches the set threshold. If so, the training of the stem label classification network is qualified.
[0024] Preferably, the steps of making the dataset include:
[0025] First, irradiate the cigarettes with an X-ray device to obtain cigarette fluoroscopic images and manually annotate them. 100 images are collected for each brand of cigarettes, with 50 images with stem labels and 50 images without stem labels, for a total of 2000 images;
[0026] Then, the trained SinGAN model is used to perform data augmentation on the images with and without stem tags for each category at training rates of 20% and 50% respectively, and the remaining 80% and 50% are used for subsequent testing;
[0027] The image augmentation ratio is 1:20. When the training rate is 20%, the augmented dataset has 20 categories, with 400 images in each category, for a total of 8000 images;
[0028] When the training rate is 50%, the augmented dataset has 20 categories, with 1000 images in each category, for a total of 20000 images.
[0029] Preferably, the use of the X-ray device to irradiate the detection object with X-rays and obtain the corresponding cigarette fluoroscopic image includes:
[0030] Based on the difference in the X-ray transmission imaging characteristics of tobacco shreds and stem tags, X-ray transmission imaging forms an X-ray black and white image;
[0031] Perform noise filtering preprocessing on the X-ray black and white image to remove background noise in the image;
[0032] Use the region growing method to perform image segmentation on the preprocessed X-ray black and white image with noise filtering to segment it into stem pixels and background pixels;
[0033] Adopt the fuzzy C-clustering algorithm to obtain the membership degree of each segmented stem pixel pair to the stem center, so as to filter out interference information and make attribution judgments;
[0034] After the fuzzy C-clustering algorithm is processed, perform shape analysis on the pixels gathered together, and calculate the area and aspect ratio of the segmented region to perform shape recognition.
[0035] The present invention provides a stem tag detection method based on X-ray vision. Data of the detection cigarette is collected through an X-ray device and manually labeled; then, a generative adversarial network is used to augment the data, and the proposed screening index is used to screen the generated samples to determine the final augmented samples. Finally, the augmented samples are used for the training of the classification network, and the trained network is further fine-tuned with real samples. It solves the problems of low efficiency and low accuracy in the stem tag detection of existing cigarette products, can improve the stem tag detection efficiency, and reduce the false detection rate or missed detection rate of stem detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the specific embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below.
[0037] Figure 1This is a schematic diagram of a method for detecting stem tags based on X-ray vision provided by the present invention.
[0038] Figure 2 This is a schematic diagram of the stem tag detection process provided by the present invention. Specific implementation manner
[0039] In order to enable those skilled in the art to better understand the solutions of the embodiments of the present invention, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings and implementation manners.
[0040] Aiming at the problems of low efficiency and low accuracy in the detection of stem tags in current cigarettes, the present invention provides a method for detecting stem tags based on X-ray vision, which solves the problems of low efficiency and low accuracy in the detection of stem tags in existing cigarette products, can improve the efficiency of stem tag detection, and reduce the false detection rate or missed detection rate of tobacco stems.
[0041] As Figure 1 and Figure 2 shown, a method for detecting stem tags based on X-ray vision includes:
[0042] S1: Randomly select cigarettes of different brands as the detection objects, irradiate the detection objects with X-rays using an X-ray device, and obtain corresponding cigarette fluoroscopic images.
[0043] S2: Use a generative adversarial network to generate multiple groups of pseudo-labeled samples from the cigarette fluoroscopic images, and screen the pseudo-labeled samples according to screening indicators to determine the final expanded labeled samples.
[0044] S3: Obtain the manually labeled samples of the detection objects, input the expanded labeled samples into a preset stem tag classification network for pre-training, and use the manually labeled samples to adjust the training network.
[0045] S4: Use the trained stem tag classification network to detect stem tags in the tested cigarette samples.
[0046] Specifically, cigarettes of different brands are randomly selected as the objects to be inspected. First, industrial X-ray equipment is used to collect data for each cigarette one by one, obtaining the original perspective images and numbering them one by one; then the cigarettes are manually peeled to detect whether they contain stem tags, and the perspective images corresponding to each cigarette are manually labeled according to the results. However, since the collection of original data is time-consuming and laborious, a generative adversarial network (GAN) is used to generate a batch of pseudo-labeled samples that can be used for subsequent network training based on the original images; then the proposed screening method is used to screen the generated data to determine the final augmented samples; finally, the screened augmented samples are used for the training of the subsequent stem tag classification network. This method can solve the problems of low efficiency and low accuracy in the stem tag detection of existing cigarette products, improve the stem tag detection efficiency, and reduce the false detection rate or missed detection rate of tobacco stems.
[0047] The method further includes: taking the overall classification accuracy as the evaluation index of the trained stem tag classification network, and the overall classification accuracy is calculated according to the formula where OA is the overall classification accuracy, N is the total number of samples, and Z is the number of samples correctly classified in all categories.
[0048] The method further includes: taking the screening index as the evaluation index of the trained stem tag classification network, and the screening index is calculated according to the formula SDF n =αNFID n +βTR n , n∈[0,N], where SDF n is the evaluation score of the nth group of generated pseudo-labeled samples, NFID n ∈[0,1] is the normalized FID score, and TR n ∈[0,1] is the normalized training evaluation score, α is the weight coefficient of NFID n , β is the weight coefficient of TR n , and α + β = 1.
[0049] Furthermore, a SinGAN model based on an improved loss function is used to generate multiple groups of the pseudo-labeled samples, and sample training is performed based on the SinGAN model. The loss function of the discriminator of the SinGAN model is:
[0050] The loss function of the generator of the SinGAN model is:
[0051] where D n is the nth discriminator, G n is the nth generator, and χ is xn and x n in the joint sampling space is the gradient penalty term, and μ is the weight coefficient is the pseudo-image generated by the (n + 1)-th generator is the pseudo-image generated by the n-th generator, and x n is the corresponding real image at each scale, and z * is the randomly selected value before training, and L div is the ratio of the distance between generated images to the distance between noises
[0052] In practical applications, SinGAN is an unconditional generative model that can learn from a single natural image, capable of capturing the internal block distribution information of the image and generating high-quality and diverse samples with the same visual content. Different from traditional GANs that only have one generative model (G) and one discriminative model (D), SinGAN has multiple generative models and discriminative models respectively, which can be considered as a cascade of multiple GANs, presenting a pyramid structure as a whole. Each GAN is responsible for learning the distribution information of the image at different scales, so new samples with arbitrary sizes and aspect ratios can be generated. These samples have obvious variations while maintaining the overall structure and fine texture features of the training image. Compared with previous single-image GAN solutions, this method is not limited to texture images and is unconditional (i.e., generating samples from noise). At the same time, different from other GANs that can only be applied to a single task, SinGAN can be used for tasks such as image generation, image segmentation, super-resolution tasks, drawn image conversion, image editing, and image harmonization. The generation process of SinGAN is from bottom to top, from rough to fine. All Gs and Ds have the same structure, consisting of 5 groups of 3×3 fully convolutional layers, so both G and D have a receptive field of 11×11. The setting of the same receptive field allows each layer of GAN to focus on the overall layout of the image and the global structure of the target
[0053] SinGAN learns the distribution of the image from a single image, which is the most important feature compared with other GAN models. Since SinGAN is a pyramid structure, it is trained layer by layer during the training process, from bottom to top. After each layer of GAN is trained, it is fixed and its network parameters are no longer changed
[0054] Corresponding to the generative models {G 0 …G N} are the discriminative models {D 0 …D N}, and each layer of the discriminative model is used to distinguish the x n downsampled from the real image x and the pseudo-image generated by the generative model True or false. Among them, the discriminator D n The loss function of adopts WGAN-GP (Wassertein Distance Generative Adversarial Networks-Gradient Penalty), which can increase the stability of network training, as shown in (1):
[0055]
[0056] Among them: D(x n ) is the generator of x n , is 's generator, χ is the joint sampling space of x n and . The third term is the gradient penalty term, and μ is the weight coefficient. WGAN-GP aims to solve the problem of over-concentration of parameters caused by weight constraints in WGAN and the problems of gradient explosion and disappearance during the training process. It proposes a gradient penalty method to set a threshold, and when the sample gradient exceeds this threshold, a penalty is imposed. This method effectively solves the above problems and stabilizes the training of the GAN network.
[0057] G n 's loss function is also called the reconstruction loss. The purpose of establishing this loss function is to hope that there is a set of randomly generated noises as input, so that the finally output image is the original image, thereby increasing the stability of training. Therefore, the author selects specific random noises here. As follows:
[0058]
[0059] Among them, z * is a value randomly selected before training and will not be changed later.
[0060] Therefore, the loss function of G n is as follows:
[0061]
[0062] Among them: is the pseudo-image generated by the (n + 1)-th generator using the above fixed noise, and x n is the corresponding real image at each scale.
[0063] The generator takes noise as input, and once the noise is selected, it will not change. At the same time, it is noted that GAN is prone to mode collapse when generating images, only generating images of several categories. Mapping to the distribution means that the distribution of these several categories of sample data is wide and the data peak is large, while the other categories are the opposite. Therefore, most of the data during the generation process falls on the categories with wide data distribution, thus reducing the diversity of the generated samples. Therefore, in order to further increase the diversity of the generated samples, a regularization term is designed and added at the generator end in this method. This regularization term intuitively increases the diversity of the generated images by maximizing the ratio of the distance between the generated images and the distance between the noises. The distance between the noises is fixed, so maximizing the ratio of the distance between the generated images and the distance between the noises can directly widen the distance between the generated images, forcing the data distribution of the generated images to fall on the categories with small peaks and narrow ranges. As follows:
[0064]
[0065] where z 1 , z 2 are different samplings from the same noise space, G() represents the generated fake samples, and d() represents the distance.
[0066] So the generator G n The final loss function is as follows:
[0067]
[0068] From the comparison of the generation effects before and after the improvement of SinGAN, it can be seen that the images generated by the original SinGAN are indistinguishable from the real images, indicating that the generated images have high authenticity but low diversity; while the fake labeled samples generated by the improved SinGAN have obvious changes in thickness, length, and the shape of the stem labels, indicating that the improved SinGAN can generate samples with higher diversity, which also increases the robustness of the subsequent network.
[0069] Furthermore, the screening of the fake labeled samples according to the screening indicators includes:
[0070] After the training of the SinGAN model is completed, N + 1 groups of fake images {PS N , …, PS n , …, PS 0} are generated, and the SinGAN model generates the fake labeled samples according to the fake images,
[0071] According to the selected screening indicator SDF n evaluate the authenticity and diversity of the fake labeled samples, and use the formula to evaluate the quality of the generated images, where F RS and respectively represent the real image RS and the nth group of generated images PS n the average value of the feature vectors, CF RS and respectively represent the covariance matrices calculated from the feature vectors of RS and PS n The Tr(·) represents the trace of the matrix.
[0072] In practical applications, SinGAN will generate N + 1 groups of pseudo-images after training, namely {PS N , …, PS n , …, PS 0}. However, not every group of pseudo-images is suitable for model training. Therefore, this method proposes an improved quantitative index for pseudo-sample screening, SDF ∈ [0, 1]. This index can comprehensively evaluate the authenticity and diversity of pseudo-samples. Specifically as follows:
[0073] SDF n = αNFID n + βTR n , n ∈ [0, N]; (6)
[0074] where SDF n represents the evaluation score of PS n , NFID n ∈ [0, 1] and TR n ∈ [0, 1] respectively represent the normalized FID (Fréchet Inception Distance) score and the training evaluation score of PS n . α and β represent the weight coefficients of NFID n and TR n respectively, and α + β = 1. At the same time, in order to make NFID n and TR n have the same contribution rate to SDF n , we set both α and β to 0.5.
[0075] The NFID n in the formula is the standardized version of FID n , specifically as follows:
[0076]
[0077] where FID n is the FID score of PS n . Min() represents the minimization operation because the quality of pseudo-samples is inversely proportional to the FID score.
[0078] FID is a metric proposed in 2017 for evaluating the quality of generated images and is specifically used to evaluate the performance of generative adversarial networks. Due to its excellent measurement method, FID can well measure the authenticity and diversity of generated images. The formula is as follows:
[0079]
[0080] where F RS and represent the average of the feature vectors of the real image RS and the nth group of generated images PS n respectively, CF RS and represent the covariance matrices calculated from the feature vectors of RS and PS n respectively, and Tr(·) represents the trace of the matrix. The feature vectors used in the formula for calculation are all extracted by the Inception V3 network pre-trained on the ImageNet dataset.
[0081] Although FID can directly calculate the distance between the generated image and the real image from the "inside" of the image to evaluate the quality of the generated image, it does not evaluate the generated samples from the perspective of improving the training quality. And improving the classification accuracy of the model and enhancing the training performance of the model are the core motivations for generating a large number of pseudo-samples. Therefore, we propose to combine TR n and FID for the evaluation of generated samples. TR n consists of two parts:
[0082]
[0083] where SIM n represents the similarity between RS and PS n , DIV n represents the relative diversity of PS n relative to RS, NSIM n ∈[0,1] and NDIV n ∈[0,1] are the normalized versions of SIM n and DIV n respectively, λ and η represent the weight coefficients of NSIM n and NDIV n , and λ + η = 1. Since the authenticity and diversity of pseudo-samples are equally important, we set the values of λ and η to 0.5.
[0084] Since the generated samples are ultimately used for model training, the quality of the generated samples is considered from the perspective of model training. Thus, if the pseudo-samples are similar to the real samples, then using the pseudo-samples in the test phase of a deep neural network trained with real samples will result in a relatively high score, and this score will not be worse than the score obtained by using real samples for testing; on the other hand, if the diversity of the pseudo-samples is not high, the pseudo-samples cannot fully cover the data distribution of the real samples, and the deep neural network trained on the pseudo-samples cannot obtain a high-precision classification result when testing real samples, that is, DIV n is very low. Therefore, the following method is adopted to calculate SIM n and DIV n :
[0085]
[0086] where DNN(RS) and DNN(PS n ) represent that the deep neural network DNN (Deep Neural Networks) is trained by RS and PS respectively n , and OA(DNN(RS),PS n ) represents the result of testing the trained network DNN(RS) with PS n , and OA(DNN(PS n ),RS) represents the result of testing the trained network DNN(PS n ) with RS
[0087] It should be noted that since the original value ranges of FID n SIM n and DIV n are different, normalization operations are necessary, which is also the reason for the normalization in equations (7) and (9). This can ensure that the final value ranges are all within [0,1]. Moreover, α + β = 1 and λ + η = 1 can ensure that the values of SDF n and TR n are restricted to [0,1].
[0088] Finally, the best pseudo-sample PS j can be determined according to the score of SDF n . The larger the SDF n , the better the quality of the pseudo-sample. That is:
[0089]
[0090] Furthermore, the use of the trained stem label classification network to detect the stem labels of the tested cigarette samples includes:
[0091] Predict and classify the target samples through a Softmax classifier with a focal loss function, and determine the loss function value according to the formula FL(p i )=-α i (1 - p i ) γ log(p i ), where FL(p i ) is the loss function value, p i represents the probability that the model predicts that the sample contains a meme label, γ represents the hyperparameter controlling "focus", and α i represents the "contribution" of controlling the positive and negative samples to the total loss.
[0092] In practical applications, in order to better solve the problem that hard example samples are difficult to identify, this method replaces the traditional cross-entropy loss function with Focal Loss. The application of this loss function can further improve the accuracy of meme label detection.
[0093] FL(p i )=-α i (1 - p i ) γ log(p i ); (12)
[0094] where FL() represents the loss function value, p i represents the probability that the model predicts that the sample contains a meme label, γ represents the hyperparameter controlling "focus", that is, controlling the model to pay more attention to hard example samples, and α i represents the "contribution" of controlling the positive and negative samples to the total loss. When α i is small, the weight of the negative sample is reduced, which means the increase of the positive sample weight, thereby reducing the impact of the negative sample on training and improving the final classification accuracy.
[0095] After the network training is completed, determine that its parameters no longer change, and then send the test samples into the deep meme label classification network to achieve meme label detection.
[0096] This method also includes: selecting cigarettes of different brands as detection objects and making the collected data into a data set, and detecting and verifying the trained meme label classification network according to the data set to determine whether the final detection rate of the meme label reaches the set threshold. If so, the training of the meme label classification network is qualified.
[0097] Further, select 20 different brands of cigarettes on the market as detection objects and make the collected data into a data set XIC-20. The steps of making the data set include:
[0098] First, irradiate cigarettes with an X-ray device to obtain cigarette fluoroscopic images and manually annotate them. 100 images are collected for each brand of cigarettes, with 50 images having stem tags and 50 images without stem tags, for a total of 2,000 images.
[0099] Then, use the trained SinGAN model to perform data augmentation on the images with and without stem tags for each category at training rates of 20% and 50% respectively. The remaining 80% and 50% are used for subsequent testing.
[0100] The image augmentation ratio is 1:20. When the training rate is 20%, the augmented dataset has 20 categories, with 400 images in each category, for a total of 8,000 images.
[0101] When the training rate is 50%, the augmented dataset has 20 categories, with 1,000 images in each category, for a total of 20,000 images.
[0102] Specifically, the parameter settings use the SinGAN default settings, i.e., N = 8, so a total of 9 groups of pseudo-samples {PS 8 ,…,PS 0} are generated. Among them, the batch size is set to 1, the learning rates of the discriminator and the generator are both 0.0005, and the optimizer is Adam; for the training of ResNet50, the batch size is set to 32, the learning rate of the last layer is 0.01, and the other layers are 0.001, and the optimizer is ASGD. At the same time, the hyperparameter settings in Focal Loss are in accordance with the original default settings, i.e., α i = 0.25, γ = 2.
[0103] In the experimental verification stage of the stem tag detection algorithm that fuses pseudo-samples and loss functions, the training rate of the dataset is the same as that in the generation stage. The experiments for each training rate of the dataset are repeated 10 times.
[0104] The workstation configuration for running the experiment is two E5-2650V4 CPUs (2.2GHz, 12×2 cores in total), 512GB, the GPU is NVIDIA TITAN RTX, and the memory is 24GB×8. Pytorch is selected as the deep learning platform.
[0105] As shown in the experimental results of the screening effectiveness comparison in Table 1, the first row of Table 1 is 9 groups of different pseudo-samples {PS 8 ,…,PS 0}, where PS 8 is generated by the bottom GAN, and PS 0 is generated by the top GAN; the second row is the quantitative screening score SDF (Formula 6); the third row is the overall classification accuracy.
[0106] Table 1
[0107]
[0108] Obviously, as shown in Table 1, the higher the SDF value, the higher the corresponding OA (Formula 13) value, which can directly verify the effectiveness of the proposed quantitative screening index. At the same time, as the generation scale increases, although the SDF value and the OA value are not much different, they are both decreasing, indicating that the quality of the generated images is gradually decreasing. Therefore, the group of pseudo-labeled samples with the highest value is selected for subsequent ResNet50 training, that is, the bottom GAN in SinGAN is used as the initial GAN to generate pseudo-samples.
[0109] As shown in Table 2 for the overall accuracy comparison experiment results, Table 2 shows the overall accuracy comparison on the dataset. RS represents the deep classification network model trained only with real samples and is also considered as the baseline method. PS uses pseudo-samples to replace real samples for training, and RS+PS jointly uses real and pseudo-samples to train the deep network classification model. Focal Loss is used to replace the traditional cross-entropy loss of RS, PS, and RS+PS respectively, obtaining RS+FL, PS+FL, and RS+FS+FL, where RS+FS+FL represents the method proposed in this paper.
[0110] Table 2
[0111]
[0112] It can be seen from the data in Table 2 that the overall performance of PS is better than that of RS, indicating that the generated pseudo-samples have good quality and can improve the performance of the deep classification network. The comparison between RS+PS and PS shows that the combination of PS and RS can further improve the performance of the deep classification network. The comparison between RS+FL, PS+FL, RS+FS+FL and RS, PS, RS+PS shows that Focal Loss can replace the traditional cross-entropy loss function and improve the classification accuracy of the network.
[0113] Furthermore, the use of the X-ray device to irradiate the detection object with X-rays and obtain the corresponding cigarette fluoroscopic image includes:
[0114] Based on the difference in the X-ray transmission imaging characteristics of cut tobacco and stem pieces, X-ray transmission imaging forms an X-ray black and white image.
[0115] Perform noise filtering preprocessing on the X-ray black and white image to remove the background noise in the image.
[0116] Use the region growing method to segment the X-ray black and white image after noise filtering preprocessing into stem pixels and background pixels.
[0117] The fuzzy C - clustering algorithm is used to obtain the membership degrees of each segmented tobacco stem pixel pair to the tobacco stem center, so as to filter out interference information and make attribution judgments.
[0118] After the fuzzy C - clustering algorithm is processed, the shape analysis is carried out on the pixels gathered together, and the area of the segmented region and the aspect ratio of the region are calculated for shape recognition.
[0119] Specifically, based on the differences in the X - ray transmission imaging characteristics between cut tobacco and stem pieces, the X - ray transmission imaging method can be effectively used for the detection and determination of cut - tobacco cigarettes. Therefore, the image recognition algorithm for tobacco stems is mainly designed for the grayscale image of X - ray transmission imaging. It includes the noise filtering pre - processing method, the image segmentation method of the image after the cigarette is transmitted by X - ray, as well as the characteristic image parameters for stem piece recognition, and the classifier algorithm for classifying and recognizing cut tobacco and stem pieces. Finally, an image recognition method for stem pieces is established. The image recognition algorithm for tobacco stems mainly includes four parts: image pre - processing, image segmentation, interference point attribution judgment, and shape feature determination. Among them, in image segmentation and interference point attribution judgment, the region - growing method with the characteristics of artificial intelligence imitation and the fuzzy C - clustering algorithm based on unsupervised machine learning function are respectively studied and adopted.
[0120] (1) Tobacco stem image pre - processing: The original image collected by the system contains a lot of background noise, which affects the quality of subsequent image segmentation and is prone to misjudgment of tobacco stems. Therefore, it is necessary to strengthen the tobacco stem information and eliminate various random noises through image pre - processing. To determine a better noise elimination algorithm, methods such as mean filtering, adaptive Wiener filtering, median filtering, and morphological filtering are respectively tested through data simulation. According to the filtering effect, the gray - scale morphological noise filter is finally adopted.
[0121] (2) Image segmentation by region - growing method: In the obtained image, the images of cut tobacco and stem pieces in the cigarette have a very high similarity in gray - level and are gathered together. The region - growing method can automatically mark the tobacco stems in the image according to the iterative rule. First, the seed pixels are determined according to the gray - level information distribution of tobacco stem imaging, and then the pixels with the same or similar properties as the seed pixels in the neighborhood around the seed pixels are merged into the region where the seed pixels are located. Then, the new pixels are regarded as new seeds for iterative operation, and further, the pixels with similar properties are gathered together to form a region.
[0122] (3) Fuzzy C - clustering membership judgment. After the image is segmented into tobacco stem pixels and background pixels, there are still many dot - like interferences in the image, which will lead to a large degree of misjudgment. Therefore, the fuzzy C - clustering algorithm based on unsupervised machine learning function is used to filter out the interference information and make membership judgment. The fuzzy C - clustering algorithm is a method of analyzing and modeling important data with fuzzy theory, which establishes an uncertain description of the sample class membership. Since the tobacco stem information is relatively complex, some tobacco stems will appear in a broken form in the segmented image. Using the traditional connectivity labeling algorithm will lead to shape judgment errors and missed recognition. By using fuzzy C - clustering, the membership degree of each segmented tobacco stem pixel to the center of the tobacco stem can be obtained, so as to determine that multiple broken modules belong to one tobacco stem information. This can effectively increase the recognition rate.
[0123] (4) Shape recognition. To further reduce the false detection rate of tobacco stems, shape recognition is added to the algorithm. The shape recognition algorithm analyzes the shape of the pixels gathered together after the fuzzy C - clustering algorithm, and calculates shape factors such as the area of the segmented region, the aspect ratio of the region, etc. The area of the region is represented by the pixel aggregation degree S of the region. The aspect ratio of the region is determined by the circumscribed rectangle method to obtain the length L and diameter D of the region, and then the aspect ratio R is calculated. When the region S is greater than 120 and R is greater than 10, the shape characteristics are met, and it is judged as a tobacco stem.
[0124] To verify the detection efficiency and accuracy of the cigarette containing - stem non - destructive detection device, ordinary cut tobacco processed by a certain manufacturer is selected. The stem pieces in the cut tobacco are manually removed. During the process of processing this part of the cut tobacco into cigarettes, it is divided into ten groups. During the process of making cigarettes in these ten groups, stem pieces are manually added to some cigarettes in each group. The proportion of the number of cigarettes with added stem pieces in each group to the total number of cigarettes in the group is 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, and 20% respectively. The cigarette containing - stem non - destructive detection device is used to sample each of the above 10 groups of cigarettes 20 times, and each sampling amount is 100 cigarettes for detection. For each sampled cigarette, its stem - containing ratio is consistent with the total stem - containing ratio of the group. After the detection is completed, the detection order of the 10 groups of cigarettes is adjusted for manual detection. Finally, the detection efficiency and accuracy of the two detection methods are compared and analyzed.
[0125] The detection results are shown in Table 3. The range of the absolute error values detected by the device is between 0.1 and 0.9, and the range of the absolute error in manual detection is between 0.8 and 1.5. The range of the standard deviation detected by the device is between 0.9468 and 1.5898, and the range of the standard deviation in manual detection is between 1.1548 and 1.5864. It can be seen from the range of the absolute error and the range of the standard deviation that the accuracy and precision of the device detection are slightly higher than those of manual detection, indicating that the non-destructive detection device for cigarette stems has high reliability in terms of detection accuracy and precision. In terms of the detection duration, the detection duration of the device is about half of that of manual detection, which greatly improves the detection efficiency.
[0126] Table 3
[0127]
[0128] Both the standard deviation and the absolute error generated by the device detection method are smaller than those of the manual detection method, and the device detection has high reliability. Compared with the manual detection method, the detection duration of the non-destructive cigarette detection device is about half of that of manual detection, which improves the detection efficiency while ensuring the detection accuracy and precision.
[0129] It can be seen that the present invention provides a method for detecting cigarette stems based on X-ray vision. Data of the tested cigarettes are collected by an X-ray device and manually labeled; then, a generative adversarial network is used to expand the data, and the generated samples are screened by the proposed screening indexes to determine the final expanded samples. Finally, the expanded samples are used for the training of a classification network, and the trained network is further fine-tuned with real samples. The problems of low efficiency and low accuracy in the detection of cigarette stems in existing cigarette products are solved, the detection efficiency of cigarette stems can be improved, and the false detection rate or missed detection rate of cigarette stem detection can be reduced.
[0130] The structure, features and function effects of the present invention have been described in detail based on the illustrated embodiments above. The above are only the preferred embodiments of the present invention, but the present invention is not limited to the scope defined by the drawings. Any changes made according to the concept of the present invention, or equivalent embodiments modified into equivalent changes, still within the spirit covered by the description and the drawings, shall fall within the protection scope of the present invention.
Claims
1. A method for detecting stem tags based on X-ray vision, characterized in that, comprising: Randomly select cigarettes of different brands as the detection objects, and use X-ray equipment to irradiate the detection objects with X-rays and obtain corresponding cigarette perspective images; Use a generative adversarial network to generate multiple groups of pseudo-labeled samples from the cigarette perspective images, and screen the pseudo-labeled samples according to screening indicators to determine the final augmented labeled samples; Obtain the manually labeled samples of the detection objects, input the augmented labeled samples into a preset stem tag classification network for pre-training, and use the manually labeled samples to adjust the training network; Use the trained stem tag classification network to detect stem tags in the tested cigarette samples; Use the overall classification accuracy as the evaluation index of the trained stem label classification network, and the overall classification accuracy is calculated according to the formula obtained, where OA is the overall classification accuracy, N is the total number of samples, and Z is the number of samples correctly classified in all categories; Use the screening metric as the evaluation metric for the trained meme tag classification network. The screening metric is calculated according to the formula where is the evaluation score of the n-th group of generated pseudo-labeled samples, is the normalized FID score sum, is the normalized training evaluation score, is the weight coefficient of is the weight coefficient of ; It consists of two parts: Represents and the similarity of, Represents the relative to the relative diversity of, and are respectively and the normalized versions of, and Represents and the weight coefficients of, and ; Use the SinGAN model based on an improved loss function to generate multiple groups of the pseudo-labeled samples, and perform sample training based on the SinGAN model. The loss function of the discriminator of the SinGAN model is: ; The loss function of the generator of the SinGAN model is as follows: ; Among them, D n is the nth discriminator, is the nth generator, is and 's joint sampling space, is the gradient penalty term, is the weight coefficient, is the th pseudo-image generated by the generator, is the pseudo-image generated by the nth generator, is the corresponding real image at each scale, is the randomly selected value before training, is the ratio of the distance between the generated images to the distance between the noises.
2. The method for detecting stem tags based on X-ray vision according to claim 1, characterized in that, The screening of the pseudo-labeled samples according to the screening indicators includes: After the training of the SinGAN model is completed, N + 1 groups of pseudo-images are generated. The SinGAN model generates the pseudo-label samples based on the pseudo-images. According to the selected screening indicators evaluate the authenticity and diversity of the pseudo-labeled samples, and use the formula to evaluate the quality of the generated images, where and represent the average feature vectors of the real images and the th group of generated images respectively, and represent the covariance matrices calculated from the feature vectors of and respectively, and represents the trace of the matrix.
3. The method for detecting stem tags based on X-ray vision according to claim 2, characterized in that, The use of the trained stem tag classification network to detect stem tags in the tested cigarette samples includes: Predict and classify the target samples through a Softmax classifier with a focal loss function, and determine the loss function value according to the formula wherein is the loss function value, represents the probability that the model predicts that the sample contains a meme tag, represents the hyperparameter that controls "focus", represents the contribution of positive and negative samples to the total loss.
4. The method for detecting stem tags based on X-ray vision according to claim 3, characterized in that, further comprising: Select cigarettes of different brands as the detection objects and make the collected data into a dataset, and use the dataset to detect and verify the trained stem tag classification network to determine whether the final detection rate of the stem tags reaches a set threshold. If so, the training of the stem tag classification network is qualified.
5. The method for detecting stem tags based on X-ray vision according to claim 4, characterized in that, The steps of making the dataset include: First, irradiate the cigarettes with X-ray equipment to obtain cigarette perspective images and perform manual labeling on them. 100 images are collected for each brand of cigarette, with 50 images with stem tags and 50 images without stem tags, for a total of 2000 images; Then use the trained SinGAN model to perform data augmentation on the images with and without stem tags in each category at training rates of 20% and 50% respectively, and the remaining 80% and 50% are used for subsequent testing; The image augmentation ratio is 1:
20. When the training rate is 20%, the augmented dataset has 20 categories, with 400 images in each category, for a total of 8000 images; When the training rate is 50%, the augmented dataset has 20 categories, with 1000 images in each category, for a total of 20000 images.
6. The method for detecting stem tags based on X-ray vision according to claim 5, characterized in that, The use of X-ray equipment to irradiate the detection objects with X-rays and obtain corresponding cigarette perspective images includes: Based on the difference in X-ray transmission imaging characteristics between tobacco shreds and stem tags, X-ray transmission imaging forms an X-ray black and white image; Perform noise filtering preprocessing on the X-ray black and white image to remove background noise in the image; Use the region growing method to segment the preprocessed X-ray black and white image after noise filtering into stem pixels and background pixels; Adopt the fuzzy C-clustering algorithm to calculate the membership degree of each segmented stem pixel to the stem center, so as to filter out interference information and make attribution judgments; After the fuzzy C-clustering algorithm is processed, perform shape analysis on the pixels gathered together, and calculate the area and aspect ratio of the segmented region for shape recognition.
Citation Information
Patent Citations
Steel plate character detection method and device based on multistage network convergence
CN113989478A
Cited By
Method for detecting and characterizing sliver in cigarette
CN116242860A
A method for detecting and characterizing the inner stem of a cigarette.
CN116242860B