Deep learning-based cut tobacco impurity detection method

By improving YOLOv8's deep learning detection method, and optimizing the backbone network using RepVGG and attention mechanism, the problem of poor accuracy and generalization in tobacco impurity detection is solved, and more efficient and accurate detection results are achieved.

CN119991641APending Publication Date: 2025-05-13HEBEI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510154456.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has low accuracy and poor generalization in the detection of tobacco impurities, especially in the tobacco processing process, which increases the detection difficulty.

Method used

Based on YOLOv8, the deep learning detection method is used to improve the convolutional network to RepVGG, and the spatial attention mechanism and channel attention mechanism are introduced, the global maximum pooling is increased, and the backbone network is optimized to improve feature extraction capabilities.

Benefits of technology

It significantly improves the accuracy and generalization ability of tobacco impurities detection, improves the robustness and detection effect of the model, reduces the amount of model parameters, and maintains efficient feature extraction and inference speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991641A_ABST
    Figure CN119991641A_ABST
Patent Text Reader

Abstract

The invention discloses a cut tobacco impurity detection method based on deep learning, and belongs to the technical field of cut tobacco impurity detection.The cut tobacco impurity detection method comprises the steps that images with impurities in the actual detection process are selected as a cut tobacco impurity data set, and target areas of various types of impurities are marked; the YOLOv8 is improved, and a convolutional network of the YOLOv8 is replaced by a re-parameterized convolution RepVGG; a space attention mechanism is introduced, and the space attention mechanism is realized by adopting a space information fusion and long-distance connection mode; a channel attention mechanism is introduced and improved, and global maximum pooling is increased; training the data set to obtain a trained detection model; and performing evaluation by using the average precision mAP and the recall rate Recall as evaluation indexes of the detection model. According to the cut tobacco impurity detection method based on deep learning provided by the invention, the efficiency and accuracy of the detection method are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tobacco impurity detection, and in particular to a tobacco impurity detection method based on deep learning. Background Art

[0002] The tobacco industry plays an important role in the global economy. From the perspective of product quality, too much impurities or poor quality tobacco will cause tobacco flavor variation and smoke quality degradation, seriously affecting the brand's market competitiveness and consumer reputation. Therefore, tobacco impurity removal is of great significance to the tobacco industry. However, tobacco processing has many steps, and it is inevitable that various impurities will be mixed in. Especially in the process of making tobacco, impurities are mixed with tobacco leaves, which increases the difficulty of impurity detection.

[0003] At present, tobacco impurity removal mainly relies on manual and machine vision algorithms. Manual detection is inefficient and labor-intensive. Tobacco impurity identification based on traditional machine vision mainly includes color sorting tables, pattern recognition and threshold segmentation methods. The existing technology includes a machine vision detection method based on the Bayesian algorithm, which pre-processes the image through image filtering, color compression and region of interest (ROI) cropping, compares the probabilities of positive samples and negative samples to design a classifier, and finally scans the image through sliding windows of different scales to determine the location of foreign matter. Based on machine vision, the color sorting table method is used to identify color-sensitive foreign matter, and the model is improved by the grayscale threshold method and the double threshold method to enhance the ability to identify impurities. The image cascade detection method designed based on color features and gradient energy realizes the positioning of tobacco impurities, and improves the detection stability through HOG features and cascade Adaboost classifier algorithms.

[0004] However, in actual production scenarios, tobacco and impurities continue to move forward with the vibration of the production line. Impurities are wrapped in a large amount of tobacco to varying degrees, and their positions and shapes change all the time, making detection more challenging. Traditional machine vision methods that rely on manually designed features to identify impurities cannot fully capture the complex information in the image and have poor generalization capabilities.

[0005] In recent years, the application of deep learning in impurity detection has received extensive attention and research. Many experts and scholars have applied deep learning technology to impurity detection tasks in various fields and have made a series of breakthroughs, providing a new solution for impurity detection. In order to solve the problem of low accuracy and poor generalization of impurity detection in cut tobacco, this paper proposes a detection algorithm for impurities in cut tobacco based on YOLOv8. Summary of the invention

[0006] The purpose of the present invention is to provide a tobacco impurity detection method based on deep learning to solve the problems existing in the background technology.

[0007] To achieve the above object, the present invention provides a tobacco impurity detection method based on deep learning, comprising the following steps:

[0008] S1. Select images with impurities in the actual detection process as tobacco impurity datasets, and mark the target areas of various types of impurities;

[0009] S2. Improve YOLOv8 by replacing its convolutional network with the re-parameterized convolution RepVGG.

[0010] S3, introduce the spatial attention mechanism, and use spatial information fusion and long-distance connection to realize the spatial attention mechanism;

[0011] S4, introduce the channel attention mechanism and improve it, and add global maximum pooling;

[0012] S5, training the data set to obtain a trained detection model;

[0013] S6. Use average precision (mAP) and recall (Recall) as evaluation indicators of the detection model for evaluation.

[0014] Preferably, the multi-branch structure of RepVGG in S2 includes a 3x3 convolution block, a 1x1 convolution block and an identity map, and convolutions of different sizes are used to extract local features and fuse information between channels respectively; the identity map directly transfers the input to the output to realize skip connection.

[0015] Preferably, the spatial information fusion in S3 includes multi-scale information fusion and spatial transposition, and the input feature map is processed by using four convolution kernels of sizes 1×1, 3×3, 5×5, and 7×7 to extract spatial adjustment information of different ranges. The process is represented by formula (1):

[0016] F i =Conv(k i ×k i )(X) (1)

[0017] In the formula, Conv(k i ×k i ) represents the i-th convolution, the convolution kernel size is k×k, i=1,2,3,4, k=i×2-1, and the sizes of the four feature maps are adjusted and fused; Formula (2) represents the fusion of feature maps of different scales:

[0018]

[0019] in Indicates the fusion operation. In the specific code implementation, the nearest neighbor interpolation method is used to adjust F iThe fusion of spatial information of different scales is performed by adding the size and sum;

[0020] The spatial information of the input feature map is connected over long distances to increase the network's perception range, which helps the network better capture the contextual information of the input data and enhance the network's representation ability, so that it can better adapt to the complex input data distribution rather than being limited to local details. The specific approach is to transpose the feature map that integrates spatial information of different scales twice and multiply it with the feature map to achieve long-distance connection. Formula (3) represents this process.

[0021]

[0022] Where F represents the feature map that integrates spatial information of different scales, F' is the transposed feature map, and attention_map is the obtained attention map.

[0023] Preferably, the S4 process is as follows:

[0024] Assume input feature map The numbers H, W, and C represent the height, width, and number of input channels, respectively. The global average pooling and global maximum pooling operators are expressed using formulas (4) and (5):

[0025]

[0026] M c =MaxPool(X c (i,i)) (5)

[0027] X c represents the c-th channel of the input feature map, and the weight of the c-th channel of the input feature map can be expressed as formula (6):

[0028] w c =σ(w 1 δ(w 0 (G c +M c ))) (6)

[0029] Where σ represents the Sigmoid activation function, which is used to map the output of the neural network to the probability distribution; δ represents the Relu activation function, which is used to introduce nonlinearity into the neural network so that the neural network can learn more complex functional relationships. 0 and w 1 There are two linear layers used to learn the relationship between channels of the feature map and generate the excitation weights for each channel.

[0030]

[0031] w=σ(C1Dk (y)) (8)

[0032] In order to further integrate channel information, as shown in formula (7), k represents the one-dimensional convolution C1D k The size of the convolution kernel; y is the feature map channel weight; w represents the attention weight after cross-channel fusion.

[0033] Preferably, S6 uses mAP and Recall at different confidence levels as evaluation indicators to measure model performance. mAP and Recall are defined as follows:

[0034]

[0035] In the formula, TP represents the number of impurities that are correctly detected; FN represents the number of impurities that are actually impurities but not detected; and the recall rate measures the proportion of all impurities that the model can correctly detect among all actual impurities. A high recall rate means that the model can identify most impurities.

[0036]

[0037] In the formula, N represents the number of categories; AP i The average precision of the i-th category. mAP measures the average detection accuracy of the model over all categories.

[0038] Therefore, the present invention adopts the above-mentioned tobacco impurity detection method based on deep learning, which has the following beneficial effects:

[0039] (1) The ECA attention mechanism is improved to form a hybrid attention mechanism STCA with SFTM, which aims to fuse the features of adjacent spaces and perform long connections between different spaces. The ECA module is improved to extract channel information, so that the model can better capture the correlation between different channels.

[0040] (2) Re-parameterize the backbone network to reduce the number of model parameters while improving the model feature extraction capability;

[0041] (3) SFTM implements the spatial attention mechanism through spatial information fusion and long-distance connection, which can more comprehensively model the spatial relationship in the input data and enhance the model's understanding of the entire space.

[0042] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is an overall flow chart of an embodiment of the present invention;

[0044] Figure 2The improved backbone network according to the embodiment of the present invention;

[0045] Figure 3 A schematic diagram of a spatial fusion and transposition module according to an embodiment of the present invention;

[0046] Figure 4 Schematic diagram of the channel attention mechanism according to an embodiment of the present invention;

[0047] Figure 5 This is a data set and impurity image of an embodiment of the present invention;

[0048] Figure 6 It is a line graph of the changes in mAP50 and recall of all-class during the training process of the embodiment of the present invention;

[0049] Figure 7 A visualized heat map of different attention mechanisms during the training process of an embodiment of the present invention;

[0050] Figure 8 This is a diagram of the detection results of an embodiment of the present invention. DETAILED DESCRIPTION

[0051] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0052] See also Figure 1 , a tobacco impurity detection method based on deep learning, comprising the following steps:

[0053] S1. Select images with impurities in the actual detection process as tobacco impurity datasets, and mark the target areas of various types of impurities.

[0054] S2. Improve YOLOv8 by replacing its convolutional network with the re-parameterized convolution RepVGG.

[0055] The model network architecture in computer vision can usually be divided into three general modules: backbone network, neck network and prediction head. The backbone network is used to extract general features including color, shape and texture; the neck network further processes and fuses these features; and the prediction head outputs results in different forms according to task requirements, such as category, probability, detection box, etc. The backbone network is the basis of many advanced tasks, and its performance greatly affects the upper limit of the entire network. This embodiment uses reparameterized convolution (RepVGG) to improve the backbone network of YOLOv8, and improves model performance and reasoning speed by using different network structures in the training and reasoning stages.

[0056] like Figure 2 As shown in the figure, the multi-branch structure of RepVGG consists of three parts: 3x3 convolution block, 1x1 convolution block and identity mapping. Convolutions of different sizes are used to extract local features and fuse information between channels. Identity mapping passes the input directly to the output, which is used to implement jump connections, thereby alleviating the gradient vanishing problem in deep networks. The output of each branch is connected to a batch normalization layer to stabilize the training process, accelerate convergence, and improve the generalization ability of the model. The convolution layers and batch normalization layers of different branches work independently during training to provide rich feature expressions. In the training phase of the model, the complex multi-branch structure of RepVGG works independently to provide rich feature expressions and improve the expression and generalization capabilities of the model. In the inference phase, the convolution kernel weights and biases of different branches are merged into an equivalent convolution kernel. The fused convolution kernel weights and biases are used as the convolution kernel parameters in the inference phase to replace the original multi-branch convolution structure, so that only one convolution operation is required during inference, simplifying the calculation and improving the speed.

[0057] This embodiment improves the YOLOv8 backbone network based on RepVGG, changes all the convolutions in the backbone network to RepVGG convolutions, and deepens the network depth, while the neck and detection head of the model remain unchanged.

[0058] S3. Introduce the spatial attention mechanism and use spatial information fusion and long-distance connection to implement the spatial attention mechanism.

[0059] The role of the spatial attention mechanism is to dynamically adjust the importance of features at different spatial locations in the neural network to enhance the model's attention to information at different locations in the image. Figure 3 As shown in Figure 2, SFTM mainly consists of two parts: multi-scale information fusion and spatial transposition.

[0060] Spatial information fusion. The role of spatial information fusion is to integrate feature information from different spatial locations to enhance the model's understanding and modeling capabilities of the overall spatial structure, help capture the global spatial structure and relationships, and thus improve the performance and generalization capabilities of the model. This embodiment uses four convolution kernels of sizes 1×1, 3×3, 5×5, and 7×7 to process the input feature map and extract spatial adjustment information of different ranges. The process is expressed by formula (1):

[0061] F i =Conv(k i ×k i )(X) (1)

[0062] In the formula, Conv(k i ×k i) represents the i-th convolution, and the convolution kernel size is k×k, i=1,2,3,4. k=i×2-1. The four feature maps are resized and fused. Formula (2) represents the fusion of feature maps of different scales.

[0063]

[0064] in Indicates the fusion operation. In the specific code implementation, the nearest neighbor interpolation method is used to adjust F i The fusion of spatial information of different scales is performed by adding and sizing them. Multi-scale feature fusion can improve the model's multi-scale perception ability of spatial features.

[0065] Long-distance connection. Long-distance connection is performed on the spatial information of the input feature map to increase the network's perception range, which helps the network better capture the contextual information of the input data and enhance the network's representation ability, so that it can better adapt to the complex input data distribution rather than being limited to local details. The specific approach is to transpose the feature map that integrates spatial information of different scales twice and multiply it with the feature map to achieve long-distance connection. Formula (3) represents this process.

[0066]

[0067] Where F represents the feature map that integrates spatial information of different scales, F' is the transposed feature map, and attention_map is the obtained attention map. In this way, the long connection of the diagonal of the feature map is realized, which helps the model better capture contextual information and expand the scope of attention.

[0068] SFTM implements the spatial attention mechanism through spatial information fusion and long-distance connection, which can more comprehensively model the spatial relationship in the input data and enhance the model's understanding of the entire space.

[0069] S4. Introduce and improve the channel attention mechanism and add global maximum pooling.

[0070] The channel attention mechanism dynamically adjusts the weight of channel features by learning the importance of each channel, thereby improving the performance of the model on specific tasks. ECA proposes a spatial attention mechanism to enhance the representation ability of deep neural networks. In the process of activating channel weights, adaptive convolution kernels are used to cross-fuse channel weights to avoid dimensionality reduction, which reduces the number of parameters and improves model performance. Figure 4 As shown, in order to better utilize channel information, this embodiment improves ECA by adding global maximum pooling and naming this part ECAblock. The numbers H, W, and C represent the height, width, and number of input channels, respectively. The global average pooling and global maximum pooling operators can be expressed by formulas (4) and (5):

[0071]

[0072] M c =MaxPool(X c (i,j)) (5)

[0073] X c represents the c-th channel of the input feature map, and the weight of the c-th channel of the input feature map can be expressed as formula (6):

[0074] w c =σ(w 1 δ(w 0 (G c +M c ))) (6)

[0075] Where σ represents the Sigmoid activation function, which is used to map the output of the neural network to the probability distribution. δ represents the Relu activation function, which is used to introduce nonlinearity into the neural network so that the neural network can learn more complex functional relationships. 0 and w 1 There are two linear layers used to learn the relationship between channels of the feature map and generate the excitation weights for each channel.

[0076]

[0077] w=σ(C1D k (y)) (8) In order to further integrate the channel information, as shown in formula (7), where k represents the one-dimensional convolution C1D k The size of the convolution kernel, y is the feature map channel weight, and w represents the attention weight after cross-channel fusion.

[0078] S5. Train the data set to obtain a trained detection model.

[0079] S6. Use average precision (mAP) and recall (Recall) as evaluation indicators of the detection model for evaluation.

[0080] The tobacco impurity dataset used in this paper includes two types of impurities: 1. Non-metallic debris outside tobacco leaves, such as rubber, plastic, nylon, feathers, paper, etc. after cutting; 2. Blocks caused by insufficient cutting of tobacco leaves during the production process. The images were collected under different lighting conditions and contain 2,300 images and corresponding annotation boxes. All images are saved in JPG format and provide annotation files in TXT and XML formats for target detection tasks. Dataset images and common impurity types are as follows: Figure 5 shown.

[0081] Since the main goal of impurity detection is to ensure that as many impurities as possible are detected to avoid missed detections, which is crucial in practical applications, recall is usually a priority, but this does not mean that precision can be completely ignored. The ideal situation is to achieve a high recall while maintaining a high precision. mAP is considered to be the main indicator for measuring the effectiveness of target detection. It combines the precision and recall of the model at different thresholds to provide an overall performance evaluation. Therefore, this embodiment uses mAP and Recall at different confidence levels as evaluation indicators to measure model performance. mAP and Recall are defined as follows:

[0082]

[0083] In the formula, TP represents the number of impurities that are correctly detected, and FN represents the number of impurities that are actually not detected. The recall rate measures the proportion of all impurities that the model can correctly detect among all the impurities that actually exist. A high recall rate means that the model can identify most impurities.

[0084]

[0085] In the formula, N represents the number of categories, AP i The average precision of the i-th category. mAP measures the average detection accuracy of the model across all categories. It is obtained by calculating the average precision for each category and then taking the average of all categories, reflecting the overall performance of the model at different thresholds.

[0086] This embodiment uses YOLOv8n as the benchmark model, and conducts ablation experiments on the tobacco impurity dataset to evaluate the impact of each module on the model performance. In order to evaluate the importance of each module, we gradually remove different parts and observe their impact on performance. As can be seen from Table 1, after adding the Rep and STCA modules respectively, the mAP and Recall of the detection results have been improved to varying degrees, but the improvement is relatively concentrated on mAP50-95 and Recall, indicating that the robustness of the model has been improved, and the target can be accurately detected under different detection accuracy requirements, and the occurrence of missed detection has been significantly reduced. After adding the Rep and STCA modules at the same time, the model has a more significant improvement in all indicators used, indicating that the improved model has found a good balance between comprehensiveness and accuracy. Figure 6 It shows the changes of mAP50 and recall of all-class during training.

[0087] Table 1 Model ablation experiment

[0088]

[0089] The STCA module needs to be upsampled before multi-scale spatial information fusion. In order to make the model perform better, this embodiment uses several common interpolation algorithms to compare the impact on STCA. As can be seen from Table 2, using the nearest algorithm, the results of the model on various indicators are the best. One of the possible reasons is that this method does not introduce new pixel values, but directly adopts the value of its nearest neighbor, which can better retain the original feature information. Analysis of the other three interpolation algorithms shows that they all fuse other pixels to varying degrees to generate new pixels. Among them, the area algorithm fuses the most pixels, and the effect is relatively poor. Bilinear and bicubic consider the pixel values ​​of different numbers around the target pixel, and perform weighted average as a new pixel to fill the feature map.

[0090] This embodiment believes that the reason why these three methods are not as good as nearest is that they fuse the information of surrounding pixels, dilute the features of impurities, and cause the loss of feature map details. Nearest is simple to calculate, fast, does not introduce new pixel values, retains feature map details, and achieves the best effect, so it is used as the upsampling algorithm of STCA.

[0091] Table 2 The impact of different interpolation algorithms on STCA

[0092]

[0093] In order to explore the performance of different attention mechanisms and find out the differences in this task, this embodiment inserts different attention mechanisms into the second layer of the backbone network of the benchmark model, and the rest of the configuration remains the same as the benchmark model, and they are tested and compared respectively.

[0094] As can be seen from Table 3, various attention mechanisms have different degrees of improvement on different evaluation indicators of different types. First, after inserting the SE and ECA attention mechanisms, the recall rate of the model for the block category decreased by 1% and 0.3% respectively. The analysis of these two attention mechanisms may be due to the lack of spatial information enhancement, which leads to the loss of some target features during the feature extraction process. Secondly, for CBAM[], the recall and mAP50 of the block category decreased by 0.4 and 1.2 percentage points respectively. Generally speaking, the recall rate index and precision are in a trade-off relationship, but from the results, while the recall rate of CBAM decreases, the overall performance decreases instead, indicating that the model after inserting CBAM not only does not enhance the feature extraction ability, but makes the model unstable. The ECA attention mechanism also has the same problem. In addition, compared with GAM[], STCA has fewer parameters and significantly improves the overall performance of the model, and is only slightly lower than GAM in the mAP50 index of impurities. The principle of SK is to use multiple convolution kernels of different scales to perform convolution operations on input features, calculate the responses of these different convolution kernels, and determine the weight of each convolution kernel through an adaptive fusion module to dynamically select the most appropriate convolution kernel. From the experimental results, SK performs well and has improved in all indicators.

[0095] Table 3 Comparative test of attention mechanism

[0096]

[0097] Finally, after inserting the STCA module proposed in this embodiment, the model performance is significantly improved comprehensively. The reason for this result is that compared with other attention mechanisms, STCA not only has the advantages of multi-scale feature fusion and long connection in space, but also can perform cross-channel information fusion weighting, so that the model can comprehensively pay attention to the channel characteristics of the input image, help the backbone network to better extract features, and thus improve the overall performance of the model. Figure 7 It is a heat map of different attention focuses. By observing the heat map, we can see how much attention weight the model assigns to different parts of the input, revealing the main areas or features that the model focuses on during the decision-making process.

[0098] In order to verify the performance of the improved model, this embodiment uses different detection models for training on the data set. First, it can be intuitively seen from Table 4 that the model proposed in this embodiment performs best in all indicators, showing its advantages in high precision and high recall rate. Compared with the baseline model, the detection effect is significantly improved, which illustrates the effectiveness of the attention mechanism and more effective feature extraction method proposed in this embodiment on this task. Secondly, the characteristics of Faster-RCNN[] dual-stage detection, its feature extraction capability and region proposal network, make it comparable to the model of this embodiment in terms of mAP50 indicators, but it is not as good as the model proposed in this embodiment in terms of recall rate and overall performance. YOLOv3tiny[] performed the worst, with a small number of parameters and a lack of more advanced feature extraction and processing modules, which makes it have certain limitations in dealing with complex backgrounds and small object recognition. In addition, models such as RegNet[], ResNeSt50[], FCOS[], and YOLOv5[] perform well in mAP50 and recall rate of the Block category, but their performance under mAP50-95 is quite different from that of the model proposed in this embodiment, indicating that these models do not perform well under higher precision requirements. Figure 8 The detection results of the baseline model and the improved model are shown.

[0099] Table 4 Model comparison test

[0100]

[0101] Therefore, the present invention adopts the above-mentioned tobacco impurity detection method based on deep learning, combines the deep learning algorithm, takes YOLOv8n as the benchmark model, introduces RepVGG, and proposes a spatial and channel feature enhancement module to optimize the benchmark model. The improved model parameter volume is reduced from 3.16M to 2.97M, which is about 5.79%, and a significant performance improvement is achieved, which verifies the effectiveness of the model proposed in the embodiment of the present invention.

[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.

Claims

1. A tobacco impurity detection method based on deep learning, characterized in that: The following steps are involved: S1. Select images with impurities in the actual detection process as tobacco impurity datasets, and mark the target areas of various types of impurities; S2. Improve YOLOv8 by replacing its convolutional network with the re-parameterized convolution RepVGG. S3, introduce the spatial attention mechanism, and use spatial information fusion and long-distance connection to realize the spatial attention mechanism; S4, introduce the channel attention mechanism and improve it, and add global maximum pooling; S5, training the data set to obtain a trained detection model; S6. Use average precision (mAP) and recall (Recall) as evaluation indicators of the detection model for evaluation.

2. The method for detecting tobacco impurities based on deep learning according to claim 1, characterized in that: The multi-branch structure of RepVGG in S2 includes 3x3 convolution blocks, 1x1 convolution blocks and identity mapping. Convolutions of different sizes are used to extract local features and fuse information between channels respectively; the identity mapping passes the input directly to the output to realize skip connection.

3. The method for detecting tobacco impurities based on deep learning according to claim 2, characterized in that: The spatial information fusion in S3 includes multi-scale information fusion and spatial transposition. The input feature map is processed by using four convolution kernels of sizes 1×1, 3×3, 5×5, and 7×7 to extract spatial adjustment information of different ranges. The process is expressed by formula (1): F i =Conv(k i ×k i )(X) (1) In the formula, Conv(k i ×k i ) represents the i-th convolution, the convolution kernel size is k×k, i=1, 2, 3, 4, k=i×2-1, and the sizes of the four feature maps are adjusted and fused; Formula (2) represents the fusion of feature maps of different scales: in Represents the fusion operation, using the nearest neighbor interpolation method to adjust F i The fusion of spatial information of different scales is performed by adding the size and sum; The spatial information of the input feature map is connected over a long distance. The feature map that integrates spatial information of different scales is transposed twice and multiplied with the feature map to achieve long-distance connection, as shown in formula (3): Where F represents the feature map that integrates spatial information of different scales; F′ is the transposed feature map; attention_map represents the obtained attention map.

4. The method for detecting tobacco impurities based on deep learning according to claim 3, characterized in that: The S4 process is as follows: Assume input feature map The numbers H, W, and C represent the height, width, and number of input channels, respectively. The global average pooling and global maximum pooling operators are expressed using formulas (4) and (5): M c =MaxPool(X c (i,j)) (5) X c represents the c-th channel of the input feature map, and the weight of the c-th channel of the input feature map is expressed as follows: w c =σ(w1δ(w0(G c +M c ))) (6) Where σ represents the Sigmoid activation function, which is used to map the output of the neural network to the probability distribution; δ represents the Relu activation function, which is used to introduce nonlinearity into the neural network so that the neural network can learn more complex functional relationships. w0 and w1 are two linear layers, which are used to learn the relationship between channels of the feature map and generate the excitation weight of each channel: in=σ(C1D k (y)) (8) In order to further integrate channel information, as shown in formula (7), k represents the one-dimensional convolution C1D k The size of the convolution kernel; y is the feature map channel weight; w represents the attention weight after cross-channel fusion.

5. The method for detecting tobacco impurities based on deep learning according to claim 4, characterized in that: S6 uses mAP and Recall at different confidence levels as evaluation indicators to measure model performance. mAP and Recall are defined as follows: In the formula, TP represents the number of impurities that are correctly detected; FN represents the number of impurities that are actually not detected; In the formula, N represents the number of categories; AP i The average precision of the i-th class.