Lightweight plastic bottle color sorting method based on transfer learning
By improving the YOLOv7 network model, the ShuffleNet v2 and CBAM attention modules were introduced, and the EIOU loss function was adopted, the problem of low accuracy in the color classification of plastic bottles was solved, and efficient and accurate classification under complex conditions was achieved.
Patent Information
- Application Number
- CN202411905438.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art has problems with low accuracy in the color classification of plastic bottles, especially when dealing with unstable color information caused by aging, pollution or light changes, and faces the challenge of complex backgrounds and adhesions and overlaps.
The lightweight plastic bottle color sorting method based on transfer learning is adopted. By establishing and expanding the recycling plastic bottle image dataset, the YOLOv7 network model is improved, the ShuffleNet v2 and CBAM attention modules are introduced, and the EIOU loss function is adopted to improve the accuracy and classification efficiency of the model.
It improves the accuracy and accuracy of color classification of plastic bottles, and can maintain good classification results when color information is unstable, meeting the requirements of real-time and efficient.
Smart Images

Figure CN120088522A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine vision detection, and particularly relates to a lightweight plastic bottle color sorting method based on transfer learning. Background Art
[0002] With the continuous expansion of the production and use scale of plastic bottles, they consume a large amount of resources during the production process and generate a large amount of garbage after use, causing serious environmental impacts. Today, when advocating resource conservation and environmental protection, classifying and recycling plastics and reprocessing them can meet the sustainable development strategy. An automated plastic bottle color classification system has emerged. This type of system uses advanced computer vision technology and machine learning algorithms to efficiently identify and classify plastic bottles of different colors, bringing convenience to the subsequent reprocessing and utilization of plastic bottles.
[0003] Target classification methods based on deep learning have now been widely applied in various fields, such as industrial classification detection and classification of agricultural and forestry pests and diseases. The task of target classification is to divide the input image or data into different categories and label the category names and location information.
[0004] Currently, the use of deep learning technology for plastic bottle classification and recycling in the prior art mainly includes classification methods based on image processing and methods based on spectral analysis. Classification methods based on image processing mainly include algorithms such as YOLOv3, FasterR-CNN, and convolutional neural network (CNN). These methods achieve high-precision classification of features such as the color and material of plastic bottles by means of optimizing the loss function, weighted sampling strategy, and improving the network structure. However, these methods may face certain challenges in dealing with overlapping plastic bottles, complex backgrounds, etc. And methods based on spectral analysis, such as using SWIR near-infrared spectroscopy and NIR sensors, classify the material of plastic bottles through spectral reflectance or data characterization process, simplifying the algorithm requirements and improving the classification efficiency.
[0005] Plastic bottle color classification faces multiple challenges in actual operation. The colors of discarded plastic bottles can vary due to aging, pollution, or strong light, and combined with the large variety of colors of plastic bottles themselves, the color information is unstable and difficult to accurately capture. In addition, during recycling, some plastic bottles may also have phenomena such as adhesion and overlap. Therefore, there is an urgent need for a color classification method to improve the classification accuracy of plastic bottles. Summary of the Invention
[0006] Aiming at the problems mentioned in the background art, the present invention proposes a lightweight plastic bottle color sorting method based on transfer learning, which improves the accuracy and precision of the network training model and can still have a good classification effect when the color information of plastic bottles is unstable.
[0007] Technical solution: To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0008] A lightweight plastic bottle color sorting method based on transfer learning, comprising the following steps:
[0009] S1: Establish and expand the recycled plastic bottle image dataset;
[0010] S2: Annotate the data to be trained, and divide the annotated dataset into a training set and a test set;
[0011] S3: Improve the YOLOv7 network model;
[0012] S4: Train the improved YOLOv7 network model, verify the model detection effect, and output the classification result.
[0013] Preferably, in S1, the specific content of establishing and expanding the recycled plastic bottle image dataset is as follows:
[0014] Physically clean the recycled plastic bottles, then use a camera to take pictures of the plastic bottles to obtain plastic bottle images, and construct an initial dataset; classify the recycled plastic bottles according to their colors, including transparent colorless, green, blue, and milky white;
[0015] Use image processing methods to expand the dataset, including translation, rotation, cropping, and adding noise processing.
[0016] Preferably, in S2, the specific content of annotating the data to be trained and dividing the annotated dataset into a training set and a test set is as follows:
[0017] First, use the annotation tool LabelImg to annotate the plastic bottle image dataset, and save the color information and position information of the annotated plastic bottles in the image to the yolo.txt file that can be directly recognized by YOLOv7; finally, divide the annotated dataset into a training set and a test set according to 4:1.
[0018] Preferably, in S3, the specific content of improving the YOLOv7 network model is as follows:
[0019] S31: Integrate the lightweight backbone network ShuffleNet v2 into the backbone network of the YOLOv7 network model;
[0020] S32: Introduce the CBAM attention module into the lightweight backbone network ShuffleNet v2;
[0021] Based on the YOLOv7 model as the basic architecture, in the basic module and spatial downsampling module of ShuffleNet v2, the CBAM attention module is introduced after the connection layer and channel shuffle layer in the left and right branches;
[0022] S33: Use the EIOU loss function instead of the IOU loss function.
[0023] Preferably, in S31, the specific content of integrating the lightweight backbone network ShuffleNet v2 into the backbone network of the YOLOv7 network model is as follows:
[0024] The basic module of ShuffleNet v2 consists of channel-wise grouped convolution and channel shuffle;
[0025] Channel-wise grouped convolution groups the input feature map by channels, and each group performs pointwise convolution, 3×3 depthwise separable convolution, and pointwise convolution. After each convolution, BN and ReLU are used, and then the results of each group are concatenated;
[0026] Channel shuffle promotes information exchange between channels by grouping, rearranging, and concatenating the feature map in the channel dimension;
[0027] The downsampling module of ShuffleNet v2 uses grouped convolution to group the input feature map for channel-wise convolution, and the output is obtained by concatenating in the channel dimension, reducing the size of the feature map by half while reducing the number of parameters and computational complexity.
[0028] Preferably, in S32, the CBAM attention module consists of a channel attention module and a spatial attention module;
[0029] The channel attention module of CBAM first performs average pooling and max pooling on the input feature map; average pooling aggregates spatial information, and max pooling extracts important features; then the two feature maps after average pooling and max pooling are input into a shared network with a multi-layer perceptron and a hidden layer;
[0030] The output formula of the channel attention module is as follows:
[0031] M C (F) = σ(MLP(avgPool(F)) + MLP(MaxPool(F)))
[0032] where F is the input feature map, σ represents the sigmoid function, AvgPool and MaxPool represent average pooling and max pooling respectively, and MLP represents the multi-layer perceptron.
[0033] Preferably, the spatial attention module of CBAM takes the output of the channel attention module as input. First, max-pooling and average-pooling operations are respectively performed in the channel dimension to obtain the significant feature values and average feature intensities at each spatial position, resulting in two single-channel feature maps that are concatenated in the channel dimension to form a two-channel feature map. Then, a 7×7 convolutional layer is used to fuse the spatial information and learn the spatial position relationship, finally generating a spatial attention map;
[0034] The output formula of the spatial attention module is as follows:
[0035] M s (F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)]))
[0036] where F is the input feature map, σ represents the sigmoid function, AvgPool and MaxPool respectively represent average pooling and max pooling, and f 7×7 represents a 7×7 convolutional operation.
[0037] Preferably, in S33, the specific content of using the EIOU loss function instead of the IOU loss function is as follows:
[0038] The calculation formula of the EIOU loss function is as follows:
[0039]
[0040] where L IoU is the IoU loss function, L dis is the distance loss, L asp is the side length loss; ρ 2 (b, b gt ) represents the Euclidean distance between the center points of the predicted box and the ground truth box, b and b gt respectively represent the coordinates of the center points of the predicted box and the ground truth box, c represents the diagonal length of the smallest closed region that can simultaneously contain the predicted box and the ground truth box, ρ 2 (ω, ω gt ) and ρ 2 (h, h gt ) respectively represent the Euclidean distances of the widths and heights of the predicted box and the ground truth box, ω and ω gt respectively represent the widths of the predicted box and the ground truth box, h and h gt respectively represent the heights of the predicted box and the ground truth box, c ω and c h respectively represent the widths and heights of the smallest closed region that can simultaneously contain the predicted box and the ground truth box. Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0041] (1) In the basic module and spatial downsampling module of ShuffleNet v2, the CBAM attention module is introduced after the connection layer and channel shuffle layer in the left and right branches; and the improved ShuffleNet v2 network is used to replace the backbone network of YOLOv7. The design concept of ShuffleNet v2 is based on channel shuffle and separable convolution techniques. This method significantly reduces the computational complexity by rearranging channels and performing separable convolution operations, while ensuring high model accuracy. This improvement enables the model to perform faster inference when dealing with complex tasks, thus meeting the real-time requirements.
[0042] (2) The CBAM attention model of the present invention double-optimizes the input features by integrating channel attention and spatial attention, enhancing the network's ability to accurately capture key features of the target.
[0043] (3) The present invention uses EIOU Loss to replace IOU Loss; although CIOU Loss takes into account the overlapping area of the bounding box, the distance between the center points, and the aspect ratio, its formula only reflects the difference in aspect ratio and fails to comprehensively optimize the actual differences between width / height and confidence. This limitation sometimes hinders the effective optimization of the model. Therefore, based on CIOU Loss, the aspect ratio is split out to propose EIOU Loss. EIOU Loss can more accurately reflect the relationship between width / height and confidence, improving the optimization effect of the model. In addition, EIOU Loss also combines with Focal Loss to focus on high-quality anchor boxes and improve the model's performance when dealing with imbalanced datasets. Through this improvement, EIOU Loss effectively enhances the adaptability and accuracy of the model for object detection tasks, especially when dealing with datasets with significant class imbalance.
[0044] EIOU Loss includes IoU loss, distance loss, and height / width (side length) loss. The height / width loss directly minimizes the differences in the height and width between the predicted target bounding box and the ground truth bounding box, resulting in a faster convergence speed and better localization results.
[0045] (4) The improved algorithm of the present invention is applied to the color classification of plastic bottles. Compared with the original YOLOv7 algorithm, this algorithm has higher classification efficiency and accuracy in the color classification of recycled plastic bottles. Brief Description of the Drawings
[0046] Figure 1 is the flow chart of the method for classifying the colors of recycled plastic bottles by improving YOLOv7 in the present invention;
[0047] Figure 2 is the model diagram of the improved ShuffleNet v2 network in the present invention;
[0048] Figure 3 It is the CBAM structure diagram in the present invention;
[0049] Figure 4 It is the diagram of the CBAM channel attention module in the present invention;
[0050] Figure 5 It is the diagram of the spatial attention module in the present invention;
[0051] Figure 6 It is the diagram of the improved YOLOv7 network model in the present invention;
[0052] Figure 7 It is the schematic diagram of the first group of results of the color classification of partial recycled plastic bottles using the improved YOLOv7 network model in the present invention;
[0053] Figure 8 It is the schematic diagram of the second group of results of the color classification of partial recycled plastic bottles using the improved YOLOv7 network model in the present invention;
[0054] Figure 9 It is the schematic diagram of the third group of results of the color classification of partial recycled plastic bottles using the improved YOLOv7 network model in the present invention;
[0055] Figure 10 It is the schematic diagram of the fourth group of results of the color classification of partial recycled plastic bottles using the improved YOLOv7 network model in the present invention. Detailed implementation manners
[0056] The following combines specific embodiments to further clarify the present invention. The embodiments are implemented on the premise of the technical solution of the present invention. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0057] The lightweight plastic bottle sorting method based on transfer learning provided in this embodiment. First, establish and expand the recycled plastic bottle image dataset, and label the dataset images; then, in the basic module and spatial downsampling module of ShuffleNet v2, introduce the CBAM attention module after the connection layer (Concat) and channel shuffle layer (Channel Shuffle) in the left and right branches; and replace the backbone network of YOLOv7 with the improved ShuffleNet v2 network; use EIOULoss (efficient intersection over union loss) instead of IOU Loss (intersection over union loss); use the training set divided from the dataset to train the improved YOLOv7 network model to obtain the final plastic bottle color classification model, and finally output the classification result. The detailed implementation manners are achieved through the following steps (as Figure 1 shown):
[0058] S1: Establish and expand the image dataset of recycled plastic bottles;
[0059] Perform physical cleaning and other operations on the recycled plastic bottles, then use a camera to take pictures of the plastic bottles to obtain plastic bottle images, and construct an initial dataset; classify the recycled plastic bottles according to their colors, including transparent and colorless, green, blue, and milky white.
[0060] Use image processing methods to expand the dataset. The image processing methods include geometric transformations such as translation, rotation, and cropping, as well as methods such as adding noise. Through these methods, richer samples can be obtained, the performance of the model can be improved, the risk of overfitting can be reduced, and the robustness and accuracy of the training model can be enhanced.
[0061] S2: Annotate the data to be trained, and divide the annotated dataset into a training set and a test set;
[0062] Use the data annotation tool LabelImg to annotate the image dataset of recycled plastic bottles. Select the rectangular box as the annotation tool and make the annotation box completely fit the color area of the plastic bottle to be annotated; then save the color information and position information of the annotated plastic bottles in the image to the yolo.txt file that can be directly recognized by YOLOv7; finally, divide the annotated dataset into a training set and a test set according to 4:1.
[0063] In this embodiment, the colors of the plastic bottles are numbered. 0 is transparent and colorless (transparency); 1 is blue (blue); 2 is green (green); 3 is milky white (milky).
[0064] S3: Improve the YOLOv7 network model;
[0065] The improvement of YOLOv7 is carried out from the following two aspects:
[0066] (1) Introduce the CBAM attention module in ShuffleNet v2 and replace the backbone network of YOLOv7 with the improved network.
[0067] (2) Adopt EIOU Loss instead of IOU Loss. This loss function can more accurately reflect the relationship between confidences higher than a certain level and improve the optimization effect of the model. The specific content is as follows:
[0068] S31: Integrate ShuffleNet v2 into the backbone network of YOLOv7;
[0069] Integrate ShuffleNet v2 into the backbone network of YOLOv7. ShuffleNet v2 is a lightweight and efficient convolutional neural network model that greatly reduces the model complexity and computational cost while ensuring the model accuracy, and improves the computational efficiency and inference speed of the model.
[0070] The basic module of ShuffleNet v2 consists of channel-wise grouped convolution and channel shuffle. Channel-wise grouped convolution groups the input feature map by channels, and for each group, pointwise convolution, 3×3 depthwise separable convolution, and pointwise convolution are performed. After each convolution, BN (Batch Normalization) and ReLU (Rectifier Linear Unit activation function) are used, and then the results of each group are concatenated. Channel shuffle promotes information exchange between channels by grouping, rearranging, and concatenating the channels of the feature map in the channel dimension.
[0071] The downsampling module of ShuffleNet v2 uses grouped convolution to perform channel-wise convolution on the input feature map in groups, and the output is obtained by concatenating in the channel dimension, reducing the size of the feature map by half while reducing the number of parameters and computational cost. Specifically, in the basic unit of ShuffleNet v2, the left branch adopts the identity mapping method to ensure that the input and output channels of the two branches are the same, and at the same time reduces the fragmentation degree. Instead of using add for the outputs of the two branches, the Concat operation is used, and then Channel Shuffle is performed to ensure information exchange. In the ShuffleNet v2 unit for spatial downsampling, Channel Split is no longer used. Each branch performs downsampling with a convolution stride of 2, and finally the results are combined to halve the spatial size of the feature map and double the number of channels.
[0072] S32: Introduce the CBAM attention module into the ShuffleNet v2 network model;
[0073] The CBAM attention module consists of a channel attention module and a spatial attention module. By integrating channel attention and spatial attention, the input features are doubly optimized. In this way, the model can simultaneously focus on valuable channels and spatial positions, thereby improving the model representation ability and decision-making accuracy, enabling the network to accurately capture the key features of the target.
[0074] The channel attention module of CBAM first performs average pooling and max pooling on the input feature map. Average pooling aggregates spatial information, and max pooling extracts important features, thus obtaining two different feature maps. These two feature maps are input into a shared network with a multi-layer perceptron and a hidden layer. This network learns the correlations between channels and determines the importance weights of each channel. The output formula of the channel attention module is as follows:
[0075] M C (F) = σ(MLP(AvgPool)F)) + MLP(MaxPool(F)))
[0076] Where F is the input feature map, σ represents the sigmoid function, AvgPool and MaxPool represent average pooling and max pooling respectively, and MLP represents the multi-layer perceptron.
[0077] The spatial attention module of CBAM takes the output of the channel attention module as the input. First, max pooling and average pooling operations are respectively performed in the channel dimension to obtain the significant feature values and average feature intensities at each spatial position, resulting in two single-channel feature maps that are concatenated in the channel dimension into a two-channel feature map, and then the 7×7 convolutional layer is used to fuse the spatial information and learn the spatial position relationship, finally generating the spatial attention map. The output formula of the spatial attention module is as follows:
[0078] M s (F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)]))
[0079] Where F is the input feature map, σ represents the sigmoid function, AvgPool and MaxPool represent average pooling and max pooling respectively, and f 7×7 represents the 7×7 convolutional operation.
[0080] S33: Use EIOU Loss instead of IOU Loss;
[0081] Based on CIOU Loss, the aspect ratio is split out, and EIOU Loss is proposed; EIOU Loss can more accurately reflect the relationship between width, height and confidence, improving the optimization effect of the model. In addition, EIOU Loss also combines with FocalLoss to focus on high-quality anchor boxes and improve the performance of the model when dealing with unbalanced datasets. Through this improvement, EIOU Loss effectively improves the adaptability and accuracy of the model for object detection tasks, especially when dealing with datasets with significant class imbalance.
[0082] The calculation formula of the EIOU loss function is as follows:
[0083]
[0084] Among them, L IoU is the IoU (Intersection over Union) loss function, that is, the Intersection over Union loss between the predicted bounding box and the ground truth bounding box. L dis is the distance loss, which is used to measure the distance difference between the center points of the predicted bounding box and the ground truth bounding box. L asp is the side length loss, which is used to penalize the side length (or aspect ratio) of the predicted bounding box to avoid the side length being wrongly enlarged or reduced. ρ 2 (b, b gt ) represents the Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box. b and b gt represent the coordinates of the center points of the predicted bounding box and the ground truth bounding box respectively. c represents the diagonal length of the smallest closed region that can contain both the predicted bounding box and the ground truth bounding box. ρ 2 (ω, ω gt ) and ρ 2 (h, h gt ) represent the Euclidean distances of the widths and heights of the predicted bounding box and the ground truth bounding box respectively. ω and ω gt represent the widths of the predicted bounding box and the ground truth bounding box respectively. h and h gt represent the heights of the predicted bounding box and the ground truth bounding box respectively. c ω and c h represent the widths and heights of the smallest closed region that can contain both the predicted bounding box and the ground truth bounding box respectively.
[0085] S4: Train the improved YOLOv7 network model, verify the detection effect of the model, and output the classification result.
[0086] During the process of training the network model, first place the divided dataset in the data folder of YOLOv7; then configure a yaml file for the self-built dataset and set the corresponding training parameters in the train.py script of YOLOv7. After the preparation work is completed, place the configured yaml file and the improved YOLOv7 network model in the computer with the configured running environment, and start the training process using the plastic bottle image dataset with the category information already annotated. During the training, the system will output the training effect of each stage and observe the mAP (mean average precision) value of the training in real time by setting the process monitoring parameters. When the training is over, save the weights of the trained network model.
[0087] Import the dataset into the improved YOLOv7 network model and set the following parameters: set the input image size to 224*224, set the batch_size (batch size) to 4, set the number of training epochs to 100, and then train the model.
[0088] Perform performance verification on the improved YOLOv7 network model.
[0089] The experimental environment is as follows: CPU: 12th Gen Intel(R) Core(TM) i9-12900H 2.50GHz, memory 32GB, graphics card is NVIDIA GTX 3060, Cuda version 11.6, and the framework used is Pytorch.
[0090] The experimental results are as follows: The accuracy value of the improved YOLOv7 network model is 96.5%, the recall rate is 95.6%, and the mAP value is 97.7%. For the unimproved YOLOv7 network model, the accuracy value is 93.5%, the recall rate is 91.3%, and the mAP value is 93.2%. It can be seen that multiple performance indicators of the improved YOLOv7 model are better than those of the original model.
[0091] Some of the output classification results are as Figures 7 - 10 shown. It can be found that the improved YOLOv7 detection model has significant effects on the color classification of recycled plastic bottles.
[0092] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A lightweight plastic bottle color sorting method based on transfer learning, characterized by: The following steps are involved: S1: Build and expand a dataset of recycled plastic bottle images; S2: Label the data to be trained and divide the labeled data set into a training set and a test set; S3: Improve the YOLOv7 network model; S4: Train the improved YOLOv7 network model, verify the model detection effect, and output the classification results.
2. The method for color sorting of lightweight plastic bottles based on transfer learning according to claim 1 is characterized in that: In S1, the specific contents of establishing and expanding the recycled plastic bottle image dataset are as follows: The recycled plastic bottles are physically cleaned, and then photographed with a camera to obtain images of the plastic bottles and construct an initial data set; the recycled plastic bottles are classified according to color, including transparent and colorless, green, blue, and milky white; Image processing methods are used to expand the data set, including translation, rotation, cropping, and adding noise processing.
3. The method for color sorting of lightweight plastic bottles based on transfer learning according to claim 1 is characterized in that: In S2, the data to be trained is labeled, and the labeled data set is divided into a training set and a test set. The specific contents are as follows: First, the plastic bottle image dataset is annotated using the labeling tool LabelImg, and the color and position information of the annotated plastic bottles in the image are saved in the yolo.txt file that can be directly recognized by YOLOv7; finally, the annotated dataset is divided into a training set and a test set.
4. The method for color sorting of lightweight plastic bottles based on transfer learning according to claim 1 is characterized in that: In S3, the specific contents of improving the YOLOv7 network model are: S31: Integrate the lightweight backbone network ShuffleNet v2 into the backbone network of the YOLOv7 network model; S32: Introducing the CBAM attention module into the lightweight backbone network ShuffleNet v2; Taking the YOLOv7 model as the basic architecture, in the basic module and spatial downsampling module of ShuffleNet v2, the CBAM attention module is introduced after the left and right branches pass through the connection layer and channel shuffle layer; S33: Use EIOU loss function instead of IOU loss function.
5. The method for color sorting of lightweight plastic bottles based on transfer learning according to claim 4 is characterized in that: In S31, the specific contents of integrating the lightweight backbone network ShuffleNet v2 into the backbone network of the YOLOv7 network model are as follows: The basic module of ShuffleNet v2 consists of channel-by-channel grouped convolution and channel shuffle; Channel-by-channel grouped convolution groups the input feature maps by channel, performs point-by-point convolution, 3×3 depth-wise separable convolution, and point-by-point convolution on each group, uses BN and ReLU after each convolution, and then concatenates the results of each group; Channel shuffling is to promote information exchange between channels by grouping, rearranging, and splicing feature maps in the channel dimension; The downsampling module of ShuffleNet v2 uses grouped convolution to group the input feature maps and perform channel-by-channel convolution, and then concatenate the output in the channel dimension, thereby halving the size of the feature map and reducing the number of parameters and computation.
6. The method for color sorting of lightweight plastic bottles based on transfer learning according to claim 4 is characterized in that: In S32, the CBAM attention module consists of a channel attention module and a spatial attention module; The channel attention module of CBAM first performs average pooling and maximum pooling on the input feature map; average pooling aggregates spatial information, and maximum pooling extracts important features; then the two feature maps after average pooling and maximum pooling are input into a shared network with a multi-layer perceptron and one hidden layer; The output formula of the channel attention module is as follows: M C (F) = σ(MLP(AvgPool(F))+MLP(MaxPool(F))) where F is the input feature map, σ represents the sigmoid function, AvgPool and MaxPool represent average pooling and maximum pooling respectively, and MLP represents multi-layer perceptron.
7. The method for color sorting of lightweight plastic bottles based on transfer learning according to claim 6 is characterized in that: The spatial attention module of CBAM takes the output of the channel attention module as input, first performs maximum pooling and average pooling operations in the channel dimension to obtain the significant feature value and average feature strength of each spatial position, obtains two single-channel feature maps, and then splices them into a dual-channel feature map in the channel dimension. Then, the spatial information is fused through a 7×7 convolutional layer and the spatial position relationship is learned, and finally a spatial attention map is generated; The output formula of the spatial attention module is as follows: M s (F)=σ(f 7×7 ([AvgPool(F);MaxPool(F)])) Among them, F is the input feature map, σ represents the sigmoid function, AvgPool and MaxPool represent average pooling and maximum pooling respectively, and f 7×7 Represents a 7×7 convolution operation.
8. The method for color sorting of lightweight plastic bottles based on transfer learning according to claim 4, characterized in that: In S33, the specific content of using the EIOU loss function instead of the IOU loss function is: The calculation formula of the EIOU loss function is as follows: Among them, L IoU is the IoU loss function, L dis is the distance loss, L asp is the edge length loss; ρ 2 (b,b gl ) represents the Euclidean distance between the center points of the predicted box and the true box, b and b gt Represent the coordinates of the center points of the predicted box and the real box respectively, c represents the diagonal length of the minimum closure area that can contain both the predicted box and the real box, ρ 2 (ω,ω gt ) and ρ 2 (h,h gt ) represent the Euclidean distances of the width and height of the predicted box and the true box, ω and ω respectively. gt Represent the width of the predicted box and the real box, h and h respectively gt Indicates the height of the predicted box and the real box, c ω and c h They represent the width and height of the minimum closure area that can contain both the predicted box and the true box.
Citation Information
Cited By
Multi-category plastic automatic backflow sorting method and system and storage medium
CN120680655A