Garbage classification method and system based on improved loss function and medium

By introducing improved loss functions into the garbage classification model and combining the feature extraction method of the CLIP model and the Transformer module, the problems of insufficient generalization capabilities and data imbalance in the existing model are solved, and higher accuracy and faster training efficiency are achieved.

CN120147731APending Publication Date: 2025-06-13NANJING FORESTRY UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510226303.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When handling complex and diverse garbage images, the existing garbage classification model lacks generalization ability and is prone to overfitting. Due to data imbalance, traditional loss functions converge slowly, making it impossible to effectively distinguish garbage categories with low sample sizes.

Method used

The improved loss function is adopted, combining cross entropy loss, focus loss and KL divergence, and the loss weight is dynamically adjusted to enhance the generalization ability and robustness of the model. At the same time, the pre-trained CLIP model and Transformer module are used to extract deep features of the image and enhance the model's feature expression ability.

Benefits of technology

It significantly improves the accuracy and efficiency of the model in garbage classification tasks, enhances the generalization ability and robustness of the model, can capture effective features more quickly, and reduces training time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147731A_ABST
    Figure CN120147731A_ABST
Patent Text Reader

Abstract

The invention provides a garbage classification method and system based on an improved loss function and a medium. The method comprises the steps that 1, a marked garbage classification data set is collected; 2, dividing a data set, and setting a verification set; step 3, preprocessing the trained data set; 4, constructing a classification network model; 5, setting an optimizer and a loss function, preparing to load a data set into the network, and carrying out model training and verification; wherein the loss function is an improved loss function; step 6, storing the training model with the best verification effect for test evaluation; and 7, carrying out actual garbage classification by utilizing the model. According to the invention, the weight of loss can be dynamically adjusted according to the actual condition of each sample. The loss function can more effectively solve the problems of class imbalance and label noise, so that the generalization ability and robustness of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of environmental protection technologies, and particularly relates to a garbage classification method, system and medium based on an improved loss function. Background Art

[0002] In recent years, with the gradual improvement of garbage classification and environmental protection awareness, garbage classification methods based on artificial intelligence have become a research hotspot. In particular, the progress of deep learning technology has made garbage classification methods based on computer vision have broad application prospects. Garbage classification methods mainly rely on the analysis of image data, and common technical means include convolutional neural networks (CNNs), self-attention mechanisms, transformers, and image feature extraction models such as the CLIP (Contrastive Language-Image Pretraining) model.

[0003] The technical means of garbage classification mainly focus on the following directions:

[0004] Methods based on convolutional neural networks (CNNs): Most traditional garbage classification methods use CNNs as the basic model, and classify by extracting features, performing convolutional operations and pooling operations on images. However, when dealing with complex and diverse garbage images, CNN models often have difficulty capturing deep features in the images, resulting in insufficient generalization ability.

[0005] Vision models based on transformers: The transformer architecture was initially applied to natural language processing (NLP) tasks and has also been successfully introduced into the field of computer vision in recent years, especially showing outstanding performance in tasks such as image classification and object detection. The self-attention mechanism can effectively capture long-range dependency information in images, making the model more robust when dealing with complex images. However, a pure transformer model requires a large amount of data and computing resources and has a slow processing speed.

[0006] Methods based on pre-trained models: The CLIP (Contrastive Language-Image Pretraining) model has mastered a general image-text correspondence relationship through contrastive learning of a large number of images and texts, and can extract image features with high semantic associations. However, although the CLIP model has achieved excellent results in many vision tasks, its application in specific tasks such as garbage classification is still in the exploration stage.

[0007] Most of the existing solutions to the garbage classification problem use traditional CNN models and common deep learning models. However, traditional models lack strong context understanding and semantic analysis capabilities when faced with complex image classification tasks, especially when dealing with garbage classification tasks with complex backgrounds and high noise interference.

[0008] Existing models are usually trained using simple loss functions (such as cross entropy loss), which lack adaptability to imbalanced categories, difficult categories, and label noise, and can easily lead to model overfitting or reduced recognition accuracy in practical applications. At the same time, due to the imbalance in the acquisition of domestic waste data, there is a large gap between the number of different waste image categories. Traditional loss functions converge slowly and cannot distinguish between waste categories with a low number of samples. Summary of the invention

[0009] To solve the above problems, the present invention provides a garbage classification method, system and medium based on an improved loss function, the purpose of which is to solve the problems in the prior art that the model is easily overfitted or the recognition accuracy is reduced in practical applications. At the same time, due to the imbalance in the data acquisition of domestic garbage, there is a large gap between the number of different garbage image categories, the traditional loss function converges slowly, and cannot distinguish between garbage categories with a low sample number.

[0010] In a first aspect, the present invention provides a garbage classification method based on an improved loss function, comprising the following steps:

[0011] Step 1: Collect labeled garbage classification data sets;

[0012] Step 2: Divide the data set and set the validation set;

[0013] Step 3, preprocessing the training data set;

[0014] Step 4, construct a classification network model;

[0015] Step 5: Set the optimizer and loss function, load the preprocessed data set into the network, and train and verify the model;

[0016] Step 6: Save the training model with the best verification effect for test evaluation;

[0017] Step 7: Use the model to perform actual garbage classification.

[0018] Furthermore, the final loss function in step 5 is:

[0019]

[0020] Among them, L totalrepresents the final loss function, i represents the i-th sample, N is the total number of samples, ω and b are learnable parameters, p i (y i ) represents the correct label information y i probability of, y i represents the label information, represents the smoothed label information, c is the c-th category, C is the total number of categories, θ represents the balance factor, γ is the adjustment factor, and CE represents the standard cross-entropy loss function, p ref,i represents the reference distribution, p i,c represents the predicted probability of the c-th category of the i-th sample; ∈ represents the smoothing weight.

[0021] Further, the step 3 includes: picture cropping, data augmentation, and picture normalization processing. Data augmentation includes flipping and rotating at different angles.

[0022] Further, the normalization formula for picture normalization processing is as follows:

[0023] Among them, x is the normalized data; x is the original data; μ is the mean of the original data, and σ is the variance of the original data.

[0024] Further, the step 4 includes:

[0025] Using the pre-trained CLIP model to first perform preliminary feature extraction to obtain a feature vector as the shallow feature vector, and then perform deep extraction on the feature vector;

[0026] Input the feature vector into the transformer module, and use four attention heads and three layers of attention networks for deep extraction. After completion, reshape it into a matrix form together with the original shallow feature for vector splicing;

[0027] Use the convolutional block to perform feature fusion on the two layers of features and input it into the next layer of convolutional network. Finally, flatten the vector and input it into the two-layer fully connected network to obtain the classification result of the specific garbage and determine the category of the final garbage.

[0028] Further, the main body of the CLIP model includes the following steps:

[0029] Divide the input three-channel picture into several squares of 32×32 size;

[0030] Flatten the pixels of each square into a feature vector;

[0031] The obtained feature vector is embedded into a high-dimensional space through a linear layer, usually with a dimension of 512;

[0032] Add position encoding to each image patch to facilitate capturing the spatial relationship between patches;

[0033] Add a CLS token to each image patch feature vector to record the high-level information of the overall image category;

[0034] Input the CLS token and the image patch features into the encoder of the Transformer module, and re-encode the features through the self-attention mechanism, and finally output a feature vector containing the image category and local information.

[0035] Furthermore, the Transformer module includes the following steps:

[0036] Embed the input feature vector into a high-dimensional space and add fixed position encoding;

[0037] Through the core self-attention mechanism, the input elements can capture the relationships between various feature vectors, and adopt the attention calculation mode of key-value pairs;

[0038] Pass the obtained feature vector through a feed-forward network for non-linear transformation and layer normalization;

[0039] Output the final feature vector through a deep residual connection.

[0040] Furthermore, the self-attention calculation formula is as follows: Among them, the formula of the softmax function is:

[0041]

[0042] Among them, Q represents the query vector, K represents the key vector, and V represents the value vector. is the probability value after softmax, i is the i-th dimension of the sample, n is the dimension of the sample, x is the sample data, and d k is the set hyperparameter used to scale the dot product of the attention scores.

[0043] Furthermore, feature fusion includes the following steps:

[0044] Reshape the obtained shallow features and deep features into matrix vectors and concatenate them to obtain a feature map;

[0045] Pass the obtained feature map through a convolutional kernel to convolve the feature maps of two channels into a feature fusion map of one channel;

[0046] Then perform BN layer normalization and Relu activation on the feature fusion map, and input it into the next layer of the convolutional network again;

[0047] Through a convolutional network; perform pooling processing on the feature map using a MaxPooling layer;

[0048] Perform BN layer normalization and Relu activation processing again to obtain the final highly abstract feature map;

[0049] Flatten the obtained feature map to obtain a feature vector;

[0050] Pass through two fully connected layers again to obtain the final probability value. First, perform softmax processing for subsequent calculation of the loss function.

[0051] In a second aspect, the present invention provides a garbage classification system based on an improved loss function, which is characterized by including:

[0052] A data module for collecting a labeled garbage classification data set; dividing the data set and setting a validation set; preprocessing the training data set;

[0053] A construction module for constructing a classification network model;

[0054] A training module for setting an optimizer and a loss function, loading the data set into the network for model training and validation;

[0055] A testing module for saving the training model with the best validation effect for testing and evaluation;

[0056] A classification module for performing actual garbage classification using the model.

[0057] In a third aspect, the present invention provides a computer-readable storage medium including a computer program, and when the computer program is executed by a processor, the steps of the above classification method are implemented.

[0058] Advantageous effects: Through the Transformer module and the pre-trained CLIP model, the powerful semantic understanding ability of the latter is fully utilized, and combined with the transformer module, the feature expression ability of the model is enhanced, enabling the model to process more complex visual information, thereby improving the accuracy of the model. In addition, an adaptive loss function combining cross-entropy loss, focal loss, and KL divergence is proposed, which can dynamically adjust the weight of the loss according to the actual situation of each sample. This loss function can more effectively handle the problems of class imbalance and label noise, thereby improving the generalization ability and robustness of the model.

[0059] By fusing the adaptive loss function with the multi-layer feature extraction of the deep network, the present invention not only improves the stability of the model during the training process, but also significantly improves the accuracy and efficiency of the model in the actual garbage classification task through the optimization of each layer of features. Description of the Drawings

[0060] Figure 1 It is a diagram for model construction;

[0061] Figure 2 It is a diagram showing the change of model training loss;

[0062] Figure 3 It is a diagram showing the change of model training loss and accuracy;

[0063] Figure 4 It is a preview diagram of prediction effect. Specific implementation manners

[0064] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention; obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0065] The present invention will be further described below in conjunction with the accompanying drawings.

[0066] Refer to Figures 1-4 , the implementation steps of the present invention are as follows:

[0067] S1. Collect a labeled garbage classification data set (a total of 14,000 pictures, 40 classifications);

[0068] S2. Divide the data set according to the ratio of 7:1.5:1.5. Among them, a validation set is set to record the generalization ability of the model during training;

[0069] S3. Preprocess the training data set, including picture cropping, data augmentation, and picture normalization processing, where

[0070] The size of the cropped picture is 224×224. Data augmentation includes flipping and rotating at different angles. The three-channel parameters for picture normalization are std = [0.229, 0.224, 0.225], mean = [0.485, 0.456, 0.406], which are the same as the three-channel normalization parameters of ImageNet, facilitating the preliminary processing of the CLIP model; the normalization formula is as follows:

[0071]

[0072] S4. First, use the pre-trained CLIP model to perform preliminary feature extraction to obtain a feature vector as the shallow feature vector, and then perform deep extraction on the feature vector. Input the feature vector into the transformer module, and use four attention heads and three layers of attention networks for deep extraction. After completion, reshape it into a matrix form together with the original shallow feature and perform vector splicing. Then use a convolutional block to fuse the features of the two layers and input them into the next convolutional network. Finally, flatten the vector and input it into two fully connected networks to obtain the classification result of the specific garbage, and then refer to the garbage category table to determine the category of the final garbage.

[0073] The main body of the CLIP model includes the following steps:

[0074] S4.1. Divide the input three-channel image into several squares of size 32×32;

[0075] S4.2. Flatten the pixels of each square into a feature vector;

[0076] S4.3. Embed the obtained feature vector into a high-dimensional space through a linear layer, usually with a dimension of 512;

[0077] S4.4. Add position encoding to each image patch to facilitate capturing the spatial relationship between patches;

[0078] S4.5. Add a CLS token to each image patch feature vector to record the high-level information of the overall image category;

[0079] S4.6. Input the CLS token and the image patch features into the encoder of the transformer module, and re-encode the features through the self-attention mechanism, and finally output a feature vector containing the image category and local information;

[0080] The above-mentioned transformer module includes the following steps:

[0081] Embed the input feature vector into a high-dimensional space and add fixed position encoding;

[0082] Through the core self-attention mechanism, the input elements can capture the relationships between various feature vectors, and adopt the attention calculation mode of key-value pairs; The self-attention calculation formula is as follows: Where Q, K, and V represent the query vector, key vector, and value vector respectively, and the formula of the softmax function is:

[0083]

[0084] The obtained feature vectors are subjected to non-linear transformation through a feed-forward network and layer normalization is performed.

[0085] Through deep residual connections, the final feature vectors are output.

[0086] The feature fusion module includes the following steps:

[0087] Both the obtained shallow features and deep features are reshaped into matrix vectors of size 32×32 and processed to obtain a feature map of size (batch_size, 2, 32, 32).

[0088] The obtained feature map is convolved through a convolutional kernel with kernel_size of 3 and padding of 1, and the feature maps of two channels are convolved into a feature fusion map of one channel.

[0089] Then, BN layer normalization and Relu activation are performed on the feature fusion map, and it is input into the next layer of the convolutional network again.

[0090] Similarly, through a convolutional network with input and output channels of 1, kernel_size of 3, and padding of 1;

[0091] Pooling processing of the feature map is performed through the MaxPooling layer;

[0092] BN layer normalization and Relu activation processing are performed again to obtain the final highly abstract feature map;

[0093] The final network output is as follows:

[0094] The obtained feature map is flattened by Flatten to obtain feature vectors of (batch size, 256);

[0095] It passes through two fully connected layers nn.Linear(256, 128) and nn.Linear(128, 40) again;

[0096] The obtained probability values of the final 40 categories are first subjected to softmax processing for subsequent calculation of the loss function;

[0097] The optimizer and loss function are set. The Adam adaptive learning rate optimizer is selected to update the gradient. For the setting of the loss function, the label smoothing cross-entropy function and the focal loss are combined to innovate a new loss function:

[0098]

[0099] Among them, label smoothing processing is performed on the one-hot encoding of the labels, and at the same time, the focal loss is added to pay attention to the difficult-to-classify garbage categories;

[0100] Among them, α and β represent the weights of different loss functions (by default, α = 0.7 and β = 0.3), and the formula for the smoothed label is:

[0101]

[0102] For the first part of the loss function, it is an improved label smoothing loss function. ∈ represents the smoothing weight. For one-hot

[0103] encoded labels, such as [1, 0, 0], are smoothed to [0.8, 0.1, 0.1]. This makes the model not overly confident in its correct category but pay attention to the losses of other categories and combines with the original cross-entropy loss function.

[0104] The second part of the loss function is the focal loss function, where θ represents the balance factor used to control the weights between different categories, γ is the adjustment factor used to control the influence of easy and hard samples, and CE represents the standard cross-entropy loss function. Adding the focal loss function solves the problem of class imbalance during training, balances the weights for each category, and enhances the generalization ability of the model;

[0105] However, simply combining the two still has many problems. One is that the values of the weights α and β of the two loss functions cannot be controlled and need to be adjusted through multiple experiments. The other is that the gradient descent methods of the two loss functions cannot be made consistent. Therefore, adaptive weights and KL divergence are introduced to solve these two problems respectively.

[0106] Among them, the formula for the adaptive weight is as follows:

[0107]

[0108] Both ω and b are learnable parameters that can be adjusted in a timely manner during training to ensure the balance of the two loss functions. Thus, the new L main is:

[0109]

[0110] Then, using the adaptive parameter λ i construct the reference distribution p ref,i , as follows:

[0111] p ref,i = (1 - λ i )p i (c) + λ i y i (c)

[0112] Based on this, the KL divergence term is defined as:

[0113]

[0114] The final loss function is defined as:

[0115]

[0116] In summary:

[0117]

[0118] Among them, L total represents the final loss function, i represents the i-th sample, N is the total number of samples, ω and b are learnable parameters, p i (y i ) represents the probability of the correct label information y i of, y i represents the label information, represents the smoothed label information, c is the c-th category, C is the total number of categories, θ represents the balance factor, γ is the adjustment factor, and CE represents the standard cross-entropy loss function, p ref,i represents the reference distribution, p i,c represents the predicted probability of the c-th category of the i-th sample; ∈ represents the smoothing weight.

[0119] The improved loss function not only saves the time for adjusting the hyperparameter weights, but also ensures that the gradient descent directions of the two loss functions are consistent, greatly improving the convergence speed and also enhancing the prediction performance of the model.

[0120] S6. Prepare and load the dataset into the network for model training and validation;

[0121] S7. Save the training model with the best validation effect for test evaluation;

[0122] S8. Use the model for actual garbage classification.

[0123] The garbage classification model of the present invention has significant technical advantages and effects. Compared with the prior art, it shows higher accuracy, faster convergence speed, and higher training efficiency. The specific advantages are as follows:

[0124] Note: Training uses a 4060 graphics card and sets the batch size to 32.

[0125] By introducing the combination of the CLIP model and the Transformer architecture, the present invention can achieve a validation accuracy of 92.5% in the garbage classification task. This significantly improves the classification performance compared to the 81.5% accuracy of using the YOLO model. The CLIP model provides strong image semantic representation ability, and the Transformer further enhances the model's understanding of image context information, thereby improving the classification accuracy.

[0126] During the training process, the model of the present invention only needs about 15 rounds to converge, while the YOLO model needs 40 rounds to converge. This shows that the model of the present invention can capture effective features more quickly during the training process, reducing the training time, improving the efficiency, and greatly saving computing power resources and power resources.

[0127] Experiments show that using the model of the present invention only takes about 1.5 minutes for each round of training on a 4060 graphics card, while the YOLO model takes a longer time to complete each round of training. This significant reduction in training time not only improves the overall efficiency but also reduces the consumption of hardware resources, having high practical value.

[0128] As described above, it is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its improved concept, making equivalent substitutions or changes, should be covered by the protection scope of the present invention.

Claims

1. A garbage classification method based on an improved loss function, characterized in that: The following steps are involved: Step 1: Collect labeled garbage classification data sets; Step 2: Divide the data set and set the validation set; Step 3, preprocessing the training data set; Step 4, construct a classification network model; Step 5: Set the optimizer and loss function, load the preprocessed data set into the network, and train and verify the model; Step 6: Save the training model with the best verification effect for test evaluation; Step 7: Use the model to perform actual garbage classification.

2. A garbage classification method based on improved loss function according to claim 1, characterized in that: The final loss function in step 5 is: Among them, L total represents the final loss function, i represents the i-th sample, N is the total number of samples, ω and b are learnable parameters, and p i (y i ) indicates the correct label information y i The probability of y i Indicates label information. represents the smoothed label information, c is the cth category, C is the total number of categories, θ represents the balance factor, γ is the adjustment factor, CE represents the standard cross entropy loss function, p ref,i represents the reference distribution, p i,c represents the predicted probability of the cth category of the ith sample; ∈ represents the smoothing weight.

3. A garbage classification method based on improved loss function according to claim 2, characterized in that: The standardization formula for image standardization is as follows: Among them, x, is the standardized data; x is the original data; μ is the mean of the original data, and σ is the variance of the original data.

4. The garbage classification method based on improved loss function according to claim 1, characterized in that: The step 4 comprises: Use the pre-trained CLIP model to perform preliminary feature extraction to obtain a feature vector as a shallow feature vector, and then perform deep extraction on the feature vector; The feature vector is input into the transformer module, and four attention heads and a three-layer attention network are used for deep extraction. After the extraction is completed, it is reshaped into a matrix form together with the original shallow features, and the vector is spliced; The convolutional blocks are used to fuse the features of the two layers and input them into the convolutional network of the next layer. Finally, the vectors are flattened and input into the two-layer fully connected network to obtain the classification results of the specific garbage and determine the final category of the garbage.

5. A garbage classification method based on improved loss function according to claim 4, characterized in that: The CLIP model body includes the following steps: Divide the input three-channel image into several 32×32 blocks; Expand the pixels of each block into a feature vector through Flatten; The obtained feature vector is embedded into a high-dimensional space through a linear layer, usually with a dimension of 512; Add position encoding to each image block to capture the spatial relationship between blocks; Add a CLS tag to each image block feature vector to record the high-level information of the overall image category; The CLS tag and image block features are input into the encoder of the transformer module together, and the features are re-encoded through the self-attention mechanism, and finally a feature vector containing image category and local information is output.

6. A garbage classification method based on improved loss function according to claim 4, characterized in that: The transformer module contains the following steps: Embed the input feature vector into a high-dimensional space and add a fixed positional encoding; Through the core self-attention mechanism, the input elements can capture the relationship between each feature vector, using the key-value pair attention calculation mode; The obtained feature vector is transformed nonlinearly through a feedforward network and then normalized at the layer; Through deep residual connection, the final feature vector is output.

7. A garbage classification method based on improved loss function according to claim 6, characterized in that: The self-attention calculation formula is as follows: The formula of the softmax function is: Among them, Q represents the query vector, K represents the key vector, and V represents the value vector. is the probability value after softmax, i is the i-th dimension of the sample, n is the dimension of the sample, x is the sample data, d k is a hyperparameter set to scale the dot product of the attention scores.

8. The garbage classification method based on improved loss function according to claim 6, characterized in that: Feature fusion includes the following steps: The shallow features and deep features obtained are reshaped into matrix vectors and concatenated to obtain feature maps; The obtained feature map is passed through the convolution kernel to convolve the feature maps of the two channels into a feature fusion map of one channel; The feature fusion map is then normalized by the BN layer and activated by Relu, and then input into the convolutional network of the next layer again; Through convolutional networks; Perform pooling of the MaxPooling layer on the feature map; Perform BN layer normalization and Relu activation again to obtain the final highly abstract feature map; Flatten the obtained feature map to obtain the feature vector; After passing through two fully connected layers again, the final probability value is first processed by softmax to facilitate the subsequent loss function calculation.

9. A garbage classification system based on an improved loss function, characterized in that: include: Data module, used to collect labeled garbage classification data sets; Divide the data set and set the validation set; Preprocess the training data set; Building module, used to build classification network model; The training module is used to set the optimizer and loss function, prepare the data set to be loaded into the network, and train and verify the model; The test module is used to save the best training model for test evaluation; The classification module is used to use the model to perform actual garbage classification.

10. A computer-readable storage medium storing one or more programs, characterized in that: The one or more programs include instructions, which, when executed by a computing device, cause the computing device to perform the method of any one of claims 1 to 8.