A garbage recognition and classification method

By building a comprehensive garbage image dataset and hypernetwork technology, combining Mixup enhancement and E-Mix modules, the accuracy and automation level of the garbage classification model are improved, and the problem of identifying unknown and abnormal samples by the garbage classification model is solved, real-time management of garbage and resource recycling are realized.

CN119229194BActive Publication Date: 2025-07-18CHINA RAILWAY 14TH BUREAU GRP NO 3 ENG CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411317374.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2025-07-18
Estimated Expiration
2044-09-20

AI Technical Summary

Technical Problem

When faced with complex and diverse image data, existing garbage classification models are difficult to effectively identify and classify unknown or abnormal garbage samples, and their generalization capabilities are insufficient.

Method used

Hypernetwork technology is used to generate classifier parameters that quickly adapt to new tasks, combine Mixup enhancement technology to improve sample diversity, and enhance abnormal sample detection capabilities through the E-Mix module, build a comprehensive spam image data set for high-quality annotation, use the feature extraction network of the ResNet-12 architecture for training, and combine supervision and self-supervised learning to deploy an automated image acquisition system for real-time recognition.

Benefits of technology

It significantly improves the accuracy and automation level of garbage classification, can quickly adapt to new tasks and effectively identify unknown types of garbage, improves the detection ability of abnormal samples, and realizes real-time management of garbage and resource recycling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229194B_ABST
    Figure CN119229194B_ABST
Patent Text Reader

Abstract

The present invention proposes a garbage recognition and classification method, belonging to the field of machine learning. By creating a comprehensive garbage image dataset and performing high-quality annotation, the accuracy of model training is ensured. Using the HyperTune module and hypernetwork technology, the model can quickly adapt to new tasks and generate classifier parameters. Combining the P-Mix and E-Mix modules, the sample diversity is enhanced through the Mixup enhancement technology, and the detection ability for abnormal samples is enhanced. In addition, the hypernetwork is further optimized in the meta-training stage through the MixMaster module to improve the model generalization. The feature extraction network adopts the ResNet-12 architecture and is pre-trained by combining supervised learning and self-supervised learning strategies to optimize the model performance. Finally, real-time image recognition and classification are realized through an automated image acquisition system, effectively improving the real-time and automation levels of garbage classification, and having the advantages of quickly adapting to new tasks and being robust to abnormal samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of machine learning, and particularly relates to a method for garbage recognition and classification. Background Art

[0002] With the rapid development of artificial intelligence technology, especially in the field of machine learning, deep learning technology has made breakthrough progress in image recognition; this technological progress provides a new solution for the automatic recognition and classification of garbage, which is of great significance for promoting the automation of garbage classification and the effective recycling of resources; by constructing an image dataset containing a wide range of samples and using a deep learning model for training, efficient recognition and classification of garbage can be achieved.

[0003] However, the image data in the real world has high complexity and diversity, which poses higher requirements for the generalization ability of the garbage classification model; to address this challenge, researchers have developed various strategies, such as using hypernetwork technology to quickly adapt to new tasks, adopting Mixup enhancement technology to improve the adaptability of the model to sample diversity, and enhancing the model's ability to quickly learn from small samples through meta-learning strategies; the combined application of these technologies enables the model to not only accurately identify common garbage types but also effectively detect and classify unknown or uncommon garbage samples.

[0004] Out-of-distribution detection is an important research direction in the field of machine learning, which focuses on the performance of the model when facing samples not encountered during the training process; in the practical application of garbage classification, the model may encounter various atypical or abnormal garbage samples; to improve the model's recognition ability for these out-of-distribution samples, researchers have introduced a variety of advanced technologies, including Mixup enhancement of abnormal samples, label assignment using the principle of maximum entropy, and generating adaptable classifier weights through hypernetworks; the combined use of these technologies not only enhances the classification performance of the model for normal samples but also significantly improves its detection and generalization ability for out-of-distribution samples, providing strong technical support for the effective implementation of garbage recognition and classification. Summary of the Invention

[0005] The present invention proposes a method for garbage recognition and classification, aiming to achieve efficient classification and accurate recognition of garbage images by combining hypernetwork, Mixup enhancement, and out-of-distribution detection technologies, while enhancing the model's detection ability for unknown category garbage to improve the accuracy and automation level of garbage classification.

[0006] A method for garbage recognition and classification proposed by the present invention includes the following steps:

[0007] S1. Create a comprehensive dataset of garbage images, covering common recyclables, hazardous waste, wet waste, and dry waste in garbage; perform high-quality annotation on the images to provide accurate supervision signals for subsequent model training;

[0008] S2. Build a HyperTune module, a classifier generated by hypernetwork technology, to quickly generate classifier parameters adapted to new tasks; this module uses meta-learning to quickly learn from a small number of samples and generate efficient classifier weights, laying the foundation for accurate classification;

[0009] S3. Build a P-Mix module, use Mixup to enhance the support set, and use weighted aggregation to calculate classifier weights. By enhancing samples in the weight space, the model can better adapt to the diversity of intra-class samples, thereby improving the classification effect;

[0010] S4. Build an E-Mix module to improve the model's ability to detect abnormal samples; OOE-Mix increases the diversity of abnormal samples by performing Mixup between out-of-task OOE samples and in-task INE samples;

[0011] S5. Build a MixMaster module, which uses the Mixup technique in two sub-modules, P-Mix and E-Mix, to train the hypernetwork in the meta-training stage;

[0012] S6. Training of the classifier model, build a feature extraction network and apply the generated classifier weights above to perform classification training on query samples; through iterative training, the model parameters are continuously optimized until the model can accurately classify new samples, and at the same time, regularization techniques are used to prevent overfitting;

[0013] S7. Real-time image acquisition and classification, deploy an automated image acquisition system to obtain image data of garbage in real time and input it into the trained model for recognition and classification; the model will quickly and accurately classify the garbage in the image based on the learned features to achieve effective management of garbage.

[0014] Preferably, in the HyperTune module in step S2, the hypernetwork is trained through meta-learning, that is, learning on multiple small-sample classification tasks; each task set contains a support set and a query set, and the model needs to learn on the support set and make predictions on the query set; it is specifically composed of a multi-layer perceptron with two fully connected layers, each layer with a size of 256, and the weights of the classifier are generated through a formula. The weight generation process is as follows:

[0015]

[0016] where, w nThe classifier weight for the n-th class, which is obtained by weighting the samples belonging to this class in the support set and is used for the final classifier; K represents the number of samples of each class in the support set. In a K-shot task, usually there are only K samples for each class; KN represents the total number of samples in the support set, including N classes with K samples in each class; 1[y s = n] is an indicator function that has a value of 1 when the sample x s belongs to class n and a value of 0 when the sample x s does not belong to class n. This is used to filter the samples belonging to class n;

[0017] In H(F(x s ), F(x s ) represents the feature vector extracted by the feature extractor from the sample x s ; H(*) is a hypernetwork used to map the feature vector to the classifier weight.

[0018] Preferably, in the P-Mix module in step S3, Mixup is used to enhance the support set, and weighted aggregation is used to calculate the classifier weight; the mixed support samples and the corresponding labels are denoted as x-s and y-s respectively; the mixing parameter PM is sampled from a fixed Beta distribution, and weighted aggregation of the sample-specific weight codes is used to calculate the extracted classifier parameters, where the aggregation weights are determined according to the probabilities of the mixed samples. The specific process is as follows:

[0019] For the samples x1 and x2 in the support set and their labels y1 and y2, generate mixed samples:

[0020]

[0021]

[0022] Then, use the weighted sum of the support set embeddings to generate the classifier weight:

[0023]

[0024] where w n represents the classifier weight for the n-th class; S represents the total number of samples in the enhanced support set; represents the probability that the s-th sample belongs to the n-th class represents the normalization coefficient to ensure that the sum of the weights w n for each class is normalized.

[0025] Preferably, in the E-Mix module in step S4, it is used to improve the model's ability to detect abnormal samples; 00E-Mix increases the diversity of abnormal samples by performing Mixup between out-of-task 00E samples and in-task INE samples. The specific process is as follows:

[0026] First, during each task training, the author samples from the unselected classes as out-of-distribution samples and mixes them with in-distribution samples as follows:

[0027]

[0028] where, represents the new samples generated by mixing in-task samples and out-of-task samples; λ is the weighting coefficient of Mixup, which controls the proportion of in-task samples and out-of-task samples;

[0029] Then, the labels of the out-of-distribution samples are set to the maximum entropy distribution, that is, the uniform distribution labels:

[0030]

[0031] where, represents the label distribution of the out-of-task samples ; represents the probability of each category under the uniform distribution;

[0032] Finally, the final mixed labels are calculated:

[0033]

[0034] where, represents the label distribution of the newly generated mixed samples ; represents the label distribution of the in-task samples ;

[0035] Preferably, in step S6, the feature extractor F adopts a ResNet-12 architecture, which includes 4 residual blocks, each residual block has 3 convolutional layers, and the numbers of convolutional kernels are 64, 160, 320, and 640 respectively, and the convolutional kernel size is 3×3; after the first 3 residual blocks, there is a 2×2 average pooling layer, and the finally obtained feature vector size is 640; in the pre-training stage, the feature extractor is trained by combining supervised learning and self-supervised learning strategies; and it is trained through two linear classifiers: one predicts the category, and the other predicts the rotation angle; the loss function combines the categorical cross-entropy loss and the binary cross-entropy loss.

[0036] A garbage recognition and classification method proposed by the present invention, compared with the prior art, ensures the accuracy and comprehensiveness of model training by constructing a comprehensive garbage image dataset and adopting high-quality annotation; by using the HyperTune module and P-Mix technology, this method can quickly adapt to new tasks and enhance sample diversity, significantly improving the classification effect; at the same time, the introduction of the E-Mix module effectively improves the model's detection ability for abnormal samples, and the meta-training strategy of the MixMaster module further enhances the generalization of the model; in addition, through the automated image acquisition system to achieve real-time image recognition and classification, this method not only improves the real-time and automation level of garbage classification, but also helps to achieve the effective management of garbage and the recycling of resources, demonstrating the advantages of fast adaptation to new tasks and robustness to abnormal samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a flowchart of a garbage recognition and classification method.

[0038] Figure 2 It is a flowchart of the construction and training of the hypernetwork model.

[0039] Figure 3 It is a network structure diagram of the hypernetwork model.

[0040] Figure 4 It is a schematic diagram of the inference classification result of the trained model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the scope of protection of the present invention.

[0042] Please refer to Figures 1-4 , the present invention provides a technical solution.

[0043] The present invention proposes a garbage recognition and classification method, aiming to achieve efficient classification and accurate recognition of garbage images by combining hypernetworks, Mixup enhancement and out-of-distribution detection technologies, while enhancing the model's detection ability for garbage of unknown categories to improve the accuracy and automation level of garbage classification, especially the accuracy of domestic waste classification.

[0044] A garbage recognition and classification method proposed by the present invention, the flowchart is as shown in the appendix Figure 1 as follows, specifically including the following steps:

[0045] S1. Create a comprehensive dataset of garbage images, covering common recyclables, hazardous waste, wet waste, and dry waste in garbage; perform high-quality annotation on the images to provide accurate supervision signals for subsequent model training;

[0046] S2. Build a HyperTune module, a classifier generated by hypernetwork technology, to quickly generate classifier parameters adapted to new tasks; this module uses meta-learning to quickly learn from a small number of samples and generate efficient classifier weights, laying the foundation for accurate classification;

[0047] S3. Build a P-Mix module, use Mixup to enhance the support set, and use weighted aggregation to calculate classifier weights. By enhancing samples in the weight space, the model can better adapt to the diversity of intra-class samples, thereby improving the classification effect;

[0048] S4. Build an E-Mix module to improve the model's ability to detect abnormal samples; OOE-Mix increases the diversity of abnormal samples by performing Mixup between out-of-task OOE samples and in-task INE samples;

[0049] S5. Build a MixMaster module, which uses the Mixup technique with two sub-modules, P-Mix and E-Mix, to train the hypernetwork in the meta-training stage;

[0050] S6. Training of the classifier model, build a feature extraction network and apply the generated classifier weights above to perform classification training on query samples; through iterative training, the model parameters are continuously optimized until the model can accurately classify new samples, and at the same time, regularization techniques are used to prevent overfitting;

[0051] S7. Real-time image acquisition and classification, deploy an automated image acquisition system to obtain real-time image data of garbage and input it into the trained model for recognition and classification; the model will quickly and accurately classify the garbage in the image based on the learned features to achieve effective management of garbage.

[0052] Furthermore, in the Figure 2 The HyperTune module mentioned in step S2 is specifically trained by the hypernetwork through meta-learning, that is, learning on multiple small-sample classification tasks; each task set contains a support set and a query set, and the model needs to learn on the support set and make predictions on the query set; it is specifically composed of a multi-layer perceptron with two fully-connected layers, each layer with a size of 256, and the weights of the classifier are generated through the formula. The weight generation process is as follows:

[0053]

[0054] where, w nThe classifier weight for the n-th class, which is obtained by weighting the samples in the support set belonging to this class and is used for the final classifier; K represents the number of samples of each class in the support set. In a K-shot task, usually there are only K samples for each class; KN represents the total number of samples in the support set, including N classes with K samples for each class; 1[y s = n] is an indicator function that takes the value 1 when the sample x s belongs to class n and 0 when the sample x s does not belong to class n. This is used to filter the samples belonging to class n;

[0055] In H(F(x s ), F(x s ) represents the feature vector extracted by the feature extractor from the sample x s ; H(*) is a hypernetwork used to map the feature vector to the classifier weight.

[0056] Furthermore, in the appendix Figure 2 and the P-Mix module mentioned in step S3, specifically, it uses Mixup to enhance the support set and uses weighted aggregation to calculate the classifier weight; Denote the mixed support samples and the corresponding labels as x-s and y-s respectively; The mixing parameter PM is sampled from a fixed Beta distribution, and the weighted aggregation of the sample-specific weight codes is used to calculate the extracted classifier parameters, where the aggregation weight is determined according to the probability of the mixed samples. The specific process is as follows:

[0057] For the samples x1 and x2 in the support set, and their labels y1 and y2, generate mixed samples:

[0058]

[0059]

[0060] Then, generate the classifier weight using the weighted sum of the support set embeddings:

[0061]

[0062] where w n represents the classifier weight for the n-th class; S represents the total number of samples in the enhanced support set; represents the probability that the s-th sample belongs to the n-th class represents the normalization coefficient to ensure that the sum of the weights w n for each class is normalized.

[0063] Furthermore, in the appendix Figure 2The E-Mix module mentioned in step S4 is specifically used to improve the model's ability to detect abnormal samples. 00E-Mix increases the diversity of abnormal samples by performing Mixup between out-of-task 00E samples and in-task INE samples. The specific process is as follows:

[0064] First, during each task training, the author samples from the unselected categories as out-of-set samples and mixes them with in-set samples as follows:

[0065]

[0066] where, represents the new sample generated by mixing in-task samples and out-of-task samples; λ is the weighting coefficient of Mixup, which controls the ratio of in-task samples and out-of-task samples;

[0067] Then, set the label of the out-of-set sample to the maximum entropy distribution, that is, the uniform distribution label:

[0068]

[0069] where, represents the label distribution of the out-of-task sample ; represents the probability of each category under the uniform distribution;

[0070] Finally, calculate the final mixed label:

[0071]

[0072] where, represents the label distribution of the newly generated mixed sample ; represents the label distribution of the in-task sample .

[0073] Furthermore, the feature extractor F mentioned in Appendix Figure 3 and step S6 specifically adopts the ResNet-12 architecture, which contains 4 residual blocks, each residual block has 3 convolutional layers, and the number of convolutional kernels is 64, 160, 320, and 640 respectively, and the convolutional kernel size is 3×3; there is a 2×2 average pooling layer after the first 3 residual blocks, and the final obtained feature vector size is 640; in the pre-training stage, the feature extractor combines supervised learning and self-supervised learning strategies for training; and is trained through two linear classifiers: one predicts the category, and the other predicts the rotation angle; the loss function combines the categorical cross-entropy loss and the binary cross-entropy loss.

Claims

1. A garbage recognition and classification method, characterized in that, It includes the following steps: S1. Create a comprehensive garbage image dataset covering common recyclables, hazardous waste, wet waste, and dry waste in garbage; perform high-quality annotation on the images to provide accurate supervision signals for subsequent model training; S2. Construct a HyperTune module, a classifier generated by hypernetwork technology, to quickly generate classifier parameters adapted to new tasks; this module learns rapidly from a small number of samples through meta-learning and generates efficient classifier weights, laying a foundation for accurate classification; S3. Construct a P-Mix module, use Mixup to enhance the support set, and use weighted aggregation to calculate classifier weights. By enhancing samples in the weight space, the model can better adapt to the diversity of intra-class samples, thereby improving the classification effect; S4. Construct an E-Mix module to improve the model's ability to detect abnormal samples; OOE-Mix increases the diversity of abnormal samples by performing Mixup between out-of-task OOE samples and in-task INE samples; S5. Construct a MixMaster module, which uses the Mixup technique with two sub-modules, P-Mix and E-Mix, to train the hypernetwork in the meta-training stage; S6. Training of the classifier model, construct a feature extraction network and apply the generated classifier weights above to perform classification training on query samples; through iterative training, the model parameters are continuously optimized until the model can accurately classify new samples, and at the same time, regularization techniques are adopted to prevent overfitting; S7. Real-time image acquisition and classification, deploy an automated image acquisition system to obtain image data of garbage in real time and input it into the trained model for recognition and classification; The model will quickly and accurately classify the garbage in the image based on the learned features, realizing the effective management of garbage.

2. The garbage recognition and classification method according to claim 1, characterized in that For the HyperTune module in step S2, the hypernetwork is trained through meta-learning, that is, learning on multiple small-sample classification tasks; each task set contains a support set and a query set, and the model needs to learn on the support set and make predictions on the query set; it is specifically composed of a multi-layer perceptron with two fully connected layers, each layer with a size of 256, and the weights of the classifier are generated through a formula. The weight generation process is as follows: Among them, w n represents the classifier weight of the nth class, which is obtained by weighting the samples belonging to this class in the support set and is used for the final classifier; K represents the number of samples of each class in the support set. In the K-shot task, usually there are only K samples for each class; KN represents the total number of support set samples, including N classes with K samples for each class; 1[y s = n] is an indicator function, which is 1 when the sample x s belongs to class n and 0 when the sample x s does not belong to class n. This is used to filter the samples belonging to class n; H(F(x s )) where F(x s ) represents the feature vector extracted by the feature extractor from the sample x s ; H(*) is a hypernetwork used to map the feature vector to the weights of the classifier.

3. A garbage recognition and classification method according to claim 1, characterized in that, The P-Mix module in step S3 uses Mixup to augment the support set and uses weighted aggregation to calculate the classifier weights; the mixed support samples and the corresponding labels are denoted as and The mixing parameter PM is sampled from a fixed Beta distribution, and the weighted aggregation of the sample-specific weight codes is used to calculate the extracted classifier parameters, where the aggregation weights are determined according to the probabilities of the mixed samples. The specific process is as follows: For samples x1 and x2 in the support set, and their labels y1 and y2, generate a mixed sample: Then, use the weighted sum of the support set embeddings to generate classifier weights: where, w n represents the classifier weight of the n-th class; S represents the total number of support set samples after augmentation; represents the probability that the s-th sample belongs to the n-th class represents the normalization coefficient to ensure that the sum of the weights w n for each class is normalized.

4. A garbage recognition and classification method according to claim 1, characterized in that The E-Mix module in step S4 is used to improve the model's detection ability for abnormal samples; OOE-Mix increases the diversity of abnormal samples by performing Mixup between out-of-task OOE samples and in-task INE samples. The specific process is as follows: First, during each task training, the author samples from unselected categories as out-of-set samples and mixes them with in-set samples for mixing: Among them, represents the new samples generated by mixing in-task samples and out-of-task samples; λ is the weighting coefficient of Mixup, which controls the proportion of in-task samples and out-of-task samples; Then, set the labels of out-of-set samples to the maximum entropy distribution, that is, the uniform distribution label: Wherein, represents the label distribution of out-of-task samples ; represents the probability of each category under the uniform distribution; Finally, calculate the final mixed label: Among them, represents the label distribution of the newly generated mixed samples ; represents the label distribution of the in-task samples .

5. A garbage recognition and classification method according to claim 1, characterized in that, The feature extractor F in step S6 adopts a ResNet-12 architecture, which contains 4 residual blocks. Each residual block has 3 convolutional layers, and the numbers of convolutional kernels are 64, 160, 320, and 640 respectively, with the convolutional kernel size of 3×3. The first 3 residual blocks are followed by 2×2 average pooling layers, and the finally obtained feature vector has a size of 640. In the pre-training stage, the feature extractor is trained by combining supervised learning and self-supervised learning strategies, and is trained through two linear classifiers: one predicts the category, and the other predicts the rotation angle. The loss function combines the categorical cross-entropy loss and the binary cross-entropy loss.

Citation Information

Patent Citations

  • Neural network construction method and device

    CN111931904A

  • Training set preparation method, training set preparation device, training set preparation equipment and medium

    CN115134589A