Junk recognition method based on concept knowledge graph
By constructing a conceptual knowledge graph of garbage images and conventional image categories, combined with the transfer learning of ViT network and Graph Transformer, the problem of scarcity and overfitting of labeled data in the garbage recognition method is solved, and the accuracy and performance of garbage image recognition are improved.
Patent Information
- Application Number
- CN202510476607.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing garbage recognition methods only focus on learning deep recognition models from a large number of annotated garbage image data. There is a problem of scarce label data and low recognition accuracy when the garbage image classifier is overfitted.
By constructing a conceptual knowledge graph between spam images and conventional image categories, pre-training and fine-tuning of Graph Transformer using ViT network, combining graph prior knowledge and self-attention mechanisms, transfer learning is performed to improve the performance of spam image recognition model.
It effectively alleviates the scarcity of labeled data, improves the parameter estimation accuracy and recognition performance of the garbage image recognition model, and improves the recognition accuracy under limited garbage image labeled data.
Smart Images

Figure CN120375074A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of garbage recognition, and in particular to a garbage recognition method based on a concept knowledge graph. Background Art
[0002] With the rapid development of artificial intelligence technology, especially in the field of machine learning, deep learning technology has made breakthrough progress in image recognition; this technological progress provides a new solution for the automatic recognition and classification of garbage, which is of great significance for promoting the automation of garbage classification and the effective recycling of resources; by constructing an image data set containing a wide range of samples and using a deep learning model for training, efficient recognition and classification of garbage can be achieved.
[0003] For example, Dong Ziyuan et al. proposed a garbage image recognition model based on a convolutional neural network. First, an attention mechanism was designed to complete local and global feature extraction, and then a feature fusion mechanism was used to fuse features of different levels and sizes to improve the garbage image recognition performance. Wu Jian et al. proposed a waste garbage analysis and recognition method based on computer vision, which realizes the recognition of waste garbage by adopting the joint technology of image feature analysis and image segmentation. Kang et al. proposed a garbage image recognition method based on ResNet-34, which optimizes its network structure from the multi-feature fusion of the input image, the feature reuse of the residual unit, and a new activation function to improve the garbage image recognition performance. Considering the interference of the background information of the garbage image, Wu et al. designed a garbage image recognition model based on YOLOv5, which reduces the interference of the background information of the garbage image by introducing candidate box monitoring, thereby improving the garbage image recognition performance. Recently, Yang et al. designed a garbage image recognition model and framework based on incremental learning, which aims to learn a sustainable updated garbage image recognition model from the stream data, thereby improving the generalization performance of the garbage recognition model for new types of garbage.
[0004] However, the existing garbage recognition methods only focus on learning a deep recognition model from a large number of labeled garbage image data, and have the disadvantages of scarce labeled data and low recognition accuracy when the garbage image classifier is overfitted. Summary of the Invention
[0005] The purpose of the present invention is to provide a garbage recognition method based on a concept knowledge graph, so as to solve the problem that the existing garbage recognition methods only focus on learning a deep recognition model from a large number of labeled garbage image data, and have the disadvantages of scarce labeled data and low recognition accuracy when the garbage image classifier is overfitted.
[0006] To achieve the above object, the present invention provides a garbage recognition method based on a concept knowledge graph. The garbage recognition method based on the concept knowledge graph includes the following steps:
[0007] Construct a concept knowledge graph between garbage images and conventional image categories based on the Internet;
[0008] According to the constructed concept knowledge graph, collect and label garbage image and conventional image data, and divide the garbage image recognition data set into a training set, a validation set and a test set;
[0009] Input the conventional natural image data set and the garbage image training set data into the constructed deep learning model for transfer training;
[0010] After the training is completed, use the validation set to evaluate the performance of the model, and make necessary adjustments and optimizations to the model;
[0011] Conduct a comprehensive performance evaluation of the final deep learning model through the test set to obtain a garbage image recognition model;
[0012] Use the formed garbage image recognition model to recognize the garbage in the image to be recognized.
[0013] Among them, in the step of "constructing a concept knowledge graph between garbage images and conventional image categories based on the Internet", the concept knowledge graph aims to describe the semantic association between garbage images and conventional image categories.
[0014] Among them, in the step of "inputting the conventional natural image data set and the garbage image training set data into the constructed deep learning model for transfer training", the transfer training stage of the deep learning model is specifically divided into a pre-training stage and a fine-tuning stage.
[0015] Among them, the specific steps of the pre-training stage are as follows:
[0016] Adopt the ViT network to learn conventional image categories;
[0017] Construct a prior knowledge structure based on the concept knowledge graph, denoted as the concept knowledge graph prior G;
[0018] Use the garbage image classification head prediction network based on Graph Transformer to model the concept knowledge graph, and calculate the high-dimensional representation vectors of each category node in the graph;
[0019] Through the above learning method based on the concept knowledge graph, generate dynamic classification head weights and combine them with the output of the image representation network.
[0020] Among them, the difference between the fine-tuning stage and the pre-training stage is that in the pre-training stage, the training data learned by the network is regular images, and in the fine-tuning stage, the training data learned by the network is replaced with garbage images.
[0021] Among them, the garbage image classification head prediction network based on Graph Transformer in both the pre-training stage and the fine-tuning stage consists of an input layer, a hidden layer, and an output layer.
[0022] Among them, the input layer is designed to transform the semantic information of the knowledge graph into the input representation space, the hidden layer is responsible for using the attention mechanism to explore the representation of each category node, and the output layer is responsible for mapping the hidden space into the feature space to obtain the classification head weights.
[0023] Among them, the operation process of the garbage image classification head prediction network based on Graph Transformer is described as:
[0024]
[0025] Among them, H represents the initial matrix composed of all category semantic vectors; A represents the adjacency matrix of the graph; Attention(·) represents the feature update process based on self-attention; W 0 , W 1 , and W 2 respectively represent the Graph Transformer parameters of the input layer, the hidden layer, and the output layer.
[0026] A garbage recognition method based on a concept knowledge graph according to the present invention. By introducing the hierarchical relationship between common image categories and garbage images and the existing common image recognition data for garbage image classification, the technical solution of the present invention can effectively alleviate the problem of scarce labeled data in existing methods; at the same time, by using the hierarchical relationship knowledge graph between common image categories and garbage image categories and relying on the structured embedding representation technology of graph neural networks, it can effectively improve the estimation accuracy of the parameters of the garbage image recognition model, and thus greatly improve the garbage image recognition performance under limited labeled garbage image data. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0028] Figure 1It is the flowchart of the steps of the garbage recognition method based on the concept knowledge graph provided by the present invention.
[0029] Figure 2 It is the schematic diagram of an example of the concept knowledge graph provided by the present invention.
[0030] Figure 3 It is the schematic diagram of the transfer training of the deep learning model provided by the present invention.
[0031] Figure 4 It is the schematic diagram of the garbage image classification head prediction network based on Graph Transformer provided by the present invention. Detailed implementation manners
[0032] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0033] Please refer to Figures 1 to 4 , the present invention provides a garbage recognition method based on a concept knowledge graph. The garbage recognition method based on the concept knowledge graph includes the following steps:
[0034] S1: Construct a concept knowledge graph between garbage images and conventional image categories based on the Internet;
[0035] S2: According to the constructed concept knowledge graph, collect and label garbage image and conventional image data, and divide the garbage image recognition data set into a training set, a validation set and a test set;
[0036] S3: Input the conventional natural image data set and the garbage image training set data into the constructed deep learning model for transfer training;
[0037] S4: After the training is completed, use the validation set to evaluate the performance of the model, and make necessary adjustments and optimizations to the model;
[0038] S5: Conduct a comprehensive performance evaluation of the final deep learning model through the test set to obtain a garbage image recognition model;
[0039] S6: Use the formed garbage image recognition model to recognize the garbage in the image to be recognized.
[0040] In this embodiment, the technical solution classifies garbage images by introducing explicit prior knowledge (i.e., the hierarchical relationship between common image categories and garbage images) and existing common image recognition data, which can effectively alleviate the problem of scarce labeled data in existing methods; at the same time, by using the knowledge graph of the hierarchical relationship between common image categories and garbage image categories and leveraging the structured embedding representation technology of graph neural networks, the estimation accuracy of the parameters of the garbage image recognition model can be effectively improved, thereby greatly enhancing the garbage image recognition performance under limited garbage image labeled data.
[0041] Further, in the step of "constructing a conceptual knowledge graph between garbage images and conventional image categories based on the Internet", the conceptual knowledge graph aims to describe the semantic association between garbage images and conventional image categories.
[0042] In this embodiment, Figure 2 A brief example of a conceptual knowledge graph of garbage images and conventional image categories is shown. For example, in the classification of recyclable garbage, categories such as "metal garbage", "rubber garbage", "paper garbage", "plastic garbage", "textile garbage", "glass garbage", etc. respectively correspond to "metal products", "rubber products", "paper products", "plastic products", "textile products", "glass products" in the category of artificial products in nature, which are a special waste form of natural images. This semantic association between garbage and natural object images provides a connection for garbage image recognition and conventional image category recognition, and lays a foundation for knowledge fusion and migration between garbage images and conventional images.
[0043] Further, in the step of "inputting the conventional natural image dataset and the garbage image training set data into the constructed deep learning model for transfer training", the stage of transfer training for the deep learning model is specifically divided into a pre-training stage and a fine-tuning stage.
[0044] Among them, the specific steps of the pre-training stage are as follows:
[0045] Use the ViT network to learn conventional image categories;
[0046] Construct a prior knowledge structure based on the conceptual knowledge graph, denoted as the conceptual knowledge graph prior G;
[0047] Use the garbage image classification head prediction network based on Graph Transformer to model the conceptual knowledge graph and calculate the high-dimensional representation vectors of each category node in the graph;
[0048] Through the above learning method based on the conceptual knowledge graph, generate dynamic classification head weights and combine them with the output of the image representation network.
[0049] In this embodiment, the core objective of the pre-training stage is to perform pre-training using a large amount of conventional image data, so as to provide a powerful image representation and classification ability for the garbage image recognition model. In this stage, the focus is on the training of the image representation network and the classification head weight prediction network. Specifically, a more generalizable image feature is learned through a rich dataset of conventional image categories, thus laying a solid foundation for the subsequent garbage image classification task.
[0050] For this purpose, this technical solution adopts the ViT (Vision Transformer) network to learn conventional image categories. ViT is an image classification model based on the Transformer architecture, which can effectively process large-scale image data and learn efficient feature representations. In this way, the ViT network can fully explore the complex features of conventional images, thus providing strong support for subsequent tasks in the pre-training stage. At the same time, in order to combine the prior information of the concept knowledge graph, a prior knowledge structure based on the graph, denoted as G, is constructed. This graph reflects the semantic relationship between garbage image categories and conventional image categories, and can help the model understand the connection between different categories. On this basis, the garbage image classification head prediction network based on GraphTransformer is used to model the concept knowledge graph, and the high-dimensional representation vectors of each category node in the graph are calculated, which provides richer semantic information for each category and makes the connection between image features and categories closer.
[0051] Finally, through this learning method based on the concept knowledge graph, dynamic classification head weights (denoted as z k ) are generated and combined with the output of the image representation network. These dynamic weights can be automatically adjusted according to the content of the image and the prior information of the graph, so as to improve the classification ability and accuracy of the model. In this way, the pre-training stage not only effectively learns large-scale conventional image data through the ViT network, but also enhances the knowledge expression ability of the model by combining the prior information of the graph. This stage provides a good starting point for the subsequent fine-tuning stage, enabling the model to perform better in the garbage image recognition task.
[0052] Furthermore, the difference between the fine-tuning stage and the pre-training stage is that the training data for network learning in the pre-training stage is conventional images, while the training data for network learning in the fine-tuning stage is replaced by garbage images.
[0053] In this embodiment, the fine-tuning stage is a crucial step in the garbage image recognition process. Its main objective is to further optimize and adjust the model obtained in the pre-training stage using garbage image data, thereby achieving efficient classification of garbage images. Specifically, the fine-tuning process is very similar to the pre-training stage, except that the source of the training data has changed. In the pre-training stage, the network learns data of regular image categories, while in the fine-tuning stage, the training data of regular images is replaced with training data of garbage images. This adjustment enables the network to learn from the features of garbage images and better understand the category features and classification boundaries of garbage images. During the fine-tuning process, the ViT network is still used to process the visual features of garbage images. This network learns the global context information of the image through the self-attention mechanism, captures the subtle differences in the image, and generates a more accurate image representation. The Graph Transformer, on the other hand, continues to leverage its advantages in the conceptual knowledge graph to conduct detailed category reasoning on garbage images through the node information in the knowledge graph, further enhancing the model's understanding of garbage images.
[0054] As the fine-tuning process progresses, the model gradually adapts to the features of garbage images, and the weights of its classification head are dynamically adjusted according to the changes in the features of garbage images. This process ensures that the weights of the classification head can more accurately reflect the features of garbage image categories, enabling the model to give highly accurate classification results when faced with garbage images. When the fine-tuning process is completed, the model is ready to test and classify new garbage images. At this stage, the test images are recognized and predicted through the already fine-tuned network, and the model classifies them according to the features of the garbage images and the learned knowledge graph to achieve the category prediction of the garbage images.
[0055] Furthermore, the garbage image classification head prediction network based on Graph Transformer in both the pre-training stage and the fine-tuning stage consists of an input layer, a hidden layer, and an output layer. The input layer is designed to transform the semantic information of the knowledge graph into the input representation space. The hidden layer is responsible for exploring the representations of each category node using the attention mechanism. The output layer is responsible for mapping the hidden space to the feature space to obtain the weights of the classification head. The operation process of the garbage image classification head prediction network based on GraphTransformer is described as follows:
[0056]
[0057] Among them, H represents the initial matrix composed of all category semantic vectors; A represents the adjacency matrix of the graph; Attention(·) represents the feature update process based on self-attention; W 0 , W 1 , and W 2Graph Transformer parameters representing the input layer, hidden layer, and output layer respectively.
[0058] In this embodiment, for the classification head weight prediction network based on Graph Transformer, its main objective is to utilize the semantic information of the knowledge graph to obtain the semantic association relationship between garbage images and regular images, so as to expand the dataset scale of garbage image samples by using the relatively large number of regular image samples. To achieve this goal, this technical solution takes the conceptual knowledge graph prior G as a graph-structured data and uses it as the input of the Graph Transformer, and then effectively predicts the classification head weight by exploring the semantic information of categories. As Figure 4 shown, the classification head weight prediction network based on Graph Transformer designed by this technical solution consists of three parts: an input layer, a multi-head attention layer, and an output layer.
[0059] In summary, this technical solution classifies garbage images by introducing explicit prior knowledge (i.e., the hierarchical relationship between common image categories and garbage images) and existing common image recognition data, which can effectively alleviate the problem of scarce labeled data in existing methods; at the same time, by using the knowledge graph of the hierarchical relationship between common image categories and garbage image categories and leveraging the structured embedding representation technology of graph neural networks, it can effectively improve the estimation accuracy of the parameters of the garbage image recognition model, and thus greatly improve the garbage image recognition performance under limited labeled garbage image data.
[0060] The above-disclosed is only a preferred embodiment of the present invention, and of course, it cannot be used to limit the scope of the rights of the present invention. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. A garbage recognition method based on a concept knowledge graph, characterized in that, It includes the following steps: Construct a conceptual knowledge graph between garbage images and regular image categories based on the Internet; According to the constructed conceptual knowledge graph above, collect and label garbage image and regular image data, and divide the garbage image recognition data set into a training set, a validation set and a test set; Input the regular natural image data set and the garbage image training set data into the constructed deep learning model for transfer training; After the training is completed, use the validation set to evaluate the performance of the model, and make necessary adjustments and optimizations to the model; Conduct a comprehensive performance evaluation on the final deep learning model through the test set to obtain a garbage image recognition model; Use the formed garbage image recognition model to recognize the garbage in the image to be recognized.
2. The garbage recognition method based on a conceptual knowledge graph according to claim 1, characterized in that In the step of "constructing a conceptual knowledge graph between garbage images and regular image categories based on the Internet", the conceptual knowledge graph aims to describe the semantic association between garbage images and regular image categories.
3. The garbage recognition method based on a conceptual knowledge graph according to claim 2, characterized in that In the step of "inputting the regular natural image data set and the garbage image training set data into the constructed deep learning model for transfer training", the stage of transfer training the deep learning model is specifically divided into a pre-training stage and a fine-tuning stage.
4. The garbage recognition method based on a conceptual knowledge graph according to claim 3, characterized in that The specific steps of the pre-training stage are as follows: Adopt a ViT network to learn regular image categories; Construct a prior knowledge structure based on the conceptual knowledge graph, denoted as the conceptual knowledge graph prior G; Use a garbage image classification head prediction network based on Graph Transformer to model the conceptual knowledge graph, and calculate the high-dimensional representation vectors of each category node in the graph; Through the above learning method based on the conceptual knowledge graph, generate dynamic classification head weights and combine them with the output of the image representation network.
5. The garbage recognition method based on a conceptual knowledge graph according to claim 4, characterized in that The difference between the fine-tuning stage and the pre-training stage is that the training data for network learning in the pre-training stage is regular images, and the training data for network learning in the fine-tuning stage is replaced with garbage images.
6. The garbage recognition method based on a conceptual knowledge graph according to claim 5, characterized in that The garbage image classification head prediction network based on Graph Transformer in both the pre-training stage and the fine-tuning stage consists of an input layer, a hidden layer and an output layer.
7. The garbage recognition method based on a conceptual knowledge graph according to claim 6, characterized in that The input layer is designed to transform the semantic information of the knowledge graph into the input representation space, the hidden layer is responsible for using the attention mechanism to explore the representation of each category node, and the output layer is responsible for mapping the hidden space into the feature space to obtain the classification head weights.
8. The garbage recognition method based on a conceptual knowledge graph according to claim 7, characterized in that The operation process of the garbage image classification head prediction network based on the Graph Transformer is described as follows: Among them, H represents the initial matrix composed of all category semantic vectors; A represents the adjacency matrix of the graph; Attention(·) represents the feature update process based on self-attention; W 0 , W 1 , and W 2 respectively represent the GraphTransformer parameters of the input layer, hidden layer, and output layer.
Citation Information
Cited By
Deodorization method and system for a waste station
CN122789083A