Garbage classification multistage verification method based on dynamic confidence threshold and storage medium
By using a multi-level verification method for waste classification with dynamic confidence thresholds, the problem that fixed confidence thresholds cannot adapt to the differences in the difficulty of waste category identification is solved. This method achieves efficient and reliable waste classification, reduces model iteration and maintenance costs, and is adaptable to various input methods and lightweight deployment.
Patent Information
- Application Number
- CN202510596031.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-10-21
AI Technical Summary
In existing waste sorting technologies, fixed confidence thresholds cannot adapt to the differences in the difficulty of identifying different waste categories, resulting in low recognition reliability and efficiency, high model iteration and upgrade costs, and poor compatibility.
A multi-level verification method for garbage classification using dynamic confidence thresholds is adopted. After the recognition results are output by the trained classification model, the softmax function is used to convert them into a probability distribution to determine the confidence and probability of garbage types. High confidence is directly output, while low confidence outputs the top number of garbage types and their probabilities. Rapid iteration is achieved through a modularly designed CNN model.
It improves the reliability of waste sorting identification and user experience, reduces model iteration and maintenance costs, adapts to different waste identification difficulties, avoids the risk of misjudgment, and supports multiple input methods and lightweight deployment.
Smart Images

Figure CN120823429A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the intersection of artificial intelligence and environmental protection technology, and specifically to a multi-level verification method and storage medium for garbage classification based on dynamic confidence thresholds. Background Art
[0002] In existing technologies, after garbage identification and detection, a fixed confidence threshold is typically assigned, which fails to adapt to the varying difficulty levels of different garbage classifications. For example, patent application number CN202411750272.3 describes a community garbage sorting platform based on ResNet, which relates to the field of household garbage sorting technology. This platform utilizes transfer learning technology to fine-tune a ResNet model based on an ImageNet pre-trained model to automatically sort community garbage. However, its main drawback is that it lacks a confidence-level decision-making process for garbage identification, nor does it employ a multi-level hierarchical recognition engine. This results in a lack of reliability and efficiency in identifying difficult garbage. In particular, the patented model lacks a modular design and a unified CNN model interface design. This results in high technical upgrade costs, long maintenance cycles, and a relatively high workload when CNN models with better performance and higher accuracy are available. Some existing technologies primarily employ single-model recognition and processing for garbage classification and identification, without a dual-stage recognition engine. Consequently, they are unable to achieve both good real-time performance and improved recognition reliability. This significantly reduces the reliability of the garbage identification information provided, especially for difficult garbage identification. At the same time, the existing garbage classification model has poor compatibility. The existing garbage classification system uses a fixed-architecture CNN model (such as ResNet18). If you want to replace it with a more advanced model (such as EfficientNet, Vision Transformer, etc.) or a model with deeper layers (such as ResNet101 and ResNet152, etc.), you must modify a large amount of code or even rewrite the entire system; the development cost is high, and the cost of model replacement and adaptation development increases. It cannot adapt to long-term technology iteration and updates, and iterative maintenance and upgrades are difficult, and it is impossible to maintain the technological advancement and accessibility of the core algorithm. Summary of the Invention
[0003] In view of the above problems, the present application provides a multi-level verification method and storage medium for garbage classification based on dynamic confidence thresholds, which solves the problem that the existing garbage classification technology uses a fixed confidence threshold and cannot adapt to the differences in recognition difficulty of different garbage classifications.
[0004] To achieve the above objectives, the inventors provide a multi-level verification method for garbage classification based on a dynamic confidence threshold, comprising the following steps:
[0005] Obtaining image information to be recognized;
[0006] Perform image preprocessing on the image information to be recognized;
[0007] Use the trained classification model to classify the processed image to be identified as garbage and output the recognition result;
[0008] After converting the recognition results output by the classification model into a probability distribution through the Softmax function, the confidence level and probability of the identified garbage type are obtained;
[0009] Determine whether the highest confidence level among the obtained garbage types exceeds a preset threshold;
[0010] If it exceeds, the garbage type corresponding to the highest confidence level is output;
[0011] If it does not exceed the limit, the preset number of garbage types and their probabilities ranked before the confidence level will be output.
[0012] In some embodiments, the multi-level verification method for garbage classification based on dynamic confidence thresholds includes the following steps:
[0013] The image information to be recognized is obtained through the camera real-time recognition module, the single image input interface, or the multiple image batch input channel.
[0014] In some embodiments, the image preprocessing of the image information to be identified specifically includes the following steps:
[0015] Scale the image information to be recognized uniformly to a preset resolution;
[0016] Convert to PyTorch tensor and automatically normalize to the range [0,1];
[0017] The ImageNet standardized mean (0.485, 0.456, 0.406) and standard deviation (0.229, 0.224, 0.225) are applied to adjust the data distribution and obtain the processed image information to be recognized.
[0018] In some embodiments, the classification model is trained by:
[0019] Get an existing dataset of junk images;
[0020] After pre-processing the acquired garbage image data, it is divided into training set, validation set and test set according to the preset ratio;
[0021] Determine the expected loading of pre-trained classification models based on the backbone network type;
[0022] Get the number of input features of the original model's fully connected layer and replace the original output layer with a new linear layer so that its output dimension matches the number of garbage categories;
[0023] Transferring classification model parameters to GPU devices to accelerate calculations;
[0024] Optimize the classification model parameters through the optimizer;
[0025] Establishing a step-wise learning rate decay strategy for the classification model, wherein the step-wise learning rate decay strategy reduces the learning rate by a preset reduction ratio during each preset training cycle;
[0026] The cross entropy loss function is selected as the loss function of the classification model;
[0027] According to the core training mechanism, the training set, validation set, and test set are input into the classification model for training;
[0028] The core training mechanism includes:
[0029] Forward propagation stage: Automatically select GPU acceleration through the intelligent data routing unit, and dynamically optimize the calculation path according to the input tensor dimension through the adaptive calculation graph builder;
[0030] Backpropagation phase: Gradient zeroing controller is used to isolate gradients between batches, and gradient safety monitor is used to monitor NaN / INF abnormal values in real time;
[0031] Parameter update stage: The StepLR and ReduceLROnPlateau strategies are combined through a hybrid learning rate regulator, and the calculation accuracy is automatically selected according to the GPU model through the hardware-aware optimizer.
[0032] In some embodiments, the following steps are also included:
[0033] All recognition results are stored in a single database file in a structured manner.
[0034] Another technical solution is provided, a storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are performed:
[0035] Obtaining image information to be recognized;
[0036] Perform image preprocessing on the image information to be recognized;
[0037] Use the trained classification model to classify the processed image to be identified as garbage and output the recognition result;
[0038] After converting the recognition results output by the classification model into a probability distribution through the Softmax function, the confidence level and probability of the identified garbage type are obtained;
[0039] Determine whether the highest confidence level among the obtained garbage types exceeds a preset threshold;
[0040] If it exceeds, the garbage type corresponding to the highest confidence level is output;
[0041] If it does not exceed the limit, the preset number of garbage types and their probabilities ranked before the confidence level will be output.
[0042] In some embodiments, the multi-level verification method for garbage classification based on dynamic confidence thresholds includes the following steps:
[0043] The image information to be recognized is obtained through the camera real-time recognition module, the single image input interface, or the multiple image batch input channel.
[0044] In some embodiments, the image preprocessing of the image information to be identified specifically includes the following steps:
[0045] Scale the image information to be recognized uniformly to a preset resolution;
[0046] Convert to PyTorch tensor and automatically normalize to the range [0,1];
[0047] The ImageNet standardized mean (0.485, 0.456, 0.406) and standard deviation (0.229, 0.224, 0.225) are applied to adjust the data distribution and obtain the processed image information to be recognized.
[0048] In some embodiments, the classification model is trained by:
[0049] Get an existing dataset of junk images;
[0050] After pre-processing the acquired garbage image data, it is divided into training set, validation set and test set according to the preset ratio;
[0051] Determine the expected loading of pre-trained classification models based on the backbone network type;
[0052] Get the number of input features of the original model's fully connected layer and replace the original output layer with a new linear layer so that its output dimension matches the number of garbage categories;
[0053] Transferring classification model parameters to GPU devices to accelerate calculations;
[0054] Optimize the classification model parameters through the optimizer;
[0055] Establishing a step-wise learning rate decay strategy for the classification model, wherein the step-wise learning rate decay strategy reduces the learning rate by a preset reduction ratio during each preset training cycle;
[0056] The cross entropy loss function is selected as the loss function of the classification model;
[0057] According to the core training mechanism, the training set, validation set, and test set are input into the classification model for training;
[0058] The core training mechanism includes:
[0059] Forward propagation stage: Automatically select GPU acceleration through the intelligent data routing unit, and dynamically optimize the calculation path according to the input tensor dimension through the adaptive calculation graph builder;
[0060] Backpropagation phase: Gradient zeroing controller is used to isolate gradients between batches, and gradient safety monitor is used to monitor NaN / INF abnormal values in real time;
[0061] Parameter update stage: The StepLR and ReduceLROnPlateau strategies are combined through a hybrid learning rate regulator, and the calculation accuracy is automatically selected according to the GPU model through the hardware-aware optimizer.
[0062] In some embodiments, the following steps are also included:
[0063] All recognition results are stored in a single database file in a structured manner.
[0064] Different from the existing technology, the above technical solution, when obtaining the image information to be identified, performs image preprocessing on the image information to be identified, and then inputs the processed image information to be identified into the trained classification model to obtain the output result. Then, after converting the output result into a probability distribution through the Softmax function, the confidence of the garbage type and its probability are obtained, and it is judged whether the highest confidence among the garbage types exceeds the preset threshold. If it exceeds, the garbage type corresponding to the highest confidence is output; if it does not exceed, the garbage types and their probabilities ranked before the preset number in confidence are output; by directly outputting the results when the confidence is high, the user experience is improved. When the confidence is low, the garbage types and their probabilities ranked before the preset number in confidence are displayed to avoid the risk of misjudgment. This dynamic confidence strategy can adapt to the differences in the difficulty of identifying different types of garbage.
[0065] The above-mentioned records related to the content of the invention are only an overview of the technical solution of this application. In order to enable ordinary technicians in this field to understand the technical solution of this application more clearly, and then implement it according to the text of the specification and the contents recorded in the drawings, and to make the above-mentioned purposes and other purposes, features and advantages of this application easier to understand, the following is an explanation in combination with the specific implementation methods and drawings of this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The accompanying drawings are only used to illustrate the principles, implementation methods, applications, characteristics and effects of the specific embodiments of this application and other related contents, and are not to be considered as limiting this application.
[0067] In the drawings of the specification:
[0068] Figure 1 A flowchart of a multi-level verification method for garbage classification based on dynamic confidence thresholds according to a specific implementation method;
[0069] Figure 2 A schematic diagram of a garbage classification model described in a specific embodiment;
[0070] Figure 3 A schematic diagram of the gradient norm training monitoring table described in the specific implementation method;
[0071] Figure 4 A schematic diagram of a dynamic learning rate scheduling diagram described in a specific embodiment;
[0072] Figure 5 A schematic diagram of a graph showing a change in model training accuracy described in a specific embodiment;
[0073] Figure 6 A schematic diagram of the mechanism of the training and validation loss change diagram described in the specific implementation method.
[0074] Figure 7 A schematic diagram of a mechanism of the verification accuracy rate change diagram described in the specific implementation method;
[0075] Figure 8 This is a schematic diagram of the structure of the multi-level verification method for garbage classification based on dynamic confidence thresholds described in the specific implementation method;
[0076] Figure 9 A schematic structural diagram of a storage medium described in a specific embodiment;
[0077] Figure 10 A schematic diagram of a garbage classification and identification system according to a specific embodiment;
[0078] Figure 11 A schematic diagram of the front-end key technologies of the garbage classification and identification system described in the specific implementation method;
[0079] Figure 12 A schematic diagram of the back-end key technologies of the garbage classification and identification system described in the specific implementation method;
[0080] Figure 13 This is a schematic diagram of an interface of the garbage classification and identification system described in the specific implementation method;
[0081] Figure 14 This is another interface diagram of the garbage classification and identification system described in the specific implementation method;
[0082] Figure 15This is another schematic diagram of the interface of the garbage classification and identification system described in the specific implementation method;
[0083] Figure 16 A schematic diagram of an interface for recording garbage classification history of the garbage classification identification system according to a specific embodiment;
[0084] Figure 17 A schematic diagram of an interface for real-time garbage classification using a camera in the garbage classification and identification system according to a specific embodiment;
[0085] Figure 18 This is another schematic diagram of an interface for real-time garbage classification using a camera in the garbage classification and identification system according to the specific implementation method; DETAILED DESCRIPTION
[0086] In order to explain in detail the possible application scenarios, technical principles, specific solutions that can be implemented, and the purpose and effects of this application, the following is a detailed description of the specific embodiments listed in conjunction with the accompanying drawings. The embodiments described herein are only used to more clearly illustrate the technical solutions of this application and are therefore only examples and are not intended to limit the scope of protection of this application.
[0087] References to "embodiments" herein mean that the specific features, structures, or characteristics described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the word "embodiment" in various places in the specification does not necessarily refer to the same embodiment, nor does it particularly limit its independence or relevance to other embodiments. In principle, in this application, as long as there are no technical contradictions or conflicts, the various technical features mentioned in the embodiments can be combined in any manner to form a corresponding implementable technical solution.
[0088] Unless otherwise defined, the technical terms used herein have the same meanings as those generally understood by those skilled in the art to which this application belongs; the use of relevant terms herein is only for describing specific embodiments and is not intended to limit this application.
[0089] See also Figure 1 This embodiment provides a multi-level verification method for garbage classification based on a dynamic confidence threshold, comprising the following steps:
[0090] Step S110: Obtain image information to be identified;
[0091] Step S120: performing image preprocessing on the image information to be recognized;
[0092] Step S130: performing garbage classification recognition on the processed image to be recognized using the trained classification model, and outputting the recognition result;
[0093] Step S140: Convert the recognition result output by the classification model into a probability distribution through the Softmax function to obtain the confidence level and probability of the identified garbage type;
[0094] Step S150: determining whether the highest confidence level among the obtained garbage types exceeds a preset threshold;
[0095] If it exceeds, then execute step S160: output the garbage type corresponding to the highest confidence level;
[0096] If not, step S170 is executed: outputting a preset number of garbage types and their probabilities ranked before confidence.
[0097] When the image information to be identified is obtained, the image information to be identified is preprocessed, and then the processed image information to be identified is input into the trained classification model to obtain the output result. Then, after the output result is converted into a probability distribution through the Softmax function, the confidence of the garbage type and its probability are obtained, and it is judged whether the highest confidence among the garbage types exceeds the preset threshold. If it exceeds, the garbage type corresponding to the highest confidence is output; if it does not exceed, the garbage types and their probabilities ranked before the preset number of confidence levels are output; by directly outputting the results when the confidence level is high, the user experience is improved. When the confidence level is low, the garbage types and their probabilities ranked before the preset number of confidence levels are displayed to avoid the risk of misjudgment. This dynamic confidence strategy can adapt to the different difficulties in identifying different types of garbage.
[0098] In some embodiments, the multi-level verification method for garbage classification based on dynamic confidence thresholds includes the following steps:
[0099] The image information to be recognized is obtained through the camera real-time recognition module, the single image input interface, or the multiple image batch input channel.
[0100] Support real-time camera shooting, single image, multiple Figure 3 Multiple input methods, unified interface processing, and improved system compatibility. The camera real-time recognition module supports real-time video stream input (such as the RTSP protocol), with a frame rate ≥ 30fps and a resolution ≥ 1080p. The single image input interface supports JPEG / PNG / BMP formats, with a single transmission delay of < 100ms. The batch input channel for multiple images supports batch upload, with a maximum concurrent processing capacity of ≥ 16 images.
[0101] In some embodiments, the image preprocessing of the image information to be identified specifically includes the following steps:
[0102] Scale the image information to be recognized uniformly to a preset resolution;
[0103] Convert to PyTorch tensor and automatically normalize to the range [0,1];
[0104] The ImageNet standardized mean (0.485, 0.456, 0.406) and standard deviation (0.229, 0.224, 0.225) are applied to adjust the data distribution and obtain the processed image information to be recognized.
[0105] First, the input image is uniformly scaled to a 224×224 resolution to ensure size compatibility. It is then converted to a PyTorch tensor and automatically normalized to the range [0, 1]. Finally, the ImageNet-standardized mean (0.485, 0.456, 0.406) and standard deviation (0.229, 0.224, 0.225) are applied to adjust the data distribution. This entire process is implemented through the torchvision.transforms pipeline. The processed image forms a 3×224×224 tensor with a batch dimension automatically added to meet model input requirements. This preprocessing preserves key image features while eliminating size and lighting variations, achieving a balanced balance between efficiency and adaptability.
[0106] In some embodiments, the classification model is trained by:
[0107] Get an existing dataset of junk images;
[0108] After pre-processing the acquired garbage image data, it is divided into training set, validation set and test set according to the preset ratio;
[0109] Determine the expected loading of pre-trained classification models based on the backbone network type;
[0110] Get the number of input features of the original model's fully connected layer and replace the original output layer with a new linear layer so that its output dimension matches the number of garbage categories;
[0111] Transferring classification model parameters to GPU devices to accelerate calculations;
[0112] Optimize the classification model parameters through the optimizer;
[0113] Establishing a step-wise learning rate decay strategy for the classification model, wherein the step-wise learning rate decay strategy reduces the learning rate by a preset reduction ratio during each preset training cycle;
[0114] The cross entropy loss function is selected as the loss function of the classification model;
[0115] According to the core training mechanism, the training set, validation set, and test set are input into the classification model for training;
[0116] The core training mechanism includes:
[0117] Forward propagation stage: Automatically select GPU acceleration through the intelligent data routing unit, and dynamically optimize the calculation path according to the input tensor dimension through the adaptive calculation graph builder;
[0118] Backpropagation phase: Gradient zeroing controller is used to isolate gradients between batches, and gradient safety monitor is used to monitor NaN / INF abnormal values in real time;
[0119] Parameter update stage: The StepLR and ReduceLROnPlateau strategies are combined through a hybrid learning rate regulator, and the calculation accuracy is automatically selected according to the GPU model through the hardware-aware optimizer.
[0120] This method combines deep learning technology and big data technology, and uses quantitative technologies such as single image recognition detection, camera real-time recognition detection, and batch image recognition detection to intelligently combine multiple scenarios in the form of detection reports to identify and classify common types of garbage in life. It greatly assists and enhances the understanding and quantitative research of garbage by environmental protection personnel, researchers, and the general public interested in environmental protection.
[0121] The main specific implementation is: a garbage classification model (trash_classifier.pt) trained based on the PyTorch framework, using a typical convolutional neural network structure (such as ResNet variant), including convolutional layers, batch normalization layers and fully connected layers, such as Figure 2 shown.
[0122] The training data is sourced from public garbage image datasets (e.g., TrashNet), which include categories such as recyclables, kitchen waste, hazardous waste, and other garbage. The dataset path is set. Data preprocessing includes image resizing, normalization, and enhancement (e.g., rotation and flipping), as well as training set augmentation. The training / test set is typically split into 70% training, 15% validation, and 15% test sets to ensure balanced classification.
[0123] The classification model training module includes replaceable pre-trained model components (ResNet18 / EfficientNet / ConvNeXt, etc.). This design solves the upgrade difficulty caused by the rigid architecture of traditional garbage classification models. It adopts: dynamic classification head adaptation interface (automatically matching feature dimensions through nn.Linear(num_features, num_classes)); weight inheritance mechanism (retaining all parameters of the pre-trained model except the last layer)". The design advantages are: (1) Plug-and-play architecture: seamless switching between different models is achieved through the backbone network type judgment statement (if 'ResNet' in backbone_type); (2) Dimension self-adaptation: automatically obtain the original model feature dimension (num_features) and match it to the new classification layer.
[0124] Specifically, the classification model configuration is implemented as follows:
[0125] Model construction and initialization: This paper uses the pre-trained ResNet18 model as the basic network structure and performs secondary development through transfer learning technology. The specific implementation steps are as follows:
[0126] 1. Model Construction: 1) Loading a pre-trained model: Use the torchvision.models module of the PyTorch framework to load the ResNet18 model pre-trained on the ImageNet dataset. This model has good image feature extraction capabilities.
[0127] 2) Modify the output layer to 6 categories: Get the number of input features of the original model's fully connected layer (num_features = model.fc.in_features), and replace the original output layer with a new linear layer (nn.Linear) so that its output dimension matches the number of garbage categories (6 categories). The implementation is as follows:
[0128] model = models.resnet18(pretrained = True) # Load pre-trained ResNet18
[0129] num_features = model.fc.in_features # Get the number of input features of the fully connected layer
[0130] model.fc = nn.Linear(num_features, 6) #Modify the output layer to 6 categories;
[0131] 3) Device migration: Transfer model parameters to GPU devices to accelerate calculations (model = model.to('cuda')).
[0132] 2. Optimizer Configuration: The Adam optimizer is used for parameter optimization, which combines the advantages of the momentum method and adaptive learning rate. The optimizer configuration parameters are as follows: base learning rate (lr): 0.001; weight decay (weight_decay): 1e-5 (L2 regularization coefficient); momentum parameters: use the default values (β1 = 0.9, β2 = 0.999). The implementation code is as follows:
[0133] optimizer=optim.Adam(model.parameters(),lr=0.001,weight_decay=1e-5);
[0134] 3. Learning rate scheduling strategy: A step-wise learning rate decay strategy (StepLR) is used, reducing the learning rate to 10% of the original value every five training cycles (epochs). This strategy helps the model adjust parameters more finely in the later stages of training. The implementation code is as follows:
[0135] scheduler=optim.lr_scheduler.StepLR(optimizer, step_size=5, gamma=0.1);
[0136] 4. Loss function selection: Select the cross entropy loss function (CrossEntropyLoss), which is suitable for multi-classification problems and can effectively measure the difference between the predicted probability distribution and the true label. The implementation code is as follows:
[0137] criterion=nn.CrossEntropyLoss()
[0138] Classification model training loop implementation:
[0139] 1. Forward propagation stage
[0140] The forward propagation subsystem includes:
[0141] Smart Data Routing Unit: Automatically select GPU acceleration via the .to('cuda') command
[0142] Adaptive computation graph builder: dynamically optimizes computation paths based on input tensor dimensions
[0143] The implementation code is as follows:
[0144] inputs = inputs.to('cuda') #Automatic device selection
[0145] labels = labels.to('cuda')
[0146] outputs = model (inputs) # dynamic graph construction
[0147] 2. Backpropagation phase
[0148] The back propagation control module includes:
[0149] Gradient zeroing controller: implement inter-batch gradient isolation via optimizer.zero_grad()
[0150] Gradient Safety Monitor: Real-time detection of NaN / INF outliers
[0151] The implementation code is as follows:
[0152] optimizer.zero_grad()#Batch gradient isolation
[0153] loss=criterion(outputs,labels)
[0154] loss.backward()#Automatic differentiation safety monitoring
[0155] 3. Parameter update phase
[0156] The parameter update system includes:
[0157] Hybrid learning rate scheduler: combining StepLR and ReduceLROnPlateau strategies
[0158] Hardware-aware optimizer: automatically selects calculation precision (FP16 / FP32) based on GPU model
[0159] The implementation code is as follows:
[0160] optimizer.step() #Hardware adaptation parameter update
[0161] scheduler.step() #Dynamic learning rate adjustment
[0162] The above implementation process constitutes the core training mechanism of the junk image classification system of this application. Through reasonable training configuration and efficient training cycle, the model can converge quickly and obtain good classification performance.
[0163] The classification model training performance parameter analysis is as follows:
[0164] 1. If Figure 3The gradient norm training monitoring table shown records the changing trend of the gradient norm (Gradient Norm) with the number of training steps (Step) during the neural network training process, which is a training process monitoring indicator unique to this application. The horizontal axis in the figure represents the number of training steps, and the vertical axis represents the L2 norm value of the gradient vector, reflecting the intensity of the model parameter update. It can be seen from the figure: the gradient norm change curve during the training process, where the horizontal axis is the number of training steps (Step), and the vertical axis is the L2 norm value (Gradient Norm) of the gradient vector. The curve shows an oscillating downward trend, indicating that the optimization strategy adopted can effectively maintain the gradient within a reasonable range and avoid the gradient disappearance or explosion phenomenon.
[0165] 2. If Figure 4 The dynamic learning rate scheduling diagram shown shows the dynamic learning rate scheduling strategy adopted, where the horizontal axis is the training cycle (Epoch) and the vertical axis is the learning rate value (logarithmic scale). The curve shows a step-by-step decline feature, with a 10-fold attenuation occurring at every 5 training cycles (such as Epoch5 / 10 / 15), achieving a three-stage optimization process of coarse adjustment → fine adjustment → fine adjustment. The chart shows the step learning rate decay strategy (Step Learning Rate Decay) adopted by the present invention. The changing process within 20 training cycles (Epoch).
[0166] In the figure, the horizontal axis is the training cycle (Epoch 1-20); the vertical axis is the learning rate value (decreasing from the initial value of 0.001 to 1e-7). The key point is that the learning rate decreases to 10% of the previous value every 5 cycles (0.001→0.0001→0.00001→...).
[0167] Technical effect:
[0168] (1) Optimization control in stages, achieved through the step-by-step descent shown in the figure:
[0169] Early stage (Epoch 1-5): A higher learning rate (0.001) quickly approaches the optimal solution
[0170] Mid-term (Epoch 6-15): Gradually reduce the learning rate (0.0001-0.000001) and fine-tune the parameters
[0171] Late stage (Epoch 16-20): extremely low learning rate (1e-7) for stable convergence
[0172] (2) Avoiding sudden changes in the local optimal periodic learning rate (such as the steep drops at Epoch 5 / 10 / 15 in the figure) helps the model escape the local optimal state:
[0173] 3. If Figure 5The model training accuracy graph shown in the figure shows the changes in model accuracy during training. The horizontal axis represents the number of training epochs, from epoch 1 to 20; the vertical axis represents the accuracy (%). As can be seen from the graph, the model's training accuracy fluctuates between 94.4% and 95.8%. Despite some fluctuation, the model's performance on the training set is relatively stable, with high accuracy. This demonstrates that the model has good learning ability on the training data and is able to correctly predict the majority of samples.
[0174] 4. If Figure 6 The training and validation loss graph shows how the model's training and validation losses change during training. The horizontal axis represents the number of training epochs, from 1 to 20, while the vertical axis represents the loss. The blue line represents the loss on the training set, and the orange line represents the loss on the validation set. As can be seen from the graph, the training loss gradually decreases with the number of training epochs, indicating that the model's predictions on the training data are becoming increasingly accurate. Furthermore, the validation loss levels off after an initial decline, indicating that the training process is stabilizing.
[0175] 5. If Figure 7 The validation accuracy graph shows how the model's validation accuracy changes during training. The horizontal axis represents the number of training epochs (epochs), from 1 to 20; the vertical axis represents the accuracy (%). The green line represents the accuracy on the validation set. As can be seen from the graph, the validation accuracy rises rapidly initially, then levels off after approximately epoch 10, ultimately stabilizing at 85%. This indicates that the model's prediction accuracy on the validation data gradually improves and stabilizes.
[0176] In some embodiments, the following steps are also included:
[0177] All recognition results are stored in a single database file in a structured manner.
[0178] The SQLite lightweight database is used to achieve efficient data persistence and query management. All recognition results (including image storage path, prediction category, confidence, Top3 candidates and timestamp) are stored in a structured manner in a single database file, where the Top3 data is stored in JSON serialization. The database design uses auto-incrementing primary keys and time indexes to optimize query efficiency, and supports fast filtering of historical records by date range and category labels (such as "query all plastic waste records in 2025"). The front end dynamically generates SQL query statements through Flask routing, and the results are displayed in reverse chronological order while maintaining the association with the original image, so that users can go back and verify at any time. The entire storage process does not require additional database services, while ensuring data integrity and perfectly adapting to lightweight deployment requirements.
[0179] In some embodiments, a multi-level verification method for garbage classification based on dynamic confidence thresholds is provided, which provides multi-scenario usage with the help of deep learning algorithms, including single image detection, batch image detection, camera real-time detection, etc.; garbage image information can be conveniently and quickly quantitatively identified and processed, including classification, recognition confidence probability, TOP3 confidence probability list, historical records, PDF detection report and other multiple quantitative information; TOP prob3 confidence probability provides the credibility of garbage intelligent classification and recognition information, especially for difficult garbage classification detection, ensuring the completeness and reliability of recognition and classification information data; the modular design of the core algorithm effectively improves the efficiency and economy of technology iteration, reduces development costs, simplifies maintenance, and innovatively and economically improves the research process and methods of intelligent garbage classification, and the access compatibility of new models in the future will be stronger.
[0180] It uses the following Figure 8 The structure shown is composed of a hardware input layer and a software processing layer, which realizes the functions of real-time image recognition, confidence assessment and grading result display.
[0181] Hardware layer
[0182] Camera real-time recognition module: supports real-time video stream input (such as RTSP protocol), frame rate ≥ 30fps, resolution ≥ 1080p.
[0183] Single image input interface: supports JPEG / PNG / BMP formats, single transmission delay <100ms.
[0184] Multiple image batch input channel: supports batch uploading, with a maximum concurrent processing number of ≥16 images.
[0185] Data transmission: All inputs are transmitted to the server via TCP / IP protocol. The server is deployed using the Flask lightweight framework to ensure data integrity and low latency.
[0186] Software Layer
[0187] Image preprocessing module: Image preprocessing uses a standardized process to quickly adapt the ResNet18 model. First, the input image is uniformly scaled to a 224×224 resolution to ensure size compatibility. Next, it is converted to a PyTorch tensor and automatically normalized to the range [0, 1]. Finally, the ImageNet-standardized mean (0.485, 0.456, 0.406) and standard deviation (0.229, 0.224, 0.225) are applied to adjust the data distribution. The entire process is implemented through the torchvision.transforms pipeline. The processed image forms a 3×224×224 tensor with a batch dimension automatically added to meet the model input requirements. This preprocessing maintains key image features while eliminating size and lighting differences, achieving a balanced balance between efficiency and adaptability.
[0188] ResNet18 model: The pre-trained ResNet18 model is used for transfer learning, and the final fully connected layer is replaced to adapt it to the 6 main garbage classification tasks. When the model is loaded, the ImageNet pre-trained weights are retained as the feature extractor, and only the newly added output layer is trained (6 neurons correspond to 6 types of garbage), which greatly improves the training efficiency of small sample data. In the inference stage, the model.eval() mode is used to turn off the randomness of Dropout and BatchNorm, and the gradient calculation is disabled with torch.no_grad(). After the model output is converted to a probability distribution by Softmax, it can output the highest confidence category and extract the Top3 candidate results, balancing the recognition accuracy and fault tolerance. The entire model is only 43MB in size, taking into account both lightweight and high precision.
[0189] Confidence Decision Judgment Module: This module utilizes a dynamic confidence decision mechanism for intelligent result display. After the model's raw scores are converted into a probability distribution using the Softmax function, the system simultaneously extracts the highest-confidence result and its probability value. When the highest confidence score exceeds a preset threshold (e.g., 0.95), a single classification result is directly output; if it falls below the threshold, the top three candidate categories and their corresponding probabilities are displayed, ranked from highest to lowest confidence. Technically, the PyTorch torch.topk() method is used to efficiently retrieve the three most probable categories. The results are then encapsulated as structured data containing the category name and confidence value (e.g., [("plastic",0.72),("metal",0.15),("glass",0.08)]). This design ensures concise output in high-confidence scenarios while providing fallback options when the model is uncertain, effectively reducing the risk of misclassification. All results are stored in a database along with information such as the original image path and timestamp, enabling subsequent query and verification.
[0190] Data storage and query: The SQLite lightweight database is used to achieve efficient data persistence and query management. All recognition results (including image storage path, prediction category, confidence, Top3 candidates and timestamp) are stored in a structured manner in a single database file, where the Top3 data is stored serialized in JSON. The database design uses auto-incrementing primary keys and time indexes to optimize query efficiency, and supports fast filtering of historical records by date range and category labels (such as "query all plastic waste records in 2025"). The front end dynamically generates SQL query statements through Flask routing, and the results are displayed in reverse chronological order while maintaining an association with the original image, so that users can go back and verify at any time. The entire storage process does not require additional database services, which perfectly adapts to lightweight deployment requirements while ensuring data integrity.
[0191] This approach has the following advantages:
[0192] (1) Multimodal input adaptation:
[0193] Support real-time camera shooting, single image, multiple Figure 3 Multiple input methods, unified interface processing, and improved system compatibility.
[0194] (2) Dynamic confidence decision:
[0195] High-confidence results are directly output to improve user experience.
[0196] When the confidence level is low, the top 3 candidates are displayed to avoid the risk of misjudgment.
[0197] Traceable data management: All identification records are stored persistently to facilitate subsequent analysis and optimization.
[0198] (4) Unified CNN model interface design:
[0199] When replacing the ResNet18 model, you only need to replace the model file. For other models such as ResNet50, ResNet101 and ResNet152, no code changes are required. This meets the needs of AI systems that require frequent model upgrades in the future, reduces maintenance cycles, and saves maintenance costs.
[0200] The core technical advantage of this invention lies in the multi-source input adaptation (real-time camera shooting, single image, multi-image and multi-source input) + intelligent hierarchical decision-making (first level: ResNet18 fast initial screening, second level: Top3 alternative verification) + traceable data management + CNN model modular design, which improves recognition reliability and model iterative optimization while ensuring real-time performance. It is suitable for garbage classification, especially avoiding the risk of misjudgment in difficult garbage classification, and improving user satisfaction with the usage experience.
[0201] In the above example, a multi-level verification method for garbage classification based on dynamic confidence thresholds was developed for 2,527 common types of real-life garbage collected from a dataset that included paper, glass, metal, paper, plastic, and other packaging. Based on the collected garbage image dataset, garbage features were learned completely from scratch to train the ResNet18 model. Simultaneously, a multi-level verification mechanism triggered by a Top3 probability confidence threshold enabled automatic identification and classification of community waste. This method can be deployed in smart trash cans, community recycling stations, mobile apps, and other terminals, addressing the issue of fuzzy garbage classification in complex scenarios. It has applications in areas such as integrated community smart trash cans, school environmental education platforms, waste treatment plant pre-sorting systems, and shopping mall recycling points systems.
[0202] This method has the following innovations:
[0203] (1) Multimodal input adaptation:
[0204] Support real-time camera shooting, single image, multiple Figure 3 Multiple input methods, unified interface processing, and improved system compatibility.
[0205] (2) Dynamic confidence decision:
[0206] High-confidence results are directly output to improve user experience.
[0207] When the confidence level is low, the top 3 candidates are displayed to avoid the risk of misjudgment.
[0208] Traceable data management: All identification records are stored persistently to facilitate subsequent analysis and optimization.
[0209] (4) Unified CNN model interface design:
[0210] When replacing the ResNet18 model, you only need to replace the model file, such as other models such as ResNet50, ResNet101 and ResNet152, without changing the code, to meet the needs of AI systems that need frequent model upgrades in the future, reduce maintenance cycles, and save maintenance costs.
[0211] Through multi-source input adaptation (real-time camera shooting, single image, multi-image and multi-source input) + intelligent hierarchical decision-making (first level: ResNet18 fast initial screening, second level: Top3 alternative verification) + traceable data management + CNN model modular design, while ensuring real-time performance, it improves recognition reliability and model iterative optimization advancement. It is suitable for garbage classification, especially avoiding the risk of misjudgment in difficult garbage classification, and improving user satisfaction with the usage experience.
[0212] As a multimodal image intelligent recognition method, this method integrates three input channels: real-time camera video stream, single image, and multiple images. The software layer uses the ResNet18 model for image recognition. A confidence threshold judgment mechanism is used to achieve intelligent hierarchical output: when the confidence level meets the standard, the recognition result is directly displayed; when it does not, the top three probability candidates are displayed. All recognition data is stored in a database to support historical queries and can be exported as PDF query reports. This system achieves efficient processing and reliable recognition of multi-source images, ensuring real-time performance while significantly reducing the risk of misjudgment.
[0213] For scientific researchers, this method helps to accelerate the credibility of microplastic identification. Through the TOP3 identification confidence probability, it provides a list of TOP3 probability candidates, avoids the risk of misjudgment, and improves the accuracy and effectiveness of research. It improves the work efficiency of the system and reduces scientific research costs (manpower, time, management, etc.). Through real-time or near real-time processing capabilities, it reduces the error of difficult garbage identification and improves scientific research efficiency.
[0214] For the general public who are interested in environmental issues, this method enhances model compatibility and unifies interface design. All models (ResNet18, ResNet53, EfficientNet, ViT, etc.) follow the same input and output formats. In this way, when changing models, only the model file needs to be replaced without changing the code. This will facilitate subsequent project maintenance, reduce development and maintenance costs, and provide an easy access interface for the emergence of new models in the future.
[0215] The modular unified interface design of the training model is more suitable for small and medium-sized enterprises, avoiding the need to rebuild the system due to technological iteration and reducing development costs. At the same time, in the fields of garbage classification and environment, AI systems that require frequent model upgrades greatly maintain the advanced nature and real-time updates of scientific research and work, and have strong future compatibility.
[0216] By providing multiple functions such as single image recognition and detection, real-time camera recognition and detection, and batch recognition and detection, the user needs of different application scenarios are met, making the method and device of the present invention have greater scene diversity and universality.
[0217] Compared with other existing patents, this technology can provide more detection information in terms of the confidence and credibility of detection results, and provide users with detailed data on spam detection, classification information, detection reports, confidence probability, TOP3 confidence probability list and other quantitative research information to help users better understand the information statistics of spam detection.
[0218] Compared with other existing patents and methods, this method has multiple advantages such as intelligence, precision, and modularity, and can be widely used in the monitoring and research of garbage classification information.
[0219] Compared with existing patents, the unified CNN modular design makes it more maintainable, easier to upgrade technology, and more advanced in the future.
[0220] By adopting a modular design of the core algorithm as the core detection algorithm design method, the high accuracy and efficiency of the model can be maintained in the future.
[0221] Through end-to-end solutions, the quantitative operating procedures for intelligent garbage classification are standardized and unified, and the research process and technical accessibility of intelligent garbage classification are innovated and improved.
[0222] Especially for difficult or unknown spam, the TOP3 confidence probabilities and confidence probability lists can be given to provide more reference information for identification, thereby improving the completeness and credibility of spam detection information.
[0223] like Figure 9 As shown, in some embodiments, a storage medium 910 stores a computer program, and when the computer program is executed by a processor 920, the following steps are performed:
[0224] Obtaining image information to be recognized;
[0225] Perform image preprocessing on the image information to be recognized;
[0226] Use the trained classification model to classify the processed image to be identified as garbage and output the recognition result;
[0227] After converting the recognition results output by the classification model into a probability distribution through the Softmax function, the confidence level and probability of the identified garbage type are obtained;
[0228] Determine whether the highest confidence level among the obtained garbage types exceeds a preset threshold;
[0229] If it exceeds, the garbage type corresponding to the highest confidence level is output;
[0230] If it does not exceed the limit, the preset number of garbage types and their probabilities ranked before the confidence level will be output.
[0231] When the image information to be identified is obtained, the image information to be identified is preprocessed, and then the processed image information to be identified is input into the trained classification model to obtain the output result. Then, after the output result is converted into a probability distribution through the Softmax function, the confidence of the garbage type and its probability are obtained, and it is judged whether the highest confidence among the garbage types exceeds the preset threshold. If it exceeds, the garbage type corresponding to the highest confidence is output; if it does not exceed, the garbage types and their probabilities ranked before the preset number of confidence levels are output; by directly outputting the results when the confidence level is high, the user experience is improved. When the confidence level is low, the garbage types and their probabilities ranked before the preset number of confidence levels are displayed to avoid the risk of misjudgment. This dynamic confidence strategy can adapt to the different difficulties in identifying different types of garbage.
[0232] In some embodiments, the multi-level verification method for garbage classification based on dynamic confidence thresholds includes the following steps:
[0233] The image information to be recognized is obtained through the camera real-time recognition module, the single image input interface, or the multiple image batch input channel.
[0234] In some embodiments, the image preprocessing of the image information to be identified specifically includes the following steps:
[0235] Scale the image information to be recognized uniformly to a preset resolution;
[0236] Convert to PyTorch tensor and automatically normalize to the range [0,1];
[0237] The ImageNet standardized mean (0.485, 0.456, 0.406) and standard deviation (0.229, 0.224, 0.225) are applied to adjust the data distribution and obtain the processed image information to be recognized.
[0238] In some embodiments, the classification model is trained by:
[0239] Get an existing dataset of junk images;
[0240] After pre-processing the acquired garbage image data, it is divided into training set, validation set and test set according to the preset ratio;
[0241] Determine the expected loading of pre-trained classification models based on the backbone network type;
[0242] Get the number of input features of the original model's fully connected layer and replace the original output layer with a new linear layer so that its output dimension matches the number of garbage categories;
[0243] Transferring classification model parameters to GPU devices to accelerate calculations;
[0244] Optimize the classification model parameters through the optimizer;
[0245] Establishing a step-wise learning rate decay strategy for the classification model, wherein the step-wise learning rate decay strategy reduces the learning rate by a preset reduction ratio during each preset training cycle;
[0246] The cross entropy loss function is selected as the loss function of the classification model;
[0247] According to the core training mechanism, the training set, validation set, and test set are input into the classification model for training;
[0248] The core training mechanism includes:
[0249] Forward propagation stage: Automatically select GPU acceleration through the intelligent data routing unit, and dynamically optimize the calculation path according to the input tensor dimension through the adaptive calculation graph builder;
[0250] Backpropagation phase: Gradient zeroing controller is used to isolate gradients between batches, and gradient safety monitor is used to monitor NaN / INF abnormal values in real time;
[0251] Parameter update stage: The StepLR and ReduceLROnPlateau strategies are combined through a hybrid learning rate regulator, and the calculation accuracy is automatically selected according to the GPU model through the hardware-aware optimizer.
[0252] In some embodiments, the following steps are also included:
[0253] All recognition results are stored in a single database file in a structured manner.
[0254] Specifically, in some embodiments, the core task of human-computer interaction is to demonstrate the capabilities of intelligent garbage identification. The main page is concise and clear in design, directly introducing the core functions of intelligent garbage classification and identification, including single image recognition and detection, real-time recognition and detection, batch recognition and detection, and other multi-scenario functions. Figure 10 As shown, the intelligent garbage identification and detection technology is mainly divided into: garbage classification identification system, front-end key technology and back-end key technology.
[0255] Front-end key technologies:
[0256] The front-end interface selects / takes photos of garbage and displays the recognition results; the user clicks the "Recognize" button, and the front-end packages the photos (one / multiple photos) into an HTTP request and sends it to the back-end; the back-end service parses the request sent by the front-end, hands the photos to the AI model for execution, and generates a recognition report; the AI model checks the photo features (color / shape / texture), compares them with the knowledge base (training data), and gives classification suggestions (recyclable / hazardous / kitchen waste, etc.); the database is used to save user history records (time + pictures + classification results), system operation logs, etc.; the browser local storage remembers the user's recent operations and stores temporary data.
[0257] like Figure 11 The flowchart of the key front-end technologies shown in the figure includes file drag and drop upload, a visual drag and drop area, and support for batch file selection; real-time preview, instant display of file name and quantity, and support for automatic line wrapping of overlong file names; asynchronous submission, data submission without refreshing the page, and no blocking of user operations when uploading large files; dynamic rendering, visualization of confidence values, intuitive display of the credibility of recognition results, and assistance for users in determining whether re-shooting is necessary; PDF generation triggering, front-end assembly of report data, automatic triggering of the download process, and provision of detection and recognition reports for environmental protection departments or users.
[0258] Backend key technologies:
[0259] like Figure 12 The diagram of the back-end key technical process shown in the figure includes: multi-file processing: parallel processing mechanism (supports uploading 50+ pictures at the same time), adopts streaming processing mechanism (avoids memory overflow) and error recovery system (supports breakpoint resumption and automatic isolation of damaged files); image preprocessing module: adopts intelligent cropping algorithm and metadata cleaning to unify images taken by different cameras into standardized input, which is convenient for improving the model recognition accuracy in the later stage; model inference module: uses dual-model collaboration, in which the main model (ResNet18) performs classification and dynamic threshold mechanism, and gives the main result confidence and TOP3 ranking results; result structuring module: mainly supports the data source of front-end generation of PDF report output, and adopts traceable data structure and tamper-proof design to ensure the reliability and structuring of data source.
[0260] The effect is achieved through an intelligent garbage identification GUI to illustrate the process of garbage classification and identification in this application:
[0261] Uploaded junk images (single / multiple) identification example 1:
[0262] The garbage classification intelligent recognition system selects a single or multiple garbage images and the recognition process is as follows:
[0263] like Figure 13 In the picture selection interface shown, click "Select Picture" (multiple selections are allowed);
[0264] Among them, such as Figure 14 As shown, two junk image files are selected, and the file names support files of any length;
[0265] After clicking "Start Recognition", the garbage classification recognition results are displayed, such as Figure 15 As shown, here, the recognition TOP3 probability is displayed simultaneously.
[0266] Click "Export Report" to generate a spam image detection report for users to view specific spam detection information reports;
[0267] Click "View History" in the upper right corner to view the history of past garbage classification. Figure 16 As shown, you can view the garbage classification history information by year, month, category, etc.
[0268] Real-time camera shooting recognition example 2:
[0269] 1. If Figure 17 The camera real-time garbage classification interface shown is connected to an external real-time camera USB device and initialized;
[0270] 2. Click the "Photo Identification" button and take a picture of the garbage to identify it immediately. Figure 18 As shown, the TOP3 recognition confidence probability recommendation ranking is given.
[0271] 3. Click the "Switch Camera" button to reset the camera, re-enter the shooting process, and go to step 1.
[0272] Finally, it should be noted that although the above embodiments have been described in the specification and drawings of this application, this does not limit the scope of patent protection of this application. All technical solutions generated by replacing or modifying equivalent structures or equivalent processes based on the essential concepts of this application using the contents recorded in the specification and drawings of this application, as well as directly or indirectly implementing the technical solutions of the above embodiments in other related technical fields, are included in the scope of patent protection of this application.
Claims
1. A multi-level verification method for garbage classification based on dynamic confidence threshold, characterized in that: The following steps are involved: Obtaining image information to be recognized; Perform image preprocessing on the image information to be recognized; Use the trained classification model to classify the processed image to be identified as garbage and output the recognition result; After converting the recognition results output by the classification model into a probability distribution through the Softmax function, the confidence level and probability of the identified garbage type are obtained; Determine whether the highest confidence level among the obtained garbage types exceeds a preset threshold; If it exceeds, the garbage type corresponding to the highest confidence level is output; If it does not exceed the limit, the preset number of garbage types and their probabilities ranked before the confidence level will be output.
2. The multi-level verification method for garbage classification based on dynamic confidence threshold according to claim 1 is characterized in that: The acquisition of the image information to be identified comprises the following steps: The image information to be recognized is obtained through the camera real-time recognition module, the single image input interface, or the multiple image batch input channel.
3. The multi-level verification method for garbage classification based on dynamic confidence threshold according to claim 1 is characterized in that: The image preprocessing for the image information to be identified specifically comprises the following steps: Scale the image information to be recognized uniformly to a preset resolution; Convert to PyTorch tensor and automatically normalize to the range [0,1]; The mean and standard deviation of ImageNet standardization are applied to adjust the data distribution to obtain the processed image information to be recognized.
4. The multi-level verification method for garbage classification based on dynamic confidence threshold according to claim 1 is characterized in that: The training method of the classification model includes the following steps: Get an existing dataset of junk images; After pre-processing the acquired garbage image data, it is divided into training set, validation set and test set according to the preset ratio; Determine the expected loading of pre-trained classification models based on the backbone network type; Get the number of input features of the original model's fully connected layer and replace the original output layer with a new linear layer so that its output dimension matches the number of garbage categories; Transferring classification model parameters to GPU devices to accelerate calculations; Optimize the classification model parameters through the optimizer; Establishing a step-wise learning rate decay strategy for the classification model, wherein the step-wise learning rate decay strategy reduces the learning rate by a preset reduction ratio during each preset training cycle; The cross entropy loss function is selected as the loss function of the classification model; According to the core training mechanism, the training set, validation set, and test set are input into the classification model for training; The core training mechanism includes: Forward propagation stage: Automatically select GPU acceleration through the intelligent data routing unit, and dynamically optimize the calculation path according to the input tensor dimension through the adaptive calculation graph builder; Backpropagation phase: Gradient zeroing controller is used to isolate gradients between batches, and gradient safety monitor is used to monitor NaN / INF abnormal values in real time; Parameter update stage: The StepLR and ReduceLROnPlateau strategies are combined through a hybrid learning rate regulator, and the calculation accuracy is automatically selected according to the GPU model through the hardware-aware optimizer.
5. The multi-level verification method for garbage classification based on dynamic confidence threshold according to claim 1 is characterized in that: The following steps are also included: All recognition results are stored in a single database file in a structured manner.
6. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the following steps are performed: Obtaining image information to be recognized; Perform image preprocessing on the image information to be recognized; Use the trained classification model to classify the processed image to be identified as garbage and output the recognition result; After converting the recognition results output by the classification model into a probability distribution through the Softmax function, the confidence level and probability of the identified garbage type are obtained; Determine whether the highest confidence level among the obtained garbage types exceeds a preset threshold; If it exceeds, the garbage type corresponding to the highest confidence level is output; If it does not exceed the limit, the preset number of garbage types and their probabilities ranked before the confidence level will be output.
7. The storage medium according to claim 6, wherein: The acquisition of the image information to be identified comprises the following steps: The image information to be recognized is obtained through the camera real-time recognition module, the single image input interface, or the multiple image batch input channel.
8. The storage medium according to claim 6, wherein: The image preprocessing for the image information to be identified specifically comprises the following steps: Scale the image information to be recognized uniformly to a preset resolution; Convert to PyTorch tensor and automatically normalize to the range [0,1]; The mean and standard deviation of ImageNet standardization are applied to adjust the data distribution to obtain the processed image information to be recognized.
9. The storage medium according to claim 6, wherein: The training method of the classification model includes the following steps: Get an existing dataset of junk images; After pre-processing the acquired garbage image data, it is divided into training set, validation set and test set according to the preset ratio; Determine the expected loading of pre-trained classification models based on the backbone network type; Get the number of input features of the original model's fully connected layer and replace the original output layer with a new linear layer so that its output dimension matches the number of garbage categories; Transferring classification model parameters to GPU devices to accelerate calculations; Optimize the classification model parameters through the optimizer; Establishing a step-wise learning rate decay strategy for the classification model, wherein the step-wise learning rate decay strategy reduces the learning rate by a preset reduction ratio during each preset training cycle; The cross entropy loss function is selected as the loss function of the classification model; According to the core training mechanism, the training set, validation set, and test set are input into the classification model for training; The core training mechanism includes: Forward propagation stage: Automatically select GPU acceleration through the intelligent data routing unit, and dynamically optimize the calculation path according to the input tensor dimension through the adaptive calculation graph builder; Backpropagation phase: Gradient zeroing controller is used to isolate gradients between batches, and gradient safety monitor is used to monitor NaN / INF abnormal values in real time; Parameter update stage: The StepLR and ReduceLROnPlateau strategies are combined through a hybrid learning rate regulator, and the calculation accuracy is automatically selected according to the GPU model through the hardware-aware optimizer.
10. The storage medium according to claim 6, wherein: The following steps are also included: All recognition results are stored in a single database file in a structured manner.
Citation Information
Patent Citations
Community garbage classification platform based on ResNet
CN119672423A
Cited By
Kitchen garbage intelligent classification method and system based on AI image recognition
CN122066997A
Intelligent classification method and system for kitchen waste based on AI image recognition
CN122066997B