Garbage classification method and system based on bidirectional long-short-term memory network and convolutional neural network

By combining an improved convolutional neural network and a bidirectional long short-term memory network, the problem of difficulty in identifying similar-looking garbage in garbage classification is solved, achieving higher classification accuracy and model expression capabilities.

CN120808033APending Publication Date: 2025-10-17GUANGZHOU NANFANG COLLEGE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510980747.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing garbage classification methods based on deep learning have difficulty in accurately identifying garbage items with similar appearance but belonging to different categories, resulting in low classification accuracy.

Method used

An improved convolutional neural network and a bidirectional long short-term memory network are combined to perform garbage classification through feature extraction, multi-level memory enhancement, feature fusion and double threshold verification.

Benefits of technology

The accuracy of garbage classification is improved. By combining the comprehensive capture of spatial and temporal features and the complementary advantages of multiple features, the expressive power of the model and the reliability of the classification results are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808033A_ABST
    Figure CN120808033A_ABST
Patent Text Reader

Abstract

The invention discloses a garbage classification method and system based on a bidirectional long-short-term memory network and a convolutional neural network, and relates to the technical field of garbage classification, and the method comprises the steps: carrying out the feature extraction of a to-be-classified garbage image through an improved convolutional neural network, and generating a convolutional output feature; recombining the convolution output features into time sequence features, performing multi-level memory enhancement on the time sequence features by adopting an improved bidirectional long-short-term memory network, and determining memory combination output features; performing weighted fusion on the convolution output features and the memory combination output features based on a feature fusion network to construct fusion features; and performing feature dimension reduction on the fused features by adopting a classification network, then calculating category probability distribution, and performing dual threshold verification based on the category probability distribution to determine a garbage classification result. Based on the above scheme, the expression ability of the model for complex features is enhanced, dual threshold verification is performed after advantage complementation among multiple features is realized through feature fusion interaction, and the garbage classification accuracy is integrally improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of garbage classification, and particularly relates to a garbage classification method and system based on a bidirectional long short-term memory network and a convolutional neural network. BACKGROUND

[0002] With the acceleration of urbanization and the improvement of people's living standards, garbage classification and treatment has become an important issue in modern city management. The traditional garbage classification method mainly relies on manual identification and classification, which has problems such as low efficiency, unstable accuracy and high labor cost. Therefore, a garbage classification method based on deep learning is proposed.

[0003] In the existing garbage classification method based on deep learning, a single model structure such as a convolutional neural network or a recurrent neural network is usually used for feature extraction before Softmax classification, which has limited feature expression ability and is difficult to accurately identify garbage items with similar appearance but belonging to different categories, resulting in low garbage classification accuracy. SUMMARY

[0004] The present application provides a garbage classification method and system based on a bidirectional long short-term memory network and a convolutional neural network, which solves the technical problem of limited feature expression ability of the existing garbage classification method based on deep learning, which is difficult to accurately identify garbage items with similar appearance but belonging to different categories, resulting in low garbage classification accuracy.

[0005] The present application provides a garbage classification method and system based on a bidirectional long short-term memory network and a convolutional neural network, which solves the technical problem of limited feature expression ability of the existing garbage classification method based on deep learning, which is difficult to accurately identify garbage items with similar appearance but belonging to different categories, resulting in low garbage classification accuracy.

[0006] The garbage classification method based on a bidirectional long short-term memory network and a convolutional neural network provided by the present application comprises:

[0007] The improved convolutional neural network is used to extract features from the garbage image to be classified to generate convolutional output features.

[0008] After the convolutional output features are reorganized into time sequence features, the improved bidirectional long short-term memory network is used to perform multi-level memory enhancement on the time sequence features to determine memory combination output features.

[0009] The convolutional output features and the memory combination output features are weighted and fused based on the feature fusion network to construct fusion features.

[0010] The classification network is used to calculate the class probability distribution after the feature dimension reduction of the fusion features, and the double threshold verification is performed based on the class probability distribution to determine the garbage classification result.

[0011] Optionally, the improved convolutional neural network is used for feature extraction on the garbage image to be classified to generate a convolutional output feature, including:

[0012] The garbage image to be classified is subjected to convolutional feature extraction by a ResNet50 backbone network to generate a plurality of hierarchical features;

[0013] A feature pyramid module is used to perform multi-scale fusion based on the hierarchical features to construct a plurality of pyramid features;

[0014] The pyramid features are input into a spatial attention module for feature enhancement, spliced, and subjected to pooling processing by a global average pooling layer to output the convolutional output feature.

[0015] Optionally, the improved bidirectional long short-term memory network is used for multi-level memory enhancement on the time sequence feature to determine a memory combination output feature, including:

[0016] A bidirectional long short-term memory network is used for feature extraction on the time sequence feature to output a plurality of initial hidden states of time steps;

[0017] Short-term memory units are used to respectively perform short-term memory enhancement on the initial hidden states as inputs to generate corresponding short-term cell states and short-term hidden states;

[0018] Medium-term memory units are used to respectively perform medium-term memory enhancement on the associated short-term cell states and initial hidden states to construct corresponding medium-term hidden states;

[0019] Long-term memory units are used to respectively perform long-term memory enhancement on the associated initial hidden states and medium-term hidden states to output corresponding long-term hidden states;

[0020] The associated short-term hidden states, medium-term hidden states, and long-term hidden states are weighted and fused according to the time steps to generate corresponding enhanced hidden states;

[0021] The enhanced hidden states are spliced to output the memory combination output feature.

[0022] Optionally, the feature fusion network is used to perform weighted fusion on the convolutional output feature and the memory combination output feature to construct a fusion feature, including:

[0023] The convolutional output feature is mapped into a query matrix, and the memory combination output feature is projected into a value matrix and a key matrix, and then transmembrane state attention calculation is performed to output a fusion attention weight;

[0024] The fusion attention weight is subjected to global average pooling to generate a compressed attention weight;

[0025] determine a convolution feature weight of the convolution output feature and a memory feature weight of the memory combined output feature based on the attention weight respectively;

[0026] multiply the convolution output feature with the convolution feature weight element by element, and multiply the memory combined output feature with the memory feature weight element by element, and then perform feature fusion to output a fusion feature.

[0027] Optionally, the classification network is used to calculate a category probability distribution after feature dimension reduction of the fusion feature, and double threshold verification is performed based on the category probability distribution to determine a garbage classification result, including:

[0028] The fusion feature is continuously dimension-reduced to determine a comprehensive dimension-reduced feature.

[0029] A category probability distribution of the comprehensive dimension-reduced feature is calculated by a temperature-adjusted softmax classifier.

[0030] The maximum probability and the information entropy corresponding to the category probability distribution are used for double threshold judgment to determine a predicted category result.

[0031] Optionally, the training process of the trained garbage recognition classification model includes:

[0032] An original garbage image is preprocessed to determine a sample garbage image.

[0033] The sample garbage image is used to train the garbage recognition classification model based on a preset multi-task loss function to determine a trained garbage recognition classification model.

[0034] The second aspect of the present application provides a garbage classification system based on a bidirectional long short-term memory network and a convolutional neural network, including:

[0035] A preprocessing module is configured to determine a trained garbage recognition classification model according to a training garbage image and a multi-task loss function, and input a garbage image to be classified into the trained garbage recognition classification model. The trained garbage recognition classification model includes an improved convolutional neural network, an improved bidirectional long short-term memory network, a feature fusion network and a classification network.

[0036] A convolution extraction module is configured to extract features from the garbage image to be classified by the improved convolutional neural network to generate a convolution output feature.

[0037] A time sequence extraction module is configured to recombine the convolution output feature into a time sequence feature, and then use the improved bidirectional long short-term memory network to perform multi-level memory enhancement on the time sequence feature to determine a memory combined output feature.

[0038] The feature fusion module is configured to perform weighted fusion on the convolution output feature and the memory combination output feature based on a feature fusion network to construct a fusion feature.

[0039] The classification module is configured to calculate a category probability distribution after performing feature dimension reduction on the fusion feature by using a classification network, and perform double threshold verification based on the category probability distribution to determine a garbage classification result.

[0040] The third aspect of the present application provides a computer device, which comprises a memory and a processor, the memory stores a computer program, and the computer program is executed by the processor to make the processor execute the steps of the garbage classification method based on the bidirectional long short-term memory network and the convolutional neural network.

[0041] The fourth aspect of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed to implement the garbage classification method based on the bidirectional long short-term memory network and the convolutional neural network.

[0042] The fifth aspect of the present application provides a computer program product, which comprises computer programs / instructions, and the computer programs / instructions are executed by a processor to implement the garbage classification method based on the bidirectional long short-term memory network and the convolutional neural network.

[0043] From the above technical solutions, the present application has the following advantages:

[0044] The scheme provides a garbage classification method based on a bidirectional long short-term memory network and a convolutional neural network, which comprises the following steps: inputting a garbage image to be classified into a trained garbage recognition classification model; the trained garbage recognition classification model comprises an improved convolutional neural network, an improved bidirectional long short-term memory network, a feature fusion network and a classification network; the improved convolutional neural network is used for feature extraction of the garbage image to be classified to generate convolutional output features; after the convolutional output features are reorganized into time sequence features, the improved bidirectional long short-term memory network is used for multi-level memory enhancement of the time sequence features to determine memory combination output features; the convolutional output features and the memory combination output features are weighted and fused based on the feature fusion network to construct fusion features; and the classification network is used for feature dimension reduction of the fusion features to calculate a category probability distribution, and double threshold verification is performed based on the category probability distribution to determine a garbage classification result. According to the spatial feature extraction capability of the improved convolutional neural network and the extraction of long-distance dependent features and the dynamic integration capability of context information of the improved bidirectional long short-term memory network, the garbage image features are comprehensively captured, the spatial information and the time sequence information in the image can be more effectively captured, the expression capability of the model for complex features is enhanced, the double threshold verification is performed after the advantages of the multiple features are complemented through feature fusion interaction, and the garbage classification accuracy is improved as a whole. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0046] Figure 1 A step flow chart of a garbage classification method based on a bidirectional long short-term memory network and a convolutional neural network provided by the embodiment of the present application;

[0047] Figure 2 A structure schematic diagram of an improved convolutional neural network provided by the embodiment of the present application;

[0048] Figure 3 A sample confidence distribution histogram based on a dynamic temperature parameter in different training stages provided by the embodiment of the present application;

[0049] Figure 4 A gradient norm comparison between an improved bidirectional long short-term memory network and a baseline model provided by the embodiment of the present application;

[0050] Figure 5 An image processing flowchart schematic diagram of a model training process provided by the embodiment of the present application;

[0051] Figure 6 A structure block diagram of a garbage classification system based on a bidirectional long short-term memory network and a convolutional neural network is provided for the embodiments of the present application. DETAILED DESCRIPTION

[0052] The embodiments of the present application provide a garbage classification method and system based on a bidirectional long short-term memory network and a convolutional neural network, which are used to solve the technical problem that the feature expression capability of the existing garbage classification method based on deep learning is limited, it is difficult to accurately identify garbage items with similar appearance but belonging to different categories, and the garbage classification accuracy is low.

[0053] In order to make the application purpose, features and advantages of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the following described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0054] Please refer to Figure 1 , Figure 1 A step flowchart of a garbage classification method based on a bidirectional long short-term memory network and a convolutional neural network is provided for the embodiments of the present application.

[0055] The garbage classification method based on a bidirectional long short-term memory network and a convolutional neural network provided by the present embodiment comprises:

[0056] Step 101, input the garbage image to be classified into the trained garbage recognition classification model; the trained garbage recognition classification model comprises an improved convolutional neural network, an improved bidirectional long short-term memory network, a feature fusion network and a classification network.

[0057] It should be noted that the present embodiment designs a garbage recognition classification model for garbage classification task, which comprises an improved convolutional neural network, an improved bidirectional long short-term memory network, a feature fusion network and a classification network. After model training of the built garbage recognition classification model to be trained, the trained garbage recognition classification model can be determined, and when the garbage image to be classified is obtained, the trained garbage recognition classification model is input for processing.

[0058] Step 102, feature extraction of the garbage image to be classified is performed through the improved convolutional neural network to generate convolution output features.

[0059] Step 102 comprises the following sub-steps:

[0060] The ResNet50 backbone network is used for convolution feature extraction of the garbage image to be classified, and multiple hierarchical features are generated.

[0061] The feature pyramid module is used for multi-scale fusion based on the hierarchical features, and multiple pyramid features are constructed.

[0062] The pyramid features are input into the spatial attention module for feature enhancement, splicing, and global average pooling layer for pooling processing, and the convolution output feature is output.

[0063] It should be noted that in the present embodiment, as shown in Figure 2 The improved convolutional neural network, i.e. the improved ResNet50, includes a ResNet50 backbone network, a feature pyramid module (FPN), a spatial attention module (SAM), and a global average pooling layer (GlobalAvgPool). The improved convolutional neural network uses a deep residual network as the basic architecture, which can effectively alleviate the gradient vanishing problem of deep networks. At the same time, by introducing the feature pyramid and spatial attention mechanism, the model can adaptively focus on the key regions in the image, improving the relevance of feature extraction. The differences between the improved ResNet50 and the standard ResNet50 are shown in Table 1:

[0064] Table 1 Comparison of structure composition of improved ResNet50 and standard ResNet50

[0065]

[0066] As shown in Table 1 and Figure 2 In the improved convolutional neural network, first, the garbage image to be classified is input into the ResNet50 backbone network to extract multiple hierarchical features , is the hierarchical index , is the hierarchical feature of the layer, which is fused with multi-scale information by the feature pyramid module to output , is the pyramid feature of the layer, which is enhanced by the spatial attention module to enhance the key region features, and finally spliced and pooled by the global average pooling layer to the same dimension feature output as the convolution output feature.

[0067] The determination process of the convolution output feature includes:

[0068] ;

[0069] ;

[0070] Where, For hierarchical index , For the The spatial attention weights of the layer, is the sigmoid activation function, is a 7×7 convolution unit, is the global maximum pooling, is the global average pooling, For the Layer pyramid features, is the convolution output feature, and the dimension of the convolution output feature is C (number of channels) × H (height) × W (width).

[0071] Step 103: After the convolution output features are reorganized into time series features, an improved bidirectional long short-term memory network is used to perform multi-level memory enhancement on the time series features to determine the memory combination output features.

[0072] It should be noted that in order to facilitate the use of the improved bidirectional long short-term memory network for feature processing, it is necessary to transform the convolution output features into a sequence form by feature reorganization methods such as spatial flattening operations, thereby obtaining the corresponding time series features. The convolution output features and the time series features meet the requirements. , is the time step, At the same time, this embodiment introduces a multi-level memory mechanism in the improved bidirectional long short-term memory network to enhance the output memory combination output features. This design enables the model to capture feature information of different time scales at the same time, significantly improving the expression ability of time series features.

[0073] In a specific implementation of this embodiment, an improved bidirectional long short-term memory network is used to perform multi-level memory enhancement on the temporal features to determine the memory combination output features, including:

[0074] A bidirectional long short-term memory network is used to extract time series features and output the initial hidden states of multiple time steps;

[0075] Through the short-term memory unit, each initial hidden state is used as input to perform short-term memory enhancement, generating the corresponding short-term cell state and short-term hidden state;

[0076] The medium-term memory unit is used to perform medium-term memory enhancement based on each associated short-term cell state and the initial hidden state, and the corresponding medium-term hidden state is constructed;

[0077] Perform long-term memory enhancement based on the initial hidden state and the mid-term hidden state of each association through the long-term memory unit, and output the corresponding long-term hidden state;

[0078] The associated short-term hidden state, the medium-term hidden state and the long-term hidden state are fused by weighting according to each time step to generate a corresponding enhanced hidden state;

[0079] The enhanced hidden states are spliced to output a memory combined output feature.

[0080] It should be noted that in the processing process of the improved bidirectional long short-term memory network, the bidirectional long short-term memory network BiLSTM is first used to process the time sequence feature step by step, then the short-term memory unit, the medium-term memory unit and the long-term memory unit form a multi-level memory enhancement mechanism for feature enhancement processing and time step hidden state splicing, and finally the output feature of the improved bidirectional long short-term memory network is output as a memory combined output feature. This design enables the model to capture feature information of different time scales at the same time, significantly improving the expression ability of the time sequence feature;

[0081] The BiLSTM includes a forward LSTM and a backward LSTM. The initial hidden state is obtained by splicing the corresponding forward hidden state and backward hidden state obtained by performing feature processing on the LSTM mechanism through the forward and backward LSTMs. For details, refer to: , , In the formula, is the time step index , is the forward hidden state of the time step , is the forward LSTM, is the input of the bidirectional long short-term memory network of the time step , is the forward hidden state of the time step , is the backward hidden state of the time step , is the backward LSTM, is the backward hidden state of the time step , is the initial hidden state of the time step (the feature dimension is D);

[0082] It can be understood that the processing process of the LSTM mechanism includes: , , , , , In the formula, is the forget gate of the time step , is a sigmoid activation function, is the weight matrix of the forget gate, is the time step The hidden state of is the time step LSTM input, is the bias term of the forget gate, is the time step The input gate, is the weight matrix of the input gate, is the bias term of the input gate, is the time step Candidate cell states, is the tanh activation function, is the weight matrix of the candidate cell state, is the bias term of the candidate cell state, is the time step The cell state, is the time step The cell state, is the time step The output gate, is the weight matrix of the output gate, is the bias term of the output gate, is the time step The hidden state of

[0083] The short-term memory unit inherits the standard LSTM structure, takes the initial hidden state as input and performs short-term memory enhancement according to the LSTM mechanism, thereby outputting the corresponding short-term cell state and short-term hidden state, including: , , where is the time step The short-term cell state, is the time step The short-term forget gate, is the time step The short-term cell state, is the time step The short-term input gate, is the time step Short-term candidate cell states, is the time step The short-term output gate, is the time step The short-term hidden state of is element-by-element multiplication; the process of determining the short-term cell state and the short-term hidden state can be specifically referred to the processing process of the aforementioned LSTM mechanism;

[0084] The medium-term memory unit adds cross-time step connections, and the medium-term hidden state determination process includes: , wherein, is the mid-term hidden state at time step , is the mid-term hidden state at time step , is layer normalization, is the mid-term weight matrix, is the mid-term gating value, is the weight matrix of the mid-term gating value;

[0085] The long-term memory unit introduces a periodical update mechanism, and the long-term hidden state determination process comprises: , wherein, is the long-term hidden state at time step , is the time step index , is the long-term hidden state at time step , is the long-term weight matrix, is the long-term gating value, is the weight matrix of the long-term gating value;

[0086] The enhanced hidden state at each time step is the weighted sum of the three levels of memory, i.e. , is the learnable short-term attention weight, is the learnable mid-term attention weight, is the learnable long-term attention weight, is the enhanced hidden state at time step , and finally the enhanced hidden states at all time steps are spliced to obtain the memory combined output feature with a dimension of TxD.

[0087] Step 104, the convolution output feature and the memory combined output feature are weighted fused based on the feature fusion network to construct a fusion feature.

[0088] Step 104 comprises the following sub-steps:

[0089] After mapping the convolution output feature into a query matrix and projecting the memory combined output feature into a value matrix and a key matrix, cross-membrane state attention calculation is performed to output a fusion attention weight;

[0090] The fusion attention weight is globally averaged pooled to generate a compressed attention weight;

[0091] Based on the attention weight, a convolution feature weight of the convolution output feature and a memory feature weight of the memory combined output feature are respectively determined;

[0092] The convolution output feature and the convolution feature weight are multiplied element by element, and the memory combination output feature and the memory feature weight are multiplied element by element, and then the feature fusion output fusion feature is performed.

[0093] It should be noted that, compared with relying on a single model structure for feature extraction, the embodiment based on the convolutional neural network and the bidirectional long short-term memory network extracts features and then effectively fuses the multi-modal features in the feature fusion network, which can simultaneously consider the spatial features and the time sequence dependent relationship of the image;

[0094] At the same time, compared with the traditional method which mainly uses simple feature splicing or average fusion strategy, this way ignores the importance difference of different features and lacks effective feature selection mechanism, which is easy to introduce redundant information and cause information loss. In the feature fusion network, the embodiment performs weight compression after deeply fusing the convolution output feature and the memory combination output feature through the multi-head attention mechanism, and realizes feature adaptive fusion through the dynamic weighting mechanism to construct the fusion feature:

[0095] Firstly, the multi-head attention mechanism is used to calculate the fusion attention weight of the convolution output feature and the memory combination output feature, including: the convolution output feature is mapped into a query matrix , the memory combination output feature is projected into a value matrix and a key matrix , wherein is the weight matrix of the learnable query matrix, is the weight matrix of the learnable key matrix, is the weight matrix of the learnable value matrix, and the query matrix, the value matrix and the key matrix are used for cross-membrane attention calculation, that is , is the fusion attention weight, is the multi-head attention, is the transpose of the key matrix, is the softmax activation function, is the dimension of the key matrix;

[0096] Then, the fusion attention weight is compressed, including: , is the compressed attention weight, is the global average pooling;

[0097] Finally, the dynamic weighting mechanism is used for feature adaptive fusion, including: the convolution feature weight of the convolution output feature is determined based on the attention weight, that is , is the convolution feature weight, is the sigmoid activation function, a weight matrix of the learnable convolution output feature, a weight matrix of the learnable compression attention weight; the memory feature weight of the memory combination output feature is determined based on the attention weight, , a memory feature weight, a weight matrix of the learnable memory combination output feature; the fusion feature is outputted by performing feature fusion on the convolution output feature, the convolution feature weight, the memory combination output feature and the memory feature weight , that is ; it can be understood that if is larger, the model will rely more on the attention weight to adjust the feature fusion, if tends to 0, the model will degenerate into the original dynamic weighting, which enhances the flexibility of the model, and and are dynamically learned weights, which can be used to balance the convolution output feature and the memory combination output feature.

[0098] Step 105: After the fusion feature is dimensionally reduced by using the classification network, the category probability distribution is calculated, and the double threshold verification is performed based on the category probability distribution to determine the garbage classification result.

[0099] Step 105 includes the following sub-steps:

[0100] The fusion feature is continuously dimensionally reduced to determine the comprehensive dimensionally reduced feature;

[0101] The category probability distribution of the comprehensive dimensionally reduced feature is calculated by the temperature-adjusted softmax classifier;

[0102] The maximum probability and the information entropy corresponding to the category probability distribution are double-thresholded to determine the predicted category result.

[0103] It should be noted that, in the classification decision link, compared with the commonly used standard softmax classifier, this method is more sensitive to the class imbalance problem and lacks the necessary confidence evaluation mechanism, and often shows poor classification effect. In this embodiment, the category probability distribution of the dimensionally reduced fusion feature is calculated in the classification network, and then the double threshold verification is performed to determine the final garbage classification result:

[0104] First, the fusion feature is continuously dimensionally reduced by using the full connection layer, the batch normalization layer, the ReLU activation function and the random inactivation operation, that is , , wherein is the fusion feature, is a full connection layer with 512 neurons, is a batch normalization layer, is the ReLU activation function, is the random dropout operation, is the intermediate dimension reduction feature, is a fully connected layer of 256 neurons, is the comprehensive dimension reduction feature;

[0105] Then, a temperature-adjusted softmax classifier is used to calculate the class probability distribution of the integrated dimensionality reduction features:

[0106] ;

[0107] Where, Index for junk categories , is a collection of garbage categories, For category The weight matrix, For category The bias term, is the temperature parameter, Index for junk categories , For category The weight matrix, For category The bias term, is the true category of the garbage image to be classified, is the garbage image to be classified, is the category probability distribution; it can be understood that and Used to adjust classification boundaries;

[0108] Among them, the temperature parameter It is used to dynamically adjust the generalization and utilization confidence of the balance model. In this embodiment, the temperature parameter can be dynamically adjusted according to the training stage. For training rounds:

[0109] ;

[0110] The sample confidence distribution histogram of this mechanism based on dynamic temperature parameters at different training stages is as follows: Figure 3As shown, in the early stage of training (τ = 1.5, Epoch≤10), the proportion of samples with low confidence interval (0.0-0.6) is high, indicating that the model's initial learning stage is conservative in judging data features, and the high temperature parameter alleviates the prediction deviation caused by unstable initial parameters by softening the probability distribution; as the training enters the middle stage (τ = 1.0, 10<Epoch≤30), the proportion of samples with confidence interval of 0.6-0.8 significantly increases, reflecting that the model gradually captures effective features and improves its discrimination ability; in the late stage of training (τ = 0.5, Epoch>30), the proportion of samples with high confidence interval (0.8-1.0) jumps to a peak, and the proportion of samples with low confidence interval decreases sharply, which confirms that the introduction of the low temperature parameter strengthens the model's focusing ability on high confidence samples, and through the dynamic temperature adjustment strategy, the gradual optimization from coarse-grained exploration to fine-grained confidence calibration is realized, which effectively balances the training stability and prediction certainty.

[0111] Finally, a double threshold judgment is made based on the category probability distribution:

[0112] , ;

[0113] wherein, is the maximum probability in the category probability distribution (used to reflect the confidence of the model on the most likely category), is the probability of the category , and is the information entropy (used to measure the uncertainty of prediction), is the garbage classification result; according to the above formula, the smaller the entropy value (such as ), the more concentrated the probability distribution (the more certain the model is about the prediction result), the larger the entropy value, the more dispersed the probability distribution (the more uncertain the model), and only when the confidence of the model about the prediction result is high enough and the prediction distribution is concentrated enough , the prediction result is accepted, otherwise it is marked as rejecting the prediction result , so as to avoid false classification caused by low confidence or ambiguous prediction. This design can adaptively adjust the parameters of different features according to the characteristics of the input data, improving the flexibility and robustness of the model.

[0114] In one specific embodiment of the present embodiment, the training process of the trained garbage recognition classification model includes:

[0115] performing image preprocessing on the original garbage image to determine a sample garbage image;

[0116] using the sample garbage image based on a preset multi-task loss function to train the garbage recognition classification model to be trained, to determine the trained garbage recognition classification model.

[0117] It should be noted that in the model training stage, a complete and systematic strategy is formulated to ensure that the model can stably and efficiently converge to the global optimal solution;

[0118] First, after obtaining the original garbage image I from the data source, image preprocessing is performed to obtain a sample garbage image for model training. In specific implementation, image preprocessing can include size adjustment, data enhancement, and global normalization. Size adjustment can unify the format and scale of input data to improve the stability and accuracy of subsequent processing. Data enhancement can expand the training samples to improve the generalization ability of the model. Global normalization processing can accelerate model convergence:

[0119] The original garbage image is resized according to a preset image size, which can be set according to the image input size of the preset garbage recognition classification model, for example, adjusted to 224x224 pixels, to ensure that the image details are preserved while matching the image input size of the preset garbage recognition classification model, thereby facilitating transfer learning. In one implementation, a bilinear interpolation algorithm can be used for size adjustment, which not only ensures the uniformity of image resolution but also effectively preserves image texture features.

[0120] Data enhancement can include random rotation, flipping, brightness and contrast adjustment, and other transformations, so that the model can still extract stable features and maintain good classification performance when facing differences caused by angle, position and lighting conditions in actual shooting. The transformation process can be referred to as follows:

[0121] ;

[0122] ;

[0123] In the formula, is the transformation matrix, is the random rotation angle, is the x-direction translation parameter and the y-direction translation parameter, is the image pixel value after brightness and contrast adjustment, is the original garbage image pixel value, is the brightness adjustment factor, is the contrast adjustment factor;

[0124] Normalization includes zero-mean normalization and standard deviation global normalization of the image, which satisfies: In the formula, is the sample garbage image, is the normalized input image, is the mean of all normalized input image pixel values, a standard deviation of all normalized input image pixel values;

[0125] Secondly, the model training configuration is initialized, including:

[0126] 1) A multi-task loss function is designed to calculate the loss function value, in which the weighted cross-entropy, L2 regularization, feature consistency loss and auxiliary task loss are organically combined, and through a dynamic weight adjustment mechanism, the loss weight is automatically allocated according to the training progress of each task, so as to realize the dynamic balance between multi-task; the multi-task loss function can include:

[0127] ;

[0128] ;

[0129] ;

[0130] ;

[0131] ;

[0132] In the formula, is a multi-task loss function, is a weighted cross-entropy, is a weight coefficient of the weighted cross-entropy, is an L2 regularization term, is a weight coefficient of the L2 regularization term, is a feature consistency loss, is a weight coefficient of the feature consistency loss, is an auxiliary task loss, is a weight coefficient of the auxiliary task loss, which is used to adjust the proportion of different loss terms in the total loss; is a garbage class index, is a garbage class set, is a weight of the class , is a true label of a sample belonging to the class , is a probability of the model predicting that the sample belongs to the class , is a weight matrix of the model, is a weight value of the row and the column of the weight matrix of the model, is an index of a feature vector, is an auxiliary classification task class index , a set of all auxiliary classification task categories, a true label of the sample belonging to a category in the auxiliary task, a probability of the model predicting that the sample belongs to a category in the auxiliary task,

[0133] 2) Initialize the optimizer (such as the Adam optimizer), set the learning rate, momentum parameter, and weight decay hyperparameters, combine the warmup learning rate scheduling strategy, batch processing strategy, gradient clipping mechanism, and early stopping mechanism, input the sample garbage image into the garbage recognition classification model to be trained for model training, thereby determining the trained garbage recognition classification model; wherein:

[0134] dynamically adjust the time step as follows: , the training phase randomly inactivates neurons with a probability of 0.5;

[0135] The Adam optimizer is used for training process optimization, the parameter setting is the decay rate of the first-order moment estimate of the calculated gradient (i.e., momentum) , the decay rate of the second-order moment estimate of the calculated gradient , the stability constant , and the learning rate adaptive adjustment is:

[0136] ;

[0137] wherein, the initial learning rate , the learning rate decay , is the decay step number, is the current training step number;

[0138] The warmup learning rate scheduling strategy can smoothly start at the beginning of training and gradually reduce the learning rate in the later period, ensuring the smoothness and final convergence effect of the training process; throughout the process, the learning rate is first linearly increased to the preset learning rate in the learning rate warm-up phase, and then the learning rate is decayed according to the cosine function law in the learning rate cosine annealing phase, including:

[0139] ;

[0140] ;

[0141] In the formula, is , the total training step number in the learning rate warm-up phase (which can be set to 10% of the total model training step number), is the initial learning rate, and the current training step number is , the total model training step number, ​​is the current learning rate;

[0142] In terms of batch processing strategy, by setting a reasonable base batch size and the number of gradient accumulation steps, the effective batch size is amplified, and the mixed precision training technology is used to further improve the training efficiency under the premise of ensuring the calculation accuracy:

[0143] ;

[0144] The gradient clipping mechanism is defined as:

[0145] ;

[0146] In the formula, is the original gradient, is the clipped gradient, is the clipping threshold (which can be set to 5.0 in this embodiment);

[0147] In order to avoid overfitting, an early stopping mechanism is introduced in the training process, which is defined as: when the validation set loss does not improve for consecutive training epochs, stop training; specifically, let the current training epoch be e, and the historical optimal validation loss be , then the training termination condition is: , wherein is the preset tolerance number of epochs, is the index of the training epoch, is the optimal validation loss of the training epoch, when the above condition is met, the training is immediately stopped and the previously saved optimal checkpoint (corresponding to the model parameters of ) is loaded, so as to ensure the generalization ability of the model;

[0148] In addition, a model ensemble strategy can also be used, that is, the last 10 checkpoint models are saved during training, and the final prediction result is the soft voting integration of these models.

[0149] In order to better illustrate the model training process, referring to Figure 4 , the overall framework diagram of the model training process of the present application is shown, which mainly includes four core steps: image preprocessing (S1), feature extraction (S2), time series feature processing (S3), and feature fusion and classification (S4);

[0150] S1, image preprocessing, preprocessing the input original garbage image, including: S11, uniformly adjusting the original garbage image to a preset size of 224x224 pixels; S12, performing data enhancement processing on the image, including random rotation, flipping, brightness and contrast adjustment, etc. transformation; S13, performing zero mean and standard deviation normalization processing on the image to obtain a sample image matrix;

[0151] S2, feature extraction, including: S21, extracting hierarchical features through an improved convolutional neural network; S22, fusing multi-scale information through a feature pyramid network, and then enhancing key region features through a spatial attention mechanism; S23, pooling the features after the same dimension to obtain convolution output features;

[0152] S3, time sequence feature processing, including: S31, processing sequence features using BiLSTM; S32, implementing a multi-level memory enhancement mechanism; S33, concatenating the final hidden states of all time steps to obtain a memory combined output feature;

[0153] S4, feature fusion and classification, including: S41, realizing deep fusion of the convolution output features and the memory combined output features through a multi-head attention mechanism; S42, realizing feature adaptive fusion through a dynamic weighting mechanism; S43, after dimension reduction, a temperature-adjusted Softmax classifier is used to classify and calculate the category probability distribution and perform double threshold verification, and output the predicted garbage classification result.

[0154] It should be noted that only the general process of the garbage classification method is briefly described here, and the specific implementation process of each step can be understood by referring to the related contents in the foregoing embodiments, which will not be repeated here. It can be understood that the present application does not limit this.

[0155] To illustrate the effect of the model training process, refer to Figure 5The gradient norm contrast result shown, during the experimental training process, the traditional LSTM sharply rises to a peak of 4.8 in the early stage (100 to 1000 steps), exposing obvious gradient instability phenomenon, and then gradually decreases but always accompanied by sharp fluctuations, and finally stabilizes at a relatively high level (0.7), and the parameter update process has a risk of avoiding oscillation; the traditional BiLSTM effectively alleviates the limitations of one-way structure by introducing a bidirectional information transmission mechanism, and the gradient norm overall decreases and the peak reduces to 2.2, but still presents fluctuations of 0.3 to 0.6 in the later stage, indicating that its optimization process has not completely overcome the inherent defects of unstable deep network training; in contrast, the improved BiLSTM of the embodiment relies on the dual optimization of dynamic attention weight distribution and gradient clipping strategy, and the gradient norm starts to maintain a steady downward trend from the initial value of 1.2, without abnormal fluctuations and finally converges to 0.2, and its curve slope and absolute value are significantly better than the traditional group. This stable gradient decay feature not only verifies the inhibitory effect of the improved mechanism on the gradient explosion and disappearance problem, but also directly reflects the high robustness and convergence efficiency of the enhanced BiLSTM model in the parameter update of the complex training scene, providing strong support for the reliability of the model in long sequence tasks.

[0156] To verify the effectiveness of the present scheme, comparative experiments are also carried out in the embodiment, under the condition that the other structures of the model are the same, the advantages of the standard ResNet50 and the improved ResNet50 in the embodiment are compared, as shown in Table 2:

[0157] Table 2 Improved advantages of improved ResNet50

[0158]

[0159] It can be understood that the method of the embodiment can be realized by software, and can be programmed in Python language, and the model can be constructed and trained using deep learning frameworks such as PyTorch. The system can be deployed on a server with GPU acceleration capability, or can be deployed on an edge computing device after model compression and optimization, and has strong practicality and expansibility.

[0160] In the embodiment of the present application, according to the spatial feature extraction capability of the improved convolutional neural network and the long-distance dependence feature extraction and dynamic integration capability of context information of the improved bidirectional long short-term memory network, the comprehensive capture of the garbage image features is realized, the spatial information and time sequence information in the image can be more effectively captured, the expression capability of the model for complex features is enhanced, the advantage complementation between multiple features is realized through feature fusion interaction, the misjudgment rate is greatly reduced through the dual verification of confidence and entropy, and through a large number of experiments, compared with the current widely used ResNet50 benchmark model, the classification accuracy of the embodiment on the standard test data set is improved by about 15 percentage points, especially in the differentiation of similar appearance garbage categories, and the recognition effect is still stable in complex environments such as light, angle and noise, which reflects its excellent robustness and practicality; from the perspective of model generalization capability, through the innovative architecture design and feature fusion strategy, the model not only performs well on the training data, but also maintains a high recognition accuracy when facing new scenes and variable environments, in the small sample learning scene, due to the efficient data enhancement and feature extraction mechanism, the training samples required by the embodiment are greatly reduced compared with the traditional deep learning method, and the embodiment has good incremental learning capability and can flexibly adapt to the subsequent changing data distribution, at the same time, in view of the strict requirements for inference speed, memory occupation and power consumption in the actual deployment process, the embodiment can be optimized on the model compression and quantization technology, so that the real-time performance and efficiency can be guaranteed on different hardware platforms, and the application requirements of garbage classification in various edge devices and embedded systems can be met, further promoting the popularization and landing application of the garbage classification system, and providing strong technical support and practical basis for environmental protection and resource recycling.

[0161] Please refer to Figure 6 , Figure 6 The structure block diagram of the garbage classification system based on the bidirectional long short-term memory network and the convolutional neural network provided by the embodiment of the present application is shown in the figure.

[0162] The garbage classification system based on the bidirectional long short-term memory network and the convolutional neural network provided by the embodiment of the present application comprises:

[0163] The preprocessing module 601 is configured to determine a trained garbage recognition classification model according to a training garbage image and a multi-task loss function, and input a garbage image to be classified into the trained garbage recognition classification model; the trained garbage recognition classification model comprises an improved convolutional neural network, an improved bidirectional long short-term memory network, a feature fusion network and a classification network.

[0164] The convolution extraction module 602 is configured to perform feature extraction on the garbage image to be classified through the improved convolutional neural network, and generate convolution output features.

[0165] The timing extraction module 603 is configured to perform multi-level memory enhancement on the timing features by using an improved bidirectional long short-term memory network after the convolution output features are reorganized into the timing features, and determine memory combination output features;

[0166] The feature fusion module 604 is configured to perform weighted fusion on the convolution output features and the memory combination output features based on a feature fusion network, and construct fusion features.

[0167] The classification module 605 is configured to calculate a category probability distribution after performing feature dimension reduction on the fusion features by using a classification network, and perform double threshold verification based on the category probability distribution to determine a garbage classification result.

[0168] Preferably, the convolution extraction module 602 is specifically configured to:

[0169] perform convolution feature extraction on the garbage image to be classified by using a ResNet50 backbone network to generate a plurality of hierarchical features;

[0170] perform multi-scale fusion on the hierarchical features based on a feature pyramid module to construct a plurality of pyramid features;

[0171] input the pyramid features into a spatial attention module for feature enhancement, splice the pyramid features after the feature enhancement, and perform pooling processing on the pyramid features by using a global average pooling layer to output convolution output features.

[0172] Preferably, the timing extraction module 603 is specifically configured to:

[0173] perform feature extraction on the timing features by using a bidirectional long short-term memory network to output a plurality of initial hidden states of time steps;

[0174] perform short-term memory enhancement on the initial hidden states by using a short-term memory unit to generate corresponding short-term cell states and short-term hidden states;

[0175] perform mid-term memory enhancement on the associated short-term cell states and initial hidden states by using a mid-term memory unit to construct corresponding mid-term hidden states;

[0176] perform long-term memory enhancement on the associated initial hidden states and mid-term hidden states by using a long-term memory unit to output corresponding long-term hidden states;

[0177] perform weighted fusion on the associated short-term hidden states, mid-term hidden states and long-term hidden states according to the time steps to generate corresponding enhanced hidden states;

[0178] splice and output the enhanced hidden states to obtain memory combination output features.

[0179] Preferably, the feature fusion module 604 is specifically used for:

[0180] After mapping the convolution output feature into a query matrix and projecting the memory combination output feature into a value matrix and a key matrix, cross-membrane state attention calculation is performed to output fusion attention weights;

[0181] Global average pooling is performed on the fusion attention weights to generate compressed attention weights;

[0182] Based on the attention weights, convolution feature weights of the convolution output feature and memory feature weights of the memory combination output feature are determined respectively;

[0183] The convolution output feature is multiplied element by element with the convolution feature weights, and the memory combination output feature is multiplied element by element with the memory feature weights, and then feature fusion is performed to output fusion features.

[0184] Preferably, the classification module 605 is specifically used for:

[0185] Continuous dimension reduction is performed on the fusion features to determine comprehensive dimension reduction features;

[0186] The class probability distribution of the comprehensive dimension reduction features is calculated by a temperature-adjusted softmax classifier;

[0187] Based on the maximum probability and the information entropy corresponding to the class probability distribution, a double threshold judgment is performed to determine a predicted class result.

[0188] Preferably, the training process of the trained garbage recognition classification model comprises:

[0189] Image preprocessing is performed on the original garbage image to determine a sample garbage image;

[0190] Based on a preset multi-task loss function, the sample garbage image is used to train the garbage recognition classification model to be trained to determine a trained garbage recognition classification model.

[0191] The embodiment of the application also provides a computer device comprising a memory and a processor, and the memory stores a computer program; when the computer program is executed by the processor, the processor executes the steps of the garbage classification method based on the bidirectional long short-term memory network and the convolutional neural network according to any one of the above embodiments.

[0192] The embodiment of the application also provides a computer readable storage medium, which stores a computer program / instruction, and when the computer program / instruction is executed by a processor, the steps of the garbage classification method based on the bidirectional long short-term memory network and the convolutional neural network according to any one of the above embodiments are implemented.

[0193] The embodiment of the present application further provides a computer program product comprising computer programs / instructions, which, when executed by a processor, implement the steps of the garbage classification method based on the bidirectional long short-term memory network and the convolutional neural network according to any of the above embodiments.

[0194] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system and modules can refer to the corresponding processes in the foregoing method embodiments, and will not be described here.

[0195] In several embodiments provided in the present application, it should be understood that the disclosed system and method can be implemented in other manners. For example, the above-described system embodiments are merely schematic, and the division of the units is merely a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.

[0196] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments.

[0197] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be in the form of hardware or in the form of software functional units.

[0198] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the entire or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0199] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A garbage classification method based on bidirectional long short-term memory network and convolutional neural network, characterized in that: include: Input the garbage image to be classified into the trained garbage recognition and classification model; The trained garbage recognition and classification model includes an improved convolutional neural network, an improved bidirectional long short-term memory network, a feature fusion network, and a classification network; Extracting features of the garbage image to be classified by using an improved convolutional neural network to generate convolution output features; After reorganizing the convolution output features into time series features, an improved bidirectional long short-term memory network is used to perform multi-level memory enhancement on the time series features to determine a memory combination output feature; Based on the feature fusion network, the convolution output feature and the memory combination output feature are weightedly fused to construct a fusion feature; A classification network is used to perform feature dimensionality reduction on the fusion features and then calculate the category probability distribution. Double threshold verification is performed based on the category probability distribution to determine the garbage classification result.

2. The garbage classification method based on bidirectional long short-term memory network and convolutional neural network according to claim 1 is characterized in that: The improved convolutional neural network is used to extract features from the garbage image to be classified to generate convolution output features, including: Perform convolutional feature extraction on the garbage image to be classified through the ResNet50 backbone network to generate multiple hierarchical features; A feature pyramid module is used to perform multi-scale fusion based on the hierarchical features to construct multiple pyramid features; The pyramid features are input into the spatial attention module for feature enhancement and then spliced, and pooled through the global average pooling layer to output the convolution output features.

3. The garbage classification method based on bidirectional long short-term memory network and convolutional neural network according to claim 1 is characterized in that: The improved bidirectional long short-term memory network is used to perform multi-level memory enhancement on the time series features to determine the memory combination output features, including: A bidirectional long short-term memory network is used to extract the time series features and output the initial hidden states of multiple time steps; Performing short-term memory enhancement using each of the initial hidden states as input through a short-term memory unit to generate corresponding short-term cell states and short-term hidden states; The medium-term memory unit is used to perform medium-term memory enhancement based on each associated short-term cell state and the initial hidden state, and the corresponding medium-term hidden state is constructed; Perform long-term memory enhancement based on the initial hidden state and the mid-term hidden state of each association through the long-term memory unit, and output the corresponding long-term hidden state; Perform weighted fusion of the associated short-term hidden state, medium-term hidden state, and long-term hidden state at each time step to generate the corresponding enhanced hidden state; Each of the enhanced hidden states is concatenated and output as a memory combination output feature.

4. The garbage classification method based on bidirectional long short-term memory network and convolutional neural network according to claim 1 is characterized in that: The feature fusion network uses the convolution output feature and the memory combination output feature for weighted fusion to construct a fusion feature, including: Mapping the convolution output features into a query matrix, and projecting the memory combination output features into a value matrix and a key matrix, performing cross-membrane state attention calculation to output fusion attention weights; Performing global average pooling on the fused attention weights to generate compressed attention weights; Determining, based on the attention weight, a convolution feature weight of the convolution output feature and a memory feature weight of the memory combination output feature; The convolution output feature is multiplied by the convolution feature weight element by element, and the memory combination output feature is multiplied by the memory feature weight element by element, and then feature fusion is performed to output a fusion feature.

5. The garbage classification method based on bidirectional long short-term memory network and convolutional neural network according to claim 1 is characterized in that: The method of using a classification network to perform feature dimensionality reduction on the fused features and then calculating the category probability distribution, and performing double threshold verification based on the category probability distribution to determine the garbage classification result includes: Continuously reducing the dimensionality of the fused features to determine comprehensive dimensionality reduction features; Calculating the class probability distribution of the comprehensive dimensionality reduction feature by a temperature-regulated softmax classifier; A double threshold judgment is performed based on the maximum probability and information entropy corresponding to the category probability distribution to determine the predicted category result.

6. The garbage classification method based on bidirectional long short-term memory network and convolutional neural network according to claim 1 is characterized in that: The training process of the trained garbage identification and classification model includes: Performing image preprocessing on the original garbage image to determine a sample garbage image; Based on a preset multi-task loss function, the sample garbage image is used to perform model training on the garbage recognition and classification model to be trained, and a trained garbage recognition and classification model is determined.

7. A garbage classification system based on bidirectional long short-term memory network and convolutional neural network, characterized in that: include: A preprocessing module is used to determine a trained garbage recognition and classification model based on the training garbage images and the multi-task loss function, and input the garbage images to be classified into the trained garbage recognition and classification model; the trained garbage recognition and classification model includes an improved convolutional neural network, an improved bidirectional long short-term memory network, a feature fusion network, and a classification network; A convolution extraction module, configured to extract features of the garbage image to be classified by using an improved convolutional neural network to generate convolution output features; A time series extraction module is used to reorganize the convolution output features into time series features, perform multi-level memory enhancement on the time series features using an improved bidirectional long short-term memory network, and determine a memory combination output feature; A feature fusion module, configured to perform weighted fusion of the convolution output features and the memory combination output features based on a feature fusion network to construct a fusion feature; The classification module is used to use a classification network to perform feature dimensionality reduction on the fusion features, calculate the category probability distribution, and perform double threshold verification based on the category probability distribution to determine the garbage classification result.

8. A computer device, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor performs the steps of the garbage classification method based on a bidirectional long short-term memory network and a convolutional neural network as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by the processor, the steps of the garbage classification method based on the bidirectional long short-term memory network and the convolutional neural network are implemented as described in any one of claims 1 to 6.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by the processor, the steps of the garbage classification method based on the bidirectional long short-term memory network and the convolutional neural network are implemented as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Heat treatment cross shaft sleeve quality inspection method based on deep learning

    CN121937464A