Garbage classification detection method and system based on closed-loop adjustment and mixed attention mechanism
By using closed-loop adjustment and mixed attention mechanism methods in garbage classification detection, a dual-path backbone network and a cross-scale feature fusion network are built, which solves the problem of insufficient accuracy and efficiency of traditional methods when processing complex garbage images, and achieves high-precision and robust garbage classification detection.
Patent Information
- Application Number
- CN202510257017.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-03
AI Technical Summary
The existing garbage classification detection methods are insufficient in accuracy and efficiency, especially when dealing with complex and changeable garbage images, traditional machine learning and convolutional neural network methods are difficult to achieve ideal detection results, and a single backbone network cannot efficiently extract large-scale features and detailed features at the same time, which can easily lead to gradient disappearance or gradient explosion, and model training is unstable.
The garbage classification detection method based on closed-loop adjustment and mixed attention mechanism is adopted. By building a dual-path backbone network and a cross-scale feature fusion network, the feature extraction process is dynamically adjusted, and the most useful feature channels are adaptively selected through the channel and spatial attention mechanism to achieve fine-grained feature selection and focus.
It improves the accuracy and robustness of garbage classification detection, avoids error amplification or feature loss of single-scale features, enhances the model's anti-interference ability of noise, occlusion and complex backgrounds, and realizes accurate identification of recyclable garbage, kitchen waste, hazardous garbage and other garbage.
Smart Images

Figure CN120088573A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of object detection in the direction of computer vision, and particularly relates to a garbage classification detection method and system based on closed-loop calibration and hybrid attention mechanism. Background Art
[0002] With the acceleration of the urbanization process and the enhancement of environmental protection awareness, garbage classification has become one of the important tasks of urban management. However, the existing garbage classification detection methods still have deficiencies in accuracy and efficiency. Especially when dealing with complex and changeable garbage images, traditional machine learning and convolutional neural network methods are difficult to achieve ideal detection effects. During the object detection process of these methods, the error propagation between the feature fusion network and the detection head often leads to feature loss or error amplification. At the same time, a single backbone network cannot efficiently extract large-scale features and detailed features simultaneously. Especially in deep networks, the uneven error propagation is prone to cause gradient disappearance or gradient explosion, resulting in unstable model training or even falling into local optima. The object detection model is prone to performance degradation when facing noise, occlusion, or complex backgrounds. A single feature fusion strategy is difficult to effectively integrate features of different scales and cannot control the accumulation of error propagation, which limits the overall performance of the model. In addition, the existing attention mechanisms are not yet mature in garbage classification detection, and there are problems of inaccurate positioning and insufficient feature utilization. Especially when dealing with complex backgrounds and multi-object images, the existing attention mechanisms are often difficult to effectively capture key features, resulting in poor detection performance. Summary of the Invention
[0003] The purpose of the present invention is to provide a garbage classification detection method and system based on closed-loop calibration and hybrid attention mechanism.
[0004] In the first aspect, the present invention provides a garbage classification detection method based on closed-loop calibration and hybrid attention mechanism, which includes the following steps:
[0005] Step 1: Obtain image data containing garbage and construct a data set;
[0006] Step 2: Construct a garbage classification detection model; the garbage classification detection model includes a main backbone network, an auxiliary backbone network, a cross-scale feature fusion network, and a detection head; the main backbone network and the auxiliary backbone network have the same structure, and both include multiple layers of feature extraction layers connected in sequence; the cross-scale feature fusion network includes a first branch and a second branch; the first branch includes multiple layers of feature fusion layers connected in sequence, and an upsampling layer is connected in series between two adjacent feature fusion layers; the second branch includes multiple layers of feature integration layers connected in sequence, and a downsampling layer is connected in series between two adjacent feature integration layers;
[0007] Each feature fusion layer of the cross-scale feature fusion network corresponds to a feature extraction layer of the main backbone network and a feature extraction layer of the auxiliary backbone network; feature fusion layers except the last one correspond to a feature integration layer;
[0008] The working process of the garbage classification detection model is divided into a first stage and a second stage; in the first stage, the measured image is input into the main backbone network to obtain a multi-scale feature extraction map; the multi-scale feature extraction map is input into the cross-scale feature fusion network to obtain a multi-scale feature map; in the second stage, the multi-scale feature map obtained in the first stage is sequentially subjected to upsampling and convolution processing and then jointly input into the auxiliary backbone network with the measured image; the feature map output by the auxiliary backbone network is sequentially subjected to upsampling and convolution processing and then jointly input into the main backbone network with the measured image; the output of the main backbone network is sequentially processed by the cross-scale feature fusion network and the detection head to obtain the garbage classification result;
[0009] Step 3: Use the dataset to train the garbage classification detection model;
[0010] Step 4: Use the trained model to classify the garbage in the measured image.
[0011] Preferably, the number of the feature fusion layers is four; the input of the remaining feature fusion layers except the first feature fusion layer is the splicing result of the output feature of the previous feature fusion layer and the output feature of the corresponding layer feature extraction layer; the input feature of the first feature fusion layer is the output feature of the corresponding layer feature extraction layer;
[0012] The input feature of the first feature integration layer is the splicing result of the feature after downsampling the output feature of the last feature fusion layer and the output features of the corresponding layer feature extraction layer and feature fusion layer; the input feature of the second feature integration layer is the splicing result of the output feature of the previous feature integration layer and the output features of the corresponding layer feature extraction layer and feature fusion layer; the input feature of the third feature integration layer is the splicing result of the output feature of the previous feature integration layer and the output feature of the corresponding layer feature fusion layer.
[0013] Preferably, the feature extraction layer, the feature fusion layer and the feature integration layer all include a feature enhancement attention module; the feature enhancement attention module includes a first branch and a second branch; the first branch includes a feature enhancement convolutional block, a bottleneck feature optimization module and a convolutional block connected in series; the second branch uses a convolutional block; the feature maps input into the feature enhancement attention module are processed respectively through the first branch and the second branch, and the processing results are spliced and then compressed through a convolutional block to obtain the enhanced image feature output by the feature enhancement attention module.
[0014] Preferably, the bottleneck feature optimization module includes a bottleneck layer and a channel-spatial comprehensive attention mechanism module; the bottleneck layer includes two feature enhancement convolutional blocks connected in series;
[0015] The channel-spatial comprehensive attention mechanism module includes a channel attention mechanism module and a spatial attention mechanism module;
[0016] In the channel attention mechanism module, for the input feature map after adjusting the number of channels by convolution, global average pooling and global max pooling are respectively performed to obtain two pooled outputs A avg and A max ; The outputs A avg and A max are respectively passed through a multi-layer perceptron and then concatenated to obtain a fusion result; according to the weights corresponding to the fusion result and the input feature map, the output result of the channel attention mechanism module is obtained;
[0017] In the spatial attention mechanism module, for the input feature map, global pooling in the horizontal and vertical directions is respectively performed to obtain the horizontal spatial weight A H and the vertical spatial weight A W ; The horizontal spatial weight A H and the vertical spatial weight A W are concatenated and then processed successively through a convolutional block, a batch normalization layer, and an H_swish activation function, and the processed result is divided into two parts according to the height and width to obtain the feature A H1 corresponding to the horizontal direction and the feature A W1 corresponding to the vertical direction; according to the weights corresponding to the feature A H1 , the feature A W1 and the input feature map, the output result of the spatial attention mechanism module is jointly obtained.
[0018] Preferably, the number of the feature fusion layers is four; the first layer of the feature fusion layer uses a feature enhancement convolutional block; the second and third layers of the feature fusion layers both include a feature enhancement attention module and a feature enhancement convolutional block connected in series; the fourth layer of the feature fusion layer uses a feature enhancement attention module; the three layers of the feature integration layer all use a feature enhancement attention module.
[0019] Preferably, the number of the feature extraction layers is four; the first layer of the feature extraction layer of the main backbone network includes two feature enhancement convolutional blocks and a feature enhancement attention module connected in series; the second and third layers of the feature extraction layers both include a feature enhancement convolutional block and a feature enhancement attention module connected in series; the last layer of the feature extraction layer includes a feature enhancement convolutional block and a multi-scale feature enhancement block connected in series.
[0020] Preferably, the multi-scale feature enhancement block includes two convolutional blocks and three max pooling layers; the feature map input to the multi-scale feature enhancement block is sequentially processed by one convolutional block and three max pooling layers, the features output by the convolutional block and the three max pooling layers are concatenated in the channel dimension, and the concatenated features are processed by the convolutional block to obtain the output of the multi-scale feature enhancement block.
[0021] Preferably, the feature enhancement convolutional block includes a two-dimensional convolutional block, a batch normalization layer, and an activation function connected in series.
[0022] Preferably, in the auxiliary backbone network of the second stage, except for the first layer feature extraction layer, the input feature of each remaining layer feature extraction layer is the sum of the output feature of the previous layer feature extraction layer and the output feature of the corresponding layer feature integration layer in the cross-scale feature fusion network; the input of the first layer feature extraction layer is the measured image.
[0023] In the main backbone network of the second stage, except for the first layer feature extraction layer, the input feature of each remaining layer feature extraction layer is the sum of the output feature of the previous layer feature extraction layer and the output feature of the corresponding layer feature extraction layer in the auxiliary backbone network; the input of the first layer feature extraction layer is the measured image.
[0024] An upsampling layer and a convolutional block are connected in series between the output of the cross-scale feature fusion network and the input of the auxiliary backbone network, and between the output of the auxiliary backbone network and the input of the main backbone network.
[0025] In a second aspect, the present invention provides a garbage classification detection system based on closed-loop calibration and hybrid attention mechanism, which includes an image acquisition module, an image preprocessing module, and a garbage classification detection module; the garbage classification detection module includes a main backbone network, an auxiliary backbone network, a cross-scale feature fusion network, and a detection head; the working process of the garbage classification detection module is divided into a first stage and a second stage; in the first stage, the main backbone network and the cross-scale feature fusion network are sequentially used to process the measured image to generate a feedback value; in the second stage, the feedback value and the measured image are combined, and are sequentially processed by the auxiliary backbone network, the main backbone network, the cross-scale feature fusion network, and the detection head to obtain the garbage classification detection result.
[0026] The beneficial effects of the present invention are:
[0027] 1. The present invention forms an independent error propagation path through the main backbone network and cross-scale feature fusion network in the first stage, alleviates the error propagation problem, dynamically adjusts the feature extraction process according to the feedback signal, and designs a dynamic weight allocation mechanism to ensure the optimization of error distribution between the dual-path backbone networks during the backpropagation process. By implementing self-check and feedback in the model, the accuracy and robustness of the model for garbage classification detection are improved. At the same time, the dual-path backbone network of the present invention focuses on feature extraction at different scales, avoids the error amplification or feature loss of a single scale, and improves the stability of model training.
[0028] 2. The present invention enables the model to adaptively select the feature channels most useful for the target detection task through the channel and spatial attention mechanisms that fuse global and local features. By combining the target coordinate information with the channel attention mechanism, the attention weights are dynamically adjusted, enabling the model to more precisely focus on the key regions of the detection target, suppressing the interference of irrelevant regions, and achieving fine-grained feature selection and focusing.
[0029] 3. The present invention realizes the fusion between low-level features (capturing detailed information) and high-level features (capturing semantic information) by fusing the output of the main backbone network in the feature integration layer, enabling the model to perform effective feature learning and classification at different spatial scales; ensuring that each scale feature can be fully utilized, the cross-scale attention feature fusion design effectively deals with noise, occlusion, and complex background interference, optimizes error propagation, and improves the anti-interference ability of the model. At the same time, the present invention can accurately identify recyclable waste, kitchen waste, hazardous waste, and other waste, reduce the error rate of manual sorting, better recycle resources through accurate classification, and reduce environmental pollution. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is the overall flowchart of the present invention.
[0031] Figure 2 It is the structural schematic diagram of the feature enhancement convolutional block in the present invention.
[0032] Figure 3 It is the structural schematic diagram of the feature enhancement attention module in the present invention.
[0033] Figure 4 It is the structural schematic diagram of the bottleneck feature optimization module in the present invention.
[0034] Figure 5 It is the structural schematic diagram of the channel-spatial comprehensive attention mechanism module in the present invention.
[0035] Figure 6 It is the structural schematic diagram of the multi-scale feature enhancement block in the present invention.
[0036] Figure 7It is a schematic structural diagram of the cross-scale feature fusion network in the present invention.
[0037] Figure 8 It is a schematic diagram of the data transmission path of the garbage classification detection model in the first stage in the present invention.
[0038] Figure 9 It is a schematic diagram of the data transmission path of the garbage classification detection model in the second stage in the present invention.
[0039] Figure 10 It is a schematic diagram of the garbage classification detection result of other garbage categories in the present invention.
[0040] Figure 11 It is a schematic diagram of the garbage classification detection result of kitchen waste categories in the present invention.
[0041] Figure 12 It is a schematic diagram of the garbage classification detection result of recyclable garbage categories in the present invention.
[0042] Figure 13 It is a schematic diagram of the result of garbage classification detection of harmful garbage categories in the present invention. Detailed implementation manners
[0043] The present invention will be further described below with reference to the accompanying drawings.
[0044] As Figure 1 shown, a garbage classification detection method based on closed-loop tuning and hybrid attention mechanism, and the garbage classification detection system used therein includes an image acquisition module, an image preprocessing module, and a garbage classification detection module. The image acquisition module is used to acquire the measured image; the image preprocessing module is used to process the measured image; the garbage classification detection module includes a main backbone network, an auxiliary backbone network, a cross-scale feature fusion network, and a detection head; the working process of the garbage classification detection module is divided into a first stage and a second stage; in the first stage, the main backbone network and the cross-scale feature fusion network are sequentially used to process the measured image to generate a feedback value; in the second stage, the feedback value and the measured image are combined, and are sequentially processed through the auxiliary backbone network, the main backbone network, the cross-scale feature fusion network, and the detection head to obtain the garbage classification detection result.
[0045] The garbage classification detection method includes the following steps:
[0046] Step 1: Establish a dataset. Construct a garbage classification attribute dataset through two methods: using an open-source garbage classification attribute dataset and a self-made garbage classification attribute dataset. Use the image annotation technology Labelimg to annotate the images, and label different garbage in the pictures into four categories: recyclable waste, harmful waste, other waste, and food waste. Divide the dataset into a training set, a validation set, and a test set according to the ratio of 7:2:1.
[0047] Step 2: Preprocess the images in the dataset through image normalization, size adjustment, and data augmentation to ensure that the model can obtain input data for efficiently processing multi-scale targets and complex scenes, and improve the robustness and generalization ability of the model. The preprocessed image tensor is X, where 1 represents the batch size, that is, one picture is input into the backbone network each time; 3 represents the number of channels of the image, corresponding to the RGB image; 640×640 represents the spatial resolution of the image (both the height H and the width W are 640 pixels).
[0048] Step 3: Construct a garbage classification detection model; the garbage classification detection model includes a dual-path backbone network, a cross-scale feature fusion network, and a detection head.
[0049] 3-1. Dual-path backbone network
[0050] The dual-path backbone network includes a main backbone network and an auxiliary backbone network. The structures of the auxiliary backbone network and the main backbone network are the same, and both include four sequentially connected feature extraction layers; the feature extraction layers in the main backbone network and the auxiliary backbone network correspond one by one; the first feature extraction layer includes two concatenated feature enhancement convolutional blocks (CBL) and a feature enhancement attention module (C3-CSA); the second and third feature extraction layers both include concatenated feature enhancement convolutional blocks and feature enhancement attention modules; the last feature extraction layer includes concatenated feature enhancement convolutional blocks and a multi-scale feature enhancement block (SPPF).
[0051] As Figure 2 shown, the feature enhancement convolutional block includes a concatenated two-dimensional convolutional block (Conv2d), a batch normalization layer (BatchNorm2d), and a SiLu activation function; the measured image input into the feature enhancement convolutional block is first subjected to a convolution operation by the two-dimensional convolutional block, expanding the number of channels from 3 to 32, and the spatial resolution is downsampled to 1 / 2 of the original due to the convolution stride of 2, changing from 640×640 to 320×320. After passing through the batch normalization layer and the SiLu activation function, the output size remains unchanged, obtaining an intermediate quantity
[0052] As shown Figure 3 in the figure, the feature enhancement attention module includes a first branch and a second branch; the first branch includes a feature enhancement convolutional block, a bottleneck feature optimization module (Bottleneck-CSA), and a convolutional block connected in series; the second branch uses a convolutional block; the feature map input to the feature enhancement attention module is processed through the first branch and the second branch respectively, and the processing results are concatenated and then compressed through a convolutional block to adjust the number of output channels, obtaining the enhanced image features output by the feature enhancement attention module.
[0053] As shown Figure 4 in the figure, the Bottleneck-CSA module includes a bottleneck layer (Bottleneck) and a channel spatial comprehensive attention mechanism module (CSA). The bottleneck layer includes two CBL modules connected in series; the input and output of the bottleneck layer are added element by element, and the channel spatial comprehensive attention mechanism module is used to process the added result to obtain the feature map output by the Bottleneck-CSA module; the bottleneck layer can alleviate the problem of gradient disappearance or gradient explosion in the deep network, and at the same time improve the feature extraction ability of the model.
[0054] As shown Figure 5 in the figure, the channel spatial comprehensive attention mechanism module includes a channel attention mechanism module and a spatial attention mechanism module. The input feature map is processed through the channel attention mechanism module and the spatial attention mechanism module respectively; where, C is the number of input channels; H and W are the height and width of the feature map respectively. 1 respectively.
[0055] The channel attention mechanism module first adjusts the number of channels through a 1×1 convolutional block to obtain the intermediate feature A conv1 = Conv 1×1 initial (A), The intermediate feature A conv1 is respectively subjected to global average pooling and global maximum pooling processing to obtain the outputs A avg and A max after the two poolings:
[0056]
[0057] The outputs A avg and A max are respectively processed by a multi-layer perceptron MLP. The multi-layer perceptron MLP compresses the number of channels to C2 / r through a fully connected layer (FC), where C 1is the number of input channels; r is the compression ratio (hyperparameter r = 16 during design). After passing through the Relu activation function to enhance the non-linear expression ability, the number of channels is restored to C2 through a fully connected layer FC, and then the two results are concatenated to obtain the fusion result A fusion :
[0058] A fusion = Concat(A avg,Fc,Relu,Fc , A max,Fc,Relu,Fc )(3)
[0059] After passing the fusion result A fusion sequentially through a fully connected layer, a Sigmoid activation function, and a convolutional block Conv2d, the channel weight W C is obtained, and the channel weight W C is multiplied by the input feature map A to obtain the output result O of the channel attention mechanism module c = A × Wc; among them, the fully connected layer is used to generate a fusion result aligned with the number of channels of the input feature map
[0060] The spatial attention mechanism module respectively performs feature aggregation on the horizontal and vertical directions of the input feature map A to generate the horizontal spatial weight and the vertical spatial weight The representation methods of global pooling in the horizontal and vertical directions are as follows
[0061]
[0062] The global information of the input feature map A in the horizontal and vertical directions is respectively extracted through global pooling in the horizontal and vertical directions. After generating the weights, the horizontal spatial weight A H and the vertical spatial weight A W are concatenated along the spatial dimension to form a feature vector containing global information in both directions, and the concatenated feature vector is sequentially transformed through a 1×1 convolutional block, a batch normalization layer BN, and an H_swish activation function to further compress and fuse the channel dimension information. Then, the processed feature is divided into two parts according to height and width to obtain the feature A H1 corresponding to the horizontal direction and the feature A W1 corresponding to the vertical direction. The feature A H1 and the feature A W1 are respectively processed through a 1×1 convolutional block and a Sigmoid activation function to generate the horizontal spatial weight A H2 and the vertical spatial weight A W2 , respectively describing the importance in their respective directions, and their formula expressions are as follows
[0063]
[0064] Among them, σ(·) is the Sigmoid activation function.
[0065] Multiply the horizontal spatial weight A H2 , the vertical spatial weight A W2 and the input feature map A to enhance the spatial response ability of the input features, thereby obtaining the output result of the spatial attention mechanism module.
[0066] Design dynamic weight adjustment, and weight and fuse the outputs of the two attention mechanisms of channels and spaces through dynamically learned weights to obtain the feature map A of the output of the channel-spatial comprehensive attention mechanism module out :
[0067] A out = σ(W 1 )·(A·W C ) + σ(W 2 )·(A·A H1 ·A W1 ) (8)
[0068] Among them, W 1 and W 2 are weight parameters, and their initial values are both 0.5. The weight parameters W 1 and W 2 are restricted to the interval [0, 1] through the Sigmoid activation function.
[0069] As Figure 6 shown, the multi-scale feature enhancement block enhances the multi-scale expression ability of features by combining convolution and multi-level pooling operations; the feature map input to the multi-scale feature enhancement block passes through a convolution block and three max poolings of K×K (K is the convolution kernel size) in sequence, splice the features output by the convolution block and the three max poolings in the channel dimension, and use the convolution block to process the spliced features to obtain the output of the multi-scale feature enhancement block.
[0070] 3-2. Cross-scale Feature Fusion Network
[0071] As Figure 7As shown in the figure, the cross-scale feature fusion network includes a first branch and a second branch; the first branch includes four layers of feature fusion layers connected in sequence, and an upsampling layer is connected in series between adjacent feature fusion layers; the feature fusion layer connected to the last layer of feature extraction layer is used as the first layer of feature fusion layer; the first layer of feature fusion layer uses a feature enhancement convolutional block; the second and third layers of feature fusion layers both include a feature enhancement attention module and a feature enhancement convolutional block connected in series; the fourth layer of feature fusion layer uses a feature enhancement attention module; the second branch includes three layers of feature integration layers connected in sequence; a downsampling layer is connected in series between adjacent feature integration layers; all three feature integration layers use feature enhancement attention modules; the output of the three feature integration layers is used as the output of the cross-scale feature fusion network. The image is fused through the feature integration layer of the second branch to enhance the semantic fusion between the context information of the image features.
[0072] In the first branch, the four layers of feature fusion layers correspond one-to-one with the four layers of feature extraction layers in the main backbone network; except for the first layer of feature fusion layer, the input of each layer of feature fusion layer is the concatenation result of the output feature of the previous feature fusion layer and the output feature of the corresponding layer of feature extraction layer; the input feature of the first layer of feature fusion layer is the output feature of the corresponding layer of feature extraction layer.
[0073] In the second branch, the three layers of feature integration layers correspond one-to-one with the first three layers of feature fusion layers; the input feature of the first layer of feature integration layer is the concatenation result of the feature after downsampling the output feature of the last layer of feature fusion layer, and the output features of the corresponding layer of feature extraction layer and feature fusion layer; the input feature of the second layer of feature integration layer is the concatenation result of the output feature of the previous layer of feature integration layer and the output features of the corresponding layer of feature extraction layer and feature fusion layer; the input feature of the third layer of feature integration layer is the concatenation result of the output feature of the previous layer of feature integration layer and the output feature of the corresponding layer of feature fusion layer. The output of the three layers of feature integration layers is used as the output of the cross-scale feature fusion network, and the feature maps output by the cross-scale feature fusion network are corresponding to the feature values of three different scales, large, medium, and small, in the detection model respectively.
[0074] The working process of the garbage classification detection model is divided into a first stage and a second stage; as Figure 8 shown, in the first stage of the garbage classification detection model, the measured image is input into the main backbone network to obtain the preliminary feature extraction values large-scale feature extraction map medium-scale feature extraction map and small-scale feature extraction map and input into the cross-scale feature fusion network to obtain feature maps Y1, Y2, and Y3.
[0075] As Figure 9As shown, in the second stage of the garbage classification detection model, the feature maps output by the cross-scale feature fusion network in the first stage are successively subjected to upsampling and convolution processing and then jointly input into the auxiliary backbone network together with the image to be detected; and the feature maps output by the auxiliary backbone network are successively subjected to upsampling and convolution processing and then jointly input into the main backbone network together with the image to be detected; the output of the main backbone network is successively processed by the cross-scale feature fusion network and the detection head to obtain the output result of the garbage classification detection model; the specific process is as follows:
[0076] (1) In the auxiliary backbone network, except for the first layer of feature extraction layer, the input features of the remaining layers of feature extraction layers are the results of adding the features output by the previous layer of feature extraction layer and the features output by the corresponding layer of feature integration layer in the cross-scale feature fusion network; the input of the first layer of feature extraction layer is the image to be detected; the outputs of the latter three layers of feature extraction layers are used as the output of the auxiliary backbone network, and the specific process is as follows:
[0077] After successively performing upsampling and convolution processing on the feature maps Y1, Y2, and Y3 output by the cross-scale feature fusion network, feedback values are obtained The upsampling layer uses the nearest neighbor interpolation method, and the number of channels is adjusted through the upsampling layer and a 1×1 convolution block. In the nearest neighbor interpolation mode, each pixel point is filled by repeating the values of neighboring pixels, so as to double the resolution of the input feature map. Suppose the input feature map is X in , and its size is H in ×W in . After upsampling and a 1×1 Conv, its size becomes:
[0078] H out =H in ×s, W out =W in ×s
[0079] where s is the upsampling ratio, and s is set to 2; H out and W out are the height and width of the output feature map respectively.
[0080] The feedback value is added to the output feature of the previous layer of feature extraction layer in the auxiliary backbone network to obtain the calibrated values P1 1 , P2 1 , P3 1 . The auxiliary backbone network is used to process the input calibrated values P1 1 , P2 1 , P3 1 to obtain the output value of the auxiliary backbone network
[0081] (2) In the main backbone network, except for the first layer feature extraction layer, the input features of the remaining feature extraction layers are the sum of the output features of the previous layer feature extraction layer and the output features of the corresponding layer feature extraction layer in the auxiliary backbone network; the input of the first layer feature extraction layer is the measured image; the feature map output by the feature extraction layer is used as the feature map output by the main backbone network; an upsampling layer and a convolutional block are connected in series between the output of the feature extraction layer in the auxiliary backbone network and the input of the feature extraction layer in the main backbone network. The specific process is as follows:
[0082] The output values P2, P3, and P4 are respectively upsampled and processed by a 1×1 Conv to obtain the auxiliary backbone network feedback values The measured image and the feedback values M1, M2, and M3 are respectively processed by the feature extraction layer to obtain the preliminary feature extraction values Large-scale feature extraction map Medium-scale feature extraction map Small-scale feature extraction map
[0083] (3) Use the feature fusion layer in the cross-scale feature fusion network to process the preliminary feature extraction map Q1, the large-scale feature extraction map Q2, the medium-scale feature extraction map Q3, and the small-scale feature extraction map Q4 to obtain the feature map B1, the feature map B2, the feature map B3, and the feature map B4. The specific process is as follows:
[0084] a. Use the convolutional feature enhancement block to process the small-scale feature extraction map Q4 to change its number of channels; the stride of the two-dimensional convolutional block in the convolutional feature enhancement block is 1, the convolutional kernel size is 3×3, the input number of channels is 512, and the output number of channels is 256, which is used for channel compression to compress the number of channels from 512 to 256 to optimize the features and obtain the feature map
[0085] b. Perform upsampling processing on the feature map B1 to obtain the feature map And splice and fuse it with the medium-scale feature extraction map Q3 to obtain the feature map The feature map B1 2 Obtain the feature map through a feature enhancement attention block The feature map B1 3 After being processed by a convolutional feature enhancement block with a stride of 1, the number of channels is halved to 128, reducing the amount of calculation while cooperating with subsequent fusion and upsampling operations to obtain the feature map
[0086] c. Perform upsampling processing on the feature map B2 to obtain the feature map And splice it with the large-scale feature extraction map Q2 to obtain the feature map The feature map B2 2 is processed through a feature enhancement attention block and a convolutional feature enhancement block with a stride of 1 to obtain a feature map
[0087] d. Upsample the feature map B3 to obtain a feature map and concatenate it with the preliminary feature extraction map Q1 to obtain a feature map The feature map B3 2 is processed through a feature enhancement attention block to obtain a feature map
[0088] (4) Use the feature integration layer in the cross-scale feature fusion network to process the feature maps B1, B2, B3, and B4 to obtain the large target feature value Y1, the medium target feature value Y2, and the small target feature value Y3. The specific process is as follows:
[0089] a. Use a two-dimensional convolutional block with a stride of 2 to downsample the feature map B4 to obtain an image feature The image feature B4 1 , the feature map B3, and the large-scale feature extraction map Q2 are cross-scale fused through Concat to obtain a feature with original semantics and context-aware fusion The feature B4 2 is processed through a feature enhancement attention block to obtain the large target feature value
[0090] b. Use a two-dimensional convolutional block with a stride of 2 to downsample the large target feature value Y1 to obtain a feature The feature Y1 1 , the feature map B2, and the medium-scale feature extraction map Q3 are cross-scale fused through Concat to obtain a feature with original semantics and context-aware fusion The feature Y1 2 is processed through a feature enhancement attention block to obtain the medium target feature value
[0091] c. Use a two-dimensional convolutional block with a stride of 2 to downsample the medium target feature value Y2 to obtain a feature The feature Y2 1 and the feature map B1 are cross-scale fused through Concat to obtain a feature with original semantics and context-aware fusion The feature Y2 2 is processed through a feature enhancement attention block to obtain the small target feature value
[0092] (5) Use the detection head Head to predict the large target feature value Y1, the medium target feature value Y2, and the small target feature value Y3, and obtain the final detection value Yout. Among them, the tensor shapes of Y1, Y2, and Y3 are B is the batch size set to 32; C is the number of channels; H and W represent the height and width of the feature map.
[0093] In some embodiments, the specific process of the detection head processing the feature values output by the cross-scale feature fusion network is as follows:
[0094] Adjust the feature map Y through a convolutional block Conv2d with a kernel size of 1×1 and filters of no×na i to the target tensor where no is the number of output parameters for each anchor point; no = nc + 5; nc is the number of classes + 4 bounding box parameters + 1 confidence; na is the number of anchor points (3 anchor points are designed for each layer). Reshape the target tensor Y i1 into a more intuitive anchor prediction form Generate the grid offset G i and the anchor box size A i :
[0095] G i = MeshGrid(H, W)
[0096] A i = Anchors × Stride i
[0097] where MeshGrid represents a function for generating a two-dimensional or three-dimensional coordinate grid; Anchors is the number of anchor points, and Strde i is the stride.
[0098] Decode the center point coordinates (x, y), the width and height (h, w), the confidence, and the class for Y i2 , and the specific process is as follows:
[0099] a. The calculation formula for center point decoding is expressed as:
[0100] (x, y) = (σ(Y i2 , xy) · 2 + G i ) · Stride i
[0101] where represents the predicted center point offset; σ(·) is the Sigmoid function.
[0102] b. The decoding calculation formula for the image width and height is expressed as:
[0103] (w, h) = (σ(Y i2 , hw) · 2) 2 ·A i
[0104] where represents the predicted scaling factor for width and height.
[0105] c. The expression for the confidence Confidence is:
[0106]
[0107] where P(Object) represents the probability that the box contains the target object; represents the intersection over union (IoU) between the predicted box and the ground truth box. The higher the value of the IoU, the better the overlap between the predicted box and the ground truth box.
[0108] d. Calculate the class probability (classprob)
[0109] Using the nc class scores output from the anchor points as the raw scores, the Softmax function is used to convert the raw scores into a probability distribution. The role of Softmax is to convert the scores of each class into a normalized probability, such that the sum of the probabilities of each box for all classes is 1. The Softmax probability P(class i ) is expressed as:
[0110]
[0111] where score i is the raw score corresponding to the class; exp is the exponential function; is the sum of the exponentialized scores of all classes.
[0112] For each bounding box, the class with the maximum Softmax probability is selected as the final predicted class, and each bounding box is assigned to the class with the highest probability; the final class probability is the probability value of each bounding box for each class. These values are calculated by the Softmax function and represent the probability that each box belongs to each class.
[0113] Concatenate the confidence Confidence, class probability ClassProb, and the decoded bounding box parameters into the final output Y out :
[0114] Y out = Concat((x, y, w, h, Confidence, ClassProb))
[0115] In other embodiments, other existing methods can also be used to obtain the final output Y out .
[0116] Step 4: Use the training set and validation set divided in Step 1 to train the garbage classification detection model, and use the test set to test the trained neural network. After 300 rounds of training, the detection effects are shown in Table 1. Among them, for the designed detection method, the average detection accuracy mAP50 for other waste in the real world reaches 0.912, the average detection accuracy mAP50 for food waste reaches 0.863, the average detection accuracy mAP50 for recyclable waste reaches 0.844, and the average detection accuracy mAP50 for hazardous waste reaches 0.941.
[0117] Table 1 Detection effects after training
[0118] C1ass Images Instances P R mAP@50 mAP@50-95 all 548 757 0.841 0.828 0.89 0.607 RecyclableWaste 548 314 0.811 0.822 0.844 0.581 HarmfulWaste 548 152 0.873 0.908 0.941 0.715 FoodWaste 548 146 0.835 0.726 0.863 0.497 OtherWaste 548 145 0.845 0.855 0.912 0.635
[0119] P (Precision) in Table 1 is the precision, which represents the proportion of true positive samples among the prediction boxes predicted by the model as positive samples. The formula is as follows:
[0120]
[0121] where TP is the number of true positives; FP is the number of false positives.
[0122] R (Recall) is the recall rate, which represents the proportion of positive samples that the model can correctly detect. The formula is as follows:
[0123]
[0124] where FN is the number of false negatives.
[0125] MAP@50 (Mean Average Precision at IoU=0.5): The average precision (AP) calculated when the intersection over union (IoU) threshold is 0.5. MAP (Mean Average Precision) is an evaluation metric that combines precision and recall and is used to measure the overall detection performance of the model across all classes.
[0126] MAP@50-95 (Mean Average Precision from IoU=0.5 to IoU=0.95): The average precision within different IoU threshold ranges (from 0.5 to 0.95). This metric considers multiple IoU thresholds to more comprehensively evaluate the performance of the model.
[0127] Step 5: Use the present invention to perform garbage classification detection on the measured image in the four attributes of garbage classification. The detection effect is as follows Figure 10 , Figure 11 , Figure 12 , Figure 13 . Among them, the designed detection method has a detection confidence of 0.9 ( Figure 10 ) for other garbage categories, a detection confidence of 0.8 ([[]] Figure 11 ) for kitchen waste categories, a detection confidence of 0.9 ([[]] Figure 12 ) for recyclable garbage categories, and a detection confidence of 0.9 ([[]] Figure 13 ) for hazardous waste categories. The designed method has a very high confidence level in the garbage detection and classification results.
[0128] Step 6: To prove the advantages of this method, a comparative experiment between the present invention and the YoloV5 model was carried out, and an ablation experiment was also carried out; The YOLOv5 model mainly uses a Backbone for feature extraction, then performs feature fusion through the Neck network, and the three different-scale feature maps obtained are sent to the detection head to obtain the final detection result. The results of the model ablation experiment are shown in Table 2.
[0129] Table 2 Model ablation experiment
[0130] Models P R MAP@50 MAP@50-90 obj_loss cls_loss YOLOv5 0.832 0.791 0.855 0.570 0.0075 0.0084 YOLOv5+CSA 0.849 0.798 0.878 0.570 0.0050 0.0027 YOLOv5+Closed-loop Tuning Mechanism 0.865 0.814 0.867 0.581 0.0076 0.0050 YOLOv5+Cross-scale Feature Fusion Network 0.888 0.771 0.862 0.573 0.0078 0.0089 The present invention 0.841 0.828 0.89 0.607 0.0048 0.0027
[0131] The obj_loss (Objectness Loss) in Table 2 represents the object loss, which measures the loss of the model when predicting whether each box contains an object and is used to calculate the confidence loss of the box; cls_loss (Classification Loss): classification loss, which measures the loss of the model when predicting the category and is calculated through the cross-entropy loss, and is used to train the model to accurately identify the category of each object. It can be seen from Table 2 that the present invention has achieved better results in different indicators.
Claims
1. A garbage classification detection method based on closed-loop calibration and hybrid attention mechanism, characterized by: The following steps are involved: Step 1: Obtain image data containing garbage and build a data set; Step 2, constructing a garbage classification detection model; the garbage classification detection model includes a main backbone network, an auxiliary backbone network, a cross-scale feature fusion network and a detection head; the main backbone network and the auxiliary backbone network have the same structure, both including multiple layers of feature extraction layers connected in sequence; the cross-scale feature fusion network includes a first branch and a second branch; the first branch includes multiple layers of feature fusion layers connected in sequence, and an upsampling layer is connected in series between two adjacent feature fusion layers; the second branch includes multiple layers of feature integration layers connected in sequence, and a downsampling layer is connected in series between two adjacent feature integration layers; Each feature fusion layer of the cross-scale feature fusion network corresponds to a feature extraction layer of the main backbone network and a feature extraction layer of the auxiliary backbone network; each feature fusion layer except the last layer corresponds to a feature integration layer; The working process of the garbage classification detection model is divided into the first stage and the second stage. In the first stage, the image to be tested is input into the main backbone network to obtain a multi-scale feature extraction map; The multi-scale feature extraction map is input into the cross-scale feature fusion network to obtain a multi-scale feature map; in the second stage, the multi-scale feature map obtained in the first stage is successively subjected to upsampling and convolution processing, and then input into the auxiliary backbone network together with the image under test; the feature map output by the auxiliary backbone network is successively subjected to upsampling and convolution processing, and then input into the main backbone network together with the image under test; the output of the main backbone network is successively processed by the cross-scale feature fusion network and the detection head to obtain the garbage classification result; Step 3: Use the data set to train the garbage classification detection model; Step 4: Use the trained model to classify the garbage in the tested image.
2. According to claim 1, a garbage classification detection method based on closed-loop calibration and hybrid attention mechanism is characterized by: The number of the feature fusion layers is four; except for the first feature fusion layer, the input of the remaining feature fusion layers is the concatenation result of the output features of the previous feature fusion layer and the output features of the corresponding feature extraction layer; The input features of the first feature fusion layer are the output features of the corresponding feature extraction layer; The input features of the first feature integration layer are the output features of the last feature fusion layer after downsampling, and the concatenation of the output features of the corresponding feature extraction layer and feature fusion layer; The input features of the second feature integration layer are the output features of the previous feature integration layer and the concatenation of the output features of the corresponding feature extraction layer and feature fusion layer; The input features of the third feature integration layer are the concatenation of the output features of the previous feature integration layer and the output features of the corresponding feature fusion layer.
3. The garbage classification detection method based on closed-loop calibration and hybrid attention mechanism according to claim 1 is characterized in that: The feature extraction layer, feature fusion layer and feature integration layer all include a feature enhancement attention module; the feature enhancement attention module includes a first branch and a second branch; the first branch includes a feature enhancement convolution block, a bottleneck feature optimization module and a convolution block connected in series; the second branch uses a convolution block; the feature map of the input feature enhancement attention module is processed by the first branch and the second branch respectively, and the processing results are spliced and compressed through the convolution block to obtain the enhanced image features output by the feature enhancement attention module.
4. According to claim 3, a garbage classification detection method based on closed-loop calibration and hybrid attention mechanism is characterized in that: The bottleneck feature optimization module includes a bottleneck layer and a channel space comprehensive attention mechanism module; the bottleneck layer includes two layers of feature enhancement convolution blocks connected in series; The channel-space comprehensive attention mechanism module includes a channel attention mechanism module and a spatial attention mechanism module; In the channel attention mechanism module, the input feature map after the convolution adjusts the number of channels is subjected to global average pooling and global maximum pooling respectively to obtain the two pooled outputs A avg and A max ; will output A avg and A max After passing through the multi-layer perceptron, they are spliced to obtain the fusion result; the output result of the channel attention mechanism module is obtained according to the weight corresponding to the fusion result and the input feature map; In the spatial attention mechanism module, the input feature map is globally pooled in the horizontal and vertical directions to obtain the horizontal spatial weight A H and vertical spatial weight A W ; The horizontal spatial weight A H and vertical spatial weight A W After splicing, it is processed by convolution block, batch normalization layer and H_swish activation function in turn, and the processing result is divided into two parts according to height and width to obtain the corresponding horizontal feature A H1 and the feature A corresponding to the vertical direction W1 ; According to feature A H1 , Feature A W1 The corresponding weights and input feature maps are used together to obtain the output results of the spatial attention mechanism module.
5. According to claim 1, a garbage classification detection method based on closed-loop calibration and hybrid attention mechanism is characterized in that: The number of the feature fusion layers is four; The first feature fusion layer uses feature enhancement convolution blocks; The second and third feature fusion layers both include a series of feature enhancement attention modules and feature enhancement convolution blocks; the fourth feature fusion layer adopts a feature enhancement attention module; the three feature integration layers shown all adopt feature enhancement attention modules.
6. The garbage classification detection method based on closed-loop calibration and hybrid attention mechanism according to claim 1 is characterized by: The number of the feature extraction layers is four; the first feature extraction layer of the main backbone network includes two feature enhancement convolution blocks connected in series and one feature enhancement attention module; The second and third feature extraction layers both include a series of feature enhancement convolution blocks and feature enhancement attention modules; the last feature extraction layer includes a series of feature enhancement convolution blocks and multi-scale feature enhancement blocks.
7. The garbage classification detection method based on closed-loop calibration and hybrid attention mechanism according to claim 6 is characterized by: The multi-scale feature enhancement block includes two convolution blocks and three maximum pooling layers; the feature map input to the multi-scale feature enhancement block is processed by a convolution block and three maximum pooling layers in sequence, the features output by the convolution block and the three maximum pooling layers are spliced in the channel dimension, and the spliced features are processed using the convolution block to obtain the output of the multi-scale feature enhancement block.
8. The garbage classification detection method based on closed-loop calibration and hybrid attention mechanism according to claim 6 is characterized by: The feature enhancement convolution block includes a series of two-dimensional convolution blocks, a batch normalization layer and an activation function.
9. The garbage classification detection method based on closed-loop calibration and hybrid attention mechanism according to claim 1 is characterized in that: In the auxiliary backbone network of the second stage, except for the first feature extraction layer, the input features of the remaining feature extraction layers are the sum of the output features of the previous feature extraction layer and the output features of the corresponding feature integration layer in the cross-scale feature fusion network; The input of the first feature extraction layer is the image under test; In the main backbone network of the second stage, except for the first feature extraction layer, the input features of the remaining feature extraction layers are the sum of the output features of the previous feature extraction layer and the output features of the corresponding feature extraction layer in the auxiliary backbone network; The input of the first feature extraction layer is the image under test; An upsampling layer and a convolution block are connected in series between the output of the cross-scale feature fusion network and the input of the auxiliary backbone network, and between the output of the auxiliary backbone network and the input of the main backbone network.
10. A garbage classification detection system based on closed-loop calibration and hybrid attention mechanism, comprising an image acquisition module, an image preprocessing module and a garbage classification detection module; characterized in that: The garbage classification detection module includes a main backbone network, an auxiliary backbone network, a cross-scale feature fusion network and a detection head; the working process of the garbage classification detection module is divided into a first stage and a second stage; in the first stage, the main backbone network and the cross-scale feature fusion network are used in turn to process the tested image to generate a feedback value; in the second stage, the feedback value and the tested image are combined, and processed in turn through the auxiliary backbone network, the main backbone network, the cross-scale feature fusion network and the detection head to obtain the garbage classification detection result.
Citation Information
Cited By
Unmanned aerial vehicle water surface floating garbage detection method and system based on multi-scale dynamic feature fusion
CN120823532A
River channel garbage identification and positioning method, system and equipment based on deep learning
CN121708512A