AI sorting garbage treatment method and system
By adopting a boundary sense knowledge distinction model that integrates attention refinement and spatial interaction mechanism in the AI garbage sorting system, the problem of insufficient recognition capabilities in existing technology in complex scenarios is solved, and high-precision identification and precise positioning of highly similar material waste is achieved.
Patent Information
- Application Number
- CN202510407592.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-05-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing AI garbage sorting technology has weak recognition capabilities in complex scenarios, especially ineffective treatment of occlusion, overlap, and debris, and the recognition accuracy of highly similar materials is low.
An effective boundary sense knowledge distinction model integrating attention refinement and spatial interaction mechanisms is adopted, and garbage image recognition and positioning are performed through the combination of input end, backbone network, attention parallel fusion module, upsampling fusion layer and detection head.
The detection and identification accuracy of highly similar materials has been significantly improved, the model's ability to capture key features of garbage in complex scenarios has been enhanced, the precise identification and detection of garbage is achieved, and the precise positioning of garbage is ensured.
Smart Images

Figure CN119909945A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to an AI garbage sorting method and system. Background Art
[0002] In the existing technology, the traditional manual garbage sorting method has problems such as low efficiency, high cost, and high labor intensity. The AI sorting method based on computer vision technology can quickly and accurately identify and classify different types of garbage through deep learning algorithms, improve recycling rates, and reduce environmental pollution. This method combines computer vision, machine learning, and robotic automation technologies, and can use large-scale data to train target detection models, enabling them to automatically identify recyclable garbage, kitchen waste, hazardous waste, and other waste, thereby realizing an intelligent, unmanned garbage sorting and processing process.
[0003] Most existing sorting technologies use target detection algorithms to give the 2D image coordinates of the garbage to be detected, and then use multiple viewpoints through multi-view geometry methods to calculate the 3D coordinates of the garbage to be detected and control the robotic arm to grasp it. However, the target detection model of the existing technology has weak recognition capabilities for occlusion, overlap, contamination, etc., and is prone to false detection and missed detection in complex scenes. In addition, the similarity of the surface material of the garbage, such as the similar optical properties of plastic bags and paper, may also lead to errors in the calculation of 3D coordinates. Summary of the invention
[0004] In order to solve one of the above-mentioned problems in the prior art, the present invention provides an AI garbage sorting method and system.
[0005] To achieve the above objectives, the present invention provides an AI garbage sorting method, which comprises: Step S101, obtaining a garbage bag; Step S102, breaking the garbage bag to obtain the garbage to be detected; Step S103, collecting image information of the garbage to be detected; Step S104, inputting the image information into a perception recognition model for recognition, the perception recognition model comprising: an input end, a backbone network, an attention parallel fusion module, an upsampling fusion layer and a detection head, the recognition process comprising the following steps: Step S1041, receiving the image information through the input end; Step S1042, performing feature extraction through the backbone network, wherein the backbone network includes at least three feature extraction layers; Step S1043, sending the output features extracted by every three adjacent feature extraction layers to the same attention parallel fusion module for processing to obtain the attention layer output; Step S1044, sending the attention layer output to the upsampling fusion layer for processing to obtain a sampling layer output; Step S1045, sending the sampling layer output to the detection head to extract boundary features and obtain boundary semantics; Step S1046, sampling the boundary semantics to the same height and width as the sampling layer output, and performing positional multiplication and additional residual connection processing on the boundary semantics and the sampling layer output to obtain enhanced features; Step S1047, performing regression analysis of the detection frame according to the enhanced features to obtain the recognition result of the garbage to be detected; Step S1048, using a classification loss function to perform classification prediction on the recognition result to obtain category information of the garbage to be detected; Step S105, obtaining the recognition result of the garbage to be detected, performing image processing in the detection frame to obtain interpretable distance feature information, performing distance measurement according to the interpretable distance feature information, and obtaining the location information of the garbage to be detected; Step S106, capturing the garbage to be detected according to the location information; Step S107: performing harmless treatment of the garbage to be detected according to the category information.
[0006] Another aspect of the present invention provides an AI sorting garbage disposal system, comprising: an automatic bag breaking module, an automatic screening module and a garbage disposal module; The automatic bag breaking module is used to obtain a garbage bag and break the bag to obtain the garbage to be tested; The automatic screening module includes: a collection unit, an identification unit, a positioning unit and a grabbing unit; The acquisition unit is used to acquire image information of the garbage to be detected; The recognition unit is used to recognize the image information using a perceptual recognition model, wherein the perceptual recognition model includes: an input end, a backbone network, an attention parallel fusion module, an upsampling fusion layer, and a detection head; The input end is used to receive the image information; The backbone network is used to extract features from the image information, and the backbone network includes at least three feature extraction layers; The attention parallel fusion module is used to receive and process the output features extracted by every three adjacent feature extraction layers to obtain the attention layer output; The upsampling fusion layer is used to receive and process the output of the attention layer to obtain the sampling layer output; The detection head is used to receive the sampling layer output and perform boundary feature extraction to obtain boundary semantics, sample the boundary semantics to the same height and width as the sampling layer output, and use positional multiplication and additional residual connection processing on the boundary semantics and the sampling layer output to obtain enhanced features, and perform regression analysis of the detection frame based on the enhanced features to obtain the recognition result of the garbage to be detected, and use the classification loss function to classify and predict the recognition result to obtain the category information of the garbage to be detected; The positioning unit is used to receive the identification result of the garbage to be detected, perform image processing in the detection frame to obtain interpretable distance feature information, perform distance measurement according to the interpretable distance feature information, and obtain the location information of the garbage to be detected; The grabbing unit is used to grab the garbage to be detected according to the position information; The garbage processing module is used to perform harmless processing of the garbage to be detected according to the category information.
[0007] The beneficial effects of the present invention are reflected in that the AI sorting garbage processing method and system provided by the present invention uses an effective boundary perception recognition model that integrates attention refinement and spatial interaction mechanisms, which can significantly improve the accuracy of detection and recognition of highly similar material garbage. The perception recognition model design of the present invention takes into account the existence of occlusion, overlap, and contamination of garbage, and realizes accurate recognition and detection of various types of garbage by enhancing the model's ability to capture key features. After successfully obtaining the 2D detection results of the garbage, advanced image processing methods, such as depth information extraction, stereo matching algorithms, etc., are further used to accurately calculate the coordinate position of the garbage in three-dimensional space, ensuring the precise positioning of the garbage position, and providing accurate navigation information for subsequent robotic arms or automated grasping equipment, thereby realizing accurate grasping and efficient processing of garbage. The system of the present invention can realize efficient recognition and detection of garbage in a variety of complex scenarios, greatly reducing false detection and missed detection caused by diverse garbage forms and complex backgrounds, and improving the efficiency and accuracy of garbage sorting and processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 This is a schematic diagram of the structure of the AI garbage sorting system provided in Example 1 of the present invention; Figure 2 This is a schematic diagram of the architecture of the perception and recognition model provided in Example 1 of the present invention; Figure 3 Another schematic diagram of the architecture of the perception recognition model provided in Embodiment 1 of the present invention; Figure 4 This is a flow chart of the AI garbage sorting method provided in Example 1 of the present invention; Figure 5 This is a flow chart of the perception recognition model recognition process provided in Example 1 of the present invention. DETAILED DESCRIPTION
[0009] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0010] Example 1 This embodiment provides an AI garbage sorting system. Figure 1 As shown, the AI sorting garbage processing system includes: an automatic bag breaking module 10, an automatic screening module 20 and a garbage processing module 30.
[0011] The automatic bag breaking module 10 is used to obtain the garbage bag and break the bag to obtain the garbage to be detected; specifically, the automatic bag breaking module 10 is responsible for splitting the packaged garbage to expose it to the garbage conveying device (such as a conveyor belt), so as to facilitate the automatic screening module 20 to identify the exposed garbage. The automatic bag breaking module 10 can be composed of two rotating blades and a nail claw. Among them, the rotating blades can be set on the left and right sides of the conveyor belt, and the rotation mode of the blades is one clockwise rotation and the other counterclockwise rotation. The nail claw can be fixed just above the conveyor belt and located between the two rotating blades. In addition, an infrared sensor can be set at the rotating blade placement. When the bagged garbage moves with the conveyor belt to the infrared sensing range, the nail claw pierces the bagged garbage downward, and the rotating blade extends into the inside of the conveyor belt and lasts for a period of time to ensure that the garbage bag can be cut. Since the garbage bag itself is fixed by the nail claw, but because the garbage bag itself is cut, its restraint on the garbage becomes weaker. Under the action of gravity and the friction of the conveyor belt, the garbage in the garbage bag will be released from the garbage bag. When the blade lasts for a period of time, the blade retracts and the nail claws are quickly retracted upward.
[0012] The automatic screening module 20 includes: a collection unit 21 , an identification unit 22 , a positioning unit 23 and a grabbing unit 24 .
[0013] The acquisition unit 21 is used to acquire image information of the garbage to be detected; the acquisition unit 21 may be an image acquisition device (such as a CCD camera in a camera, etc.), which captures image information of the garbage on the conveyor belt through the camera.
[0014] The recognition unit 22 is responsible for identifying the garbage on the conveyor belt (the garbage has passed through the automatic bag breaking module). The collection unit 21 and the recognition unit 22 can be both set inside the camera, with the camera as a carrier. The recognition unit 22 stores a perception recognition model 220 for using the perception recognition model 220 to recognize image information. Figure 2 As shown, the perception recognition model 220 may include: an input end 221, a backbone network 222, an attention parallel fusion module 223, an upsampling fusion layer 224 and a detection head 225.
[0015] In an optional implementation, the system of this embodiment may include two sets of acquisition units 21 and two sets of identification units 22, and the two sets of acquisition units 21 and the two sets of identification units 22 are respectively set correspondingly; the two sets of acquisition units 21 respectively acquire images of the garbage to be detected, and the two sets of identification units 22 respectively identify the corresponding image information. Specifically, two cameras can be set as carriers of the two sets of acquisition units 21 and the identification units 22, respectively, and the two sets of identification systems can improve the fault tolerance rate of the identification results and improve the stability of the identification; the two cameras can also be used to assist the positioning unit 23 in locating the garbage position and improve the positioning accuracy.
[0016] The input terminal 221 is used to receive image information; after receiving the image information, the image information is input into the backbone network 222 for feature extraction.
[0017] The backbone network 222 is used to extract features from the image information. The backbone network 222 includes at least three feature extraction layers. Specifically, the backbone network 222 includes multiple feature extraction layers, depending on the type of the backbone network 222 used. The backbone network 222 of this embodiment can use ResNet34. Each feature extraction layer extracts features from the image information, and obtains feature maps that are output to the attention parallel fusion module 223 for processing.
[0018] The attention parallel fusion module 223 is used to receive and process the output features extracted by every three adjacent feature extraction layers to obtain the attention layer output; specifically, the output features extracted by the feature extraction layer can exist in the form of feature maps. There can be multiple attention parallel fusion modules 223 of the perception recognition model 220, which can be set according to the type of model and the number of feature extraction layers. Since the sizes of the feature maps output by different feature extraction layers in the backbone network are different, the sizes of the three output feature maps received by each attention parallel fusion module 223 are also different. Therefore, in the attention parallel fusion module 222, the two smaller feature maps are first upsampled to the same size as the larger feature map, and then spliced in the channel direction. After the splicing is completed, the principal component analysis algorithm (Principal Component Analysis, PCA) is used to reduce the number of channels to the preset number of channels (the purpose is to reduce the dimension and remove possible noise), and the SEWeight module (extraction channel attention module) is used to converge the feature map into a channel attention vector, which contains fine spatial information at multiple scales, so as to obtain the attention layer output. The attention parallel fusion module 222 can utilize attention refinement and spatial interaction to optimize the recognition effect, thereby achieving more accurate and efficient detection of similar material garbage.
[0019] The upsampling fusion layer 224 is used to receive and process the output of the attention layer to obtain the sampling layer output; specifically, a fully connected layer and an activation function may be used in the upsampling fusion layer 224 to process the output of the attention layer, and the activation function may be a linear rectification function, etc. The number of upsampling fusion layers 224 in this embodiment may also be multiple, which may be set according to the number of attention parallel fusion modules 223.
[0020] The detection head 225 is used to receive the sampling layer output and perform boundary feature extraction to obtain boundary semantics, sample the boundary semantics to the same height and width as the sampling layer output, and use the bitwise multiplication and additional residual connection processing on the boundary semantics and the sampling layer output to obtain the enhanced features, and perform regression analysis of the detection frame based on the enhanced features to obtain the recognition result of the garbage to be detected, and use the classification loss function to classify and predict the recognition result to obtain the category information of the garbage to be detected. Specifically, by receiving the output data from the sampling layer, the boundary features are finely extracted so as to capture the key boundary information from the sampling layer output, and then the boundary semantics are obtained for subsequent image analysis. In order to maintain the consistency of the data and facilitate subsequent processing, the system will adjust the extracted boundary semantics to the same height and width as the sampling layer output through sampling technology. After ensuring the consistency of the boundary semantics and the sampling layer output in size, the boundary semantics and the sampling layer output can be fused and complemented by bitwise multiplication, bitwise addition, etc. to enhance the expression ability of the features. At the same time, in order to improve the richness and robustness of the features, additional residual connection processing is introduced to combine the original features with the features processed by the positional multiplication and other methods, so as to obtain further enhanced features. Based on these enhanced features, the regression analysis of the detection frame can be performed to predict the recognition results of the garbage to be detected. Specifically, confidence-driven bounding box positioning can be introduced to perform the regression of the detection frame. The detection head 225 not only makes full use of the boundary semantic information, but also improves the accuracy and reliability of the detection frame regression analysis by fusing and enhancing features, thereby improving the accuracy of garbage recognition.
[0021] For garbage classification prediction, after garbage identification, the cross-entropy loss function such as Cross-Entropy Loss can be used to give the category information of garbage, so as to facilitate the classification and processing of garbage according to the garbage category. Cross-Entropy Loss can quantify the difference between the probability distribution predicted by the model and the true category distribution, thereby guiding the model to continuously optimize its prediction ability during the training process. By applying Cross-Entropy Loss, the probability of each garbage sample being correctly classified and the confidence of the model in distinguishing different garbage categories can be obtained. Design an automated classification process based on the category information calculated by Cross-Entropy Loss.
[0022] In an optional embodiment, if Figure 3In the architecture of the perception recognition model 220 shown, the perception recognition model 220 may include three attention parallel fusion modules and three upsampling fusion layers. The three attention parallel fusion modules include: a first attention parallel fusion module 2231, a second attention parallel fusion module 2232, and a third attention parallel fusion module 2233; the three upsampling fusion layers include: a first upsampling fusion layer 2241, a second upsampling fusion layer 2242, and a third upsampling fusion layer 2243; The backbone network 222 includes five feature extraction layers, and the five feature extraction layers include: a first feature extraction layer, a second feature extraction layer, a third feature extraction layer, a fourth feature extraction layer and a fifth feature extraction layer; The attention parallel fusion module 223 receives the output features extracted by every three adjacent feature extraction layers and processes them to obtain the attention layer output, specifically including: the first attention parallel fusion module 2231 receives the output features extracted by the first feature extraction layer, the second feature extraction layer and the third feature extraction layer and processes them to obtain the first attention layer output; the second attention parallel fusion module 2232 receives the output features extracted by the second feature extraction layer, the third feature extraction layer and the fourth feature extraction layer and processes them to obtain the second attention layer output; the third attention parallel fusion module 2233 receives the output features extracted by the third feature extraction layer, the fourth feature extraction layer and the fifth feature extraction layer and processes them to obtain the third attention layer output; The upsampling fusion layer 224 receives the attention layer output and processes it to obtain the sampling layer output, specifically including: the first upsampling fusion layer 2241, the second upsampling fusion layer 2242 and the third upsampling fusion layer 2243 respectively receive the first attention layer output, the second attention layer output and the third attention layer output and process them to obtain the first sampling layer output, the second sampling layer output and the third sampling layer output respectively.
[0023] Through Figure 3 The set perception recognition model architecture can refine the attention mechanism and enhance the ability to capture target features.
[0024] In an optional embodiment, before the sampling layer output is sent to the detection head for boundary feature extraction, the first sampling layer output, the second sampling layer output and the third sampling layer output are connected and sent to the detection head for boundary feature extraction. Specifically, since there are three sampling layer outputs, it is necessary to first connect the sampling layer outputs through the concat function and then perform feature extraction. Through integration, the diversity and complementary features captured by different sampling layers can be utilized, thereby achieving higher accuracy and robustness in the subsequent boundary feature extraction process.
[0025] The following is a specific example of the recognition unit 22 using the perceptual recognition model 220 to perform recognition, and the Figure 3 The given perception recognition model architecture is used for calculation. The specific steps are as follows: The image of garbage on the conveyor belt captured by the camera is received through the input terminal 221, and then input into the backbone network 222 for feature extraction. The output features of the three adjacent feature extraction layers in the backbone network 222 are sent to the same attention parallel fusion module 223. Assume that the output feature maps of each feature extraction layer of the backbone network 222 are ,but Send it to the attention parallel fusion module 2231, Send it to the attention parallel fusion module 2232, Send to the attention parallel fusion module 2233; In the attention parallel fusion module 223, the two smaller feature maps are first upsampled to the same size as the larger feature map, and then spliced in the channel direction. After the splicing is completed, PCA is used to reduce the number of channels to the preset number of channels, and the SEWeight module is used to converge the feature map into a channel attention vector, which contains fine spatial information at multiple scales. The outputs of the three attention parallel fusion modules are , that is, the output of the first attention layer, the output of the second attention layer, and the output of the third attention layer. Then, are input to three upsampling fusion layers 224 respectively, and the sampling layer output is calculated The specific formula is as follows: ; in, represents the upsampling fusion layer, express Activation function, represents the fully connected layer, is the coefficient. After passing through the upsampling fusion layer 224, The three are connected and sent to the detection head 225. The detection head 225 first extracts the boundary features: ; in, Represents boundary semantics, express Dilated convolution, express convolution, for Activation function. Then the boundary semantics Sampling to The same height and width, and then apply counter-integration and counter-integration between them, and use an additional residual connection to ensure no information loss: ; in, represents bitwise addition, represents the bitwise multiplication, h represents the enhanced feature, and H represents Convolution, U represents upsampling. Finally, according to The final detection result can be obtained by performing regression of the detection frame according to the traditional method. The regression analysis method of the detection frame belongs to the existing technology and will not be described in detail here.
[0026] The positioning unit 23 is used to receive the recognition result of the garbage to be detected, perform image processing in the detection frame to obtain interpretable distance feature information, perform distance measurement based on the interpretable distance feature information, and obtain the location information of the garbage to be detected; specifically, after the positioning unit 23 obtains the recognition result, it uses traditional image processing methods such as grayscale processing, binarization processing, edge recognition, etc. in the detection frame to obtain interpretable manual features (such as interpretable distance feature information, location features, etc.), and obtains the target distance through triangulation.
[0027] In an optional embodiment, the positioning unit 23 receives the recognition results of the garbage to be detected, performs image processing in the detection frame to obtain interpretable distance feature information, performs distance measurement based on the interpretable distance feature information, and obtains the position information of the garbage to be detected, which specifically includes: the positioning unit 23 receives the recognition results of the garbage to be detected by the two sets of recognition units 22 respectively, performs image processing in the detection frame respectively to obtain interpretable distance feature information, performs distance measurement based on the interpretable distance feature information, and obtains the position information of the garbage to be detected relative to the two sets of acquisition units 21. Specifically, when two sets of acquisition units 21 and two sets of recognition units 22 are used, since the internal parameters and positions of the cameras where each set of acquisition units 21 and recognition units 22 are located are known and fixed, after the distance is measured, the positional relationship between the garbage target to be detected and any camera can be obtained, so as to accurately locate the garbage to be detected.
[0028] The grasping unit 24 is used to grasp the garbage to be detected according to the position information; specifically, the grasping unit 24 can be composed of a robotic arm, and the relationship between the base of the robotic arm and the camera where the identification unit 22 is located remains constant, so the position relationship between the target garbage and the camera can be converted into a position relationship with the base of the robotic arm. At this time, grasping can be achieved by controlling the upper limbs of the robotic arm (the control between the upper limbs of the robotic arm and the base is measured by the internal position relationship of the robotic arm, which belongs to the prior art).
[0029] The garbage processing module 30 is used to perform harmless treatment of the garbage to be detected according to the category information. The garbage processing module 30 can perform corresponding treatment according to the identification result of the garbage to be detected, such as recycling, crushing, incineration, etc.
[0030] Different harmless treatments can be implemented for the following different types of garbage: (1) Plastic waste: garbage identified as plastic can be put into a crusher to crush the plastic, and then packaged for subsequent sale; (2) Organic waste: organic waste can be sent to an incinerator as a material for burning to generate electricity; (3) Textile waste: reusable textiles (such as old clothes) can be cleaned, disinfected, and refurbished; (4) Paper waste: paper can be pulped to remove impurities, and then compressed using hydraulic or mechanical pressure to make new paper; (5) Metal waste: recycled metals can be smelted to extract pure metals for use in industrial production.
[0031] The AI sorting garbage disposal system provided in this embodiment can significantly improve the accuracy of detection and identification of highly similar material garbage by using an effective boundary perception recognition model that integrates attention refinement and spatial interaction mechanisms. The perception recognition model design of this embodiment takes into account the existence of occlusion, overlap, and contamination of garbage. By enhancing the model's ability to capture key features, it achieves accurate recognition and detection of various types of garbage. After successfully obtaining the 2D detection results of the garbage, advanced image processing methods are further used to accurately calculate the coordinate position of the garbage in three-dimensional space, ensuring the precise positioning of the garbage position, and providing accurate navigation information for subsequent robotic arms or automated grasping equipment, thereby achieving accurate grasping and efficient processing of garbage. The system using this embodiment can achieve efficient recognition and detection of garbage in a variety of complex scenarios, greatly reducing false detections and missed detections caused by diverse garbage forms and complex backgrounds, and improving the efficiency and accuracy of garbage sorting and processing.
[0032] In an optional embodiment, before using the perceptual recognition model 220 to identify the garbage to be detected, the recognition unit 22 also uses the training samples to train the perceptual recognition model 220 according to the supervised learning method, so that the perceptual recognition model converges and obtains the training weights; in the process of inputting the image information into the perceptual recognition model 220 for recognition, no non-maximum suppression is performed to obtain fuzzy samples, and the perceptual recognition model 220 is fine-tuned using the training weights and fuzzy samples so that the perceptual recognition model 220 converges, wherein the fuzzy samples are formed by multiple detection frames with the largest intersection-and-union ratio with the detection frame with the maximum confidence. Specifically, in this embodiment, the perceptual recognition model 220 is pre-trained based on fuzzy samples to improve the performance of the recognition unit 22. The following is a specific example of a training method based on fuzzy samples and fine-tuning the model until convergence according to the training method. The specific steps can be as follows: Step (1), firstly, the perceptual recognition model 220 of the recognition unit 22 is trained according to the traditional supervised learning method, that is, the manually annotated detection box is the true value, and a large number of manually annotated label values are used to train the perceptual recognition model 220 of the recognition unit 22, and make it converge to obtain the trained weights. The loss function can use Smooth L1 Loss (smooth loss function) and Cross-Entropy Loss (cross entropy loss function, CE); Step (2), after the perceptual recognition model 220 converges, the samples that have not participated in the training are sent to the recognition unit 22. The recognition unit 22 does not perform non-maximum suppression on the output results during the recognition process, and retains the first four detection frames with the largest intersection-over-union ratio with the detection frame with the maximum confidence in each category, together with the detection frame with the maximum confidence as the true value of the fuzzy sample; Step (3), load the weights trained in step (1), use the blurred samples generated in step (2) to fine-tune the recognition unit 22, and use a new loss function during fine-tuning, which is defined as follows: ; in, Represents the initial impairment weight assigned to all samples, specifically: ; It represents the intersection-over-combination ratio between the prediction result and the true value of fuzzy sample 1. to They are the intersection and union ratios of the prediction results and the true values of other fuzzy samples, is the weight, The value of is 1, The weight value is 0.25, is the cross entropy loss function.
[0033] This embodiment also provides an AI sorting garbage processing method, which realizes garbage processing through the above-mentioned AI sorting garbage processing system. The specific situation of the AI sorting garbage processing system has been described in the above-mentioned AI sorting garbage processing system part. Here, only the steps of the AI sorting garbage processing method are briefly described. For the rest, please refer to the above-mentioned AI sorting garbage processing system.
[0034] like Figure 4 As shown, the AI garbage sorting method of this embodiment includes the following steps: Step S101, obtaining a garbage package; specifically, the garbage package may be automatically received through an external entrance.
[0035] Step S102, breaking the garbage bag to obtain the garbage to be detected; specifically, the automatic bag breaking module can be responsible for splitting the packaged garbage to expose it to the garbage conveying device (such as a conveyor belt), thereby facilitating the automatic screening module to identify the exposed garbage.
[0036] Step S103, collecting image information of the garbage to be detected; specifically, the image information of the garbage on the conveyor belt can be captured by an image acquisition device (such as a CCD camera in a camera).
[0037] Step S104, input the image information into the perception recognition model for recognition, the perception recognition model includes: an input end, a backbone network, an attention parallel fusion module, an upsampling fusion layer and a detection head; specifically, the method of this embodiment uses the perception recognition model to perform garbage recognition, such as Figure 5 As shown, the identification process may include the following steps: Step S1041, receiving image information through an input end; after receiving the image information, inputting the image information into a backbone network for feature extraction.
[0038] Step S1042, feature extraction is performed through the backbone network, and the backbone network includes at least three feature extraction layers; specifically, the backbone network includes multiple feature extraction layers, depending on the type of backbone network used, and the backbone network of this embodiment can use ResNet34. Each feature extraction layer extracts features from the image information respectively, and obtains feature maps that are output to the attention parallel fusion module for processing.
[0039] Step S1043, the output features extracted by every three adjacent feature extraction layers are sent to the same attention parallel fusion module for processing to obtain the attention layer output; specifically, the output features extracted by the feature extraction layer can exist in the form of feature maps. The perception recognition model of this embodiment can have multiple attention parallel fusion modules, which can be set according to the type of model and the number of feature extraction layers. Since the sizes of the outputs of different feature extraction layers in the backbone network are different, the sizes of the three output feature maps received by each attention parallel fusion module are also different. Therefore, in the attention parallel fusion module, the two smaller feature maps are first sampled to the same size as the larger feature map, and then spliced in the channel direction. After the splicing is completed, PCA is used to reduce the number of channels to the preset number of channels, and the SEWeight module is used to converge the feature map into a channel attention vector, which contains fine spatial information at multiple scales, so as to obtain the attention layer output. The attention parallel fusion module can optimize the recognition effect by using attention refinement and spatial interaction, thereby achieving more accurate and efficient detection of similar material garbage.
[0040] Step S1044, the output of the attention layer is sent to the upsampling fusion layer for processing to obtain the sampling layer output; specifically, in the upsampling fusion layer, a fully connected layer and an activation function can be used to process the output of the attention layer, and the activation function can be a linear rectification function, etc. The number of upsampling fusion layers in this embodiment can also be multiple, which can be set according to the number of attention parallel fusion modules.
[0041] In an optional embodiment, the perception recognition model includes three attention parallel fusion modules and three upsampling fusion layers, the three attention parallel fusion modules respectively include: a first attention parallel fusion module, a second attention parallel fusion module and a third attention parallel fusion module, the three upsampling fusion layers respectively include: a first upsampling fusion layer, a second upsampling fusion layer and a third upsampling fusion layer; the backbone network includes five feature extraction layers, the five feature extraction layers respectively include: a first feature extraction layer, a second feature extraction layer, a third feature extraction layer, a fourth feature extraction layer and a fifth feature extraction layer; Step S1043 may include: sending the output features extracted by the first feature extraction layer, the second feature extraction layer and the third feature extraction layer to the first attention parallel fusion module for processing to obtain the first attention layer output, sending the output features extracted by the second feature extraction layer, the third feature extraction layer and the fourth feature extraction layer to the second attention parallel fusion module for processing to obtain the second attention layer output, sending the output features extracted by the third feature extraction layer, the fourth feature extraction layer and the fifth feature extraction layer to the third attention parallel fusion module for processing to obtain the third attention layer output; Step S1044 may include: sending the first attention layer output, the second attention layer output and the third attention layer output to the first upsampling fusion layer, the second upsampling fusion layer and the third upsampling fusion layer for processing, respectively, to obtain the first sampling layer output, the second sampling layer output and the third sampling layer output, respectively.
[0042] Through the above-mentioned perception recognition model architecture, the attention mechanism can be refined and the ability to capture target features can be enhanced.
[0043] In an optional embodiment, sending the sampling layer output to the detection head for boundary feature extraction includes: connecting the first sampling layer output, the second sampling layer output and the third sampling layer output and sending them to the detection head for boundary feature extraction. Specifically, since there are three sampling layer outputs, it is necessary to first connect the sampling layer outputs through the concat function and then perform feature extraction. Through integration, the diversity and complementary features captured by different sampling layers can be utilized, so as to achieve higher accuracy and robustness in the subsequent boundary feature extraction process.
[0044] Step S1045, sending the sampling layer output to the detection head for boundary feature extraction to obtain boundary semantics; specifically, by receiving the output data from the sampling layer, fine extraction of boundary features is performed to capture key boundary information from the sampling layer output, and then obtain boundary semantics for subsequent image analysis.
[0045] Step S1046, the boundary semantics are sampled to the same height and width as the sampling layer output, and the boundary semantics and the sampling layer output are processed using positional multiplication and additional residual connection to obtain enhanced features; in order to maintain data consistency and facilitate subsequent processing, the system will adjust the extracted boundary semantics to the same height and width as the sampling layer output through sampling technology. After ensuring the consistency of the boundary semantics and the sampling layer output in size, the boundary semantics and the sampling layer output can be fused and complemented by positional multiplication, positional addition, etc. to enhance the expressive power of the features. At the same time, in order to improve the richness and robustness of the features, additional residual connection processing is introduced to combine the original features with the features processed by positional multiplication, etc., so as to obtain further enhanced features.
[0046] Step S1047, perform regression analysis of the detection frame based on the enhanced features to obtain the recognition result of the garbage to be detected. Performing regression analysis of the detection frame based on these enhanced features can predict the recognition result of the garbage to be detected. Specifically, confidence-driven bounding box positioning can be introduced to perform regression of the detection frame.
[0047] Step S1048, use the classification loss function to classify and predict the recognition results to obtain the category information of the garbage to be detected; specifically, after the garbage is identified, the cross-entropy loss function such as Cross-Entropy Loss can be used to give the category information of the garbage, so as to facilitate the classification and processing of the garbage according to the garbage category. Cross-Entropy Loss can quantify the difference between the probability distribution predicted by the model and the true category distribution, thereby guiding the model to continuously optimize its prediction ability during the training process. By applying Cross-Entropy Loss, the probability of each garbage sample being correctly classified can be obtained, as well as the confidence of the model in distinguishing different garbage categories. Design an automated classification process based on the category information calculated by Cross-Entropy Loss.
[0048] Step S105, obtaining the recognition result of the garbage to be detected, performing image processing in the detection frame to obtain interpretable distance feature information, performing distance measurement based on the interpretable distance feature information, and obtaining the location information of the garbage to be detected; specifically, after obtaining the recognition result, traditional image processing methods such as grayscale processing, binarization processing, edge recognition, etc. can be used in the detection frame to obtain interpretable manual features (such as interpretable distance feature information, location features, etc.), and the target distance can be obtained through triangulation.
[0049] Step S106, grabbing the garbage to be detected according to the position information; specifically, the grabbing can be achieved by setting a mechanical arm and controlling the mechanical arm.
[0050] Step S107, the garbage to be detected is subjected to harmless treatment of the corresponding category according to the category information. Specifically, the garbage to be detected can be subjected to corresponding treatment according to the identification result, such as recycling, crushing, incineration, etc.
[0051] The AI sorting garbage processing method provided in this embodiment can significantly improve the accuracy of detection and identification of highly similar material garbage by using an effective boundary perception recognition model that integrates attention refinement and spatial interaction mechanisms. The perception recognition model design of this embodiment takes into account the existence of occlusion, overlap, and contamination of the garbage shape, and realizes accurate recognition and detection of various types of garbage by enhancing the model's ability to capture key features. After successfully obtaining the 2D detection results of the garbage, advanced image processing methods are further used to accurately calculate the coordinate position of the garbage in three-dimensional space, ensuring the precise positioning of the garbage position, and providing accurate navigation information for subsequent robotic arms or automated grasping equipment, thereby realizing accurate grasping and efficient processing of garbage. The method of this embodiment can realize efficient recognition and detection of garbage in a variety of complex scenarios, greatly reducing false detection and missed detection caused by diverse garbage shapes and complex backgrounds, and improving the efficiency and accuracy of garbage sorting and processing.
[0052] In an optional embodiment, before step S104, the AI garbage sorting method further includes: step S103a, using training samples to train the perceptual recognition model according to the supervised learning method, so that the perceptual recognition model converges and obtains training weights; in the process of executing step S104, non-maximum suppression is not performed to obtain fuzzy samples, and the perceptual recognition model is fine-tuned using the training weights and fuzzy samples to make the perceptual recognition model converge, wherein the fuzzy samples are formed by multiple detection frames with the largest intersection-and-union ratio with the maximum confidence detection frame. Specifically, in this embodiment, the perceptual recognition model is pre-trained based on fuzzy samples, and the model is fine-tuned until convergence, so as to improve the recognition performance of the perceptual recognition model.
[0053] In an optional embodiment, after step S1047, the method further includes: using a classification loss function to classify and predict the recognition results to obtain category information of the garbage to be detected; processing the garbage to be detected includes: performing harmless processing of the corresponding category according to the category information of the garbage to be detected. Specifically, after the garbage is identified, a cross-entropy loss function such as Cross-Entropy Loss can be used to give the category information of the garbage, so as to facilitate the classification and processing of the garbage according to the garbage category. Cross-Entropy Loss can quantify the difference between the probability distribution predicted by the model and the true category distribution, thereby guiding the model to continuously optimize its prediction ability during the training process. By applying Cross-Entropy Loss, the probability of each garbage sample being correctly classified and the confidence of the model in distinguishing different garbage categories can be obtained. An automated classification process can be designed based on the category information calculated by Cross-Entropy Loss.
[0054] In the description of the embodiments of the present invention, it should be understood that the terms "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "center", "top", "bottom", "top", "bottom", "inside", "outside", "inner side", "outer side" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. Among them, "inside" refers to an internal or enclosed area or space. "Periphery" refers to the area surrounding a specific component or a specific area.
[0055] In the description of the embodiments of the present invention, the terms "first", "second", "third", and "fourth" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first", "second", "third", and "fourth" may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.
[0056] In the description of the embodiments of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "install", "connect", "connect", and "assemble" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0057] In the description of the embodiments of the present invention, specific features, structures, materials or characteristics may be combined in a suitable manner in any one or more embodiments or examples.
[0058] In the description of the embodiments of the present invention, it should be understood that "-" and "~" represent a range between two values, and the range includes the endpoints. For example: "AB" represents a range greater than or equal to A and less than or equal to B. "A~B" represents a range greater than or equal to A and less than or equal to B.
[0059] In the description of the embodiments of the present invention, the term "and / or" herein is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " herein generally indicates that the associated objects before and after are in an "or" relationship.
[0060] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An AI garbage sorting method, characterized in that: The method comprises: Step S101, obtaining a garbage bag; Step S102, breaking the garbage bag to obtain the garbage to be detected; Step S103, collecting image information of the garbage to be detected; Step S104: input the image information into a perception recognition model for recognition. The perception recognition model includes: an input end, a backbone network, an attention parallel fusion module, an upsampling fusion layer, and a detection head. The recognition process includes the following steps: Step S1041, receiving the image information through the input end; Step S1042, performing feature extraction through the backbone network, wherein the backbone network includes at least three feature extraction layers; Step S1043, sending the output features extracted by every three adjacent feature extraction layers to the same attention parallel fusion module for processing to obtain the attention layer output; Step S1044, sending the attention layer output to the upsampling fusion layer for processing to obtain a sampling layer output; Step S1045, sending the sampling layer output to the detection head to extract boundary features and obtain boundary semantics; Step S1046, sampling the boundary semantics to the same height and width as the sampling layer output, and performing positional multiplication and additional residual connection processing on the boundary semantics and the sampling layer output to obtain enhanced features; Step S1047, performing regression analysis of the detection frame according to the enhanced features to obtain the recognition result of the garbage to be detected; Step S1048, using a classification loss function to perform classification prediction on the recognition result to obtain category information of the garbage to be detected; Step S105, obtaining the recognition result of the garbage to be detected, performing image processing in the detection frame to obtain interpretable distance feature information, performing distance measurement according to the interpretable distance feature information, and obtaining the location information of the garbage to be detected; Step S106, capturing the garbage to be detected according to the location information; Step S107: performing harmless treatment of the garbage to be detected according to the category information.
2. The AI garbage sorting method according to claim 1 is characterized in that: Before step S104, the method further includes: step S103a, using training samples to train the perception recognition model according to a supervised learning method, so that the perception recognition model converges and obtains training weights; During the execution of step S104, non-maximum suppression is not performed to obtain blurred samples, and the perceptual recognition model is fine-tuned using the training weights and the blurred samples so that the perceptual recognition model converges, wherein the blurred samples are formed by multiple detection frames having the largest intersection-over-union ratio with the detection frame with the maximum confidence.
3. The AI garbage sorting method according to claim 1, characterized in that: The perception recognition model includes three attention parallel fusion modules and three upsampling fusion layers, the three attention parallel fusion modules respectively include: a first attention parallel fusion module, a second attention parallel fusion module and a third attention parallel fusion module, and the three upsampling fusion layers respectively include: a first upsampling fusion layer, a second upsampling fusion layer and a third upsampling fusion layer; The backbone network includes five feature extraction layers, and the five feature extraction layers include: a first feature extraction layer, a second feature extraction layer, a third feature extraction layer, a fourth feature extraction layer and a fifth feature extraction layer; The step S1043 includes: Send the output features extracted by the first feature extraction layer, the second feature extraction layer and the third feature extraction layer to the first attention parallel fusion module for processing to obtain the first attention layer output, send the output features extracted by the second feature extraction layer, the third feature extraction layer and the fourth feature extraction layer to the second attention parallel fusion module for processing to obtain the second attention layer output, send the output features extracted by the third feature extraction layer, the fourth feature extraction layer and the fifth feature extraction layer to the third attention parallel fusion module for processing to obtain the third attention layer output; The step S1044 includes: The first attention layer output, the second attention layer output and the third attention layer output are respectively sent to the first upsampling fusion layer, the second upsampling fusion layer and the third upsampling fusion layer for processing to obtain the first sampling layer output, the second sampling layer output and the third sampling layer output respectively.
4. The AI garbage sorting method according to claim 3 is characterized in that: The sending the sampling layer output to the detection head for boundary feature extraction includes: connecting the first sampling layer output, the second sampling layer output and the third sampling layer output and sending them to the detection head for boundary feature extraction.
5. The AI garbage sorting method according to claim 1, characterized in that: The image processing includes at least one of the following: grayscale processing, binarization processing and edge recognition.
6. An AI garbage sorting system, characterized in that: include: Automatic bag breaking module, automatic screening module and garbage disposal module; The automatic bag breaking module is used to obtain a garbage bag and break the bag to obtain the garbage to be tested; The automatic screening module includes: a collection unit, an identification unit, a positioning unit and a grabbing unit; The acquisition unit is used to acquire image information of the garbage to be detected; The recognition unit is used to recognize the image information using a perceptual recognition model, wherein the perceptual recognition model includes: an input end, a backbone network, an attention parallel fusion module, an upsampling fusion layer, and a detection head; The input end is used to receive the image information; The backbone network is used to extract features from the image information, and the backbone network includes at least three feature extraction layers; The attention parallel fusion module is used to receive and process the output features extracted by every three adjacent feature extraction layers to obtain the attention layer output; The upsampling fusion layer is used to receive and process the output of the attention layer to obtain the sampling layer output; The detection head is used to receive the sampling layer output and perform boundary feature extraction to obtain boundary semantics, sample the boundary semantics to the same height and width as the sampling layer output, and use positional multiplication and additional residual connection processing on the boundary semantics and the sampling layer output to obtain enhanced features, and perform regression analysis of the detection frame based on the enhanced features to obtain the recognition result of the garbage to be detected, and use the classification loss function to classify and predict the recognition result to obtain the category information of the garbage to be detected; The positioning unit is used to receive the identification result of the garbage to be detected, perform image processing in the detection frame to obtain interpretable distance feature information, perform distance measurement according to the interpretable distance feature information, and obtain the location information of the garbage to be detected; The grabbing unit is used to grab the garbage to be detected according to the position information; The garbage processing module is used to perform harmless processing of the garbage to be detected according to the category information.
7. The AI sorting garbage processing system according to claim 6 is characterized in that: Before using the perceptual recognition model to recognize the garbage to be detected, the recognition unit further trains the perceptual recognition model using training samples according to a supervised learning method, so that the perceptual recognition model converges and obtains training weights; In the process of inputting the image information into the perceptual recognition model for recognition, non-maximum suppression is not performed to obtain a fuzzy sample, and the perceptual recognition model is fine-tuned using the training weights and the fuzzy sample so that the perceptual recognition model converges, wherein the fuzzy sample is formed by multiple detection frames with the largest intersection-over-union ratio with the detection frame with the maximum confidence.
8. The AI sorting garbage processing system according to claim 6, characterized in that: The perception recognition model includes three attention parallel fusion modules and three upsampling fusion layers, the three attention parallel fusion modules respectively include: a first attention parallel fusion module, a second attention parallel fusion module and a third attention parallel fusion module, and the three upsampling fusion layers respectively include: a first upsampling fusion layer, a second upsampling fusion layer and a third upsampling fusion layer; The backbone network includes five feature extraction layers, and the five feature extraction layers include: a first feature extraction layer, a second feature extraction layer, a third feature extraction layer, a fourth feature extraction layer and a fifth feature extraction layer; The attention parallel fusion module receives the output features extracted by every three adjacent feature extraction layers and processes them to obtain the attention layer output, specifically including: The first attention parallel fusion module receives and processes the output features extracted by the first feature extraction layer, the second feature extraction layer and the third feature extraction layer to obtain the first attention layer output, the second attention parallel fusion module receives and processes the output features extracted by the second feature extraction layer, the third feature extraction layer and the fourth feature extraction layer to obtain the second attention layer output, and the third attention parallel fusion module receives and processes the output features extracted by the third feature extraction layer, the fourth feature extraction layer and the fifth feature extraction layer to obtain the third attention layer output; The upsampling fusion layer receives the attention layer output and processes it to obtain the sampling layer output, specifically including: The first upsampling fusion layer, the second upsampling fusion layer and the third upsampling fusion layer respectively receive the first attention layer output, the second attention layer output and the third attention layer output and process them to obtain the first sampling layer output, the second sampling layer output and the third sampling layer output respectively.
9. The AI sorting garbage processing system according to claim 6, characterized in that: The detection head receiving the sampling layer output and performing boundary feature extraction specifically includes: connecting the first sampling layer output, the second sampling layer output and the third sampling layer output and sending them to the detection head for boundary feature extraction.
10. The AI sorting garbage processing system according to claim 6, characterized in that: The system includes two sets of collection units and two sets of identification units, and the two sets of collection units and the two sets of identification units are respectively arranged correspondingly; The two sets of acquisition units respectively acquire images of the garbage to be detected, and the two sets of recognition units respectively recognize the corresponding image information; The positioning unit receives the identification result of the garbage to be detected, performs image processing in the detection frame to obtain interpretable distance feature information, performs distance measurement according to the interpretable distance feature information, and obtains the location information of the garbage to be detected, specifically including: The positioning unit receives the recognition results of the two sets of recognition units on the garbage to be detected, performs image processing in the detection frame to obtain interpretable distance feature information, performs distance measurement based on the interpretable distance feature information, and obtains the position information of the garbage to be detected relative to the two sets of collection units.
Citation Information
Patent Citations
Box type domestic garbage recycling product classification system
CN108421724A
Garbage detection and identification method based on improved yolov5 network
CN113158956A
Garbage can unattended system design method based on machine vision
CN114092877A
Household garbage comprehensive treatment method and system
CN114247726A
Garbage type identification method and device, computer equipment and storage medium
CN114742989A
Cited By
Garbage sorting integrated equipment
CN121178449A