Garbage recognition and classification method, computer-readable storage medium, and robot
By adopting the combination of defuzzy network and deep convolutional neural network models in robots, efficient identification and classification of garbage is achieved, and the problem of high false detection rates and missed detection rates in traditional robots in garbage inspection is solved, which improves detection speed and versatility, and reduces environmental impact.
Patent Information
- Application Number
- CN202111521559.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-13
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-13
AI Technical Summary
Traditional robots have high false detection rates and missed detection rates during garbage inspection, insufficient detection speed and versatility, and are greatly affected by the environment.
A garbage identification and classification method is adopted to defuzzle the original frame image through a preset defuzzy network, and target objects are detected and tracked in combination with a deep convolutional neural network model to reduce the error detection rate and miss detection rate, and improve detection speed and versatility.
It effectively reduces the false detection rate and missed detection rate of ground garbage identification, improves detection speed and versatility, and is relatively less affected by the environment, and has strong robustness of detection and tracking algorithms.
Smart Images

Figure CN114398950B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robots, and in particular, to a garbage recognition and classification method, a computer-readable storage medium, and a robot. Background Art
[0002] During the garbage inspection process of traditional robots, in addition to cleaning sundries such as fallen leaves and paper scraps, it is also necessary to specifically detect, classify, and collect recyclable garbage, so as to realize the recycling of recyclable garbage. Most traditional algorithms detect recyclable garbage by segmenting the road surface in the image and using algorithms such as connected component detection, grayscale histogram, and feature matching on the road surface to determine the category and position information of the objects on the road surface, and then recycle specific recyclable garbage, such as plastic bottles and aluminum cans. However, since most of the garbage on the road surface has been crushed and has different shapes, the feature matching effect is mostly poor. In addition, due to the change of the viewing position during the movement of the robot, the shape, size, etc. of the same object captured will change, and it is easy to have the situation of missing detection in multiple consecutive frames. At this time, the robot will update the path and move towards the next garbage, resulting in low operation efficiency. Summary of the Invention
[0003] The present invention aims to solve at least one of the technical problems in the related art to some extent. To this end, the first object of the present invention is to propose a garbage recognition and classification method, which can effectively reduce the false detection rate and missing detection rate of ground garbage recognition, improve the detection speed and versatility, and is less affected by the environment.
[0004] The second object of the present invention is to propose a computer-readable storage medium.
[0005] The third object of the present invention is to propose a robot.
[0006] To achieve the above object, the first aspect embodiment of the present invention proposes a garbage recognition and classification method, including: performing deblurring processing on the original frame image according to a preset deblurring network to obtain a target image;
[0007] Inputting the target image into a preset deep convolutional neural network model for target object detection to obtain a first feature vector map of the target object; tracking the target object according to the first feature vector map, and sorting the target object according to the category and position of the target object.
[0008] According to the garbage recognition and classification method of the embodiments of the present invention, first, according to a preset deblurring network, the original frame image is deblurred to obtain a target image, then the target image is input into a preset deep convolutional neural network model for target object detection to obtain a first feature vector map of the target object, and finally, the target object is tracked according to the first feature vector map, and the target object is sorted according to the category and position of the target object. Thus, this method can, according to the deblurring network, reduce the impact of robot movement on the image, make the image clearer, and by adding a tracking algorithm after the detection algorithm, it can eliminate the misrecognition of occasional frames, effectively reduce the false detection rate and missed detection rate of ground garbage recognition, improve the detection speed and versatility, and is less affected by the environment, and the detection and tracking algorithms have strong robustness.
[0009] In addition, the garbage recognition and classification method according to the above embodiments of the present invention may further have the following additional technical features:
[0010] According to an embodiment of the present invention, the above garbage recognition and classification method further includes: obtaining a first sample set, where each training sample in the first sample set includes a first original blurred video frame and a second original blurred video frame adjacent to the first original blurred video frame; inputting the first original blurred video frame and the second original blurred video frame into the deblurring network for deblurring processing respectively to obtain a first predicted clear image and a second predicted clear image; performing optical flow prediction, blurring processing, and warping processing on the first predicted clear image to obtain a first blurred image and a second warped image, and performing optical flow prediction, blurring processing, and warping processing on the second predicted clear image to obtain a second blurred image and a first warped image; training the deblurring network based on minimizing the loss between the first original blurred video frame and the first blurred image, minimizing the loss between the second original blurred video frame and the second blurred image, minimizing the loss between the first predicted clear image and the first warped image, and minimizing the loss between the second predicted clear image and the second warped image.
[0011] According to an embodiment of the present invention, the deep convolutional neural network model includes a backbone network, a neck network, and a head network, and the method further includes: obtaining a training sample set; performing standardized preprocessing on the images in the training sample set, and inputting the preprocessed images into the backbone network for feature extraction to obtain feature maps of different scales; inputting the feature maps of different scales into the neck network, performing upsampling and feature fusion to obtain tensor data of different scales; inputting the tensor data of different scales into the head network, and performing gradient update based on the loss function and backpropagation to train the deep convolutional neural network model.
[0012] According to an embodiment of the present invention, obtaining a training sample set includes: obtaining a second sample set, where the second sample set includes actual captured images and synthetic images of target objects with different shapes under different backgrounds and different visions of the same background; selecting a partial sample set from the second sample set as the training sample set; randomly selecting a preset group of training samples from the training sample set, and performing random cropping, scaling, and stitching on each group of training samples to generate new training samples, and adding the new training samples to the training sample set.
[0013] According to an embodiment of the present invention, the neck network performs upsampling using the bilinear interpolation method. The neck network includes a feature pyramid, and the feature pyramid includes an FPN (Feature Pyramid Network) structure and a PAN (Pixel Aggregation Network) structure. Both the FPN structure and the PAN structure include three layers of structures: P / 4, P / 8, and P / 16.
[0014] According to an embodiment of the present invention, the above garbage recognition and classification method further includes: parsing and codec processing the images in the video stream based on Deep Streamer to obtain the original frame images; optimizing the deep convolutional neural network model based on TensorRT.
[0015] According to an embodiment of the present invention, tracking the target object according to the first feature vector map includes: predicting the second feature vector map of the target object using the Kalman filter model; performing GIOU matching on the first feature vector map and the second feature vector map; when they are matched, updating the parameters of the Kalman filter model and saving the position of the target object as the historical trajectory of the target object; when they are not matched, saving the position of the target object as the historical trajectory of the target object and continuing to track the target object. If the first feature vector map and the second feature vector map of the target object can be matched in multiple consecutive frames, then continue to track the target object.
[0016] According to an embodiment of the present invention, tracking the target object according to the first feature vector map further includes: if the first feature vector map and the second feature vector map of the target object are not matched in multiple consecutive frames, then stop tracking the target object and delete the historical trajectory of the target object; or, if the first feature vector map and the second feature vector map of the target object are not matched in multiple consecutive frames, then gradually adjust the GIOU matching threshold, and perform GIOU matching on the first feature vector map and the second feature vector map of the target object according to the adjusted GIOU matching threshold, and when it is determined that the first feature vector map and the second feature vector map of the target object are not matched in multiple consecutive frames, stop tracking the target object and delete the historical trajectory of the target object.
[0017] To achieve the above object, an embodiment of the second aspect of the present invention provides a computer-readable storage medium, on which a garbage recognition and classification program is stored. When the garbage recognition and classification program is executed by a processor, the above-mentioned garbage recognition and classification method is implemented.
[0018] The computer-readable storage medium according to the embodiment of the present invention can effectively reduce the false detection rate and missed detection rate of ground garbage recognition, improve the detection speed and versatility, and is less affected by the environment when implementing the above-mentioned garbage recognition and classification method.
[0019] To achieve the above object, a robot provided by an embodiment of the third aspect of the present invention includes a memory, a processor, and a garbage recognition and classification program stored on the memory and executable on the processor. When the processor executes the garbage recognition and classification program, the above-mentioned garbage recognition and classification method is implemented.
[0020] The robot according to the embodiment of the present invention can effectively reduce the false detection rate and missed detection rate of ground garbage recognition, improve the detection speed and versatility, and is less affected by the environment when implementing the above-mentioned garbage recognition and classification method.
[0021] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. Description of the Drawings
[0022] Figure 1 It is a flowchart of the garbage recognition and classification method according to an embodiment of the present invention;
[0023] Figure 2 It is a flowchart of the garbage recognition and classification method according to another embodiment of the present invention;
[0024] Figure 3 It is a flowchart of the deep convolutional neural network model according to an embodiment of the present invention;
[0025] Figure 4 It is a schematic diagram of convolution calculation according to an embodiment of the present invention;
[0026] Figure 5 It is a diagram of the improved Neck part model according to an embodiment of the present invention;
[0027] Figure 6 It is an effect diagram of the use of the algorithm according to an embodiment of the present invention;
[0028] Figure 7 It is a block diagram of the robot according to an embodiment of the present invention. Detailed Embodiments
[0029] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention, but should not be construed as limiting the present invention.
[0030] A garbage recognition and classification method, a computer-readable storage medium, and a robot proposed according to an embodiment of the present invention will be described below with reference to the accompanying drawings.
[0031] Currently, image detection algorithms based on deep learning can better detect the categories and positions of specific targets and have good detection effects on semi-occluded objects. However, deep learning not only requires a large amount of image data for training, but also requires the background in the images to be diverse. Moreover, there are significant performance differences in feature extraction of targets by neural networks adopted by different deep learning methods. The pixel ratio of recyclable garbage in an image is often small, and detecting it can be classified as small target detection. If the number of downsampling layers of the neural network is too large, deeper layers will lose the feature information of small targets, and the detection performance for small targets will be poor at this time. In addition, since most of the detection targets are small objects, when enhancing the image data during model training, it is necessary to appropriately scale targets of different sizes to expand the training sample size of small targets.
[0032] The existing technical solution is to scale the input image as required after the camera acquires it, and then process it according to traditional image algorithms or deep learning algorithms. When using traditional image algorithms, the road surface is mainly segmented, and the features of the objects on the road surface are calculated through traditional image processing algorithms such as connected components, gray histograms, and feature matching to find plastic bottles and aluminum cans, and the center point coordinates are given. In the robot system, the coordinate information is sent to the path planning node through Topic. When using deep learning algorithms, high-dimensional feature vectors are mainly extracted through convolutional neural networks to obtain the categories and corresponding center point coordinates of the objects on the road surface. In the robot system, the coordinate information is sent to the path planning node through Topic, where the convolutional neural network is obtained through continuous iterative training on the training data set.
[0033] The above method is prone to false detections when the target is small or in a semi-occluded state, and there is also a problem of missed detections caused by changes in the perspective, distance, and relative movement speed of the target due to the movement of the robot during the recycling journey. For this reason, the present invention proposes a garbage recognition and classification method, which can effectively reduce the false detection rate and missed detection rate of ground garbage recognition, improve the detection speed and versatility, and is less affected by the environment.
[0034] Figure 1 It is a flowchart of the garbage recognition and classification method according to an embodiment of the present invention;
[0035] Figure 2 Flow chart of the garbage recognition and classification method according to another embodiment of the present invention;
[0036] Figure 3 Flow chart of the deep convolutional neural network model according to an embodiment of the present invention;
[0037] Figure 4 Schematic diagram of convolutional calculation according to an embodiment of the present invention;
[0038] Figure 5 Improved Neck part model diagram according to an embodiment of the present invention;
[0039] Figure 6 Effect diagram of the use of the algorithm according to an embodiment of the present invention;
[0040] Figure 7 Block diagram of the robot according to an embodiment of the present invention.
[0041] As Figure 1 shown, the garbage recognition and classification method according to the embodiment of the present invention may include the following steps:
[0042] S1, according to the preset deblurring network, perform deblurring processing on the original frame image to obtain a target image.
[0043] In an embodiment of the present invention, the image in the video stream is parsed and codec processed based on Deep Streamer to obtain the original frame image. The algorithm of the present invention is based on Nvidia hardware, and Deep Streamer is used in the data input and inference parts to improve the image input speed. Among them, Deep Streamer is developed based on GStreamer, encapsulates functions such as data parsing, codec, and preprocessing into plugins, and combines them in a certain order to form a pipeline to complete operations such as video preprocessing, detection, tracking, and communication with the cloud. This algorithm mainly utilizes its image parsing and codec functions when reading in the video.
[0044] According to an embodiment of the present invention, as Figure 2 shown, the above method includes the following steps:
[0045] S11, obtain a first sample set, and each training sample in the first sample set includes a first original blurred video frame and a second original blurred video frame adjacent to the first original blurred video frame.
[0046] S12, input the first original blurred video frame and the second original blurred video frame into the deblurring network for deblurring processing respectively to obtain a first predicted clear image and a second predicted clear image.
[0047] S13. Perform optical flow prediction, blurring processing, and warping processing on the first predicted clear image to obtain a first blurred image and a second warped image, and perform optical flow prediction, blurring processing, and warping processing on the second predicted clear image to obtain a second blurred image and a first warped image; S14. Train the deblurring network based on the minimization of the loss between the first original blurred video frame and the first blurred image, the minimization of the loss between the second original blurred video frame and the second blurred image, the minimization of the loss between the first predicted clear image and the first warped image, and the minimization of the loss between the second predicted clear image and the second warped image.
[0048] Specifically, during the actual movement of the robot, the captured video images will be blurred due to the movement. Therefore, deblurring processing is required, that is, using the deblurring network to perform deblurring processing on the original frame images to restore the images and make them clearer. Among them, the deblurring network can use SRN (Scale Recurrent Network). When training the deblurring network, first, use the camera installed in the robot to take pictures as the robot moves, and a first sampling set can be obtained. The first sampling set can include multiple training samples, and each training sample includes two adjacent original images (the first original blurred video frame and the second original blurred video frame). During the shooting process, due to the movement of the robot, the captured images are blurred. Input the two consecutive blurred images into the deblurring network for blurring processing. The deblurring network will obtain two deblurred predicted outputs, that is, two clear images can be obtained, namely the first predicted clear image and the second predicted clear image. Then, input these two predicted outputs (two clear images) into the optical flow network (such as PWC-Net, a compact and effective optical flow estimation CNN model) to predict the optical flow, that is, the first optical flow and the second optical flow can be obtained. The blurring module performs blurring processing on the first predicted clear image and the second predicted clear image obtained after the first deblurring respectively according to the first optical flow and the second optical flow to obtain a first blurred image and a second blurred image. The image warping module performs warping processing on the first predicted clear image according to the first optical flow to obtain a second warped image, and performs warping processing on the second predicted clear image according to the second optical flow to obtain a first warped image. Finally, by minimizing the loss between the outputs of SRN and PWC-Net, the self-supervised training of the SRN deblurring network can be completed, and thus an end-to-end SRN network for removing dynamic blur can be obtained. Thus, an SNN network for removing dynamic blur is trained by a self-supervised method, reducing the impact of robot movement on the image.
[0049] S2. Input the target image into a preset deep convolutional neural network model for target object detection to obtain a first feature vector map of the target object.
[0050] According to an embodiment of the present invention, the deep convolutional neural network model includes a backbone network, a neck network, and a head network. As Figure 3 shown, the method includes the following steps:
[0051] S21. Obtain a training sample set.
[0052] According to an embodiment of the present invention, obtaining a training sample set includes: obtaining a second sample set, where the second sample set includes actual captured images and synthetic images of target objects with different shapes under different backgrounds and different visual effects of the same background; selecting a part of the sample set from the second sample set as the training sample set; randomly selecting a preset group of training samples from the training sample set, and performing random cropping, scaling, and stitching on each group of training samples to generate new training samples, and adding the new training samples to the training sample set.
[0053] Specifically, the training of deep learning depends on the data set. In addition to the image data captured in the actual working scenario, it is also necessary to collect a large number of small-sized objects with different shapes under complex backgrounds and different perspectives, such as plastic bottles and aluminum cans, and add promotional images containing plastic bottles or aluminum cans as training data to improve the robustness of the model. From the large number of collected data sets, randomly select some samples as the training sample set. During the training process of the model, randomly select pictures from the training sample set and use the mosaic algorithm to enhance the image data. For example, by randomly cropping, scaling, and stitching four images, the sample quantity of small-sized objects is expanded, and at the same time, a single image data containing multiple scenarios is generated to improve the generalization ability of the model. Finally, the generated single image is put back into the training sample set again to make the data set contain various garbage distribution situations as much as possible.
[0054] S22. Perform standardized preprocessing on the images in the training sample set, and input the preprocessed images into the backbone network for feature extraction to obtain feature maps of different scales.
[0055] Specifically, in order to improve the accuracy of the deep convolutional neural network model, for the training sample set in step S21, perform standardized preprocessing on the images therein. For example, by transforming the original data, the data is transformed into a range with a mean of 0 and a standard deviation of 1. Then input the processed images into the backbone network for feature extraction to obtain feature maps of different scales.
[0056] S23. Input the feature maps of different scales into the neck network, and perform upsampling and feature fusion to obtain tensor data of different scales.
[0057] According to an embodiment of the present invention, the Neck network performs upsampling using the bilinear interpolation method. The Neck network includes a feature pyramid, and the feature pyramid includes an FPN structure and a PAN structure. Both the FPN structure and the PAN structure include three layers of structures: P / 4, P / 8, and P / 16.
[0058] Specifically, the Neck part consists of a feature pyramid FPN and a path aggregation network PAN. The feature pyramid transfers the feature vectors of large targets at high levels to low levels, and the path aggregation network transfers the feature information of small targets at low levels to high levels, achieving the fusion and complementarity of high-level features and low-level features. However, most recyclable wastes such as plastic bottles and aluminum cans on the road surface are relatively small in size. If the number of layers of the feature pyramid is set too many, it will cause the loss of feature information of small targets at low levels. Therefore, structural modifications are made to the downsampling part of the feature pyramid, so that the output of the Neck module is the downsampling results of the original features by 4, 8, and 16 times, retaining the features of small targets at low levels of the feature pyramid to a great extent. The specific method is to delete a group of CBL calculations. Among them, a group of CBL calculations includes three calculation processes: convolution, normalization, and activation. The deleted convolution operation is as Figure 4 shown, and the deleted normalization operation is as follows: For a d-dimensional input x = (x (1) ,...x (d ), the batch normalization formula for each dimension is as follows:
[0059]
[0060] where k ∈ [1, d], i ∈ [1, m], and ε is a very small constant. And represent the mean and variance of each dimension, and their calculation formulas are as follows:
[0061]
[0062]
[0063] The output calculation formula is: where γ (k) and β (k) are learnable parameters during the training process. The deleted Leaky Relu activation operation is as shown in the formula:
[0064] where α is a hyperparameter set manually.
[0065] The feature vector finally output by the original Neck structure is 32 times smaller than the original image. After deleting a group of CBL operations, it becomes 16 times smaller than the original image, thus retaining the feature information of small targets. The feature pyramid of the improved Neck part is as Figure 5 shown.
[0066] In addition, in the upsampling part of the path aggregation network, the bilinear interpolation method is used to replace the original nearest neighbor interpolation method to weaken the interference caused by outliers in the training samples to feature transmission. The calculation process of the bilinear interpolation method is as follows: First, traverse and select four pixel points in the image, and record the four pixel points as Q 11 , Q 12 , Q 21 and Q 22 in the order from left to right and from top to bottom. Then, perform bilinear interpolation on these four pixel points, and finally iterate through to complete the interpolation and magnification of the original image. Specifically, first calculate two points in the x direction:
[0067]
[0068]
[0069] Then, perform interpolation calculation in the y direction to obtain the interpolated pixel points:
[0070]
[0071] Thus, through structural improvement of the downsampling part of the feature pyramid and using the bilinear interpolation method, the output of the Neck module is the result of downsampling the original features by 4, 8, and 16 times, greatly retaining the features of small targets in the low layer of the feature pyramid and effectively improving the learning ability and detection accuracy of the neural network for small targets.
[0072] S24, input the tensor data of different scales into the head network, and perform gradient update based on the loss function and backpropagation to train the deep convolutional neural network model.
[0073] That is to say, the head is the network for obtaining the output content of the network. The head network uses the features extracted in the above steps, and based on these features, makes predictions such as performing gradient update based on the loss function and backpropagation to train the deep convolutional neural network model.
[0074] Input the target image obtained through the above step S1 into the above deep convolutional neural network model for target object detection, and a first feature vector map of the target object can be obtained.
[0075] In one embodiment of the present invention, the deep convolutional neural network model is optimized based on TensorRT. That is, after Deep Streamer reads the video stream, TensorRT can perform inference calculations on the image to obtain the detection result. It should be noted that TensorRT is an inference optimizer that can parse network models such as Caffe and TensorFlow, and then map them one by one to the corresponding layers in TensorRT. The Yolov5 model trained using the Pytorch deep learning platform in this algorithm will be mapped into TensorRT to improve the calculation speed. During the mapping process, TensorRT merges operations such as convolution, bias, and activation layers, reducing the startup time of CUDA cores in the GPU and the time for exclusive write operations for the input / output of each layer. At the same time, when there is a high demand for real-time performance, the calibration algorithm provided by TensorRT can also be used to quantize the tensors in the network, converting the original 32-bit floating-point numbers into 16-bit floating-point numbers or 8-bit integers. Thus, the detection solution that uses deep streamer to obtain video stream data and then uses TensorRT for network quantization and inference calculations significantly improves the calculation speed of the detection algorithm.
[0076] S3. Track the target object according to the first feature vector map, and sort the target object according to the category and position of the target object.
[0077] According to an embodiment of the present invention, tracking the target object according to the first feature vector map includes: predicting the second feature vector map of the target object using the Kalman filter model; performing GIOU matching on the first feature vector map and the second feature vector map; when the matching is successful, updating the parameters of the Kalman filter model and saving the position of the target object as the historical trajectory of the target object; when the matching fails, saving the position of the target object as the historical trajectory of the target object and continuing to track the target object. If the first feature vector map and the second feature vector map of the target object can be matched in multiple consecutive frames, continue to track the target object.
[0078] Specifically, although the above-improved algorithm has improved the detection accuracy and ability for small targets, there are still a small number of missed detections and false detections. To address this issue, the sort algorithm can be used to track the detected targets after detection. Currently, the mostly used one is the deep sort algorithm. Among them, the application scenarios of the deep sort algorithm are mostly in blocks, shopping malls, etc. with relatively large vehicle or pedestrian flows. Most of the tracked objects have the problem of being partially or completely occluded by other objects in a short period of time. By training a separately trained feature extraction network to save the feature vectors of the tracked objects for a period of time, when the tracked object reappears after being partially or completely occluded and causing the tracking to be lost, the feature extraction network of deep sort will compare the current feature vector of the target with the saved feature vector, so as to achieve the re-identification of the tracked object. Deep sort can improve the accuracy by using the appearance feature vectors of the targets for re-identification and has good effects when there are many tracked objects and there are many occlusions between them. However, since deep sort requires an additional feature extraction network during calculation, the time complexity of the algorithm is relatively high, and the number of road surface wastes such as plastic bottles and aluminum cans is relatively small compared to the pedestrian and vehicle flows in shopping malls and streets. Using deep sort will cause waste of computing resources. The sort algorithm mainly lacks the feature extraction network compared with the deep sort algorithm. It directly performs tracking and re-identification by comparing the IOU of the tracking box and the detection box, has a fast calculation speed, and is suitable for tracking after road surface waste detection. By optimizing the sort algorithm, formulating strategies for sort, eliminating misdetected objects that only appear in a few frames and retaining the historical trajectories of all tracked objects for re-identification when the target reappears after being briefly occluded.
[0079] Considering that during the process of the robot moving forward, affected by the environment and the moving speed of the robot, the position of the tracking target relative to the robot may change drastically in a short period of time. At this time, using the traditional IOU algorithm, the tracking prediction box will have no intersection with the detection box. Therefore, this algorithm replaces the IOU algorithm with the GIOU algorithm to evaluate the matching degree between the tracking prediction box and the detection box, that is, uses GIOU to replace the IOU in the original sort algorithm to judge the confidence of the Kalman filter prediction result. Among them, the calculation method of GIOU is shown in the following formula:
[0080]
[0081] Among them, A c represents the minimum closed interval of the tracking prediction box and the detection box, and U represents the union part of the tracking prediction box and the detection box.
[0082] That is to say, the target box containing position and size information is predicted for the tracking target by Kalman filtering and GIOU matching is performed with the detected target box. Considering that the tracked object may be occluded, when the predicted box of the tracked object suddenly fails to match the detected box in terms of GIOU at a certain moment, the algorithm in this paper will give a frame number threshold th. If GIOU can be rematched within the next frame number threshold, it is considered that the target re-identification is successful and tracking continues, otherwise it is considered that the target is completely lost and its saved historical trajectory is deleted.
[0083] When the detected target box does not match the predicted one, it is considered that a new tracking object may be detected at this time. The position information is retained as the historical trajectory and then added to the tracking candidate queue. If the target can be continuously detected and matched with GIOU in the following consecutive frames, it is moved from the tracking candidate queue to the tracking queue at this time, otherwise it is deleted, so that misdetected targets can be eliminated; when the detected target box matches the predicted one, the Kalman filter parameters are updated at this time and the position information of the tracked object is retained as the historical trajectory. According to an embodiment of the present invention, when tracking the target object according to the first feature vector map, it further includes: if the first feature vector map and the second feature vector map of the target object do not match in consecutive frames, stop tracking the target object and delete the historical trajectory of the target object; or, if the first feature vector map and the second feature vector map of the target object do not match in consecutive frames, gradually adjust the GIOU matching threshold, and perform GIOU matching on the first feature vector map and the second feature vector map of the target object according to the adjusted GIOU matching threshold, and when it is determined that the first feature vector map and the second feature vector map of the target object do not match in consecutive frames, stop tracking the target object and delete the historical trajectory of the target object.
[0084] That is to say, when tracking the target object, the tracking prediction and the detection result are matched in a real-time search range. If the number of mismatches between the tracking prediction and the detection result exceeds a set threshold th, the mismatched target is set to a semi-lost state, and a smaller GIOU domain value is set for the next frame of tracking to enlarge the search area of each semi-lost target. If the target still cannot be found after continuously enlarging the search range within the set number of frames, the target is set to a completely lost state at this time and its historical trajectory is deleted.
[0085] Thus, by replacing IOU in the original sort algorithm with GIOU to evaluate the matching degree between the tracking prediction and the detection, the problem that the tracking object has a large displacement in a short time due to the change of the machine re-identification traveling speed or being affected by the environment, and the tracking prediction result has no intersection with the detection result, resulting in tracking failure, is effectively solved.
[0086] For example,Figure 6 As shown, in the left figure, the box ① is the result output by the deep convolutional neural network model detection algorithm, representing the position where the detection target is located. The box ② is the result output by the Kalman prediction in the tracking module. The right figure is the final output result of the tracking module, where the points inside correspond to the historical trajectories of each target, and the ID represents the current tracking target. For an object, if the detection algorithm can continuously detect it within the specified minimum number of frames, it is regarded as a tracking target and initialized. First, obtain the maximum ID value. If there is none, initialize a maximum ID value equal to 0. After adding 1 to the maximum ID value, assign it to the target and update the maximum ID value. During tracking, the ID of the target remains unchanged. For Figure 6 the object in the box ① in Figure 6 , because it can be detected under the set frame number threshold, it will be regarded as a target to be tracked. The position of the previous frame is predicted through the Kalman filter of the tracking module, and the prediction result of the current frame (
[0087] the box ② shown). It is not difficult to see that for the results in detection results 2, 3 and Kalman prediction results 2, 3, there is a certain overlap between them. At this time, it is feasible to calculate the IOU by the original sort algorithm. However, for detection result 1 and the corresponding Kalman prediction result 1, there is no overlap between the detection box and the prediction box. At this time, the original sort algorithm fails, resulting in tracking failure. However, we can see that the prediction box can return to the position of the detection box after translation within a certain range. Therefore, there are certain defects in using IOU for matching evaluation. The algorithm in this paper uses GIOU to judge the matching degree. By setting the range parameter, even if the prediction box and the detection box do not overlap at all, corresponding evaluation values can be obtained. As long as the evaluation value exceeds the threshold, it is considered that the tracking is successful, and then the final tracking result is output by combining the detection result and the Kalman prediction result.
[0088] Corresponding to the above embodiments, the present invention also proposes a computer-readable storage medium.
[0089] The computer-readable storage medium of the present invention stores a garbage recognition and classification program thereon. When the garbage recognition and classification program is executed by a processor, the above-mentioned garbage recognition and classification method is implemented.
[0090] The computer-readable storage medium of the embodiments of the present invention can effectively reduce the false detection rate and missed detection rate of ground garbage recognition, improve the detection speed and versatility, and is less affected by the environment by executing the above garbage recognition and classification method.
[0091] Corresponding to the above embodiments, the present invention also provides a robot.
[0092] As Figure 7 shown, the robot 200 of the present invention may include: a memory 210, a processor 220, and a garbage recognition and classification program stored on the memory 210 and executable on the processor 220. When the processor 220 executes the garbage recognition and classification program, the above garbage recognition and classification method is implemented.
[0093] The robot of the present invention can effectively reduce the false detection rate and missed detection rate of ground garbage recognition, improve the detection speed and versatility, and is less affected by the environment by executing the above garbage recognition and classification method.
[0094] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.
[0095] It should be understood that each part of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0096] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0097] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0098] In the present invention, unless otherwise clearly defined and limited, terms such as "installed", "connected", "connected to", "fixed", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements or the interaction relationship between two elements, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0099] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A garbage recognition and classification method, characterized in that, the method includes: Performing deblurring processing on the original frame image according to a preset deblurring network to obtain a target image; Inputting the target image into a preset deep convolutional neural network model for target object detection to obtain a first feature vector map of the target object; Tracking the target object according to the first feature vector map, and sorting the target object according to the category and position of the target object; wherein, Obtaining a first sample set, each training sample in the first sample set includes a first original blurred video frame and a second original blurred video frame adjacent to the first original blurred video frame; Inputting the first original blurred video frame and the second original blurred video frame into the deblurring network respectively for deblurring processing to obtain a first predicted clear image and a second predicted clear image; Performing optical flow prediction, blurring processing and warping processing on the first predicted clear image to obtain a first blurred image and a second warped image, and performing optical flow prediction, blurring processing and warping processing on the second predicted clear image to obtain a second blurred image and a first warped image; Training the deblurring network based on minimizing the loss between the first original blurred video frame and the first blurred image, minimizing the loss between the second original blurred video frame and the second blurred image, minimizing the loss between the first predicted clear image and the first warped image, and minimizing the loss between the second predicted clear image and the second warped image.
2. The method according to claim 1, characterized in that, the deep convolutional neural network model includes a backbone network, a neck network and a head network, and the method further includes: Obtaining a training sample set; Performing standardized preprocessing on the images in the training sample set, and inputting the preprocessed images into the backbone network for feature extraction to obtain feature maps of different scales; Inputting the feature maps of different scales into the neck network, performing upsampling and feature fusion to obtain tensor data of different scales; Inputting the tensor data of different scales into the head network, and performing gradient update based on the loss function and backpropagation to train the deep convolutional neural network model.
3. The method according to claim 2, characterized in that, the obtaining of the training sample set includes: Obtaining a second sample set, the second sample set includes actual captured images and synthetic images of target objects with different shapes under different backgrounds and different visions of the same background; Selecting a part of the sample set from the second sample set as the training sample set; Randomly selecting a preset number of groups of training samples from the training sample set, randomly cropping, scaling and splicing each group of training samples to generate new training samples, and adding the new training samples to the training sample set.
4. The method according to claim 2, characterized in that, The neck network uses bilinear interpolation for upsampling. The neck network includes a feature pyramid, and the feature pyramid includes an FPN structure and a PAN structure. Both the FPN structure and the PAN structure include three layers: P / 4, P / 8, and P / 16.
5. The method according to claim 1, wherein, the method further includes: parsing and encoding / decoding the images in the video stream based on Deep Streamer to obtain the original frame images; optimizing the deep convolutional neural network model based on TensorRT.
6. The method according to claim 1, wherein, tracking the target object according to the first feature vector map includes: predicting a second feature vector map of the target object using a Kalman filter model; performing GIOU matching on the first feature vector map and the second feature vector map; when they are matched, updating the parameters of the Kalman filter model and saving the position of the target object as the historical trajectory of the target object; when they are not matched, saving the position of the target object as the historical trajectory of the target object and continuing to track the target object. If the first feature vector map and the second feature vector map of the target object can be matched in multiple consecutive frames, then continue to track the target object.
7. The method according to claim 6, wherein, tracking the target object according to the first feature vector map further includes: if the first feature vector map and the second feature vector map of the target object are not matched in multiple consecutive frames, then stop tracking the target object and delete the historical trajectory of the target object; or, if the first feature vector map and the second feature vector map of the target object are not matched in multiple consecutive frames, then gradually adjust the GIOU matching threshold, perform GIOU matching on the first feature vector map and the second feature vector map of the target object according to the adjusted GIOU matching threshold, and when it is determined that the first feature vector map and the second feature vector map of the target object are not matched in multiple consecutive frames, stop tracking the target object and delete the historical trajectory of the target object.
8. A computer-readable storage medium, wherein, a garbage recognition and classification program is stored thereon, and when the garbage recognition and classification program is executed by a processor, it implements the garbage recognition and classification method according to any one of claims 1-7.
9. A robot, wherein, it includes: a memory, a processor, and a garbage recognition and classification program stored on the memory and executable on the processor. When the processor executes the program, it implements the garbage recognition and classification method according to any one of claims 1-7.
Citation Information
Patent Citations
Construction method of automatic garbage classification model and garbage sorting method and system
CN113469264A