A Visual Recognition Method for Waste Sorting Robots Based on Deep Learning

By adopting deep learning methods and synchronous calibration technology in garbage sorting robots, the synchronous fusion of 2D and 3D camera data is achieved, solving the problem of difficulty in data synchronization in the existing technology, and improving the efficiency and accuracy of robot crawling.

CN114049557BActive Publication Date: 2025-05-27CHINA TIANYING +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111323743.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-10
Publication Date
2025-05-27
Estimated Expiration
2041-11-10

AI Technical Summary

Technical Problem

The prior art is difficult to achieve synchronous integration of 2D cameras and 3D camera data, resulting in low robot grabbing efficiency and inaccurate positioning accuracy.

Method used

The visual recognition method of garbage sorting robot based on deep learning is adopted. The 2D surface array camera and the 3D line array camera are synchronously calibrated under the same world coordinate system, and the YOLOv4 object detection model is established, the encoder pulse value is obtained in real time, the 2D and 3D image data are collected simultaneously, and the data synchronous fusion is realized through the data analysis and processing module.

Benefits of technology

The synchronous fusion of 2D cameras and 3D camera data is realized, which improves detection speed and recognition accuracy, and enhances the accuracy and robustness of real-time capture of robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114049557B_ABST
    Figure CN114049557B_ABST
Patent Text Reader

Abstract

The present invention discloses a visual recognition method for a garbage sorting robot based on deep learning, wherein a 2D array camera and a 3D linear array camera complete synchronous calibration in the same world coordinate system; a YOLOv4 target detection model is established; multithreading is started, encoder pulse values ​​are obtained in real time, and image data is collected by the 2D camera and the 3D camera at the same time; the current contour image data in the memory is obtained to obtain the actual height information and width information of the target object, and the information is sent to the garbage sorting robot together with the rectangular frame center coordinates, rectangular frame width, rectangular frame height and target type obtained by the YOLOv4 target detection model, so that the robot can realize real-time online grabbing and classification. The present invention synchronously collects images of 2D cameras and 3D cameras, and the encoder pulse values ​​correspond the target image detected by the YOLOv4 detection model to the 3D scanning data, and the target height and width information are obtained. This method has fast detection speed and high recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a visual recognition method, in particular to a visual recognition method for a waste sorting robot based on deep learning, belonging to the field of visual recognition of waste sorting robots. Background Art

[0002] Robots grasp materials mainly through visual recognition of front-end images. For two-dimensional objects, the target materials can be collected by a 2D camera. However, for three-dimensional solid materials, due to reasons such as inconsistent shapes and sizes, a 3D camera must be used. In traditional 3D cameras, the height information of an object is mainly obtained by installing a laser displacement sensor at a fixed position to obtain the height information of the object. Although the linear performance of the laser displacement sensor is very good and the accuracy is high, during the conveying process, due to many unstable factors during the movement of the object, it often leads to low grasping efficiency of the robot, and even worse, it may cause the robot to perform empty grasping and collisions. On this basis, usually, a method of fusing a 2D camera and a 3D camera is used for recognition. However, these two cameras must ensure that the field of view area is the same when the object passes through the two cameras. If the correspondence is inaccurate, it will lead to inaccurate positioning accuracy of the robot. The existing technology cannot well fuse the 2D camera and the 3D camera to achieve complete data synchronization. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a visual recognition method for a waste sorting robot based on deep learning to achieve synchronous fusion of 2D camera and 3D camera data.

[0004] To solve the above technical problem, the technical solution adopted by the present invention is:

[0005] A visual recognition method for a waste sorting robot based on deep learning, characterized by comprising the following steps:

[0006] Step 1: Synchronously calibrate a 2D area array camera and a 3D line array camera in the same world coordinate system;

[0007] Step 2: Establish a YOLOv4 target detection model;

[0008] Step 3: Start multi-threading, and obtain the encoder pulse value in real time. The 2D camera and the 3D camera simultaneously collect image data. The 2D camera collects the RGB image to be recognized and detected, and the 3D camera collects the contour and height information of the target image;

[0009] Step 4: Input the RGB image into the YOLO v4 target detection model to obtain the center coordinates of the rectangular frame, the width of the rectangular frame, the height of the rectangular frame, and the target category;

[0010] Step 5: The 3D camera cyclically acquires single contour images and sequentially stores the data in a memory with a specified size. The memory occupancy size of the single contour image is determined by the number of encoder pulses;

[0011] Step 6: Obtain the current contour image data in the memory, get the actual height information and width information of the target object, and send them together with the center coordinates of the rectangular box, the width of the rectangular box, the height of the rectangular box, and the target category obtained by the YOLOv4 target detection model to the garbage sorting robot, enabling the robot to achieve real-time online grasping and classification.

[0012] Further, the calibration method of the 2D area array camera in Step 1 is as follows:

[0013] Use the camera to acquire checkerboard calibration plates at different positions and rotation angles;

[0014] Perform internal parameter calibration using Matlab to obtain the internal parameter matrix and distortion coefficients;

[0015] Determine a world coordinate system on the checkerboard calibration plate. Use the 4-point calibration method in PNP to determine the pixel coordinates of 4 corner points on this calibration plate and the corresponding world coordinates of the 4 corner points in the determined world coordinate system. Use the solvePNP operator to perform external parameter calibration on the 4 corner points to obtain the rotation matrix and translation matrix of the external parameters, and complete the external parameter calibration.

[0016] Further, the calculation formula for converting from pixel coordinates to world coordinates is as follows:

[0017]

[0018] where: u and v are respectively the pixel abscissa and pixel ordinate in the pixel coordinate system; x w , y w , z w are respectively the abscissa, ordinate, and vertical coordinate in the world coordinate system; R is the rotation matrix; T is the translation matrix; u 0 , v 0 , f x , f y are the internal parameters of the camera, that is, u 0 and v 0 are respectively the abscissa of the image center and the ordinate of the image center, f x and f y are respectively the horizontal equivalent focal length and the vertical equivalent focal length; s is the camera coordinate in the camera coordinate system.

[0019] Further, the calibration method of the 3D line array camera is as follows:

[0020] Use 3D laser to irradiate on the calibration plate so that the laser is parallel to the x-axis of the fixed world coordinate system;

[0021] Turn off the laser, increase the exposure time, and capture a picture with clear corner points;

[0022] After the picture is captured, turn on the laser, move the conveyor belt forward a certain distance so that the laser falls on another checkerboard, and make the laser parallel to the x-axis;

[0023] Turn off the laser, increase the exposure time, and capture another picture with clear corner points;

[0024] Finally, perform calibration. After saving the data, the calibration of the 3D camera can be completed.

[0025] Further, the specific steps of step two are as follows: Use a 2D camera to capture RGB images, the annotator annotates the captured picture information, and use the YOLOv4 object detection model for model training to generate the final YOLOv4 object detection model.

[0026] Further, the specific steps of step four are as follows:

[0027] Input the captured RGB image into the YOLOv4 object detection model with an input size of 608*608 to obtain a list of all position boxes Bounding Box where target materials exist in the image. After filtering by the non-maximum suppression NMS algorithm, obtain the coordinate position information of the final target garbage points to be retained;

[0028] The non-maximum suppression (NMS) algorithm is as follows:

[0029]

[0030] Among them, Si represents the score of each border, M represents the box with the highest current score, bi represents a certain box among the remaining boxes, Nt is the set NMS threshold, and IOU is the overlapping area ratio of two recognition boxes;

[0031] When the YOLOv4 object detection model detects the target value, obtain the current encoder pulse value in real time. The obtained pulse value is the pulse value of the center coordinate of the target material. Use the height of the rectangle detected by the YOLOv4 object detection model as the movement direction of the conveyor belt, and send the current encoder pulse value, the detected target center coordinate, and the width and height information of the rectangle to the data analysis and processing module.

[0032] Further, the specific steps of step five are as follows:

[0033] The data analysis and processing module calculates the starting position of the target material in the 3D storage memory. The specific calculation formula is as follows:

[0034]

[0035] Among them, A represents the current encoder pulse value; B represents the initial start pulse value; 1216 indicates that there are 1216 scan points on one contour; 3 represents a total of 3 coordinate values for x, y, and z; 4 indicates that each of the x, y, and z values of each scan point occupies 4 bytes;

[0036] Cyclically acquire the single contour image of the 3D line array camera and sequentially store the data in the memory with a specified size. The specific calculation formula for the storage byte size is as follows:

[0037]

[0038] Among them, 1.6mm represents the world coordinate distance between two scanned contours; 1216 indicates that there are 1216 scan points on one contour; 3 represents a total of 3 coordinate values for x / y / z; 4 indicates that each of the x / y / z values of each scan point occupies 4 bytes;

[0039] According to the starting position of the target material in the 3D storage memory and the storage byte size of the target material, finally comprehensively calculate and obtain the final end position of the target material in the 3D storage memory. The specific calculation formula is as follows:

[0040] Final end position = memory initial position + storage byte size

[0041] The data analysis and processing module realizes the reading of the corresponding target 3D scan memory data, and thus can complete the reading of the x, y, and z values of the target material in the three-dimensional world in the 3D storage memory data.

[0042] Furthermore, the specific source of the world coordinate distance of 1.6mm between two contours is as follows:

[0043] The 3D camera scans one contour with 5 pulses. There are 1216 points on one contour, and each point contains x / y / z values, with each value occupying 4 bytes. The encoder pulse value for one circle is counted as 1000, and the distance traveled by the conveyor belt for one circle is 320mm, that is, the distance for one pulse is 0.32mm. Every 5 pulses, that is, 5 * 0.32 = 1.6mm, scans one contour. Therefore, the distance of the world coordinate corresponding to the distance between two contours is 1.6mm.

[0044] Furthermore, when the number of 3D acquisition contours reaches 50,000, the 3D data starts to store data again from the starting position of the memory space, overwriting the previously stored data, and continues to store sequentially backward for infinite loop storage of data. When the program stops running, the opened memory space is released to prevent memory overflow or leakage.

[0045] Furthermore, the specific content of step six is as follows:

[0046] Obtain the 3D memory data data. Use the x and y values as the row and column pixel coordinates corresponding to Mat in OpenCV. Normalize the z value from 0 to 255, and use the normalized value as the grayscale value corresponding to the x and y pixel coordinates in Mat. The height value is obtained according to the pixel value of the target center coordinate, and the corresponding grayscale value in Mat is extracted. The actual height information of the target object can be obtained through inverse normalization of the grayscale value;

[0047] Use algorithms such as OpenCV thresholding segmentation, opening and closing operations, and minimum contour processing to process the Mat image, and extract the width information of the target object;

[0048] Finally, send the center coordinate position of the target object obtained from 2D, the target type, and the target height information and width information obtained from 3D to the robot as a whole, so that the robot can achieve real-time online grasping and classification.

[0049] Compared with the prior art, the present invention has the following advantages and effects:

[0050] 1. By synchronously collecting 2D camera and 3D camera images, the present invention corresponds the target image detected by the YOLOv4 detection model and the 3D scan data through the encoder pulse value, and obtains the target height and width information. This method has a fast detection speed and high recognition accuracy;

[0051] 2. The present invention forms a cyclic 3D memory data storage system during the data reading process, which not only does not cause memory overflow, but also can combine the memory space with the encoder to read the corresponding memory space data data in real time;

[0052] 3. The present invention applies the deep learning YOLO target detection model to the 2D camera and 3D camera fusion technology, and uses OpenCV to process the memory data data, so that the robot has higher real-time grasping accuracy and stronger robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a flowchart of a method for visual recognition of a garbage sorting robot based on deep learning according to the present invention.

[0054] Figure 2 is a schematic diagram of the camera synchronization calibration process according to the present invention.

[0055] Figure 3 is a schematic diagram of the calibration board before 3D camera calibration according to the present invention.

[0056] Figure 4 is a schematic diagram of the calibration board after 3D camera calibration according to the present invention.

[0057] Figure 5The present invention processes OpenCV data and sends the processing results to the robot to realize the grasping flow chart.

[0058] Figure 6 It is a schematic diagram of the target width AB obtained by the memory data data corresponding to the OpenCV processing target of the present invention. DETAILED DESCRIPTION

[0059] In order to elaborate on the technical scheme adopted by the present invention to achieve the predetermined technical purpose, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only partial embodiments of the present invention, rather than all embodiments, and the technical means or technical features in the embodiments of the present invention can be replaced without paying creative work. The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0060] like Figure 1 As shown, a garbage sorting robot visual recognition method based on deep learning of the present invention comprises the following steps:

[0061] Step 1: The 2D area array camera and 3D line array camera are synchronously calibrated in the same world coordinate system.

[0062] like Figure 2 As shown, the calibration method of the 2D array camera is:

[0063] Use the camera to collect checkerboard calibration plates at different positions and rotation angles;

[0064] Matlab is used to perform internal parameter calibration to obtain the internal parameter matrix and distortion coefficient;

[0065] Determine a world coordinate system on the checkerboard calibration plate, use the 4-point calibration method in PNP to determine the pixel coordinates of the 4 corner points on the calibration plate and the world coordinates corresponding to the 4 corner points in the determined world coordinate system, use the solvePNP operator to perform extrinsic parameter calibration on the 4 corner points, obtain the rotation matrix and translation matrix of the extrinsic parameters, and complete the extrinsic parameter calibration.

[0066] Furthermore, the calculation formula for converting from pixel coordinates to world coordinates is as follows:

[0067]

[0068] Where: u and v are the horizontal and vertical coordinates of the pixel in the pixel coordinate system respectively; x w ,y w 、z w are the horizontal, vertical and vertical coordinates in the world coordinate system respectively; R is the rotation matrix; T is the translation matrix; u 0 、v0 , f x , f y is the camera internal parameter, that is, u 0 and v 0 are the abscissa and ordinate of the image center respectively, and f x and f y are the horizontal equivalent focal length and vertical equivalent focal length respectively; s is the camera coordinate in the camera coordinate system.

[0069] The calibration method of the 3D line array camera is as follows:

[0070] Use 3D laser to irradiate on the calibration board so that the laser is parallel to the x-axis of the fixed world coordinate system;

[0071] As Figure 3 shown, turn off the laser, increase the exposure time, and collect a picture with clear corner points;

[0072] After the picture is collected, turn on the laser, then move the conveyor belt forward a certain distance so that the laser falls on another checkerboard grid and the laser is parallel to the x-axis;

[0073] As Figure 4 shown, turn off the laser, increase the exposure time, and collect another picture with clear corner points;

[0074] Finally, perform calibration. After saving the data, the calibration of the 3D camera can be completed. The calibration process of the 4 corner points is the same as that of the 2D calibration process.

[0075] Step 2: Establish a YOLOv4 object detection model; use a 2D camera to collect RGB images, and the annotator annotates the collected picture information. Then use the YOLOv4 object detection model for model training to generate the final YOLOv4 object detection model.

[0076] Step 3: Start multi-threading to obtain the encoder pulse value in real time. The 2D camera and 3D camera perform image data collection simultaneously. The 2D camera collects the RGB image to be recognized and detected, and the 3D camera collects the contour and height information of the target image;

[0077] Step 4: Input the RGB image into the YOLO v4 object detection model to obtain the center coordinates of the rectangle box, the width of the rectangle box, the height of the rectangle box, and the target category.

[0078] Input the collected RGB image into the YOLOv4 object detection model with an input size of 608*608 to obtain a list of all position boxes Bounding Box of the target materials existing in the image. After filtering by the non-maximum suppression NMS algorithm, obtain the coordinate position information of the final target garbage points to be retained;

[0079] The non-maximum suppression (NMS) algorithm is as follows:

[0080]

[0081] Among them, Si represents the score of each bounding box, M represents the box with the highest current score, bi represents a certain box among the remaining boxes, Nt is the set NMS threshold, and IOU is the overlapping area ratio of two recognition boxes;

[0082] When the YOLOv4 object detection model detects the target value, the current encoder pulse value is obtained in real time. The obtained pulse value is the pulse value of the center coordinates of the target material. The height of the rectangular box detected by the YOLOv4 object detection model is used as the movement direction of the conveyor belt, and the current encoder pulse value, the detected target center coordinates, and the width and height information of the rectangular box are sent to the data analysis and processing module together.

[0083] Step Five: The 3D camera cyclically collects single contour images and stores the data in memory of a specified size in sequence. The memory occupancy size of the single contour image is determined by the number of encoder pulses.

[0084] The data analysis and processing module calculates the starting position of the target material in the 3D storage memory. The specific calculation formula is as follows:

[0085]

[0086] Among them, A represents the current encoder pulse value; B represents the initial start pulse value; 1216 represents that there are 1216 scanning points on one contour; 3 represents a total of 3 coordinate values of x, y, and z; 4 represents that each of the x, y, and z values of each scanning point occupies 4 bytes;

[0087] Cyclically collect single contour images of the 3D line array camera and store the data in memory of a specified size in sequence. The specific calculation formula for the storage byte size is as follows:

[0088]

[0089] Among them, 1.6mm represents the world coordinate distance between two scanned contours; 1216 represents that there are 1216 scanning points on one contour; 3 represents a total of 3 coordinate values of x / y / z; 4 represents that each of the x / y / z values of each scanning point occupies 4 bytes;

[0090] According to the starting position of the target material in the 3D storage memory and the storage byte size of the target material, the final end position of the target material in the 3D storage memory is finally obtained through comprehensive conversion. The specific calculation formula is as follows:

[0091] Final end position = memory initial position + storage byte size

[0092] In the data analysis and processing module, reading the corresponding target 3D scan memory data can complete reading the x, y, and z values of the target material in the 3D storage memory data in the three-dimensional world.

[0093] The world coordinate distance of 1.6 mm between the two contours is specifically sourced as follows:

[0094] The 3D camera scans one contour with 5 pulses. There are 1216 points on one contour, and each point contains x / y / z values, with each value occupying 4 bytes. The encoder pulse value for one full circle is counted as 1000, and the distance traveled by the conveyor belt in one full circle is 320 mm, that is, the distance for one pulse is 0.32 mm. For every 5 pulses, which is 5 * 0.32 = 1.6 mm, one contour is scanned. Therefore, the distance between the two contours corresponding to the world coordinates is 1.6 mm.

[0095] In the data reading process of the present invention, a cyclic 3D memory data storage system is formed. When the number of 3D acquisition contours reaches 50,000, the 3D data starts to be stored again from the starting position of the memory space, overwriting the previously stored data, and continues to be stored sequentially backward, performing infinite cyclic storage of the data. When the program stops running, the allocated memory space is released to prevent memory overflow or leakage.

[0096] Step six: Obtain the contour image data in the memory, obtain the actual height information and width information of the target object, and send them together with the center coordinates of the rectangular box, the width of the rectangular box, the height of the rectangular box, and the target category obtained by the YOLOv4 target detection model to the garbage sorting robot, enabling the robot to achieve real-time online grasping and classification.

[0097] As Figure 5 shown, obtain the 3D memory data data. Take the x and y values as the row and column pixel coordinates corresponding to Mat in OpenCV, perform normalization processing on the z value from 0 to 255, and take the normalized value as the grayscale value corresponding to the x and y pixel coordinates in Mat. The height value is obtained according to the pixel value of the target center coordinate, and the corresponding grayscale value in Mat is extracted. Through inverse normalization processing of the grayscale value, the actual height information of the target object can be obtained;

[0098] Use algorithms such as OpenCV thresholding segmentation, opening and closing operations, and minimum contour processing to process the Mat image, and extract the width information of the target object; As Figure 6 shown, for extracting the contour information of the corresponding garbage target, the length of points AB is the width information of this target.

[0099] Finally, send the center coordinate position of the target object obtained by 2D, the target category, and the target height information and width information obtained by 3D to the robot as a whole, enabling the robot to achieve real-time online grasping and classification.

[0100] In the present invention, by synchronously collecting images of a 2D camera and a 3D camera, the encoder pulse value corresponds the target image detected by the YOLOv4 detection model with the 3D scan data, and the height and width information of the target are obtained. This method has a fast detection speed and high recognition accuracy. In the process of data reading, the present invention forms a cyclic 3D memory data storage system, which not only will not cause memory overflow, but also can combine the memory space with the encoder to read the data data in the corresponding memory space in real time. The present invention applies the deep learning YOLO object detection model to the 2D camera and 3D camera fusion technology, and uses OpenCV to process the memory data data, so that the robot has higher real-time grasping accuracy and stronger robustness.

[0101] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the equivalent embodiments by using the above-disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the technical solution content of the present invention and is based on the technical essence of the present invention, any simple modification, equivalent replacement and improvement of the above embodiments still fall within the protection scope of the technical solution of the present invention.

Claims

1. A visual recognition method for a waste sorting robot based on deep learning, characterized in that it includes the following steps: Step 1: Synchronously calibrate the 2D area array camera and the 3D line array camera in the same world coordinate system; Step 2: Establish a YOLOv4 object detection model; Step 3: Start multi-threading, and obtain the encoder pulse value in real time. The 2D camera and the 3D camera simultaneously collect image data. The 2D camera collects the RGB image to be recognized and detected, and the 3D camera collects the contour and height information of the target image; Step 4: Input the RGB image into the YOLO v4 object detection model to obtain the center coordinates of the rectangular box, the width of the rectangular box, the height of the rectangular box, and the target category; Step 5: The 3D camera circularly collects single contour images, and stores the data in the memory with a specified size in sequence. The memory occupancy size of the single contour image is determined by the number of encoder pulses; The specific content of Step 5 is as follows: The data analysis and processing module calculates the starting position of the target material in the 3D storage memory. The specific calculation formula is as follows: where A represents the current encoder pulse value; B represents the initial start pulse value; 1216 means there are 1216 scan points on one contour; 3 represents 3 coordinate values of x, y, and z; 4 means that the x, y, and z values of each scan point each occupy 4 bytes; Circularly collect single contour images of the 3D line array camera, and store the data in the memory with a specified size in sequence. The specific calculation formula for the stored byte size is as follows: where 1.6mm represents the world coordinate distance between two scanned contours; 1216 means there are 1216 scan points on one contour; 3 represents 3 coordinate values of x / y / z; 4 means that the x / y / z values of each scan point each occupy 4 bytes; According to the starting position of the target material in the 3D storage memory and the storage byte size of the target material, finally comprehensively convert to obtain the final end position of the target material in the 3D storage memory. The specific calculation formula is as follows: Final end position = memory initial position + storage byte size The data analysis and processing module realizes the reading of the corresponding target 3D scan memory data, and then the x, y, and z values of the target material in the three-dimensional world in the 3D storage memory data can be read; Step 6: Obtain the current contour image data in the memory, obtain the actual height information and width information of the target object, and send them together with the center coordinates of the rectangular box, the width of the rectangular box, the height of the rectangular box, and the target category obtained by the YOLOv4 object detection model to the waste sorting robot, so that the robot can achieve real-time online grasping and classification.

2. The visual recognition method for a waste sorting robot based on deep learning according to claim 1, characterized in that: The calibration method of the 2D area array camera in Step 1 is as follows: Use the camera to collect checkerboard calibration plates at different positions and rotation angles; Perform internal parameter calibration using Matlab to obtain the internal parameter matrix and distortion coefficients; Determine a world coordinate system on the checkerboard calibration board. Use the 4-point calibration method in PNP to determine the pixel coordinates of 4 corner points on this calibration board and the corresponding world coordinates of the 4 corner points in the determined world coordinate system. Use the solvePNP operator to calibrate the external parameters of the 4 corner points, obtain the rotation matrix and translation matrix of the external parameters, and complete the external parameter calibration.

3. A vision recognition method for a waste sorting robot based on deep learning according to claim 1, characterized in that: The calculation formula for converting from pixel coordinates to world coordinates is as follows: Where: u and v are the pixel abscissa and pixel ordinate in the pixel coordinate system; x w , y w , z w are the abscissa, ordinate and vertical coordinate in the world coordinate system respectively; R is the rotation matrix; T is the translation matrix; u 0 , v 0 , f x , f y are the internal parameters of the camera, that is, u 0 and v 0 are the abscissa of the image center and the ordinate of the image center respectively, and f x and f y are the horizontal equivalent focal length and the vertical equivalent focal length respectively; s is the camera coordinate in the camera coordinate system.

4. A vision recognition method for a waste sorting robot based on deep learning according to claim 2, characterized in that: The calibration method of the 3D line array camera is: Use 3D laser to irradiate on the calibration board, and make the laser parallel to the x-axis of the fixed world coordinate system; Turn off the laser, increase the exposure time, and collect a picture with clear corner points; After the picture is collected, turn on the laser, and then move the conveyor belt forward a certain distance so that the laser falls on another checkerboard and the laser is parallel to the x-axis; Turn off the laser, increase the exposure time, and collect another picture with clear corner points; Finally, perform calibration. After saving the data, the calibration of the 3D camera can be completed.

5. A vision recognition method for a waste sorting robot based on deep learning according to claim 1, characterized in that: The specific content of step two is: Use a 2D camera to collect RGB images, the annotation personnel annotate the collected picture information, and use the YOLOv4 object detection model for model training to generate the final YOLOv4 object detection model.

6. A vision recognition method for a waste sorting robot based on deep learning according to claim 1, characterized in that: The specific content of step four is: Input the collected RGB image into the YOLOv4 object detection model with an input size of 608*608 to obtain a list of all position boxes Bounding Box of target materials existing in the image, and filter through the non-maximum suppression NMS algorithm to obtain the coordinate position information of the final target waste points to be retained; The non-maximum suppression (NMS) algorithm is as follows: Among them, Si represents the score of each bounding box, M represents the box with the highest current score, bi represents a certain box among the remaining boxes, Nt is the set NMS threshold, and IOU is the overlapping area ratio of two recognition boxes; When the YOLOv4 object detection model detects the target value, obtain the current encoder pulse value in real time. The obtained pulse value is the pulse value of the center coordinate of the target material. Use the height of the rectangular box detected by the YOLOv4 object detection model as the movement direction of the conveyor belt, and send the current encoder pulse value, the detected target center coordinate, and the width and height information of the rectangular box to the data analysis and processing module together.

7. A vision recognition method for a waste sorting robot based on deep learning according to claim 1, characterized in that: The specific source of the world coordinate distance of 1.6mm between two contours is as follows: The 3D camera scans a contour with 5 pulses. There are 1216 points on a contour. Each point contains x / y / z values, and each value occupies 4 bytes. The encoder pulse value for one circle is 1000. The distance of one circle of the conveyor belt is 320mm, that is, the distance of one pulse is 0.32mm. Every 5 pulses, that is, 5*0.32=1.6mm, scan a contour. Therefore, the distance of the world coordinates corresponding to the spacing between two contours is 1.6mm.

8. According to claim 1, a garbage sorting robot visual recognition method based on deep learning, Features: When the number of 3D collected contours reaches 50,000, the 3D data is re-stored from the starting position of the memory space, overwriting the previously stored data, and continues to store data in sequence, performing an infinite loop storage of data. When the program stops running, the opened memory space is released to prevent memory overflow or leakage.

9. According to claim 1, a garbage sorting robot visual recognition method based on deep learning, Features: The step six is ​​specifically as follows: Get the 3D memory data data, use the x and y values ​​as the row and column pixel coordinates corresponding to the Mat in OpenCV, normalize the z value from 0 to 255, and use the normalized value as the grayscale value corresponding to the x and y pixel coordinates in the Mat. The height value is extracted from the grayscale value in the corresponding Mat according to the pixel value of the target center coordinate. The actual height information of the target object can be obtained by denormalizing the grayscale value. Use OpenCV threshold segmentation, opening and closing operations, and minimum contour processing algorithms to process the Mat image and extract the width information of the target object; Finally, the center coordinate position of the target object obtained by 2D, the target type, and the target height and width information obtained by 3D are sent to the robot as a whole, so that the robot can realize real-time online grasping and classification.

Citation Information

Patent Citations

  • Rapid three-dimensional panoramic image stitching method

    CN108389157A

  • Garbage sorting system based on visual and deep learning and garbage sorting method

    CN110743818A