AGV-oriented tray dynamic identification and positioning method
Through the combination of RGB-D camera and deep learning model, real-time identification and positioning of AGV pallets are solved, and the problems of poor environmental adaptability and positioning error in the prior art are realized, and efficient pallet pose calculation and AGV trajectory optimization are achieved.
Patent Information
- Application Number
- CN202510505708.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-01
AI Technical Summary
The existing AGV target recognition and positioning technology relies on preset routes and markers, has poor environmental adaptability, insufficient flexibility, and is susceptible to positioning errors, resulting in low work efficiency.
The RGB-D camera is used to collect multimodal images in real time, and the tray ROI area is extracted through the CUDA acceleration model. The point cloud is divided into RANSAC and projected into two-dimensional images. The contour key points are extracted using Canny edge detection and Hough transformation, and the pose is back-projected to the three-dimensional space to calculate the position, and passed to the AGV motion control module for trajectory adjustment.
Real-time recognition and high-precision positioning of pallet positions are realized, and the accuracy and response speed of AGV's automated handling system are improved.
Smart Images

Figure CN120411232A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and detection technology, and in particular to an AGV-oriented pallet dynamic identification and positioning method. Background Art
[0002] In recent years, with the rapid development of industrial automation and intelligent manufacturing, logistics automation has become an irreversible trend, and intelligent handling equipment, represented by AGVs, has ushered in significant development opportunities. As the core component of AGV systems, target recognition and positioning technology directly determines the accuracy and efficiency of AGV operations.
[0003] The shortcoming of existing technologies is that traditional AGVs primarily rely on technologies such as magnetic point guidance, QR code guidance, RFID guidance, or laser reflector guidance to achieve target positioning. While these technologies achieve a certain degree of automated handling, they still have significant limitations, requiring pre-set routes and markers, poor environmental adaptability, and insufficient flexibility. Some manufacturers use improved AGVs using lidar or 3D cameras. While these can achieve target recognition through a single acquisition, their reliance on initial positioning accuracy makes them susceptible to positioning errors and AGV motion errors, leading to mission failure and the need to retreat a certain distance for re-identification and positioning, severely limiting work efficiency. Summary of the Invention
[0004] The purpose of the present invention is to overcome the shortcomings of the existing technology. To achieve the above purpose, a dynamic identification and positioning method of pallets for AGV is adopted to solve the problems raised in the above background technology.
[0005] A method for dynamic identification and positioning of pallets for AGV, comprising the following steps:
[0006] Step S1: First, collect pallet images to train and deploy model weights to the local project as the basis for image recognition;
[0007] Step S2: Use the RGB-D camera to synchronously collect multimodal images in real time and store them in the buffer to ensure the timeliness of the data for recognition calls;
[0008] Step S3: Call the CUDA acceleration model to extract the pallet ROI area from the latest RGB image as the target range for depth data association;
[0009] Step S4: align the depth map to the RGB coordinate system based on the camera's internal and external parameters, then extract depth data based on the pallet ROI and generate a local point cloud as the 3D data basis for subsequent point cloud analysis;
[0010] Step S5: Segment the point cloud of the front plane of the tray by RANSAC and project it to generate a two-dimensional image, converting the three-dimensional features into a two-dimensional processing target;
[0011] Step S6. Subsequently, based on Canny and Hough transforms, extract the pixel coordinates of the key points of the tray contour as the data for three-dimensional pose back-projection;
[0012] Step S7. Back-project the pixel coordinates of the key points into the three-dimensional space to calculate the pose, and complete the mapping between the physical space and the image space;
[0013] Step S8. Transmit the pose information to the AGV motion control module, and pass the pose information to the AGV motion control module for decision-making output of trajectory adjustment.
[0014] As a further solution of the present invention: The specific steps in the step S1 include:
[0015] Step S11. Use a camera to collect tray images, annotate the tray information, and then perform preprocessing to make a data set, where the tray information includes the tray type and the position information;
[0016] Step S12. Divide the data set into a training set, a validation set, and a test set according to a preset ratio, where the training set is collected to train the model, obtain the optimal weight file, and evaluate the model performance on the test set;
[0017] Step S13. After converting the format of the obtained optimal weight file, deploy it locally for dynamic detection of the tray.
[0018] As a further solution of the present invention: The specific steps in the step S2 include:
[0019] Step S21. Use an RGB-D camera to collect and obtain synchronized data and store it in the buffer;
[0020] Step S22. Monitor the survival time of the data in the buffer in real time. When any data stays in the buffer for more than the preset time threshold, it is used as the latest synchronized RGB-D data for output, otherwise it is cleared immediately.
[0021] As a further solution of the present invention: The specific steps in the step S3 are to initialize the target detection model, and then read the RGB image from the buffer; perform inference on the image and use CUDA to accelerate the inference; detect the tray, extract the tray ROI region and the category for output.
[0022] As a further solution of the present invention: The specific steps in the step S4 include:
[0023] Step S41. Read the camera internal parameters and external parameters;
[0024] Step S42. Based on the read camera internal parameters and external parameters, align the depth map to the RGB image coordinate system;
[0025] Step S43: Extract corresponding depth data from the depth map based on the pallet ROI;
[0026] Step S44: Generate a local point cloud based on the extracted depth data;
[0027] Step S45: Store the generated local point cloud data into the point cloud data structure and output it.
[0028] As a further solution of the present invention: The specific steps in Step S5 include:
[0029] Step S51: Calculate the distribution histogram of the generated local point cloud in the Z direction;
[0030] Step S52: Find the interval with the densest distribution of the local point cloud in the Z direction, and this interval is the effective range where the pallet exists;
[0031] Step S53: Use the Z values of the front and back intervals of the main interval as thresholds to perform pass-through filtering on the point cloud, removing background, ground, and suspended noise points;
[0032] Step S54: Use voxel grid downsampling to process the point cloud and reduce the point cloud density;
[0033] Step S55: Segment the point cloud of the front plane of the pallet based on the RANSAC algorithm;
[0034] Step S56: Calculate the plane parameters, construct a rotation matrix through the Rodriguez formula, and rotate the plane to be parallel to the XOY plane;
[0035] Step S57: Project the plane point cloud parallel to the XOY onto the XOY plane;
[0036] Step S58: Convert the projected plane point cloud into a two-dimensional image;
[0037] Step S59: Package the plane parameters, two-dimensional image, rotation matrix, and projection relationship into a unified data structure and output it.
[0038] As a further solution of the present invention: The specific steps in Step S6 include:
[0039] Step S61: Perform preprocessing and binary processing on the two-dimensional image obtained in Step S5;
[0040] Step S62: Use the Canny algorithm to detect the edges in the image;
[0041] Step S63: Extract the straight line segment l from the edge image based on the Hough transform i and store it in the set L;
[0042] Step S64: Use the angle filtering and length filtering methods to filter out the straight line segments in set L that are not the tray contours;
[0043] Step S65: Divide the straight line segments in set L into two types of straight line segments, L X and L Y in the X and Y directions;
[0044] Step S66: Cluster the line segments in L X by the Y coordinate to obtain the upper and lower boundaries of the tray; and cluster the line segments in L Y by the X coordinate to obtain the contour lines of each side post;
[0045] Step S67: Calculate and output the pixel coordinates of the key points of the tray based on the upper and lower boundary lines of the tray and the contour lines of each side post.
[0046] As a further solution of the present invention: The specific steps in step S7 include:
[0047] Step S71: Extract the hardware synchronization timestamp from the RGB-D data;
[0048] Step S72: Read the plane parameters, rotation matrix, and projection relationship from the data structure output in step S5;
[0049] Step S73: Back-project the pixel coordinates of the key points to the camera three-dimensional coordinate system based on the rotation matrix and projection relationship to obtain the three-dimensional coordinates of the key points;
[0050] Step S74: Calculate the rotation angles of the tray plane relative to the X and Y axes respectively based on the plane parameters;
[0051] Step S75: Package the timestamp, three-dimensional coordinates of the key points, and tray rotation angle into pose data and output.
[0052] As a further solution of the present invention: The specific steps in step S8 include:
[0053] Step S81: Send the pose information to the AGV motion control module;
[0054] Step S82: The AGV motion control module dynamically adjusts its own pose according to the relative position information of the tray.
[0055] Compared with the prior art, the present invention has the following technical effects:
[0056] With the above technical solution, real-time recognition and positioning of the tray pose are achieved through multimodal data fusion calculation. Train and deploy the tray recognition model locally to build the basis for image recognition; use an RGB-D camera to synchronously collect RGB and depth images to ensure data timeliness; call CUDA to accelerate the model to extract the tray ROI area in the RGB image and associate the depth data target range; align the depth map to the RGB coordinate system to generate a local point cloud; segment the front plane point cloud of the tray based on RANSAC and project it into a two-dimensional image; combine Canny edge detection and Hough transform to extract the pixel coordinates of the contour key points; back-project the key points into three-dimensional space to calculate the pose and establish the mapping between the physical and image spaces; transfer the pose to the AGV motion control module to drive the trajectory adjustment decision. By synchronizing multimodal data and accelerating calculations, the processing efficiency is improved. Combining three-dimensional point cloud segmentation and two-dimensional feature extraction enables high-precision pose calculation, supports real-time trajectory optimization of the AGV, and enhances the accuracy and response speed of the automated handling system. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The following describes in detail the specific embodiments of the present invention with reference to the accompanying drawings:
[0058] Figure 1 It is a flowchart of the recognition and positioning method for the disclosed embodiments of this application;
[0059] Figure 2 It is a schematic diagram of the tray under the camera coordinate system for the disclosed embodiments of this application;
[0060] Figure 3 It is a flowchart of the model deployment for the disclosed embodiments of this application;
[0061] Figure 4 It is a flowchart of the image acquisition for the disclosed embodiments of this application;
[0062] Figure 5 It is a flowchart of the tray recognition for the disclosed embodiments of this application;
[0063] Figure 6 It is an effect diagram of detecting the tray ROI area for the disclosed embodiments of this application;
[0064] Figure 7 It is a flowchart of obtaining the local point cloud for the disclosed embodiments of this application;
[0065] Figure 8 It is a scene point cloud diagram for the disclosed embodiments of this application;
[0066] Figure 9 It is a local point cloud diagram for the disclosed embodiments of this application;
[0067] Figure 10 It is a diagram of the point cloud processing steps for the disclosed embodiments of this application;
[0068] Figure 11 The planar point cloud map aligned with XOY for the disclosed embodiment of this application;
[0069] Figure 12 The two-dimensional planar image of the tray for the disclosed embodiment of this application;
[0070] Figure 13 The image processing flow chart for the disclosed embodiment of this application;
[0071] Figure 14 The effect diagram of the tray contour line and key point processing for the disclosed embodiment of this application;
[0072] Figure 15 The key point point cloud effect diagram for the disclosed embodiment of this application;
[0073] Figure 16 The flow chart for calculating the pose information of the tray for the disclosed embodiment of this application. Detailed implementation manners
[0074] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0075] Please refer to Figure 1 , in the embodiments of the present invention, a method for dynamic recognition and positioning of a tray for an AGV includes the following steps:
[0076] Step S1, model deployment: First, collect tray images, train and deploy model weights to the local project as the basis for image recognition. The specific steps include:
[0077] Step S11, use a camera to collect tray images, label the tray information, and then perform preprocessing to make a data set, where the tray information includes the tray type and position information;
[0078] Step S12, divide the data set into a training set, a validation set, and a test set according to a preset ratio. Among them, collect the training set to train the model, obtain the optimal weight file, and evaluate the model performance on the test set;
[0079] Step S13, after converting the format of the obtained optimal weight file, deploy it locally for dynamic detection of the tray.
[0080] Among them, for the "model deployment" mentioned in step (1), to achieve dynamic recognition and positioning of the target during the process of AGV picking up and placing goods, it is first necessary to be able to quickly identify and track the target. For this purpose, in this example, the deep learning object detection model is deployed in the local project. Through this model, the target ROI area in the image can be efficiently inferred, providing support for subsequent extraction of the target point cloud based on the ROI.
[0081] In this embodiment, as Figure 3 shown, the steps of model deployment are as follows:
[0082] Use a camera to collect pallet images, label the types and position information of the pallets, and then organize the image and label data to make a dataset.
[0083] Divide the dataset into a training set, a validation set, and a test set according to 7:2:1, use the dataset to train the model, save the optimal weight file, and evaluate the model performance on the test set.
[0084] Convert the model weight file to the ONNX format and deploy it in the local C++ project to achieve dynamic pallet detection.
[0085] Step S2, Image acquisition: Use an RGB-D camera to collect multi-modal images in real-time and synchronously and store them in the buffer to ensure data timeliness for recognition and call;
[0086] Specifically, an RGB-D camera is used to collect image data. This camera can synchronously output RGB images and depth images, and supports inferring the target ROI area through the RGB image and calculating the target pose through the depth data. Its specific steps include:
[0087] Step S21, Use the RGB-D camera to collect and obtain synchronous data and store it in the created buffer;
[0088] Step S22, Monitor the survival time of the buffer data in real-time. When any data stays in the buffer for more than the preset time threshold, it is used as the latest synchronous RGB-D data output, otherwise it is immediately cleared.
[0089] In this embodiment, as Figure 4 shown, the steps of image acquisition are as follows:
[0090] Turn on the RGB-D camera.
[0091] Obtain synchronous RGB-D data.
[0092] Create a buffer area and temporarily store the RGB-D data in the buffer area.
[0093] Monitor the survival time of the buffer data in real-time. When any data stays in the buffer for more than 30 ms, it is immediately cleared to ensure that only the latest frame is retained.
[0094] Extract the latest and unexpired synchronized RGB-D data from the buffer and output it to the subsequent processing module.
[0095] Step S3, Tray Recognition: Call the CUDA-accelerated model to extract the tray ROI region from the latest RGB image as the target range for depth data association; among them, the tray ROI region is mainly extracted based on the deep learning object detection algorithm. At the same time, the CUDA acceleration technology is used to accelerate the inference process, significantly shortening the inference time from 100 - 200 ms to about 20 ms.
[0096] The specific steps are as follows: Initialize the object detection model, and then read the RGB image from the buffer; perform inference on the image and use CUDA to accelerate the inference; when the tray is detected, extract the tray ROI region and category for output.
[0097] As Figure 5 shown, the specific steps of tray recognition are as follows:
[0098] Step S31, Initialize the object detection model.
[0099] Step S32, Read the RGB image from the buffer.
[0100] Step S33, Perform inference on the image and use CUDA to accelerate the inference process:
[0101] The tray recognition module uses YOLOv8 for model training and deploys inference through the OpenCV DNN module. Actual tests show that: in the CPU mode, the inference takes 100 - 200 ms; after enabling CUDA acceleration, the GPU inference time can be reduced to 20 ms, and the recognition efficiency is increased by about 5 - 10 times.
[0102] Step S34, Determine whether the tray is detected. If the tray is detected, extract the tray ROI region and category, as Figure 6 shown, and output it to the subsequent module; otherwise, execute step S32.
[0103] Step S4, Obtain the local point cloud: Align the depth map to the RGB coordinate system based on the internal and external camera parameters, and then extract the depth data according to the tray ROI and generate the local point cloud as the three-dimensional data basis for subsequent point cloud analysis; among them, mainly based on the tray ROI region obtained in step (3), obtain the depth data from the aligned depth map and generate the point cloud.
[0104] The specific steps of step S4 include:
[0105] Step S41, Read the internal and external camera parameters;
[0106] Step S42: Align the depth map to the RGB image coordinate system based on the read camera intrinsic and extrinsic parameters;
[0107] Step S43: Extract the corresponding depth data from the depth map based on the pallet ROI;
[0108] Step S44: Generate a local point cloud based on the extracted depth data;
[0109] Step S45: Store the generated local point cloud data in the point cloud data structure and output it.
[0110] As Figure 7 shown, the steps to obtain the local point cloud are as follows:
[0111] Step S41: Read the camera intrinsic and extrinsic parameters:
[0112] Depth camera intrinsic matrix K d :
[0113]
[0114] RGB camera intrinsic matrix K c :
[0115]
[0116] Extrinsic parameters from depth to color camera: (rotation matrix R, translation vector T).
[0117] Extrinsic matrix: [R|T].
[0118] Where: f x ,f y is the focal length in the x / y direction (in pixel units), c x ,c y is the optical center.
[0119] Step S42: Align the depth map to the RGB image coordinate system through coordinate transformation:
[0120] For each pixel (u d ,v d ) in the depth map, calculate its 3D coordinate P d in the depth camera coordinate system based on the depth value z d and the depth camera intrinsic K d =(x d ,y d ,z d ):
[0121]
[0122] Transform P through the calibrated extrinsic parameters (rotation matrix R and translation vector T)d Convert to the RGB coordinate system P c =(x c , y c , z c ):
[0123] P c = R·P d + T
[0124] Based on the RGB camera internal parameter matrix K c Project P c onto the pixel coordinates (u c , v c ) in the RGB camera coordinate system, and each pixel corresponds to a depth value z c :
[0125]
[0126] Step S43: Extract the corresponding depth data from the depth map based on the pallet ROI.
[0127] Step S44: Generate a local point cloud based on the depth data. The scene point cloud is as shown in Figure 8 , and the local point cloud is as shown in Figure 9 :
[0128] Generate the corresponding point cloud coordinates P = (x, y, z) through the mapping relationship between the depth image pixel coordinates and the RGB camera internal parameter matrix K c :
[0129]
[0130] Step S45: Store the local point cloud data into the point cloud data structure and output it to the subsequent module.
[0131] Step S5: Point cloud processing: Segment the point cloud of the front plane of the pallet by RANSAC and project it to generate a 2D image, converting the 3D feature into a 2D processing target; its specific steps include:
[0132] Step S51: Calculate the distribution histogram of the generated local point cloud in the Z direction;
[0133] Step S52: Find the interval with the densest distribution of the local point cloud in the Z direction, and this interval is the effective range where the pallet exists;
[0134] Step S53: Use the Z values of the front and back intervals of the main interval as thresholds to perform a pass-through filter on the point cloud to remove background, ground, and suspended noise points;
[0135] Step S54: Use voxel grid downsampling to process the point cloud to reduce the point cloud density;
[0136] Step S55: Segment the point cloud of the front plane of the tray based on the RANSAC algorithm;
[0137] Step S56: Calculate the plane parameters, construct a rotation matrix through the Rodriguez formula, and rotate the plane to be parallel to the XOY plane;
[0138] Step S57: Project the plane point cloud parallel to the XOY plane onto the XOY plane;
[0139] Step S58: Convert the projected plane point cloud into a two-dimensional image;
[0140] Step S59: Package the plane parameters, two-dimensional image, rotation matrix, and projection relationship into a unified data structure and output.
[0141] In this embodiment, as Figure 10 shown, the steps of point cloud processing are as follows:
[0142] Step S51: Calculate the distribution histogram of the local point cloud in the Z direction:
[0143] First, equally divide the effective measurement range (0 - 5000 mm) of the depth camera into 20 intervals:
[0144] Bin i = [250(i - 1), 250i) mm, i ∈ {1, 2,..., 20}
[0145] (Note: The 20th interval is a closed interval [4750, 5000] mm)
[0146] Then calculate the distribution histogram of the point cloud in the Z direction.
[0147] Step S52: Find the interval where the point cloud is most dense in the Z direction, and this interval is the effective range where the tray exists:
[0148] Count the number of points H i in each interval, and calculate the maximum number of points H k and the corresponding interval Bin k .
[0149] Furthermore, judge whether the number of points H k in the intervals Bin k-1 , Bin k+1 before and after satisfies H k-1 > ∑H k+1 / 5. Merge the intervals that meet the conditions with Bin i to obtain the main interval Bin i (Note: If neither the front nor the back interval meets the conditions, then Bin k remains unchanged) t (Note: If neither the front nor the back interval meets the conditions, then Bin k= Bin t )。
[0150] Finally, take the maximum and minimum Z values (Z t and Z max and Z min ) of the main interval Bin as the thresholds for Z-direction through filtering
[0151] Step S53: Take the front and back Z values (Z max and Z min ) of the main interval as the thresholds, and perform through filtering on the point cloud to remove background, ground, and suspended noise points.
[0152] Step S54: Use voxel grid downsampling (Voxel Grid Filter) to filter the point cloud and reduce the point cloud density.
[0153] Step S55: Based on the RANSAC algorithm, segment the point cloud of the front plane of the tray. The front plane of the tray is as shown Figure 2 . The plane model equation is: ax + by + cz + d = 0, and the normal vector is: n = (a, b, c) T .
[0154] Step S56: Calculate the plane parameters, construct a rotation matrix through the Rodriguez formula, and rotate the plane to be parallel to the XOY plane (the XOY plane is the front plane of the camera as shown Figure 2 ):
[0155] Rotate the tray plane to be parallel to the XOY plane: The normal vector of the tray plane n1 = (a, b, c) T , the normal vector of the XOY plane n2 = (0, 0, 1) T , and normalize the cross product of the two normal vectors to obtain the rotation axis axis:
[0156]
[0157] Obtain the sine and cosine values of the rotation angle through the dot product:
[0158]
[0159] Use the Rodriguez formula to construct the rotation matrix R:
[0160] R = I + sinθ·K + (1 - cosθ)·K 2
[0161] where K is the cross product matrix of the rotation axis:
[0162]
[0163] Finally, according to the rotation matrix R, the planar point cloud is rotated to be parallel to the XOY plane of the camera coordinate system. At this time, the coordinates of each point in the planar point cloud are P i =(x i , y i , Z d ). Save the rotation matrix R for subsequent inverse transformation of pixel coordinates to three-dimensional space.
[0164] Step S57: Project the planar point cloud parallel to XOY onto the XOY plane:
[0165] Project the planar point cloud onto the XOY plane. As Figure 11 shown, the coordinates of each point in the planar point cloud at this time are:
[0166] P i =(x i , y i , 0).
[0167] Save the distance Z from the planar point cloud to the XOY plane d for subsequent inverse transformation of pixel coordinates to three-dimensional space.
[0168] Step S58: Convert the projected planar point cloud into a two-dimensional image:
[0169] Create a black image. Convert the x i , y i of each point P i =(x i , y i , 0) in the projected planar point cloud into pixel values, and fill the corresponding pixels of each point with white in the image, thus obtaining the two-dimensional image corresponding to the planar point cloud. As Figure 12 shown.
[0170] Step S59: Package the point cloud plane parameters, two-dimensional image, rotation matrix, and projection relationship into a unified data structure and output it to the subsequent module.
[0171] Step S6: Image processing: Subsequently, based on Canny and Hough transforms, extract the pixel coordinates of the key points of the tray contour as the data for three-dimensional pose back-projection. The specific steps include:
[0172] Step S61: Preprocess and binarize the two-dimensional image obtained in Step S5;
[0173] Step S62: Use the Canny algorithm to detect the edges in the image;
[0174] Step S63: Extract the straight line segment l i from the edge image based on the Hough transform and store it in the set L;
[0175] Step S64. Use the angle screening and length screening methods to filter out the straight line segments in set L that are not the tray contours;
[0176] Step S65. Divide the straight line segments in set L into two types of straight line segments, L X and L Y in the X and Y directions;
[0177] Step S66. Cluster the line segments in L X by the Y coordinate to obtain the upper and lower boundaries of the tray; and cluster the line segments in L Y by the X coordinate to obtain the contour lines of each side post;
[0178] Step S67. Calculate the pixel coordinates of the tray key points according to the upper and lower boundary lines of the tray and the contour lines of each side post and output them.
[0179] As Figure 13 shown, the steps of image processing are as follows:
[0180] Read the two-dimensional image from the data structure output in step S59.
[0181] Preprocess the two-dimensional image.
[0182] Perform binary processing on the image.
[0183] Use the Canny algorithm to detect the edges in the image
[0184] Apply the Hough transform to extract the straight line segments l i from the edge image and store them in set L.
[0185] Use the angle screening and length screening methods to filter out the straight line segments in set L that are not the tray contours.
[0186] Divide the straight line segments in set L into two types of straight line segments, L X and L Y in the X and Y directions.
[0187] As Figure 14 shown, further cluster the line segments in L X by the Y coordinate to obtain the upper and lower boundaries of the tray; cluster the line segments in L Y by the X coordinate to obtain the contour lines of each side post.
[0188] Calculate the pixel coordinates of the tray key points according to the upper and lower boundary lines of the tray and the contour lines of each side post and output them to the subsequent module.
[0189] Step S7. Calculate the pose information of the tray: Back-project the pixel coordinates of the key points into the three-dimensional space to calculate the pose, and complete the mapping between the physical space and the image space; The specific steps include:
[0190] Step S71: Extract the hardware synchronization timestamp from the RGB-D data;
[0191] Step S72: Read the plane parameters, rotation matrix, and projection relationship from the data structure output in Step S5;
[0192] Step S73: Back-project the key-point pixel coordinates to the camera's three-dimensional coordinate system based on the rotation matrix and projection relationship to obtain the key-point three-dimensional coordinates;
[0193] Step S74: Calculate the rotation angles of the tray plane relative to the X and Y axes respectively based on the plane parameters;
[0194] Step S75: Pack the timestamp, key-point three-dimensional coordinates, and tray rotation angles into pose data and output.
[0195] In this embodiment, as Figure 16 shown, the steps for calculating the tray pose information are as follows:
[0196] Extract the hardware synchronization timestamp from the original RGB-D data.
[0197] Read the plane parameters, rotation matrix, and projection relationship from the data structure output in Step S59.
[0198] Back-project the key-point pixel coordinates to the camera's three-dimensional coordinate system based on the rotation matrix and projection relationship. The key points in three-dimensional space are as Figure 15 shown, and the key-point three-dimensional coordinates are obtained:
[0199] The key-point pixel coordinates are (u i , v i ), the rotation matrix is R, and the projection distance is Z d .
[0200] Since the rotation matrix is an orthogonal matrix, therefore:
[0201] R T = R -1
[0202] Through coordinate transformation, the key-point three-dimensional coordinates (x i , y i , z i ) can be obtained:
[0203]
[0204] Calculate the rotation angles of the tray plane relative to the X and Y axes respectively based on the plane parameters.
[0205] The plane normal vector is n = (a, b, c) T , and calculate the rotation angles of the X and Y axes:
[0206]
[0207]
[0208] Pack the timestamp, the three-dimensional coordinates of the key points, and the tray rotation angle into pose data and output it to the subsequent module.
[0209] Step S8: Transmit the pose information to the AGV motion control module: Transmit the pose information to the AGV motion control module for decision-making output of trajectory adjustment. The specific steps include:
[0210] Step S81: Send the pose information to the AGV motion control module;
[0211] Step S82: The AGV motion control module dynamically adjusts its own pose according to the relative position information of the tray.
[0212] In this embodiment, after all the above eight steps are executed, the pose information of the tray can be accurately obtained. Since the system takes a relatively short cycle of 60 - 100 ms to execute the above steps once, the system can update the tray pose information to the AGV at a frequency of more than 10 Hz, thus realizing the dynamic recognition and positioning of the tray by the AGV.
[0213] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents, and all should be included within the protection scope of the present invention.
Claims
1. A method for dynamic recognition and positioning of pallets for AGV, characterized in that It includes the following steps: Step S1: First, train and deploy the model weights to the local project by collecting pallet images, serving as the basis for image recognition; Step S2: Use an RGB-D camera to synchronously collect multi-modal images in real time and store them in the buffer to ensure data timeliness for recognition and invocation; Step S3: Invoke the CUDA-accelerated model to extract the pallet ROI region from the latest RGB image, serving as the target range for depth data association; Step S4: Align the depth map to the RGB coordinate system based on the internal and external camera parameters, and then extract depth data according to the pallet ROI to generate a local point cloud, serving as the three-dimensional data basis for subsequent point cloud analysis; Step S5: Segment the front plane point cloud of the pallet by RANSAC and project it to generate a two-dimensional image, converting the three-dimensional features into two-dimensional processing targets; Step S6: Subsequently, extract the pixel coordinates of the key points of the pallet contour based on Canny and Hough transforms, serving as the data for three-dimensional pose back-projection; Step S7: Back-project the pixel coordinates of the key points into the three-dimensional space to calculate the pose, completing the mapping between the physical space and the image space; Step S8: Transmit the pose information to the AGV motion control module, and transfer the pose information to the AGV motion control module for decision-making output of trajectory adjustment.
2. The method for dynamically identifying and positioning a pallet for an AGV according to claim 1, characterized in that The specific steps in step S1 include: Step S11: Use a camera to collect pallet images, annotate the pallet information, and then perform preprocessing to make a data set, where the pallet information includes the pallet type and position information; Step S12: Divide the data set into a training set, a validation set, and a test set according to a preset ratio. Train the model with the training set to obtain the optimal weight file, and evaluate the model performance on the test set; Step S13: After converting the format of the obtained optimal weight file, deploy it locally for dynamic pallet detection.
3. The method for dynamically identifying and positioning a pallet for an AGV according to claim 1, wherein The specific steps in step S2 include: Step S21: Use an RGB-D camera to collect and obtain synchronous data and store it in the buffer; Step S22: Monitor the survival time of the buffer data in real time. When any data stays in the buffer for more than the preset time threshold, it is used as the latest synchronous RGB-D data output, otherwise it is immediately cleared.
4. The method for dynamically identifying and positioning a tray for an AGV according to claim 1, wherein The specific steps in step S3 are to initialize the object detection model, and then read the RGB image from the buffer; perform inference on the image and use CUDA to accelerate the inference; detect the pallet, and extract and output the pallet ROI region and category.
5. The method for dynamically identifying and positioning a tray for an AGV according to claim 1, wherein, The specific steps in step S4 include: Step S41: Read the internal and external camera parameters; Step S42: Align the depth map to the RGB image coordinate system based on the read internal and external camera parameters; Step S43: Extract the corresponding depth data from the depth map based on the pallet ROI; Step S44: Generate a local point cloud based on the extracted depth data; Step S45: Store the generated local point cloud data in the point cloud data structure and output it.
6. The method for dynamically identifying and positioning a tray for an AGV according to claim 1, wherein The specific steps in step S5 include: Step S51: Calculate the distribution histogram of the generated local point cloud in the Z direction; Step S52: Find the interval where the local point cloud is most dense in the Z direction, and this interval is the effective range where the pallet exists; Step S53: Use the Z values of the front and rear intervals of the main interval as thresholds to perform a pass-through filter on the point cloud, removing background, ground, and suspended noise points; Step S54: Use voxel grid downsampling to process the point cloud and reduce the point cloud density; Step S55: Based on the RANSAC algorithm, segment the point cloud of the front plane of the tray; Step S56: Calculate the plane parameters, construct a rotation matrix through the Rodriguez formula, and rotate the plane to be parallel to the XOY plane; Step S57: Project the plane point cloud parallel to the XOY plane onto the XOY plane; Step S58: Convert the projected plane point cloud into a two-dimensional image; Step S59: Package the plane parameters, two-dimensional image, rotation matrix, and projection relationship into a unified data structure and output.
7. The method for dynamically identifying and positioning a tray for an AGV according to claim 1, wherein, The specific steps in Step S6 include: Step S61: Perform preprocessing and binary processing on the two-dimensional image obtained in Step S5; Step S62: Use the Canny algorithm to detect edges in the image; Step S63: Extract a straight line segment l from the edge image based on the Hough transform i and store it in the set L; Step S64: Use angle filtering and length filtering methods to filter out the straight line segments in set L that are not the tray contour; Step S65: Divide the line segments in set L into two types of line segments, namely, those in the X direction and those in the Y direction, in set L X and L Y two types of line segments; Step S66. Cluster the line segments in L X by Y coordinate to obtain the upper and lower boundaries of the tray; and cluster the line segments in L Y by X coordinate to obtain the contour lines of each side post; Step S67: Calculate and output the pixel coordinates of the key points of the tray based on the upper and lower boundary lines of the tray and the contour lines of each side post.
8. The method for dynamically identifying and positioning a pallet for an AGV according to claim 1, characterized in that, The specific steps in Step S7 include: Step S71: Extract the hardware synchronization timestamp from the RGB-D data; Step S72: Read the plane parameters, rotation matrix, and projection relationship from the data structure output in Step S5; Step S73: Based on the rotation matrix and projection relationship, back-project the pixel coordinates of the key points into the camera three-dimensional coordinate system to obtain the three-dimensional coordinates of the key points; Step S74: Based on the plane parameters, calculate the rotation angles of the tray plane relative to the X and Y axes respectively; Step S75: Package the timestamp, three-dimensional coordinates of the key points, and tray rotation angle into pose data and output.
9. The method for dynamically identifying and positioning a tray for an AGV according to claim 1, wherein, The specific steps in Step S8 include: Step S81: Send the pose information to the AGV motion control module; Step S82: The AGV motion control module dynamically adjusts its own pose according to the relative position information of the tray.
Citation Information
Cited By
Bulk cargo cabin real-time fusion detection method based on laser point cloud
CN120953342A
Method for accurately extracting key points of two types of point clouds based on RGB-D image
CN121458996A
A method for accurate extraction of key points from two types of point clouds based on RGB-D images
CN121458996B