Target detection de-weight positioning method and system for double-arm cooperative wheel-tracked robot

By performing target detection and deduplication processing on the two-arm cooperative wheel-sliding robot, combined with the unified coordinate conversion framework of three-dimensional information and radar coordinate system, the robot's accuracy and efficiency problems of target detection and positioning in complex environments are solved, and higher crawling operation reliability and dynamic response capabilities are achieved.

CN120056113APending Publication Date: 2025-05-30江淮前沿技术协同创新中心

Patent Information

Application Number
CN202510268024.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When performing complex grab tasks, the dual-arm cooperative wheel-slide robot faces problems such as overlapping view angles of object detection cameras and redundant data, low positioning accuracy, and limitations of multi-sensor coordinate conversion.

Method used

By acquiring camera data and lidar data for time alignment, object detection and deduplication processing are performed based on the aligned data, deduplication judgment is performed based on three-dimensional information, and coordinate conversion is performed to the robotic arm base coordinate system, and coordinate conversion is performed through a unified coordinate conversion framework based on the radar coordinate system.

Benefits of technology

It effectively avoids the same target being repeatedly detected by multiple sensors, reduces the amount of data, reduces the target positioning error, improves the reliability of subsequent grab operations, and enhances the accuracy and dynamic adaptability of target positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120056113A_ABST
    Figure CN120056113A_ABST
Patent Text Reader

Abstract

The invention provides a target detection deduplication positioning method and system for a double-arm cooperative wheel-tracked robot, and relates to the technical field of target detection deduplication positioning of robots, the method incorporates three-dimensional information into a deduplication algorithm, so that the target position can be judged more accurately, the same target is prevented from being repeatedly detected by multiple sensors, the data volume is reduced, and the detection efficiency is improved. The target positioning error is effectively reduced, and the reliability of subsequent grabbing operation is improved; through combination of accurate three-dimensional point cloud data provided by a laser radar, the limitation of a depth camera in a complex environment is effectively made up, and the target identification capability is significantly improved. A unified coordinate conversion framework based on a radar coordinate system is introduced, the precision and dynamic adaptability of target positioning are enhanced, the defects of a traditional coordinate conversion method in a dynamic environment are overcome, efficient grabbing operation of a mechanical arm is supported, and the dynamic response capacity of the robot is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection and duplicate removal positioning for robots, and more specifically, to a method and system for target detection and duplicate removal positioning for a dual-arm collaborative wheeled tracked robot. Background Art

[0002] With the continuous development of robot technology, especially in the application of dual-arm collaborative wheeled tracked robots, target detection and positioning technology has become an important part of their autonomous operation and environmental perception. Dual-arm collaborative wheeled tracked robots are widely used in tasks that require high-precision grasping and operation, such as complex object handling, industrial assembly, precision maintenance, etc. In these applications, the robot needs to rely on advanced sensor systems, such as cameras, radars, lidars, etc., to identify and locate objects in the environment, especially dynamic targets. In order to achieve precise dual-arm grasping tasks, how to effectively perform target detection, positioning, and duplicate removal has become a key challenge to improve the operation accuracy and efficiency of the robot.

[0003] First of all, in a complex dynamic environment, the target object may be partially occluded or move rapidly, resulting in the inability of cameras and lidars to simultaneously obtain complete target information. Although existing robot systems use multiple cameras to expand the overall perspective by covering different fields of view to improve the detection range of targets, however, in the overlapping area of multiple camera fields of view, especially when the target is at a long distance or the viewing angles are close, the problem of repeated detection of the same target by multiple cameras is likely to occur. This repeated detection usually leads to the generation of redundant data, which in turn affects subsequent operations, such as the accuracy of robotic arm grasping. Current duplicate removal algorithms only focus on the overlap degree of bounding boxes in two-dimensional images and ignore the true position relationship of targets in three-dimensional space, resulting in poor duplicate removal effects.

[0004] Secondly, in order to perform accurate grasping tasks, it is not only necessary to achieve accurate detection of the target object, but also necessary to obtain the three-dimensional position information of the object. Although in existing multi-camera combination technologies, depth cameras are used to extract depth information in images through stereo vision algorithms and then estimate the position of the target in three-dimensional space, which can achieve a certain degree of depth perception, due to factors such as the resolution of the image, the difference in camera viewing angles, and lighting changes, the depth information is often not very accurate, especially in long-distance or complex scenarios, where the depth error is large. This leads to errors in target positioning and affects the accuracy of subsequent operations.

[0005] Thirdly, compared with cameras, lidar performs more stably and accurately in long-distance perception and environmental modeling. However, there are perspective differences between lidar and cameras, and the docking of their coordinate systems is a key issue in multi-sensor fusion. Although existing camera-lidar calibration algorithms have been relatively mature and can achieve coordinate conversion between the two, in some applications, especially in the robotic arm or chassis system of wheeled-tracked robots, it is still very difficult to directly convert the targets detected by the camera from the camera coordinate system to the base coordinate system or the end coordinate system of the robotic arm. This is because the robotic arm coordinate system usually has a high degree of freedom, and as the robotic arm moves, the coordinate changes are very complex, and traditional calibration methods cannot effectively handle these dynamic changes. Therefore, there are certain limitations in the coordinate conversion of the existing technology between the robotic arm and the camera and lidar systems. Especially in multi-task execution and dynamic operations, there is still much room for improvement in the accuracy and efficiency of the conversion.

[0006] Therefore, for the problems faced by the dual-arm collaborative wheeled-tracked robot in performing complex grasping tasks, such as the overlapping of the target detection camera perspectives and data redundancy, low positioning accuracy, and limitations in multi-sensor coordinate conversion, the key lies in how to accurately obtain the position of the target object in the three-dimensional space and effectively eliminate the data redundancy caused by the overlapping of multi-sensor perspectives. At the same time, considering the dynamic coordinate changes during the movement of the robotic arm, traditional calibration methods are difficult to meet the requirements of accuracy and real-time performance in practical applications. Most current object detection frameworks usually use the IOU (Intersection over Union) value in the two-dimensional space for deduplication, but this method is prone to failure when the objects are very close or similar. Incorporating three-dimensional information into the deduplication process can effectively reduce the misrecognition rate and improve the robustness in complex environments. This deduplication method that combines two-dimensional detection and three-dimensional space judgment may be more accurate than the traditional two-dimensional-only deduplication in practical applications.

[0007] CN117593620A discloses a multi-target detection method and device based on the fusion of a camera and lidar, which is applicable to the fields of autonomous driving and robotics. By fusing lidar point cloud data with camera images, a deep learning model is used to classify and locate targets. Although the fusion of camera and lidar data is solved, the perspective difference problem still exists. Due to the different working principles and installation angles of the camera and lidar, there are large differences between the sensor data, resulting in large position deviations of the targets in the fields of view of different sensors.

[0008] CN112731371A discloses an integrated target tracking system and method for lidar and vision fusion. Although the system proposed in this patent performs target tracking through the fusion of lidar and vision information, its fusion process mainly relies on spatial registration and target state prediction. This method is prone to target positioning errors and tracking instability problems when sensor data is asynchronous or there are high dynamic changes.

[0009] In the paper "Fusion of LiDAR and Camera for 3D Object Detection in Autonomous Vehicles" published by Zhou, X., & Sun, Z. et al. in IEEE Transactions on Intelligent Transportation Systems, a 3D object detection method based on the fusion of lidar (LiDAR) and camera is proposed, aiming to solve the problems of object detection and positioning in autonomous vehicles. Although progress has been made in multi-sensor fusion, it still faces problems such as perspective differences, poor duplicate removal effect, and depth perception errors. Especially in the overlapping area of the multi-sensor field of view, repeated detection and positioning errors will affect the execution of subsequent tasks. Summary of the Invention

[0010] To solve the above problems, an embodiment of the present invention provides a method for duplicate removal and positioning of object detection for a two-armed collaborative wheeled tracked robot. The method includes: obtaining camera data and lidar data, and performing time alignment on the camera data and the lidar data; performing object detection based on the aligned camera data to obtain a plurality of two-dimensional detection frames; performing distance detection on the plurality of two-dimensional detection frames of the same detection category based on the aligned lidar data, and performing duplicate removal processing on the plurality of two-dimensional detection frames with a distance less than a distance threshold to obtain duplicate-removed two-dimensional detection frames; converting the spatial coordinates of the duplicate-removed two-dimensional detection frames to the base coordinate system of the robotic arm to obtain spatial coordinates in the base coordinate system of the robotic arm; and sending the spatial coordinates in the base coordinate system of the robotic arm to the robotic arm so that the robotic arm performs a grasping operation.

[0011] The target detection and duplicate removal positioning method for the dual-arm collaborative wheeled tracked robot provided by the embodiments of the present invention incorporates three-dimensional information into the duplicate removal algorithm, which can more accurately determine the target position, avoid repeated detection of the same target by multiple sensors, reduce the data volume, effectively reduce the target positioning error, and improve the reliability of subsequent grasping operations; by combining the precise three-dimensional point cloud data provided by the lidar, it effectively makes up for the limitations of the depth camera in complex environments and significantly improves the target recognition ability; introducing a unified coordinate transformation framework based on the radar coordinate system enhances the accuracy and dynamic adaptability of target positioning, solves the deficiencies of traditional coordinate transformation methods in dynamic environments, supports the efficient grasping operation of the robotic arm, and improves the dynamic response ability of the robot.

[0012] Optionally, the distance detection of multiple two-dimensional detection frames of the same detection category according to the aligned lidar data includes: converting the three-dimensional coordinates of the two-dimensional detection frame in the camera coordinate system into three-dimensional coordinates in the radar coordinate system based on the transformation matrix between the camera coordinate system and the radar coordinate system; the three-dimensional coordinates include the center point coordinates and depth values of the two-dimensional detection frame; determining the radar point cloud coordinates closest to the center point in the aligned lidar data according to the three-dimensional coordinates of the center point of the two-dimensional detection frame in the radar coordinate system; calculating the distance between any two two-dimensional detection frames corresponding to the radar point cloud coordinates of the same detection category.

[0013] The embodiments of the present invention incorporate three-dimensional information into the duplicate removal algorithm, which can more accurately determine the target position and avoid the situation of repeated detection of the same target by multiple sensors.

[0014] Optionally, before performing duplicate removal processing on multiple two-dimensional detection frames with a distance less than a preset threshold, the method further includes: performing overlap detection on multiple two-dimensional detection frames of the same detection category; performing duplicate removal processing on multiple two-dimensional detection frames with an overlap degree greater than the overlap degree threshold and a distance less than the preset threshold.

[0015] The embodiments of the present invention fuse the two-dimensional detection frame with the three-dimensional point cloud data of the lidar, improving the detection accuracy.

[0016] Optionally, the duplicate removal processing includes: retaining the two-dimensional detection frame with the highest confidence among multiple two-dimensional detection frames with an overlap degree greater than the overlap degree threshold and a distance less than the preset threshold.

[0017] The embodiments of the present invention optimize the duplicate removal algorithm by combining three-dimensional space positioning, depth information, and radar point cloud data, and accurately obtain the position of the target object in three-dimensional space.

[0018] Optionally, the conversion of the spatial coordinates of the deduplicated two-dimensional detection box to the base coordinate system of the robotic arm to obtain the spatial coordinates in the base coordinate system of the robotic arm includes: determining the transformation relationship from the radar coordinate system to the base coordinate system of the robotic arm based on the installation position relationship between the robotic arm and the radar sensor; and converting the spatial coordinates of the deduplicated two-dimensional detection box in the radar coordinate system to the base coordinate system of the robotic arm based on the transformation relationship to obtain the spatial coordinates in the base coordinate system of the robotic arm.

[0019] The embodiment of the present invention solves the problem of difficult coordinate conversion between a traditional camera and a robotic arm coordinate system through a unified coordinate conversion framework based on the radar coordinate system, enhancing the accuracy and dynamic adaptability of target positioning.

[0020] Optionally, the time alignment of the camera data and the lidar data includes: performing time alignment on the camera data and the lidar data based on the ROS message synchronization mechanism, and aligning the camera data and the lidar data with the same timestamp or a timestamp difference less than a set time difference.

[0021] The embodiment of the present invention can eliminate the time differences of different sensors through time alignment, ensuring that the data can be accurately corresponding to the same time point.

[0022] Optionally, the method further includes: visualizing the deduplicated two-dimensional detection box; the visualization includes: marking the deduplicated two-dimensional detection box in the point cloud data, or marking the deduplicated two-dimensional detection box on the image, as well as marking the class label and confidence level of the deduplicated two-dimensional detection box.

[0023] The embodiment of the present invention visualizes the results to verify the fusion and deduplication effects.

[0024] Optionally, the method further includes: adjusting the distance threshold and the overlap degree threshold according to the visualization result.

[0025] The embodiment of the present invention can improve the detection accuracy by adjusting the threshold.

[0026] Optionally, the method further includes: adjusting the distance threshold and the overlap degree threshold according to the usage environment of the dual-arm collaborative wheeled tracked robot and the spatial distribution of the target object.

[0027] The embodiment of the present invention can adjust the deduplication threshold according to the dynamic changes of the environment and the target object, improve the accuracy, and adapt to the dynamically changing environment.

[0028] The embodiment of the present invention provides a target detection deduplication and positioning system for a dual-arm collaborative wheeled tracked robot, which is used to execute the method described in any one of the above.

[0029] The target detection and duplicate removal positioning system for a dual-arm collaborative wheeled-tracked robot provided by an embodiment of the present invention can achieve the same technical effects as the above-mentioned target detection and duplicate removal positioning method for a dual-arm collaborative wheeled-tracked robot. Description of the Drawings

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.

[0031] Figure 1 It is a schematic flowchart of a target detection and duplicate removal positioning method for a dual-arm collaborative wheeled-tracked robot provided by an embodiment of the present invention;

[0032] Figure 2 It is a specific schematic flowchart of the target detection and duplicate removal positioning method for a dual-arm collaborative wheeled-tracked robot provided by an embodiment of the present invention. Detailed Embodiments

[0033] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will give a detailed description of the specific embodiments of the present invention with reference to the drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0034] Aiming at the problems faced by a dual-arm collaborative wheeled-tracked robot during the execution of complex grasping tasks, such as the overlapping of target detection camera views and data redundancy, low positioning accuracy, and limitations in multi-sensor coordinate transformation, the key lies in how to accurately obtain the position of the target object in three-dimensional space and effectively eliminate the data redundancy caused by the overlapping of multi-sensor views.

[0035] Considering the change of the dynamic coordinate system during the movement of the robotic arm, the traditional calibration method is difficult to meet the requirements of accuracy and real-time performance in practical applications. By combining the multi-modal data of the camera and the radar, through coordinate system conversion and duplicate removal algorithms, the coordinate information of the target object can be accurately obtained, effectively solving the problem of insufficient target positioning accuracy, realizing the conversion from the camera coordinate system to the radar coordinate system, and further improving the positioning accuracy and operability of the dual-arm collaborative wheeled tracked robot when performing grasping tasks, and enhancing the reliability and flexibility of the system. Most current object detection frameworks, including YOLO (You Only Look Once) and SSD (Single Shot Multibox Detector), usually use the IOU value in the two-dimensional space for duplicate removal, but this method is prone to failure when the objects are very close or similar. Incorporating three-dimensional information into the duplicate removal process can effectively reduce the misrecognition rate and enhance the robustness in complex environments. This duplicate removal method that combines two-dimensional detection and three-dimensional space judgment is more accurate in practical applications than the traditional two-dimensional-only duplicate removal.

[0036] Based on this, the embodiments of the present invention propose a target detection duplicate removal and positioning method and system based on the fusion of multi-camera combination and radar point cloud data. By combining camera and radar data and applying a duplicate removal algorithm in the overlapping area of the multi-sensor perspective, the position of the target object in the three-dimensional space can be accurately obtained, avoiding the positioning accuracy problems caused by the perspective differences and data redundancy of a single sensor in the prior art. At the same time, through a unified coordinate conversion framework based on the radar coordinate system, the problem of difficult coordinate conversion between the traditional camera and the robotic arm is solved. This method can perform target positioning and duplicate removal in real time and accurately, improve the accuracy and stability in the grasping tasks of the dual-arm collaborative robot, realize the seamless conversion from the camera coordinate system of the target object to the radar coordinate system, and significantly improve the operation efficiency and reliability of the robot in a dynamic environment.

[0037] Figure 1 The flowchart of the target detection duplicate removal and positioning method for the dual-arm collaborative wheeled tracked robot provided by the embodiments of the present invention includes:

[0038] S102, obtaining camera data and lidar data, and performing time alignment on the camera data and the lidar data.

[0039] The camera data can be images, depth data, etc. collected by the camera, and the lidar data can be three-dimensional point cloud data collected by the lidar. Considering that there may be a situation where the timestamps of the camera and the lidar are not synchronized in time, in this embodiment, time alignment is performed on the above data to eliminate the time difference of the data of different sensors and ensure that the data can be accurately corresponding to the same time point.

[0040] Specifically, based on the ROS (Robot Operating System) message synchronization mechanism, the camera data and lidar data are time-aligned, and the camera data and lidar data with the same timestamp or a timestamp difference less than the set time difference are aligned. For example, when the data sampling frequencies of the camera and the lidar are the same, precise synchronization can be performed to align the data of the above-mentioned camera and lidar at the same moment, that is, their timestamps are exactly the same; when the data sampling frequencies of the camera and the lidar are different, if the gap between the timestamps of the two data is less than the set time difference, approximate alignment can be performed. Since the camera and radar data have been aligned under the same timestamp, it ensures that accurate results can be obtained when performing data processing and fusion in subsequent steps.

[0041] S104. Perform object detection based on the aligned camera data to obtain multiple two-dimensional detection boxes.

[0042] For camera data, target extraction can be performed to obtain the two-dimensional detection box (bounding box) of the target. Exemplarily, it includes steps such as camera initialization, image acquisition, image preprocessing, object detection, and extracting the two-dimensional detection box from the detection results.

[0043] S106. Perform distance detection on multiple two-dimensional detection boxes of the same detection category based on the aligned lidar data, and perform duplicate removal processing on multiple two-dimensional detection boxes with a distance less than the distance threshold to obtain the two-dimensional detection boxes after duplicate removal.

[0044] Traditional depth cameras often have large errors when obtaining three-dimensional position information, resulting in low accuracy of target detection. In this embodiment, three-dimensional information fusion and duplicate removal positioning are combined with the accurate three-dimensional point cloud data provided by the lidar, which can more accurately judge the target position, avoid the situation of repeated detection of the same target by multiple sensors, reduce misidentification, reduce the data volume, effectively reduce the target positioning error, and improve the reliability of subsequent grasping operations.

[0045] For the two-dimensional detection boxes obtained in the foregoing steps, only the two-dimensional detection boxes of the same type are compared for duplication. Specifically, the two-dimensional detection boxes obtained based on the camera data are converted to the radar coordinate system, and the corresponding points (which can be the same position or the nearest position points) of the two-dimensional detection boxes in the lidar point cloud data are found. Based on the coordinates of this point in the lidar point cloud data, the spatial distance between any two two-dimensional detection boxes can be calculated. If the spatial distance is less than a certain threshold, it is determined that these two two-dimensional detection boxes are duplicates. Only one of the multiple two-dimensional detection boxes determined to be duplicates is retained, and the rest are deleted, and then the two-dimensional detection box set is updated to ensure the uniqueness and precise positioning of each object in three-dimensional space.

[0046] Exemplarily, for multiple two-dimensional detection boxes with an overlap degree greater than the overlap degree threshold and a distance less than the preset threshold, only the two-dimensional detection box with the highest confidence can be retained.

[0047] Exemplarily, distance detection is performed based on the following method:

[0048] First, based on the transformation matrix between the camera coordinate system and the radar coordinate system, the three-dimensional coordinates of the two-dimensional detection box in the camera coordinate system are converted into the three-dimensional coordinates in the radar coordinate system. The three-dimensional coordinates in the camera coordinate system can include the center point coordinates and depth values of the two-dimensional detection box. Second, according to the three-dimensional coordinates of the center point of the above two-dimensional detection box in the radar coordinate system, the radar point cloud coordinates closest to the center point in the aligned lidar data are determined. Then, the distance between any two two-dimensional detection boxes is calculated based on the radar point cloud coordinates corresponding to the two two-dimensional detection boxes of the same detection category. In the above aligned point cloud data, points that completely coincide with or are closest to the center points of each two-dimensional detection box can be found, and the distance between any two two-dimensional detection boxes is calculated based on these points in the point cloud.

[0049] In this embodiment, two-dimensional information and three-dimensional information can also be fused for duplicate removal and positioning. Based on this, the above method can also include the step of performing overlap detection on the two-dimensional detection boxes.

[0050] Specifically, overlap detection can be performed on multiple two-dimensional detection boxes of the same detection category; then, duplicate removal processing is performed on multiple two-dimensional detection boxes with an overlap degree greater than the overlap degree threshold and a distance less than the preset threshold.

[0051] Exemplarily, based on the transformation matrix between the camera coordinate system and the lidar coordinate system, the three-dimensional coordinates of the center point of the detection box of the detected target in the radar coordinate system can be obtained. The coordinate distances of the center points of multiple detection boxes of the same detection category in the radar coordinate system are calculated. If it is less than the set threshold, they are considered to be the same target. Second, for each two-dimensional detection box of the detected target, the overlapping area between it and the detection boxes of other detected targets is calculated. If the overlap degree of the two boxes in the image is relatively high (higher than the set threshold), they are considered to be possibly the same target. Finally, combining the overlapping degree of the two-dimensional bounding boxes and the distance error in the three-dimensional space, it is determined whether the two detection boxes are duplicates. If both of these conditions are met, it is considered a duplicate detection.

[0052] S108, convert the spatial coordinates of the above duplicate-removed two-dimensional detection box to the base coordinate system of the robotic arm to obtain the spatial coordinates in the base coordinate system of the robotic arm.

[0053] The spatial coordinates of the duplicate-removed two-dimensional detection box are relative to the radar coordinate system, and it is also necessary to convert them from the radar coordinate system to the base coordinate system of the robotic arm.

[0054] In this embodiment, by introducing a unified coordinate transformation framework based on the radar coordinate system, the accuracy and dynamic adaptability of target positioning are enhanced. This transformation framework can accurately complete the transformation between the camera, lidar, and robotic arm coordinate systems, solve the deficiencies of traditional coordinate transformation methods in dynamic environments, support the efficient grasping operation of the robotic arm, and improve the dynamic response ability of the robot.

[0055] S110. Send the spatial coordinates in the above-mentioned robotic arm base coordinate system to the robotic arm so that the robotic arm performs a grasping operation.

[0056] Send the above-mentioned transformed spatial coordinates to the robotic arm. The robotic arm takes the transformed object position as input, performs the grasping operation task of the robotic arm, or processes it according to the control algorithm of the robotic arm.

[0057] After obtaining the above-mentioned deduplicated two-dimensional detection frames, visualization can be performed to verify the fusion and deduplication effects. This visualization process may include: marking the deduplicated two-dimensional detection frames in the point cloud data, or marking the deduplicated two-dimensional detection frames on the image, as well as marking the class labels and confidence levels of the deduplicated two-dimensional detection frames.

[0058] Furthermore, after visualization, the above-mentioned distance threshold and overlap threshold can also be adjusted according to the visualization results. Users can visually verify and adjust the deduplication criteria, such as the distance threshold, overlap threshold, etc., to improve the detection accuracy.

[0059] The target detection deduplication and positioning method for a dual-arm collaborative wheeled tracked robot provided by an embodiment of the present invention incorporates three-dimensional information into the deduplication algorithm, which can more accurately determine the target position, avoid repeated detection of the same target by multiple sensors, reduce the data volume, effectively reduce the target positioning error, and improve the reliability of subsequent grasping operations; by combining the precise three-dimensional point cloud data provided by the lidar, it effectively makes up for the limitations of the depth camera in complex environments and significantly improves the target recognition ability; by introducing a unified coordinate transformation framework based on the radar coordinate system, it enhances the accuracy and dynamic adaptability of target positioning, solves the deficiencies of traditional coordinate transformation methods in dynamic environments, supports the efficient grasping operation of the robotic arm, and improves the dynamic response ability of the robot.

[0060] The following embodiments will detail the target detection deduplication and positioning method for a dual-arm collaborative wheeled tracked robot provided by an embodiment of the present invention. Figure 2 The specific process schematic diagram of this method is shown, and this method can be carried out according to the following steps:

[0061] Step 1. Align the camera data and lidar data in terms of time.

[0062] To ensure that data from different sensors (such as cameras and lidars) can be correctly fused and subsequently processed, it is first necessary to ensure the alignment of their timestamps. Different sensors may have different timestamps due to factors such as sampling frequency and data transmission delay. Timestamp alignment can eliminate these differences and ensure that data can be accurately mapped to the same time point. The detailed description of this step is as follows:

[0063] Step 1.1 Data acquisition, including synchronous data collection, obtaining the raw data from the camera and lidar, and extracting their timestamp information (header.stamp).

[0064] First, ensure that the camera and lidar sensors are synchronously started at the hardware level to ensure that they can collect data within a similar time frame. Second, the output of each sensor will carry a timestamp (header.stamp), which marks the specific time of data collection and is the basis for timestamp alignment. The timestamps of the camera and lidar are usually generated by their respective hardware systems, but due to differences between the hardware, the timestamps often vary. For subsequent deduplication and localization operations, it is necessary to first extract these timestamps from the outputs of the two sensors and ensure that subsequent synchronization operations can be based on these timestamps.

[0065] Step 1.2 Use the ROS message synchronization mechanism for timestamp alignment.

[0066] First, create sensor subscribers to receive the corresponding sensor data. Then configure the synchronizer, that is, select different synchronization methods and pass the subscribers of each sensor to the synchronizer for processing. For example, when the data sampling frequencies of the two sensors are basically the same, message_filters::TimeSynchronizer can be used for precise time synchronization. It will ensure that at each callback function call, the data from the two sensors come from the same moment, that is, the timestamps are exactly the same. When the sampling frequencies of the two sensors are not exactly the same, message_filters::ApproximateTimeSynchronizer is used to handle it. It will synchronize the data with approximately matching timestamps according to the set maximum time difference (such as 0.1 seconds). If the difference between the timestamps of the two data is less than the set tolerance time window, the callback function will be triggered.

[0067] Step 1.3 Callback processing. After using the synchronization mechanism, ROS will pass the data with aligned timestamps to the callback function.

[0068] In this callback function, developers can simultaneously process data from the camera and lidar, and perform further calculations and data fusion according to requirements. At this time, since the camera and radar data are aligned under the same timestamp, it ensures accurate results when performing data processing and fusion in subsequent steps. For example, subsequent operations such as deduplication and coordinate system conversion can be based on data with the same timestamp, avoiding potential errors caused by data collected at different time points.

[0069] Step 2: Camera Data Processing and 2D Detection Box Extraction

[0070] In this step, the goal is to process the images captured by the camera and extract the 2D detection boxes of the targets. This process usually includes camera initialization, image acquisition, image preprocessing, object detection, and extracting 2D detection boxes from the detection results. The detailed description of this step is as follows:

[0071] Step 2.1 Camera Initialization. Before performing image processing or object detection operations, it is first necessary to initialize the camera device and perform necessary configurations. This usually includes specifying the camera width, height, frame rate, setting the depth stream and color stream, etc. If the configuration is successful, proceed to the next step; if the configuration fails, stop the program.

[0072] Step 2.2 Image and Depth Data Acquisition.

[0073] For example, acquire image and depth data through a RealSense camera and perform necessary preprocessing. Obtain the image stream (RGB image and depth image) from the RealSense camera, and use the CvBridge library in ROS to convert the ROS image message into an OpenCV image. Before processing, the depth image can be synchronized with the RGB (RED, BLUE, GREEN) image to facilitate subsequent 3D positioning (for example, depth information helps identify the distance between the object and the camera). The preprocessing steps for the original image mainly include resizing, color conversion, etc.

[0074] Step 2.3 Object Detection and 2D Detection Box Extraction.

[0075] After completing image preprocessing, it is necessary to perform object detection on the RGB image obtained from the camera through a pre-trained object detection model, such as an ONNX (OpenNeural Network Exchange) model, and input the image into the object detection model. The output of the object detection model is usually an array containing all detection results, each of which includes the object's category ID and the coordinates of the bounding box (i.e., the object's detection box) and confidence. The bounding box is represented by the coordinates of the upper left corner and the lower right corner, or by the center point and the width and height. For example, [x, y, w, h], where x, y are the coordinates of the center point of the bounding box, and w, h are the width and height of the bounding box. This coordinate is in the image coordinate system, that is, two-dimensional, and will be converted using the above coordinates later and then used for deduplication judgment.

[0076] In order to facilitate subsequent processing and visualization, in this step, the detection box coordinates can be converted into standard [x mon ,y mon ,x max ,y max ] format, ensuring that subsequent deduplication and spatial judgment can use these standardized coordinates, where x min ,y min is the coordinate of the upper left corner, x max ,y max is the coordinate of the lower right corner. Finally, draw the bounding box and display the category label and confidence. Visualize the detection box on the image. Each box corresponds to a detected object.

[0077] Step 3: Fusion of 3D information to remove relocation

[0078] In order to optimize the accuracy of target recognition, reduce the probability of misidentification, and avoid the situation where the same target may be detected multiple times due to overlapping views of multiple cameras, the two-dimensional detection frame obtained in step 2 is fused with the three-dimensional point cloud data of the lidar to improve the detection accuracy. The specific steps are as follows:

[0079] Step 3.1 obtains the depth information and calculates the 3D coordinates of each detection box.

[0080] Using the depth value of each pixel in the depth map, the 2D image coordinates are mapped to the 3D space to obtain the 3D coordinates (i.e. X, Y, depth Z) of each detection frame. Combined with the camera intrinsic parameters (focal length, principal point, etc.), the spatial position of each object is inferred through the depth information. For example, the pixel coordinates (x, y) are converted to the three-dimensional coordinates (X, Y, Z) in the camera coordinate system using the following formula:

[0081]

[0082] where (x, y) are pixel coordinates in the image, (c x , c y ) are the coordinates of the camera principal point, f x , f y is the focal length of the camera, and Z is the depth value obtained from the depth map.

[0083] Step 3.2 LiDAR and point cloud data acquisition and preprocessing.

[0084] Obtain the original point cloud data from the LiDAR. Each point in the point cloud contains position coordinates (X, Y, Z) and intensity values. Then, preprocess the point cloud data, including operations such as denoising, filtering, and downsampling, to remove irrelevant or noisy points and reduce the computational burden.

[0085] Step 3.3 Align the camera coordinate system with the LiDAR coordinate system. Use the extrinsic parameters (rotation matrix and translation vector) between the camera and the LiDAR to transform the data in the camera coordinate system to the LiDAR coordinate system, or map the LiDAR point cloud to the camera coordinate system.

[0086] Assume the extrinsic parameter matrices of the camera and the LiDAR are R and T. Then, points can be transformed from one coordinate system to another through the following formula:

[0087] P new = R·P old + T

[0088] where P old is the point in the original coordinate system, and P new is the point in the transformed coordinate system.

[0089] Step 3.4 Remove duplicate detection boxes and determine whether objects are duplicates by combining 2D and 3D information.

[0090] First, by comparing the center points of the camera detection boxes with the spatial positions of the LiDAR point cloud, calculate the 3D distance between them. If the distance is less than the set threshold, they can be considered the same object. Second, for each 2D detection box of the camera, calculate the overlapping area with other detection boxes. If the overlap degree of the two boxes in the image is high, they may be considered the same object. Finally, combine the overlapping degree of the 2D bounding boxes and the distance error in 3D space (for example, the 3D position distance of the two boxes is less than a certain threshold) to determine whether the two detection boxes are duplicates. If both conditions are met simultaneously, it is considered a duplicate detection.

[0091] Based on the transformation matrix between the camera coordinate system and the lidar coordinate system, the three-dimensional coordinates of the center point of the detection box of the detected target in the lidar coordinate system can be obtained through the matrix calculation in step 3.3. Calculate the coordinate distances of the center points of multiple detection boxes under the same detection category in the lidar coordinate system. If it is less than the set threshold, it is considered the same target. Secondly, for the 2D detection box of each detected target, calculate its overlapping area with the detection boxes of other detected targets. If the overlap degree of the two boxes in the image is high (higher than the set threshold), it is considered that they may be the same target. Finally, combine the overlapping degree of the 2D bounding boxes and the distance error in the 3D space to jointly judge whether the two detection boxes are duplicates. If both conditions are met simultaneously, it is considered a duplicate detection.

[0092] Step 3.5 Once a detection box is determined to be a duplicate, one of the detection boxes can be removed, keeping the box with a higher confidence level, or keeping the most accurate box according to certain rules (such as preferentially selecting lidar data).

[0093] Step 3.6 Update the bounding boxes according to the new 3D deduplication results to ensure the uniqueness and precise positioning of each object in the 3D space.

[0094] Step 4, Visualization and Deduplication Result Publishing

[0095] To verify the fusion and deduplication effects, the results can be visualized.

[0096] Step 4.1 Mark the deduplicated targets in the point cloud data, or mark the deduplicated 2D detection boxes on the image, and mark the class labels and confidence levels.

[0097] In addition, through visual verification, the deduplication criteria, such as distance threshold and confidence threshold, can be further adjusted to improve the detection accuracy. The deduplication threshold (in step 3.4) can also be adjusted according to the dynamic changes of the environment and target objects. For example, a smaller distance threshold is used for deduplication in a low-density environment, while a larger threshold is used in a dense environment. Adaptive threshold adjustment can be achieved through environmental perception (such as sensor density, spatial distribution of targets, etc.).

[0098] Step 4.2 In a robot system, it is a very common and efficient practice to use ROS for data exchange and communication. It enables each module of the robot system to share data efficiently and in real time, without the need to couple complex communication protocols and with high real-time performance. Therefore, in order to facilitate the precise grasping and target operation of the subsequent robotic arm, first use CvBridge to convert the OpenCV image format to the ROS message format, and then publish the bounding box information through the ROS topic, including information such as the coordinates, class, and confidence level of the bounding box, so that other nodes (such as the control system or the grasping system) can subscribe and perform corresponding actions.

[0099] In this step, publishing the object detection results (including images, bounding box information, and 3D coordinates) through ROS not only simplifies the communication between modules but also enhances the scalability and real-time performance of the system. At the same time, to avoid a complex calibration process, converting the data to the lidar coordinate system first can simplify the coordinate transformation with the robotic arm control system and improve the accuracy.

[0100] Step 5: Coordinate Transformation

[0101] In the deduplicated data, the 3D spatial coordinates and 2D detection boxes of the object can be used for the grasping and positioning of the robotic arm, and the robotic arm will perform object positioning and grasping based on this information. However, at this time, the spatial position of the object is relative to the radar coordinate system, and it is necessary to convert the 3D position of the object from the radar coordinate system to the base coordinate system of the robotic arm.

[0102] Step 5.1 Obtain the transformation relationship from the radar coordinate system to the base coordinate system of the robotic arm through the installation position relationship between the robotic arm and the radar sensor, that is, obtain the rotation matrix R and the translation vector T through extrinsic calibration.

[0103] Step 5.2 In ROS, developers do not need to manually handle the calculation of matrices and vectors, but can automatically query and apply these coordinate transformations through the tf2 library. tf2 (Transform Library) provides a mechanism for broadcasting and receiving coordinate transformations. Developers can use tf2 to listen for the transformations between sensors (such as lidar and the base of the robotic arm).

[0104] Step 5.3 Use the transformed object position as input to execute the grasping or operation tasks of the robotic arm, or process it according to the control algorithm of the robotic arm.

[0105] Step 6: Object Grasping Feedback and Real-Time Correction

[0106] Steps 1 to 5 will be cycled in each control cycle, continuously performing processes such as object detection, deduplication and positioning, coordinate transformation, and sending object grasping commands.

[0107] When performing the grasping operation, judge whether the grasping is successful through the feedback information of the robotic arm (such as grasping force, grasping angle, etc.). If the grasping fails, re-attempt to grasp through real-time correction strategies (such as adjusting the grasping posture, force, etc.).

[0108] Based on this, the above method can also include: adjusting the distance threshold and overlap threshold according to the usage environment of the dual-arm collaborative wheeled tracked robot and the spatial distribution of the target objects.

[0109] This step can achieve closed-loop control, enhancing the robustness and flexibility of the grasping task. This adjustment and correction process is carried out continuously, ensuring that the robotic arm can adapt to the dynamically changing environment and ensuring successful grasping. Once the robotic arm successfully grasps the target, the control system will send a confirmation signal indicating the end of the grasping process; if the grasping fails, the system will decide whether to re-grasp or terminate the task based on the reason for the failure, and perform appropriate logging and error feedback.

[0110] The embodiment of the present invention reduces the generation of redundant data through an accurate de-duplication and localization method. In the area where the multi-sensor perspectives overlap, traditional two-dimensional de-duplication methods are prone to redundant data, while incorporating three-dimensional information into the de-duplication algorithm can more accurately determine the target position, avoiding the situation where the same target is repeatedly detected by multiple sensors. This not only reduces the data volume but also effectively reduces the target localization error and improves the reliability of subsequent grasping operations.

[0111] The embodiment of the present invention improves the accuracy and robustness of target detection by fusing camera and lidar data. Compared with traditional two-dimensional image-based detection methods, traditional depth cameras often have large errors when obtaining three-dimensional position information, resulting in low accuracy of target detection. This embodiment effectively compensates for the limitations of depth cameras in complex environments by combining the accurate three-dimensional point cloud data provided by lidar, significantly enhancing the target recognition ability. Especially in the case of target occlusion and fast movement, it can reduce misidentifications, thereby improving the overall detection accuracy of the robot.

[0112] The embodiment of the present invention enhances the accuracy and dynamic adaptability of target localization by introducing a unified coordinate transformation framework based on the radar coordinate system. This method can accurately complete the transformation between the camera, lidar, and robotic arm coordinate systems, solving the deficiencies of traditional coordinate transformation methods in dynamic environments, supporting the efficient grasping operation of the robotic arm, and improving the dynamic response ability of the robot.

[0113] The embodiment of the present invention also provides a target detection de-duplication and localization system for a dual-arm collaborative wheeled and tracked robot, which can achieve the same technical effects as the above-mentioned target detection de-duplication and localization method for a dual-arm collaborative wheeled and tracked robot.

[0114] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be determined by the scope defined by the claims.

[0115] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0116] The various embodiments in this specification are described in a progressive manner, with each embodiment highlighting the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the above embodiments, the description is relatively simple, and reference can be made to the description in the method part for related parts.

[0117] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A target detection and de-relocation method for a dual-arm collaborative wheeled robot, characterized in that: The method comprises: Acquire camera data and lidar data, and perform time alignment on the camera data and the lidar data; Perform object detection based on the aligned camera data to obtain multiple two-dimensional detection frames; Performing distance detection on the multiple two-dimensional detection frames of the same detection category based on the aligned laser radar data, and performing deduplication processing on the multiple two-dimensional detection frames whose distances are less than a distance threshold to obtain deduplicated two-dimensional detection frames; Convert the spatial coordinates of the two-dimensional detection frame after deduplication to a robot base coordinate system to obtain the spatial coordinates in the robot base coordinate system; The spatial coordinates in the robot base coordinate system are sent to the robot so that the robot performs a grasping operation.

2. The method according to claim 1, characterized in that The performing distance detection on the multiple two-dimensional detection frames of the same detection category according to the aligned laser radar data includes: Based on a conversion matrix between the camera coordinate system and the radar coordinate system, converting the three-dimensional coordinates of the two-dimensional detection frame in the camera coordinate system into the three-dimensional coordinates in the radar coordinate system; the three-dimensional coordinates include the center point coordinates and the depth value of the two-dimensional detection frame; Determine, according to the three-dimensional coordinates of the center point of the two-dimensional detection frame in the radar coordinate system, the radar point cloud coordinates closest to the center point in the aligned laser radar data; The distance between the two two-dimensional detection frames is calculated based on the radar point cloud coordinates corresponding to any two two-dimensional detection frames of the same detection category.

3. The method according to claim 2, characterized in that Before performing deduplication processing on the plurality of two-dimensional detection frames whose distance is less than a preset threshold, the method further includes: Performing overlapping detection on a plurality of the two-dimensional detection frames of the same detection category; Deduplication processing is performed on the multiple two-dimensional detection frames whose overlap is greater than the overlap threshold and whose distance is less than the preset threshold.

4. The method according to claim 3, characterized in that The deduplication process comprises: The two-dimensional detection frame with the highest confidence among the multiple two-dimensional detection frames whose overlap is greater than the overlap threshold and whose distance is less than the preset threshold is retained.

5. The method according to claim 1, characterized in that The step of converting the spatial coordinates of the two-dimensional detection frame after deduplication into a robot base coordinate system to obtain the spatial coordinates in the robot base coordinate system includes: The transformation relationship from the radar coordinate system to the robot arm base coordinate system is determined by the installation position relationship between the robot arm and the radar sensor; Based on the transformation relationship, the spatial coordinates of the deduplicated two-dimensional detection frame in the radar coordinate system are converted to the robotic arm base coordinate system to obtain the spatial coordinates in the robotic arm base coordinate system.

6. The method according to claim 1, characterized in that The time aligning the camera data and the laser radar data includes: The camera data and the lidar data are time aligned based on the ROS message synchronization mechanism, and the camera data and the lidar data with the same timestamp or a timestamp difference less than a set time difference are aligned.

7. The method according to claim 3, characterized in that The method further includes: visualizing the two-dimensional detection frame after deduplication; the visualization includes: Mark the deduplicated 2D detection frame in the point cloud data, or mark the deduplicated 2D detection frame on the image, as well as the category label and confidence of the deduplicated 2D detection frame.

8. The method according to claim 7, characterized in that The method further comprises: The distance threshold and the overlap threshold are adjusted according to the visualization result.

9. The method according to claim 3, characterized in that: The method further comprises: The distance threshold and overlap threshold are adjusted according to the use environment of the dual-arm collaborative wheeled-track robot and the spatial distribution of the target object.

10. A target detection and relocation system for a dual-arm collaborative wheeled robot, characterized in that: Used to execute the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Laser radar and vision fusion integrated target tracking system and method

    CN112731371A

  • Multi-target detection method and device based on camera and laser radar fusion

    CN117593620A

Cited By

  • Target duplicate removal method and device, electronic equipment and storage medium

    CN122391676A