Multi-point cloud view angle robotic arm grabbing system based on deep learning
Through the multi-view point cloud splicing technology based on deep learning and GraspNet crawling estimation network, the problem of incomplete point cloud data in complex crawling scenarios is solved, high-precision three-dimensional modeling and crawling pose estimation are realized, and the grab accuracy and success rate of the robot arm are significantly improved.
Patent Information
- Application Number
- CN202510269580.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-06
AI Technical Summary
In complex crawling scenarios, traditional point cloud data acquisition methods have incompleteness and low precision, resulting in limited success rate of robot crawling tasks.
Using multi-view point cloud stitching technology based on deep learning, point cloud data is collected from multiple perspectives through a depth camera, and deep learning models such as RPMNet are used for high-precision stitching to generate a complete three-dimensional object model. At the same time, GraspNet grab estimation network calculates the optimal grab pose of an object and passes it to the robot arm through the ROS service communication mechanism.
It significantly improves the splicing accuracy of point cloud data and the reliability of grab pose estimation, improves the grab accuracy and success rate of the robot arm in complex dynamic environments, and solves the bottleneck problems of traditional methods in high-precision point cloud generation and real-time feedback.
Smart Images

Figure CN120107366A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of three-dimensional robotic arm grasping, and specifically relates to a multi-view point cloud registration method based on deep learning, which generates a complete three-dimensional point cloud model of an object through high-precision splicing, thereby more accurately describing the spatial information of the object and improving the accuracy and reliability of subsequent three-dimensional grasping. Background Art
[0002] At present, with the rapid development of deep learning technology, the ability of intelligent robots in environmental perception and task execution has been significantly improved. Through the deep learning-based visual model, the robot can efficiently and accurately detect, identify and semantically segment the target objects in the environment. At the same time, with the help of multimodal data fusion, such as the combination of visual information and depth information, the robot's scene understanding ability is further enhanced. This enables robots to show higher efficiency and flexibility when performing complex tasks such as object grasping, assembly, audio-visual navigation, etc., and promotes the widespread application of intelligent robot technology.
[0003] However, when faced with unstructured complex scenes, target objects often have overlapping, occluded or complex geometric features, which makes the point cloud data obtained through a single fixed perspective often incomplete and difficult to meet the needs of high-precision 3D modeling. Due to the defects of point cloud data, the accuracy of robots in tasks such as target detection and posture estimation will be seriously affected, thus limiting the success rate of robot grasping tasks. In this context, how to improve the integrity and accuracy of point cloud data, and then enhance the perception ability of robots, has become an urgent problem to be solved in the field of robotics research.
[0004] In order to solve the above problems, multi-view point cloud registration methods have gradually attracted widespread attention. By dynamically adjusting the position and posture of the visual sensor, active vision can obtain point cloud data of the target object from multiple perspectives, thereby significantly improving the visibility and detection accuracy of the target. Compared with single-view point cloud acquisition, multi-view acquisition can effectively make up for the shortcomings caused by object occlusion or blind spots, and obtain more comprehensive point cloud information. Despite this, traditional point cloud registration methods still face several challenges: First, the registration accuracy of these methods is highly dependent on parameter settings, and these parameters usually need to be adjusted based on experience. The optimal parameters in different application scenarios vary greatly, resulting in a lack of unified standards; second, traditional methods have high requirements for computing resources, especially in dynamic environments, processing point cloud data from multiple perspectives has real-time problems, which often leads to calculation delays; finally, with the increasing complexity of application scenarios, the efficiency and accuracy of traditional registration methods often cannot meet actual needs, especially in complex grasping or detection tasks with high-precision alignment and real-time feedback, it is difficult to ensure stable performance. Therefore, traditional methods face significant performance bottlenecks in the efficient fusion of multi-view point cloud data, and more efficient and intelligent algorithms are urgently needed to meet these challenges.
[0005] In summary, traditional point cloud registration methods have significant deficiencies in dealing with complex grasping scenarios, especially in terms of the accuracy and efficiency of point cloud data fusion. In order to solve this problem, combining multi-view point cloud stitching technology with deep learning methods can provide strong technical support for high-precision 3D modeling and significantly improve the accuracy and stability of the robot arm in actual grasping tasks. How to break through these technical bottlenecks and optimize point cloud data processing and grasping posture prediction algorithms has become one of the core challenges in current intelligent robot research. Summary of the invention
[0006] The purpose of the present invention is to provide a multi-point cloud view robotic arm grasping system based on deep learning, aiming to solve the problems of incomplete point cloud data, insufficient registration accuracy, and low grasping estimation accuracy in complex grasping scenarios in the prior art. By introducing multi-view point cloud stitching technology and deep learning grasping estimation method, the present invention can significantly improve the grasping accuracy and reliability of the system in complex dynamic environments, thereby effectively improving the success rate of grasping tasks of the robotic arm.
[0007] In order to achieve the above object, the present invention provides the following technical solutions:
[0008] A multi-point cloud perspective robotic arm grasping system comprises the following steps:
[0009] S1: Use the depth camera to obtain point cloud data from multiple perspectives according to the graph optimization algorithm
[0010] S2: Use deep learning methods to stitch point cloud data with high precision to generate a complete 3D model of the object;
[0011] S3: Use the grasp estimation network to calculate the optimal grasping pose of the object;
[0012] S4: The grasping posture information is transmitted to the robot arm through the Robot Operating System (ROS) service communication mechanism;
[0013] S5: Plan and optimize the robot's grasping path based on the MoveIt! package;
[0014] S6: Subscribe to joint status information and drive the robot arm to complete the actual grasping action.
[0015] The present invention is a multi-point cloud perspective robotic arm grasping system based on deep learning. In step S1, a depth camera is fixedly installed at the end of the robotic arm, the initial position of the camera is set, and point cloud data of the scene is collected from different positions. The posture of each camera is used as a node, and the matching results between point clouds are used as edges to construct a graph containing nodes and edges. The graph optimization algorithm is applied to optimize the posture of the camera by minimizing the errors of all edges. Finally, the optimized camera posture is obtained, so that the camera scans the scene from different angles and covers the area of all stacked objects.
[0016] After processing, these multi-view data are converted into corresponding point cloud data, minimizing the data loss caused by object occlusion or blind spots. This method effectively reduces the risk of grasping failure in traditional single-view point cloud acquisition in complex scenes and improves the success rate of subsequent grasping tasks.
[0017] In step S2, the present invention inputs point cloud data collected from different perspectives into the deep learning model RPMNet, and uses its excellent point cloud feature extraction and alignment capabilities to accurately implement high-precision point cloud stitching operations. Compared with traditional point cloud stitching methods, the RPMNet model can more accurately identify and align point cloud features, reduce stitching errors, and thus provide higher-quality three-dimensional point cloud input for subsequent grasping estimation. This high-precision three-dimensional point cloud model provides a more reliable foundation for subsequent grasping tasks, greatly improving the grasping accuracy and the overall stability of the system.
[0018] In steps S3 and S4, the system uses the powerful point cloud feature extraction capability and deep reasoning function of the GraspNet grasp estimation network to analyze the complete 3D point cloud model and calculate the optimal grasping posture of the object. In order to improve the flexibility of the system, the grasping posture calculation process is completed on the server side deployed on the Windows system. The calculation results are transmitted to the robot arm through the ROS service communication mechanism, ensuring stable and efficient cross-platform data interaction. This mechanism effectively avoids the problems caused by the distribution of computing resources or performance differences between devices, and realizes efficient and accurate grasping instruction generation and execution.
[0019] In step S5, the present invention makes full use of the powerful path planning and dynamic optimization capabilities of the MoveIt! package, combines the 3D point cloud model of the object and the kinematic parameters of the robot arm, and plans the most suitable grasping path. During the path planning process, the system can identify and bypass potential obstacles through a dynamic optimization algorithm, while taking into account the joint limitations of the robot arm to ensure the safety of the grasping path and the accuracy of the grasping action. Even in complex environments, the system can effectively adjust the path, optimize the grasping action, and avoid accidental collisions and grasping failures.
[0020] In step S6, the system monitors the motion status of each joint of the robot arm in real time by subscribing to the joint status information in ROS to ensure that the planned path is highly consistent with the actual grasping action performed by the robot arm. At the same time, the system introduces a feedback mechanism for the state of the object after grasping, and determines whether the grasping is successful by detecting the state of the grasped object. If the grasping fails, the system can quickly feedback and provide a basis for improvement, providing an important reference for the optimization of subsequent grasping tasks. This mechanism further improves the stability and adaptability of the system, enabling it to cope with different grasping scenarios and task requirements.
[0021] The present invention significantly improves the stitching accuracy of point cloud data and the reliability of grasping pose estimation by combining multi-view point cloud stitching technology with deep learning grasping estimation method. Through cross-platform data processing and communication mechanism, the system can be flexibly applied in a variety of complex grasping scenarios, solving the bottleneck problems of traditional technologies in high-precision point cloud generation, dynamic environment processing and real-time feedback, thereby significantly improving the grasping accuracy, stability and adaptability of the robot arm, and meeting the actual application needs in diverse and complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is the system flow of the present invention;
[0023] Figure 2 It is the multi-view point cloud fusion process of the present invention;
[0024] Figure 3 It is the network structure of the point cloud registration of the present invention;
[0025] Figure 4 is a system block diagram of the present invention;
[0026] Figure 5 It is a node information diagram of ROS and Windows server of the present invention; Figure (5) is the OFDM signal processing flow; Specific implementation plan
[0027] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings.
[0028] This design adopts a point cloud registration method based on deep learning, using neural networks to automatically learn the features of point cloud data, and autonomously optimizes the registration parameters according to the characteristics of the data, fundamentally solving the problem of manually setting parameters in traditional point cloud registration algorithms. Unlike traditional methods, deep learning models can dynamically adjust parameters based on the input point cloud data without human intervention, and automatically identify the optimal matching relationship between point clouds. This not only reduces the need for human intervention, but also greatly improves the automation and robustness of the registration process, especially when facing complex or dynamic environments, and can better adapt to changes. The specific implementation process is described as follows:
[0029] First, by fixing the depth camera at the end of the robot arm, setting the initial position of the camera, collecting point cloud data of the scene from different positions, taking the pose of each camera as a node and the matching results between point clouds as edges, a graph containing nodes and edges is constructed. The graph optimization algorithm is applied to optimize the camera pose by minimizing the error of all edges. Finally, the optimized camera pose is obtained, so that the camera scans the scene from different angles and covers the area of all stacked objects. Finally, depth maps and RGB images of multiple perspectives are obtained, such as Figure 1 As shown. By extracting depth information and color information from these images, they can be converted into corresponding point cloud data. Next, point cloud filtering technology is used to process these point cloud data from different perspectives to remove noise and improve the quality of the point cloud, making the point cloud data from each perspective more accurate and reliable. Then, the high-quality point cloud data from all perspectives are summarized and input into the RPMNet network, and its powerful feature extraction and alignment capabilities are used to effectively splice and fuse these point cloud data. Finally, a complete and high-precision 3D model is obtained, which realizes the seamless splicing and reconstruction of point cloud data from multiple perspectives.
[0030] In the process of point cloud registration, RPMNet is the key technology. Figure 2is the network structure of RPMNet. RPMNet is a point cloud registration method based on deep learning. Compared with traditional methods such as ICP, it can still achieve high-precision point cloud alignment in the presence of noise, occlusion or sparse point clouds. The network structure of RPMNet is mainly composed of three core parts: point cloud feature extraction, parameter prediction and matching matrix calculation. Each part plays a vital role in the point cloud registration process to ensure the accuracy and stability of the final registration.
[0031] The first module of RPMNet is responsible for extracting effective features from the input point cloud data. This module is based on PointNet++, an efficient point cloud feature learning network that can capture global structural information while retaining local geometric details. PointNet++ uses multi-level convolution operations and pooling layers to first extract local features in a smaller neighborhood, and then gradually expand to a larger range to learn more comprehensive point cloud distribution characteristics. The neighborhood information of each point cloud is encoded, and feature extraction is performed through a shared multi-layer perceptron (MLP), and finally the global features of the point cloud as a whole are generated through maximum pooling. Finally, the local and global features are combined into a high-dimensional feature vector F_X and F_Y, which represent the features of the two input point clouds respectively, as the basis for subsequent point cloud matching calculations.
[0032] After feature extraction is completed, the second module of RPMNet further processes these features to predict the key parameters required for point cloud matching. This module merges the features F_X and F_Y of the two point clouds through feature stitching and inputs them into a shared multi-layer perceptron (MLP) network for further processing. At this stage, the network automatically learns how to calculate the transformation parameters required by the traditional RPM algorithm, mainly including key parameters such as rotation matrix and translation vector. These parameters are the core of point cloud registration and directly determine how to perform rigid alignment between point clouds.
[0033] Once the best matching parameters are obtained, the last module of RPMNet will use these parameters to calculate the final point cloud matching matrix. This module first combines the previously extracted features F_X and F_Y with the predicted parameters and calculates the initial matching matrix, which is then optimized using the Sinkhorn normalization method to ensure the numerical stability and convergence of the matching matrix. Ultimately, these optimized matching parameters will be used to calculate the transformation matrix of the point cloud, achieving high-precision alignment of the point cloud.
[0034] Compared with the traditional RPM algorithm, RPMNet no longer relies on manually set parameters, but automatically learns how to optimize these parameters through deep neural networks. In this way, the system can maintain high adaptability and accuracy in the face of complex and dynamically changing environments, and can adapt to different environments and object changes, significantly improving the accuracy and efficiency of point cloud registration.
[0035] After obtaining the complete 3D point cloud model, the system uses the GraspNet deep learning network to estimate the grasping posture of the target object. The overall process is as follows: Figure 3 As shown in the figure. GraspNet is a deep learning network designed specifically for object grasping tasks. It can accurately extract the geometric information of an object from point cloud data and predict the most suitable grasping posture based on the shape and features of the object. Using the reasoning ability of the deep learning model, GraspNet can automatically calculate the optimal grasping posture of an object, including the grasping position, direction, and appropriate contact surface. Through this process, the system greatly reduces the steps that require manual adjustment in traditional grasping methods, and provides more efficient and accurate grasping posture estimation.
[0036] Once the grasping posture information is calculated, the system transmits this information to the robot arm control node through the ROS service mechanism. ROS (Robot Operating System) provides an efficient message transmission and service call mechanism to ensure that the grasping posture data can be seamlessly transmitted to the robot arm control system, realizing cross-platform data interaction and real-time command transmission. The introduction of this mechanism greatly improves the flexibility and real-time response capability of the system, ensuring that the grasping task can be executed quickly.
[0037] The system uses the MoveIt! package to plan and optimize the path. MoveIt! is a mature open source robot motion planning framework that can plan an optimal grasping path based on the robot arm's kinematic model, the object's 3D point cloud information, and the target grasping posture. During the path planning process, MoveIt! not only takes into account the robot arm's kinematic constraints, but also adjusts the path planning in real time so that the robot arm can smoothly reach the optimal grasping position. In this process, the system also integrates a dynamic optimization algorithm that can identify potential obstacles in real time and dynamically adjust the path to avoid collision risks. Whether there are static obstacles in the operating environment or the robot arm encounters a dynamically changing environment when performing a grasping action, the optimization algorithm can respond in time to ensure the safety and effectiveness of the path.
[0038] Through this series of processing, the system can achieve efficient, accurate and safe grasping path planning, ensuring that the robot arm can smoothly and accurately perform grasping tasks. This not only improves the grasping success rate, but also enhances the robustness and flexibility of the system in complex environments.
[0039] The specific node diagram is as follows Figure 4 As shown in the figure, the workflow of each node in the system is shown. Data transmission and control are carried out between each node through ROS communication. Specifically, the depth camera node is responsible for acquiring RGB images, depth images and mask images, and outputting point cloud data from multiple perspectives. The point cloud stitching node receives the point cloud data from the depth camera node, stitches it using the deep learning model, and generates a complete three-dimensional point cloud model of the object. The grasping estimation node receives the three-dimensional point cloud model and performs grasping pose estimation, and outputs the optimal grasping pose of the object. The path planning node receives the three-dimensional point cloud model of the object and performs path planning and dynamic optimization through the MoveIt! function package, and finally outputs the optimized grasping path. The robot arm control node receives the grasping path information from the path planning node, controls the robot arm to perform grasping actions, and monitors the execution status of the robot arm in real time by subscribing to the status information of each joint of the robot.
Claims
1. A multi-point cloud perspective robotic arm grasping system based on deep learning, characterized in that: The system is implemented through the following steps. S1. Obtain point cloud data of objects from multiple perspectives through posture planning based on graph optimization. S2. Use deep learning methods to stitch the acquired point cloud data with high precision to generate a complete three-dimensional model of the object. S3. On the server side of the Windows system, use the deep learning model to infer the grasping posture of the point cloud data and calculate the optimal grasping posture of the object. Through ROSBridge, the calculated grasping posture information is transmitted from the Windows system to the robot arm control system running on the Linux system. S4. After receiving the grasping posture information, the robot arm control system drives the robot arm to complete the grasping action.
2. The multi-point cloud perspective robotic arm grasping system according to claim 1, characterized in that: In step S1, the depth camera is installed at the end of the robot arm to construct a graph model. The nodes in the graph represent different camera perspectives or postures. i Represents an x i The camera pose is usually expressed as (T i ), where T i is the camera's transformation matrix, including position and orientation. ij Represents the constraint relationship between two camera views, usually including relative pose (T ij ) and error term. Define an optimization objective to minimize the square error. The specific formula is as follows: The goal is to find a set of poses that minimizes the error of all constraints, e ij (T i ,T j ) is the error between the two camera poses. The camera poses are adjusted by the Gauss-Newton method. During the optimization process, a set of optimal camera poses will be found so that each camera view can align the point cloud data relative to other viewpoints as accurately as possible, thereby maximizing the coverage of the scene and avoiding blind spots. Ultimately, the optimization result will give the precise pose of each camera node, including the optimal position and orientation of the camera. In this way, RGB images and depth images of objects captured from multiple viewpoints are obtained, and these image data are converted into high-quality point cloud data.
3. The multi-point cloud perspective robot arm grasping system according to claim 1 or 2, characterized in that: In step S2, the RPMNet deep learning model is used to perform efficient feature extraction and alignment on the acquired point cloud data from multiple viewpoints, so as to achieve accurate stitching of the multi-viewpoint cloud data and generate a complete three-dimensional point cloud model of the object.
4. The multi-point cloud perspective robot arm grasping system according to claims 1 to 3, characterized in that: In step S3, the deep learning model performs grasping posture reasoning in the Windows system, uses the trained network to analyze the point cloud data, and calculates the optimal grasping posture of the object. The communication between the Windows system and the Linux system is realized through ROSBridge, and the calculated grasping posture information is transmitted in real time to the robot arm control system running in the Linux system.
5. The multi-point cloud perspective robot arm grasping system according to claims 1 to 4, characterized in that: In step S4, the robot arm control system drives the robot arm to perform the grasping action according to the received grasping posture information, and adjusts the operation path according to the real-time feedback information of the grasping action to ensure the successful execution of the grasping task.
Citation Information
Cited By
Robot visual guidance method and system
CN120852519A
Part surface point cloud acquisition method based on 6D pose estimation
CN120876738A
Space robot sensing system and method based on active vision strategy
CN121200028A
A space robot perception system and method based on active vision strategy
CN121200028B
Robot visual servo control method based on depth point cloud
CN122274993A