A part grasping method, device, equipment and storage medium

CN122606658APending Publication Date: 2026-08-21LANZHOU UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611113990.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-27
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

二维视觉方案难以获取零件的深度与姿态信息,在堆叠遮挡场景下易出现识别漏检、定位偏差,无法满足三维抓取的精度需求;纯三维视觉方案虽能提供完整的空间位姿信息,但现有算法普遍计算复杂度高、算力负载大,难以在工业边缘设备上实现实时运行,且在密集堆叠、严重遮挡的极端工况下,位姿估计易出现歧义,导致抓取成功率下降,无法兼顾识别精度与部署实时性

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122606658A_ABST
    Figure CN122606658A_ABST
Patent Text Reader

Abstract

The application discloses a part grabbing method, device, equipment and storage medium. The method comprises the following steps: acquiring color depth image data and mechanical arm motion state data of a disordered stacked part scene, and obtaining coordinate system conversion parameters through multi-coordinate system calibration and time synchronization processing; completing part identification through a lightweight target detection network, and outputting detection bounding boxes and grabbability labels of each stacked part; extracting corresponding area depth point clouds, combining a part three-dimensional model to perform six-degree-of-freedom pose estimation to obtain a target part space pose; performing online grabbing decision based on an offline grabbing template library to obtain an optimal grabbing pose, and completing trajectory planning optimization in combination with dynamics and scene collision constraints; and controlling a mechanical arm to perform grabbing and integrating to generate an intelligent grabbing control scheme. The application effectively improves the grabbing precision and real-time performance in a stacked occlusion scene, and is suitable for industrial part sorting, feeding and discharging and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of industrial robot vision grasping technology, and in particular to a part grasping method, device, equipment and storage medium. Background Technology

[0002] With the continuous advancement of intelligent transformation in the manufacturing industry, the application of automated operations such as robotic sorting and loading / unloading in industrial production scenarios is becoming increasingly widespread. In actual production sites, industrial parts are often stored in a disordered stacked form. In such scenarios, parts occlude each other and their postures are random and changeable, which places high demands on the robot's target recognition, pose estimation, and stable grasping capabilities. Achieving efficient and reliable grasping of disordered stacked parts is a key link in improving the level of automation and production efficiency of production lines.

[0003] Current solutions for grasping disordered parts typically employ 2D vision or pure 3D vision to locate the parts. 2D vision solutions struggle to acquire depth and pose information of the parts, and are prone to missed detections and positioning errors in stacked and occluded scenarios, failing to meet the accuracy requirements of 3D grasping. While pure 3D vision solutions can provide complete spatial pose information, existing algorithms generally have high computational complexity and heavy computing load, making it difficult to achieve real-time operation on industrial edge devices. Furthermore, in extreme conditions of dense stacking and severe occlusion, pose estimation is prone to ambiguity, leading to a decrease in grasping success rate, making it impossible to balance recognition accuracy and real-time deployment.

[0004] Therefore, how to achieve high-precision identification and stable pose estimation of disordered stacked parts in industrial scenarios with limited computing power, and improve the success rate of grasping and operation efficiency in stacked occlusion environments, is an urgent problem to be solved. Summary of the Invention

[0005] In view of this, the part grasping method, apparatus, device, and storage medium provided in this application embodiment can achieve high-precision identification and stable pose estimation of disordered stacked parts in industrial scenarios with limited computing power, improving the grasping success rate and operational efficiency in stacked occlusion environments. The part grasping method, apparatus, device, and storage medium provided in this application embodiment are implemented as follows: This application provides a part gripping method, including: Acquire color depth image data of a scene with disordered stacked parts and robotic arm motion state data. Perform multi-coordinate system calibration and time synchronization processing on the color depth image data and the robotic arm motion state data to obtain coordinate system transformation parameters. Based on the coordinate system transformation parameters, the color depth image data is input into a lightweight target detection network for part recognition processing to obtain the detection bounding box and graspability label data of each stacked part; Extract the depth point cloud data of the region corresponding to the detection bounding box, and input the depth point cloud data into the three-dimensional model of the part for six-degree-of-freedom pose estimation to obtain the spatial pose data of the target part. The spatial pose data is processed online based on a pre-built offline crawling template library to obtain the optimal crawling pose data. Based on the dynamic constraints of the robotic arm and the scene collision constraints, the optimal grasping pose data is processed for trajectory planning and optimization to obtain grasping motion trajectory data. Based on the grasping motion trajectory data, the robotic arm is controlled to perform part grasping processing to obtain robotic arm pose data and grasping status data; The robotic arm pose data, scene 3D point cloud information, robotic arm operating status, and grasping status data are integrated and processed to obtain an intelligent grasping control scheme.

[0006] In some embodiments, the step of inputting the color depth image data into a lightweight object detection network for part recognition processing based on the coordinate system transformation parameters to obtain detection bounding boxes and graspability label data for each stacked part includes: Data augmentation and annotation segmentation are performed on images of stacked parts with different degrees of occlusion to obtain a parts stacking detection dataset; A lightweight target detection network is obtained by performing cross-stage feature fusion structure reconstruction, lightweight transformation of convolutional units, and attention mechanism embedding on the baseline target detection network. The part stack detection dataset is input into the lightweight object detection network for training and performance verification, resulting in a converged lightweight object detection network. Spatial alignment and input normalization are performed on the coordinate system transformation parameters and the color depth image data to obtain the image data to be detected. The image data to be detected is input into the lightweight object detection network that has been trained and converged for part identification and graspability evaluation, so as to obtain the detection bounding box and graspability label data of each stacked part.

[0007] In some embodiments, the step of extracting depth point cloud data of the region corresponding to the detection bounding box and inputting the depth point cloud data into the 3D model of the part for six-degree-of-freedom pose estimation processing to obtain the spatial pose data of the target part includes: The detection bounding box and the color depth image are subjected to target region cropping and depth back projection processing to obtain the depth point cloud data of the region of interest corresponding to the target part; The three-dimensional model of the part is processed by neural implicit representation construction to obtain renderable part model data; Multiple initial pose sampling processes are performed on the six-DOF pose space to obtain a set of initial pose assumptions corresponding to different postures; The initial pose hypothesis set, the renderable part model data and the depth point cloud data of the region of interest are subjected to rendering comparison and iterative optimization processing to obtain multiple sets of optimized candidate pose data. The candidate pose data are subjected to hierarchical comprehensive scoring and sorting filtering to obtain the spatial pose data of the target part.

[0008] In some embodiments, the online crawling decision processing of the spatial pose data based on a pre-built offline crawling template library to obtain optimal crawling pose data includes: The pre-built offline grasping template library and the spatial pose data are subjected to coordinate system mapping and transformation to obtain a set of candidate grasping poses in the robot arm base coordinate system; The candidate grasping pose set and the scene depth point cloud data are subjected to collision safety verification, grasping width adaptation verification and inverse kinematics reachability verification to obtain a valid candidate pose set that meets the constraints. Multi-objective comprehensive scoring processing is performed on the set of effective candidate poses to obtain comprehensive score data corresponding to each effective pose; The comprehensive scoring data and the set of effective candidate poses are sorted and filtered to obtain the optimal capture pose data.

[0009] In some embodiments, the optimal grasping pose data is processed by trajectory planning and optimization based on the dynamic constraints and scene collision constraints of the robotic arm to obtain grasping motion trajectory data, including: The current joint state data of the robotic arm and the optimal grasping pose data are subjected to sampled collision-free motion planning to obtain joint space geometric path data. The joint space geometric path data and the robotic arm dynamics constraints are subjected to time-optimized parameterization to obtain the initial grasping trajectory data; The initial grasping trajectory data and scene collision constraints are subjected to multi-objective iterative optimization processing to obtain grasping motion trajectory data.

[0010] In some embodiments, controlling the robotic arm to perform part grasping processing based on the grasping motion trajectory data to obtain robotic arm pose data and grasping state data includes: The grasping motion trajectory data is processed by joint space control command conversion to obtain real-time motion control data of each joint of the robotic arm; The real-time motion control data and the robotic arm body are subjected to motion tracking and driving processing to obtain the real-time pose data of the robotic arm end effector. The gripping parameters of the robotic arm's end effector are adapted and the opening and closing control is processed to obtain the part gripping status data. The real-time pose data and the part gripping and holding status data are synchronized and integrated in time to obtain the robot arm pose data and gripping status data.

[0011] In some embodiments, acquiring color depth image data of a scene with disordered stacked parts and robotic arm motion state data, and performing multi-coordinate system calibration and time synchronization processing on the color depth image data and the robotic arm motion state data to obtain coordinate system transformation parameters, includes: Intrinsic parameter calibration and lens distortion correction are performed on the depth camera to obtain the camera's internal imaging parameters; The robotic arm is modeled to obtain the coordinate transformation parameters of each link of the robotic arm; The robotic arm is controlled to move to multiple different spatial poses, and the calibration plate image data and the robotic arm end pose data under the corresponding poses are collected simultaneously to obtain calibration sample data. The calibration sample data is subjected to linear closed-form solution and nonlinear iterative optimization to obtain the rigid body transformation parameters between the camera coordinate system and the robot arm end coordinate system. The image acquisition timing sequence of the depth camera and the motion control timing sequence of the robotic arm are timestamped to obtain multi-system time synchronization parameters. The camera's internal imaging parameters, the coordinate system transformation parameters of each link, the calibration sample data, the rigid body transformation parameters, and the multi-system time synchronization parameters are integrated and processed to obtain the coordinate system transformation parameters.

[0012] This application provides a parts gripping device, comprising: The acquisition module is used to acquire color depth image data of a scene of disordered stacked parts and robotic arm motion state data, and to perform multi-coordinate system calibration and time synchronization processing on the color depth image data and the robotic arm motion state data to obtain coordinate system transformation parameters. The processing module is used to input the color depth image data into a lightweight target detection network for part recognition processing based on the coordinate system transformation parameters, so as to obtain the detection bounding box and graspability label data of each stacked part; The processing module is also used to extract the depth point cloud data of the region corresponding to the detection bounding box, input the depth point cloud data into the three-dimensional model of the part for six-degree-of-freedom pose estimation processing, and obtain the spatial pose data of the target part. The processing module is also used to perform online crawling decision processing on the spatial pose data based on a pre-built offline crawling template library to obtain the optimal crawling pose data. The processing module is also used to perform trajectory planning and optimization processing on the optimal grasping pose data based on the dynamic constraints of the robotic arm and the scene collision constraints, so as to obtain grasping motion trajectory data. The processing module is also used to control the robotic arm to perform part grasping processing based on the grasping motion trajectory data, so as to obtain robotic arm pose data and grasping state data. The integration module is used to integrate and process the robotic arm pose data, scene 3D point cloud information, robotic arm operating status and grasping status data to obtain an intelligent grasping control scheme.

[0013] The computer device provided in this application includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements the method described in this application.

[0014] The computer-readable storage medium provided in this application embodiment stores a computer program thereon, which, when executed by a processor, implements the method described in this application embodiment.

[0015] This application provides a part grasping method, apparatus, device, and storage medium. It acquires color depth image data of a disordered stacked part scene and robotic arm motion state data, and obtains coordinate system transformation parameters through multi-coordinate system calibration and time synchronization processing. Part recognition is achieved through a lightweight target detection network, outputting detection bounding boxes and graspability labels for each stacked part. Depth point clouds of the corresponding region are extracted, and six-degree-of-freedom pose estimation is performed using the part's 3D model to obtain the target part's spatial pose. Online grasping decision-making is performed based on an offline grasping template library to obtain the optimal grasping pose, and trajectory planning optimization is completed by combining dynamics and scene collision constraints. The robotic arm is controlled to perform grasping and an intelligent grasping control scheme is generated. This enables high-precision recognition and stable pose estimation of disordered stacked parts in industrial scenarios with limited computing power, improving the grasping success rate and operational efficiency in stacked occlusion environments, and solving the technical problems mentioned in the background art. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A schematic diagram illustrating the implementation process of a part gripping method provided in an embodiment of this application; Figure 2A schematic diagram illustrating the implementation process of obtaining detection bounding boxes and graspability tag data of each stacked part, provided in an embodiment of this application; Figure 3 This is a schematic diagram of a parts gripping device provided in an embodiment of this application. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0019] The following description of some technologies involved in the embodiments of this application is provided to aid understanding and should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, some descriptions of well-known functions and structures are omitted in the following description.

[0020] This embodiment discloses a parts grasping method, specifically applied to the automated sorting and loading / unloading of disordered stacked linkage parts in industrial production sites. The hardware system upon which this application relies mainly consists of three parts: a vision perception unit, a robotic arm execution unit, and a central control unit. The vision perception unit employs a depth camera with active infrared stereo imaging capabilities to collect color and depth information of the scene. The robotic arm execution unit uses a six-degree-of-freedom collaborative robotic arm with a two-finger parallel adaptive gripper at its end. The depth camera is rigidly fixed to the flange at the end of the robotic arm, forming a visual configuration where the "eye" is on the hand. The central control unit carries a robot operating system, responsible for visual data processing, grasping decision calculation, and motion control command issuance, enabling the coordinated operation of all units.

[0021] Figure 1 This is a schematic diagram of a part gripping method provided in an embodiment of this application, including steps 101 to 107. Figure 1 This is merely one execution order shown in the embodiments of this application and does not represent the only execution order for a part gripping method. Where the final result can be achieved, [the following may be observed]. Figure 1 The steps shown can be performed in parallel or in reverse order.

[0022] Step 101: Obtain color depth image data of the disordered stacked parts scene and robotic arm motion state data. Perform multi-coordinate system calibration and time synchronization processing on the color depth image data and robotic arm motion state data to obtain coordinate system transformation parameters.

[0023] In this embodiment, firstly, color and depth images of the scene containing disordered stacked parts are acquired in real time using a depth camera. Simultaneously, the operating parameters and end-effector pose information of each joint of the robotic arm are read to obtain the robotic arm's motion state data. Then, multi-coordinate system calibration is performed: firstly, intrinsic parameter calibration and lens distortion correction are performed on the depth camera to establish a mapping relationship from pixel coordinates to camera spatial coordinates; secondly, a modified link parameter method is used to kinematically model the robotic arm, determining the coordinate system transformation relationship between adjacent links; thirdly, the robotic arm is controlled to move to multiple different spatial poses, and at each pose, image data from the calibration board and the corresponding end-effector pose data are simultaneously acquired. Through linear closed-form solving combined with nonlinear iterative optimization, the rigid body transformation relationship between the camera coordinate system and the robotic arm end-effector coordinate system is calculated; finally, the image acquisition timing of the depth camera and the motion control timing of the robotic arm are timestamped to achieve time synchronization of multiple systems. The camera intrinsic parameters, link coordinate system transformation parameters, hand-eye rigid body transformation parameters, and time synchronization parameters are integrated to obtain complete coordinate system transformation parameters, which serve as the basis for subsequent spatial coordinate mapping.

[0024] Step 102: Based on the coordinate system transformation parameters, the color depth image data is input into the lightweight target detection network for part recognition processing to obtain the detection bounding box and graspability label data of each stacked part.

[0025] In this embodiment, the acquired color depth images are preprocessed for spatial alignment using coordinate system transformation parameters and then input into a pre-trained lightweight object detection network. This lightweight object detection network is based on a single-stage detection network architecture and adapts to stacked parts detection scenarios through three core improvements: First, a cross-stage collaborative feature fusion module is introduced to reconstruct the network's neck structure, enhancing the bidirectional interaction between shallow detail features and deep semantic features; second, spatial and channel-based convolutional optimization is used to optimize the convolutional units within the network, suppressing feature information redundancy and strengthening the extraction of effective features; third, a spatially efficient attention module is embedded in the feature extraction structure at the network's end, improving the ability to distinguish target parts in stacked occlusion scenarios. During network training, a dataset of industrial parts stacked images covering three levels of occlusion—unoccluded, slightly occluded, and heavily occluded—is pre-constructed. The sample size is expanded using data augmentation techniques, and then proportionally divided into training, validation, and test sets to complete network training, optimization, and performance verification. After the detection network runs, it outputs the detection bounding box corresponding to each stacked part in the scene, and generates a graspability label for each part to distinguish between the upper part that can be directly grasped and the lower part that is occluded and cannot be grasped.

[0026] Step 103: Extract the depth point cloud data of the region corresponding to the detection bounding box, input the depth point cloud data into the 3D model of the part for six-degree-of-freedom pose estimation processing, and obtain the spatial pose data of the target part.

[0027] In this embodiment, based on the detected bounding boxes of each part, the region of interest corresponding to the target part is cropped from the color depth image. The depth pixels within this region are then back-projected into a 3D space using camera imaging parameters to obtain the local depth point cloud data corresponding to the target part. Subsequently, six-degree-of-freedom pose estimation is performed using a pre-stored 3D model of the part: First, a neural implicit representation is constructed for the 3D model of the part, generating a renderable model representation containing geometric and appearance attributes. Multiple different initial pose hypotheses are sampled and generated within the six-degree-of-freedom pose space. The rendering results of the part model under each pose are compared with the actual observed local depth point cloud. Iterative optimization continuously reduces the matching error between the rendering results and the actual observation. All optimized candidate poses are given a hierarchical comprehensive score, with scoring dimensions covering rendering consistency, depth fit, and detection confidence. Finally, the result with the highest comprehensive score is selected as the spatial pose data of the target part.

[0028] Step 104: Based on the pre-built offline crawling template library, perform online crawling decision processing on the spatial pose data to obtain the optimal crawling pose data.

[0029] In this embodiment, an offline grasping template library is pre-built. The construction process is as follows: the three-dimensional model of the target part is voxelized, and the farthest point sampling method is used to uniformly sample the surface of the part to obtain uniformly distributed candidate grasping points; at each candidate grasping point, a local coordinate system is constructed based on the surface normal vector and the principal curvature direction, and multiple sets of candidate grasping poses are generated through translation and rotation augmentation; the mechanical stability of each set of candidate grasping poses is evaluated based on the force helix theory, the poses that meet the force closure condition are selected and the corresponding grasping quality score is calculated, and all qualified poses and corresponding scores are structured and stored to form an offline grasping template library. In the online decision-making phase, all candidate grasping poses in the offline grasping template library are transformed into the robot arm's base coordinate system by combining the spatial pose data of the target part, resulting in a set of candidate grasping poses. Three checks are then performed on all candidate poses in this set: collision safety check along the approach direction, grasping width adaptation check in the gripper opening and closing dimension, and inverse kinematics accessibility check within the robot arm's workspace. Candidate poses that do not meet the constraints are eliminated. The remaining valid candidate poses are then subjected to a multi-objective comprehensive evaluation, with evaluation dimensions including grasping quality, surface normal alignment, collision safety margin, and kinematic accessibility. The poses are then sorted from highest to lowest comprehensive score, and the pose with the highest score is selected as the optimal grasping pose data.

[0030] Step 105: Based on the dynamic constraints of the robotic arm and the scene collision constraints, the optimal grasping pose data is processed for trajectory planning and optimization to obtain grasping motion trajectory data.

[0031] In this embodiment, the optimal grasping pose data is processed for trajectory planning and optimization using dynamic constraints such as joint velocity and acceleration of the robotic arm, as well as collision constraints of scene obstacles, as boundary conditions. First, the current joint state data of the robotic arm is acquired. A sampling motion planner performs a random search in the high-dimensional joint space to generate a collision-free geometric path from the current joint state to the target grasping pose. Next, a time-optimal trajectory generation algorithm is used to parameterize the geometric path over time, generating initial grasping trajectory data while strictly satisfying the upper limits of velocity and acceleration for each joint. Finally, a covariance matrix adaptive evolution strategy is used to perform multi-objective iterative optimization of the initial trajectory, simultaneously considering four optimization objectives: total trajectory execution time, motion smoothness, collision safety margin, and end-effector positioning accuracy. Ultimately, grasping motion trajectory data that can be directly executed is obtained.

[0032] Step 106: Control the robotic arm to perform part grasping processing based on the grasping motion trajectory data to obtain robotic arm pose data and grasping status data.

[0033] In this embodiment, the generated grasping motion trajectory data is converted into real-time motion control commands for each joint of the robotic arm, driving the robotic arm to move smoothly along the planned path. During the movement, the end effector pose data of the robotic arm is collected in real time. When the end effector reaches the target grasping pose, the two-finger gripper at the end effector is controlled to perform a closing action according to the appropriate gripping force and opening / closing stroke, completing the grasping and clamping of the target part. At the same time, the gripper's clamping state data is collected. Finally, the real-time pose data of the robotic arm and the gripper's grasping and clamping state data are integrated in time synchronization to obtain the robotic arm pose data and grasping state data.

[0034] Step 107: Integrate and process the robotic arm pose data, scene 3D point cloud information, robotic arm running status, and grasping status data to obtain an intelligent grasping control scheme.

[0035] In this embodiment, the real-time acquired robotic arm pose data, the overall three-dimensional point cloud information of the scene, the current operating status of the robotic arm, and the grasping status data are integrated and structured in multiple dimensions to form a complete data closed loop that includes environmental perception results, grasping decision results, and execution status feedback. This generates an intelligent grasping control scheme that can support continuous operation for subsequent cyclic grasping operations of stacked parts.

[0036] This application's embodiments construct a complete unordered part grasping process, from environmental perception, part identification, pose estimation, grasping decision-making to trajectory execution, enabling automated grasping of stacked parts in industrial scenarios, improving production line efficiency and automation levels. A two-level positioning scheme of "lightweight detection + six-degree-of-freedom pose estimation" is adopted, ensuring real-time detection through a lightweight network and grasping accuracy through three-dimensional pose calculation, balancing computational constraints and positioning requirements, and adapting to the deployment conditions of industrial edge devices. Combining an offline grasping template library with an online decision-making mechanism, offline mechanical evaluation results are reused while adapting to actual on-site working conditions, improving the reliability of grasping decisions and ensuring the response speed of online processing. Trajectory optimization is achieved by integrating dynamic constraints and collision constraints, improving the smoothness of motion execution and positioning accuracy while ensuring safe and collision-free grasping operations, and reducing the probability of grasping failure.

[0037] In the above Figure 1 Based on the above, this application embodiment also provides a schematic diagram of the implementation process for obtaining the detection bounding box and graspability label data of each stacked part. For example... Figure 2 As shown, steps 201 to 205 are included: Step 201: Perform data augmentation and annotation segmentation on images of stacked parts with different degrees of occlusion to obtain a parts stacking detection dataset.

[0038] In this embodiment, images of stacked parts in two scenarios—stacked inside a box and scattered outside a box—are collected in an industrial setting. Based on the degree of visual occlusion, the parts are categorized into three levels: unoccluded (parts are completely separated and have intact outlines), lightly occluded (edges overlap and occlusion area is less than half), and heavily occluded (highly densely stacked and occlusion area is not less than half). The samples cover different placement postures, lighting conditions, and background interference. The collected original images undergo multi-dimensional data augmentation processing, including random cropping, translation transformation, dynamic brightness adjustment, noise injection, image rotation, and mirror flipping, to expand the sample size and improve the model's generalization ability. All augmented images are randomly divided into training, validation, and test sets according to a preset ratio. Bounding boxes and occlusion levels are labeled for each part in each image, ultimately forming a parts stacking detection dataset.

[0039] Step 202 involves reconstructing the baseline object detection network through cross-stage feature fusion, lightweighting convolutional units, and embedding attention mechanisms to obtain a lightweight object detection network.

[0040] In this embodiment, a single-stage lightweight object detection network is used as the baseline model, and structural improvements are made in three dimensions: First, a cross-stage collaborative feature fusion module is introduced to reconstruct the network's neck structure, unify the channel dimensions of multi-scale features, construct a bidirectional feature fusion path from top to bottom and bottom to top, and replace standard convolutional downsampling with a spatial and channel-decoupled downsampling method. Simultaneously, depthwise separable convolutions are introduced into the backbone network to enhance the bidirectional interaction between shallow detail features and deep semantic features while compressing the number of parameters. Second, a spatial and channel-decoupled convolutional module is used to optimize the convolutional units within the network. The spatial reconstruction unit decouples key features from redundant features, and the channel reconstruction unit adaptively calibrates the channel response, suppressing information redundancy and enhancing the extraction of effective features. Third, a spatially efficient attention module is embedded in the feature extraction structure at the network's end to adaptively recalibrate features in occluded regions, improving target discrimination capabilities in stacked occlusion scenarios. Simultaneously, the network's depth, width, and channel allocation are jointly optimized to adapt to the computational constraints of edge devices, ultimately resulting in an improved lightweight object detection network.

[0041] Step 203: Input the part stack detection dataset into the lightweight object detection network for training and performance verification to obtain a converged lightweight object detection network.

[0042] In this embodiment, a parts stack detection dataset is input into a constructed lightweight object detection network for training. The training process employs a stochastic gradient descent optimizer, setting the corresponding training epochs and batch size. Input images are uniformly scaled to a fixed size. During training, the model's convergence status is monitored in real-time using a validation set. After training, the model's detection accuracy, recall, and inference speed are verified using a test set, along with its robustness under different occlusion levels. Finally, a lightweight object detection network that has converged and meets performance targets is obtained.

[0043] Step 204: Spatial alignment and input normalization are performed on the coordinate system transformation parameters and color depth image data to obtain the image data to be detected.

[0044] In this embodiment, based on coordinate system transformation parameters, the color image and depth image acquired by the depth camera are spatially aligned at the pixel level to ensure a one-to-one correspondence between color pixels and depth values; at the same time, the color image is subjected to size normalization and numerical standardization processing to adjust it to the specifications required by the network input, thereby obtaining the image data to be detected that conforms to the network input format.

[0045] Step 205: Input the image data to be detected into the trained and converged lightweight object detection network for part recognition and graspability evaluation to obtain the detection bounding box and graspability label data of each stacked part.

[0046] In this embodiment, the preprocessed image data to be detected is input into a lightweight object detection network that has been trained and converged. The network extracts features at multiple levels and fuses them at multiple scales to output the bounding box coordinates and class confidence of each part in the scene. At the same time, the network evaluates the graspability of the parts by combining the degree of occlusion and the stacking position of the parts, and generates a binary graspability label for each part. The labels are marked as the upper part that can be grasped directly and the lower part that is occluded and cannot be grasped temporarily. Finally, the bounding box and graspability label data of each stacked part are obtained.

[0047] This application's embodiments improve the model's generalization ability to different stacking conditions by constructing a part stacking detection dataset covering multiple occlusion levels and performing data augmentation, making it adaptable to complex part placement states on-site. Through multi-dimensional network improvements such as cross-stage feature fusion reconstruction, lightweight convolutional unit modification, and attention mechanism embedding, the model size is compressed, computational requirements are reduced, and the network's feature extraction capability for occluded parts is enhanced, improving part detection accuracy in stacking scenarios. Simultaneous output of graspability labels can pre-screen out ungraspable parts that are occluded at the bottom layer, reducing invalid calculations in subsequent pose estimation and grasping decisions, further improving overall processing efficiency. Spatial alignment preprocessing combined with coordinate system transformation parameters ensures spatial consistency between detection results and depth information, improving the baseline accuracy of subsequent pose calculations.

[0048] In some embodiments, the depth point cloud data of the region corresponding to the detection bounding box is extracted, and the depth point cloud data is input into the three-dimensional model of the part for six-degree-of-freedom pose estimation processing to obtain the spatial pose data of the target part. This includes: performing target region cropping and depth back projection processing on the detection bounding box and the color depth image to obtain the depth point cloud data of the region of interest corresponding to the target part.

[0049] Specifically, based on the output detection bounding boxes and graspability labels, target parts marked as directly graspable are first selected. Using the detection bounding boxes of each target part as the scope, the original color depth image is cropped, removing background pixels outside the bounding boxes and retaining only the local area where the target part is located. Combining the internal imaging parameters of the depth camera, each depth pixel within the cropped area is back-projected into 3D space, and the corresponding 3D spatial coordinates of each pixel are calculated, ultimately forming a region-of-interest depth point cloud data containing only the local geometric information of the target part. By limiting the processing range through detection bounding boxes, the amount of data processing required for subsequent pose estimation can be significantly reduced, while also minimizing the interference of background noise on the solution process.

[0050] Furthermore, neural implicit representation construction processing is performed on the 3D model of the part to obtain renderable part model data.

[0051] Specifically, based on the 3D solid model of the target part, a corresponding neural implicit representation field is constructed. This representation field includes two parts: geometric representation and appearance representation. The geometric representation describes the surface geometry of the part and can output signed distance values ​​corresponding to any point in 3D space, thereby defining the solid boundary of the part. The appearance representation combines the material and texture features of the part's surface and can synthesize the part's appearance color information from the corresponding angle based on the viewing perspective. Based on this neural implicit representation field, it is possible to generate rendered images and depth information of the part in any spatial pose, thereby obtaining renderable part model data that supports rendering from new perspectives.

[0052] Furthermore, multiple initial pose sampling processes are performed on the six-degree-of-freedom pose space to obtain a set of initial pose assumptions corresponding to different postures.

[0053] Specifically, within the six-degree-of-freedom rigid body motion space, multiple sets of distinct initial pose parameters are generated according to a uniform sampling rule. Each set of parameters includes three-dimensional position parameters and three-dimensional orientation parameters, corresponding to the spatial position and orientation of the part in the camera coordinate system, respectively. All initial pose parameters are stored in a structured manner to form a set of initial pose assumptions, which serve as the initial solution for subsequent iterative optimization, thereby covering more possible pose states and reducing the probability of getting trapped in local optima.

[0054] Furthermore, the initial pose hypothesis set, renderable part model data, and region of interest depth point cloud data are subjected to rendering comparison and iterative optimization to obtain multiple sets of optimized candidate pose data.

[0055] Specifically, each set of pose parameters in the initial pose hypothesis set is substituted into the renderable part model data in turn to generate the corresponding part rendering depth information and rendering appearance image under that pose. The rendering results are compared point by point with the depth point cloud data of the region of interest and the color image information of the corresponding region to calculate the matching error between the rendering results and the actual observation data. The gradient descent strategy is used to iteratively update the parameters of each pose to gradually reduce the matching error. After a preset number of iterations and convergence, multiple sets of optimized candidate pose data are obtained, and the matching error value corresponding to each set of candidate poses is recorded simultaneously.

[0056] Furthermore, the candidate pose data are subjected to hierarchical comprehensive scoring and sorting filtering to obtain the spatial pose data of the target part.

[0057] Specifically, a hierarchical comprehensive scoring system is constructed, with three scoring dimensions: rendering appearance consistency, depth fit, and detection confidence. Rendering appearance consistency measures the degree of matching between the rendered image and the observed image; depth fit measures the geometric fit between the rendered depth and the measured point cloud; and detection confidence directly uses the output confidence score of the target part's target detection network. Based on this scoring system, all candidate pose data are comprehensively scored, sorted from highest to lowest score, and the candidate poses with the highest comprehensive score are selected as the final spatial pose data for the target part. This hierarchical scoring mechanism can effectively alleviate pose ambiguity in stacked occlusion scenarios and improve the robustness of pose estimation results.

[0058] This application's embodiments significantly reduce the processing scope of pose estimation by obtaining a local depth point cloud through cropping the target region based on the detected bounding box and back-projecting it. This reduces computational load, improves computation speed, and eliminates interference from background noise. A renderable part model is constructed using a neural implicit representation, which accurately represents the geometric and appearance features of the part, providing a high-precision comparison benchmark for pose optimization and improving the accuracy of pose estimation. Through iterative optimization of multiple initial pose sampling and rendering comparisons, more pose possibilities can be covered, avoiding getting trapped in local optima and improving the global accuracy of pose estimation. A hierarchical comprehensive scoring mechanism is used to select the optimal pose, integrating multi-dimensional matching indicators, which effectively alleviates pose ambiguity problems in stacked occlusion scenes and improves the robustness of pose estimation results.

[0059] In some embodiments, online grasping decision processing is performed on spatial pose data based on a pre-built offline grasping template library to obtain optimal grasping pose data, including: performing coordinate system mapping transformation processing on the pre-built offline grasping template library and spatial pose data to obtain a set of candidate grasping poses in the robot arm base coordinate system.

[0060] Specifically, the pre-built offline grasping template library is obtained through preprocessing. First, the 3D model of the part is voxelized. Then, the farthest point sampling method is used to generate uniformly distributed candidate grasping points on the surface of the part. A local coordinate system is constructed by combining the surface normal and principal curvature direction of each grasping point. After translation and rotation augmentation, multiple sets of candidate grasping poses are generated. Finally, the mechanical stability is evaluated and screened based on the force helix theory before being structured and stored.

[0061] In the online processing stage, based on coordinate system transformation parameters, all candidate grasping poses in the offline grasping template library relative to the coordinate system of the part itself are first transformed to the camera coordinate system by combining the spatial pose data of the target part; then, through the rigid body transformation parameters between the camera coordinate system and the robot arm end effector coordinate system, they are transformed to the robot arm end effector coordinate system; finally, through the robot arm kinematic model, they are transformed to the robot arm base coordinate system to obtain the pose description of all candidate grasping poses in the robot arm base coordinate system, forming a set of candidate grasping poses.

[0062] Furthermore, collision safety verification, grasping width adaptation verification, and inverse kinematics reachability verification are performed on the candidate grasping pose set and the scene depth point cloud data to obtain a valid candidate pose set that meets the constraints.

[0063] Specifically, three constraint checks are performed on each set of candidate grasping poses in sequence: The first is a collision safety check, which simulates the motion path of the gripper and the end effector of the robotic arm along the grasping approach direction and performs collision detection with the scene depth point cloud data to eliminate candidate poses with collision risks; the second is a grasping width adaptation check, which calculates the gripping width of the part corresponding to the grasping pose, compares it with the effective opening and closing stroke of the end effector, and eliminates candidate poses that exceed the gripper's stroke range; the third is an inverse kinematics reachability check, which solves the joint angle solutions corresponding to the candidate grasping poses based on the robotic arm kinematic model and eliminates candidate poses without reasonable solutions or that exceed the joint limits; the remaining candidate poses after the three checks constitute the set of valid candidate poses that satisfy all constraints.

[0064] Furthermore, a multi-objective comprehensive scoring process is performed on the set of valid candidate poses to obtain the comprehensive score data corresponding to each valid pose.

[0065] Specifically, a multi-dimensional comprehensive scoring system is constructed, comprising four dimensions: grasping quality score, surface normal alignment score, collision safety margin score, and kinematic reachability score. The grasping quality score is determined based on the grasping mechanical stability value obtained from offline evaluation; higher stability results in a higher score. The surface normal alignment score is determined based on the deviation between the grasping approach direction and the normal to the part's surface; smaller deviations result in a higher score. The collision safety margin score is determined based on the minimum distance between the pose-corresponding path and scene obstacles; larger distances result in a higher score. The kinematic reachability score is determined based on the margins of joint angles and limits; larger margins result in a higher score. After assigning corresponding weights to each score, a weighted sum is performed to obtain the comprehensive score data for each valid candidate pose.

[0066] Furthermore, the comprehensive scoring data and the set of effective candidate poses are sorted and filtered to obtain the optimal capture pose data.

[0067] Specifically, all valid candidate poses are sorted in descending order of comprehensive score data, and the pose with the highest comprehensive score is selected as the final optimal grasping pose data for subsequent trajectory planning and grasping execution.

[0068] This application's embodiments adapt the offline grasping template to the actual pose of the current part through coordinate system mapping transformation, realizing the online reuse of the offline template and avoiding the high computational load of generating grasping poses point by point online, thus improving the efficiency of grasping decisions. Triple checks are set up for collision safety, grasping width, and inverse kinematics reachability, which can eliminate invalid candidate poses that pose collision risks, exceed the gripper's travel range, or exceed the robotic arm's working range layer by layer, ensuring the feasibility and safety of the final output pose. A multi-objective comprehensive scoring system is adopted to evaluate candidate poses from multiple dimensions such as grasping stability, posture adaptability, and safety margin, which can screen out the grasping pose with the best overall performance, improving the success rate and operational stability of actual grasping.

[0069] In some embodiments, trajectory planning and optimization processing is performed on the optimal grasping pose data based on the dynamic constraints of the robotic arm and the scene collision constraints to obtain grasping motion trajectory data, including: sampling-based collision-free motion planning processing is performed on the current joint state data of the robotic arm and the optimal grasping pose data to obtain joint space geometric path data.

[0070] Specifically, the current joint state data of the robotic arm is acquired, including real-time angles and operating speeds of each joint, with the optimal grasping pose data serving as the trajectory endpoint. Discrete nodes are generated within the high-dimensional joint space of the robotic arm using random sampling. Combined with scene collision constraints corresponding to the scene depth point cloud data, collision detection is performed on the connection paths between each sampled node. Nodes and paths interfering with scene obstacles are eliminated. Feasible nodes are gradually connected to generate a continuous, collision-free path from the current joint state to the target grasping pose, ultimately yielding joint space geometric path data composed of multiple sets of discrete joint angle points.

[0071] Furthermore, the joint space geometric path data and the robotic arm dynamic constraints are subjected to time-optimized parameterization to obtain the initial grasping trajectory data.

[0072] Specifically, the dynamic constraints of the robotic arm serve as boundary conditions, including the maximum permissible speed and maximum permissible acceleration limits for each joint. The joint space geometric path obtained in the first step is parameterized in the time dimension. Under the premise that the speed and acceleration of each joint do not exceed the limits throughout the entire process, the shortest total motion time is calculated. Corresponding timestamps, joint speeds, and joint acceleration parameters are assigned to each path point on the geometric path, generating initial grasping trajectory data with temporal information.

[0073] Furthermore, the initial grasping trajectory data and scene collision constraints are subjected to multi-objective iterative optimization processing to obtain grasping motion trajectory data.

[0074] Specifically, using scene collision constraints as hard constraints, the initial grasping trajectory data undergoes multi-objective iterative optimization. The optimization objectives include four aspects: total trajectory execution time, motion smoothness, collision safety margin, and end-effector positioning accuracy. During the optimization process, collision checks are continuously performed on each point on the trajectory to ensure that scene collision constraints are met throughout the process. At the same time, the timing and joint parameters of each path point are iteratively adjusted to achieve a balance and optimality among multiple optimization objectives. After iterative convergence, the grasping motion trajectory data that can be directly sent to the robotic arm controller for execution is obtained.

[0075] This application's embodiments utilize sampling-based collision-free motion planning to generate geometric paths. This allows for rapid solution of feasible paths satisfying scene collision constraints within a high-dimensional joint space, effectively avoiding obstacles and ensuring safety during the motion process. Through time-optimized parameterization, the shortest motion time is achieved while strictly meeting the robotic arm's dynamic constraints, improving the efficiency of single-cycle grasping operations and adapting to the cycle time requirements of industrial production lines. Multi-objective iterative optimization is conducted, balancing motion smoothness, positioning accuracy, and safety margin while ensuring collision constraints. This makes the final trajectory more closely resemble the actual operating characteristics of the robotic arm, reducing operational shock and vibration, and improving the accuracy of end-effector positioning.

[0076] In some embodiments, the robotic arm is controlled to perform part grasping processing based on the grasping motion trajectory data to obtain robotic arm pose data and grasping state data, including: performing joint space control command conversion processing on the grasping motion trajectory data to obtain real-time motion control data of each joint of the robotic arm.

[0077] Specifically, the grasping motion trajectory data is used as input. This trajectory data includes the target angles, velocities, and acceleration parameters of each joint arranged in time sequence within the joint space. Combined with the fixed control cycle of the robotic arm controller, the trajectory points are interpolated to generate control commands for each joint target corresponding to each control cycle. At the same time, motion verification is performed based on the previously calibrated link coordinate system transformation parameters to ensure that the end effector pose corresponding to the command meets the grasping accuracy requirements. Finally, real-time motion control data that can be directly sent to the robotic arm servo system is obtained.

[0078] Furthermore, the real-time motion control data and the robotic arm itself are processed for motion tracking to obtain the real-time pose data of the robotic arm's end effector.

[0079] Specifically, real-time motion control data is sent to the servo drives of each joint of the robotic arm, driving each joint to move synchronously along the planned trajectory. During the movement, the actual angle data of each joint is collected in real time by the coded sensors of each joint. Combined with the forward kinematics model of the robotic arm and the transformation parameters of the link coordinate system, the three-dimensional position and attitude information of the end flange of the robotic arm is calculated in real time, that is, the real-time pose data of the end of the robotic arm, which is used to monitor the motion execution status throughout the process.

[0080] Furthermore, the gripping parameters of the robotic arm's end effector are adapted and the opening and closing control is processed to obtain the gripping status data of the part.

[0081] Specifically, based on the part gripping width corresponding to the optimal gripping pose, the parameters of the end effector two-finger parallel adaptive gripper are adapted, and the corresponding target opening and closing stroke and gripping force threshold are set. When the real-time pose of the robotic arm end effector reaches the target gripping pose, a closing control command is sent to the gripper to drive it to smoothly close and grip the target part. After gripping, the real-time opening and closing degree and gripping force data are collected by the position and force feedback sensors built into the gripper to determine whether the part gripping is stable and reliable, and finally generate part gripping state data containing gripping status, gripping force and opening and closing degree information.

[0082] Furthermore, the real-time pose data and the part gripping and holding status data are synchronized and integrated in time to obtain the robot arm pose data and gripping status data.

[0083] Specifically, based on the multi-system time synchronization parameters obtained from calibration, the real-time pose data of the robotic arm end effector and the gripper's part grasping and holding status data are timestamped according to a unified time reference to eliminate data delay deviations from different sensors. The aligned data is then structured and integrated to output robotic arm pose data containing the entire motion process and corresponding grasping status data for subsequent operation status feedback and closed-loop control.

[0084] This application's embodiments convert trajectory data into real-time joint space control commands, directly adapting to the control logic of the robotic arm servo system, ensuring accurate execution of the planned trajectory and reducing trajectory tracking errors. Real-time acquisition and calculation of the robotic arm's end-effector pose allows for continuous monitoring of the motion execution status, providing data support for operational status feedback and closed-loop adjustments. Adapting gripping parameters to the grasping pose and controlling the opening and closing of the gripper ensures stable part gripping, preventing insufficient gripping force leading to part detachment or mismatched gripping stroke causing grasping failure. Time synchronization integration of pose data and gripping status data eliminates time delay deviations between different sensing systems, improving the consistency of status data and providing a reliable status basis for subsequent closed-loop operational control.

[0085] In some embodiments, color depth image data of a scene of disordered stacked parts and motion state data of a robotic arm are acquired. Multi-coordinate system calibration and time synchronization processing are performed on the color depth image data and the motion state data of the robotic arm to obtain coordinate system transformation parameters, including: performing intrinsic parameter calibration and lens distortion correction processing on the depth camera to obtain internal imaging parameters of the camera.

[0086] Specifically, a checkerboard calibration board is used as the calibration reference. Multiple sets of calibration board images are collected at different angles and distances. The internal parameters of the depth camera, such as focal length and principal point, are solved through the calibration algorithm. At the same time, the radial and tangential distortion coefficients of the lens are solved to complete the lens distortion correction. The projection mapping relationship from the image pixel coordinates to the three-dimensional coordinates of the camera coordinate system is established, and finally the internal imaging parameters of the camera are obtained.

[0087] Furthermore, the robotic arm is modeled to obtain the coordinate transformation parameters of each link of the robotic arm.

[0088] Specifically, the improved link parameter method is used to perform kinematic modeling on a six-DOF collaborative manipulator. Corresponding body coordinate systems are established for each link of the manipulator in turn, and the homogeneous coordinate transformation relationship between adjacent links is derived to obtain the coordinate system transformation parameters of each link. These parameters can be used to solve the forward and inverse kinematics of the manipulator and are the basis for subsequent end-effector pose calculation and motion control.

[0089] Furthermore, the robotic arm is controlled to move to multiple different spatial poses, and the calibration plate image data and the robotic arm end pose data under the corresponding poses are collected simultaneously to obtain calibration sample data.

[0090] Specifically, the robotic arm is controlled to move a depth camera fixed at its end to multiple different spatial poses, maintaining the stability of the robotic arm's posture in each pose. At each pose, the depth camera is simultaneously triggered to acquire calibration plate image data, and based on the coordinate transformation parameters of each link, the robotic arm end pose data at the corresponding moment is obtained through forward kinematics calculation. The calibration plate image corresponding to each pose is stored in a one-to-one correspondence with the robotic arm end pose, finally obtaining calibration sample data.

[0091] Furthermore, the calibration sample data is subjected to linear closed-form solution and nonlinear iterative optimization to obtain the rigid body transformation parameters between the camera coordinate system and the robotic arm end effector coordinate system.

[0092] Specifically, based on the calibration sample data, the initial solution of the rigid body transformation between the camera coordinate system and the robot arm end effector coordinate system is first calculated using a linear closed-loop solution method. Then, with the minimum reprojection error as the optimization objective, the initial solution is optimized for accuracy using a nonlinear iterative optimization algorithm to reduce the influence of calibration noise, and finally, the accurate rigid body transformation parameters between the camera coordinate system and the robot arm end effector coordinate system are obtained.

[0093] Furthermore, the image acquisition timing of the depth camera and the motion control timing of the robotic arm are time-stamp aligned to obtain multi-system time synchronization parameters.

[0094] Specifically, using a unified system clock as a reference, the time deviation between the image acquisition triggering sequence of the depth camera and the motion state feedback sequence of the robotic arm is calibrated, and the fixed time delay offset between the two is calculated. Based on this time delay offset, the timestamps of the two types of data are aligned, and multi-system time synchronization parameters are generated to ensure the time consistency between subsequent visual data and robotic arm motion data.

[0095] Furthermore, the camera's internal imaging parameters, the coordinate system transformation parameters of each link, the calibration sample data, the rigid body transformation parameters, and the time synchronization parameters of multiple systems are integrated and processed to obtain the coordinate system transformation parameters.

[0096] Specifically, the internal imaging parameters of the camera, the coordinate system transformation parameters of each link of the robotic arm, the calibration sample data, the rigid body transformation parameters between the camera and the end effector, and the time synchronization parameters of multiple systems are structurally integrated to form a complete coordinate system transformation parameter system. This system can support the full-link spatial mapping from image pixel coordinates to the camera coordinate system, the end effector coordinate system of the robotic arm, and the base coordinate system of the robotic arm. It also supports the time synchronization matching of data from multiple systems. The calibration sample data can be used for parameter traceability and accuracy verification. Finally, coordinate system transformation parameters that can support the subsequent full-process processing are obtained.

[0097] This application's embodiments establish a complete mapping relationship from image pixels to the robot arm's base coordinate system through end-to-end spatial calibration, including camera intrinsic parameter calibration, robot arm kinematic modeling, and hand-eye rigid body transformation calibration. This ensures spatial consistency between visual perception results and robot arm motion control, improving the overall system's positioning accuracy. The calibration solution method, employing a combination of linear closed-form solving and nonlinear iterative optimization, balances solution efficiency and calibration accuracy, effectively reducing the impact of calibration noise and improving the accuracy of hand-eye transformation parameters. By aligning timestamps across multiple systems, the time delay deviation between visual acquisition and robot arm state feedback is eliminated, ensuring the time synchronization of multi-source data and avoiding positioning and control errors caused by timing discrepancies. Integrating multiple types of calibration parameters to form a unified coordinate system transformation parameter system provides unified benchmark support for the entire process of spatial mapping, pose calculation, and motion control, ensuring the consistency of data benchmarks at each stage.

[0098] While this application provides the method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-inventive labor. The order of steps listed in this embodiment is merely one possible execution order among many and does not represent the only execution order. In actual device or client product execution, the methods shown in this embodiment or the accompanying drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment).

[0099] like Figure 3 As shown in the figure, this application embodiment also provides a part gripping device 300. The device includes: The acquisition module 301 is used to acquire color depth image data of the disordered stacked parts scene and robotic arm motion state data, and to perform multi-coordinate system calibration and time synchronization processing on the color depth image data and robotic arm motion state data to obtain coordinate system transformation parameters.

[0100] The processing module 302 is used to input color depth image data into a lightweight target detection network for part recognition processing based on coordinate system transformation parameters, so as to obtain the detection bounding box and graspability label data of each stacked part.

[0101] The processing module 302 is also used to extract the depth point cloud data of the region corresponding to the detection bounding box, input the depth point cloud data into the three-dimensional model of the part for six-degree-of-freedom pose estimation processing, and obtain the spatial pose data of the target part.

[0102] The processing module 302 is also used to perform online grasping decision processing on spatial pose data based on a pre-built offline grasping template library to obtain the optimal grasping pose data.

[0103] The processing module 302 is also used to perform trajectory planning and optimization processing on the optimal grasping pose data based on the dynamic constraints of the robotic arm and the scene collision constraints, so as to obtain grasping motion trajectory data.

[0104] The processing module 302 is also used to control the robotic arm to perform part grasping processing based on the grasping motion trajectory data, and to obtain the robotic arm pose data and grasping status data.

[0105] The integration module 303 is used to integrate and process the robot arm pose data, scene 3D point cloud information, robot arm running status and grasping status data to obtain an intelligent grasping control scheme.

[0106] Some modules in the apparatus described in this application can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0107] The apparatus or module described in the above embodiments can be implemented by a computer chip or physical entity, or by a product with a certain function. For ease of description, the above apparatus is described by dividing it into various modules according to their functions. When implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware. Of course, a module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.

[0108] The methods, apparatus, or modules described in this application can be implemented in a computer-readable program code manner. The controller can be implemented in any suitable manner, such as a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of a memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code manner, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included within it for implementing various functions can also be considered as structures within the hardware component. Alternatively, the device used to implement various functions can be viewed as either a software module that implements the method or a structure within a hardware component.

[0109] This application also provides an apparatus, the apparatus comprising: a processor; a memory for storing processor-executable instructions; wherein, when the processor executes the executable instructions, it implements the method described in this application.

[0110] This application also provides a non-volatile computer-readable storage medium storing a computer program or instructions thereon, which, when executed, enables the method described in this application embodiment to be implemented.

[0111] Furthermore, in the various embodiments of the present invention, each functional module can be integrated into a processing module, or each module can exist independently, or two or more modules can be integrated into a single module.

[0112] The aforementioned storage media include, but are not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), Cache, Hard Disk Drive (HDD), or Memory Card. The memory can be used to store computer program instructions.

[0113] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, or it can be embodied in the process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0114] The various embodiments described in this specification are presented in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. All or part of this application can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, mobile communication terminals, multiprocessor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.

[0115] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of this application.

Claims

1. A method for gripping parts, characterized in that, include: Acquire color depth image data of a scene with disordered stacked parts and robotic arm motion state data. Perform multi-coordinate system calibration and time synchronization processing on the color depth image data and the robotic arm motion state data to obtain coordinate system transformation parameters. Based on the coordinate system transformation parameters, the color depth image data is input into a lightweight target detection network for part recognition processing to obtain the detection bounding box and graspability label data of each stacked part; Extract the depth point cloud data of the region corresponding to the detection bounding box, and input the depth point cloud data into the three-dimensional model of the part for six-degree-of-freedom pose estimation to obtain the spatial pose data of the target part. The spatial pose data is processed online based on a pre-built offline crawling template library to obtain the optimal crawling pose data. Based on the dynamic constraints of the robotic arm and the scene collision constraints, the optimal grasping pose data is processed for trajectory planning and optimization to obtain grasping motion trajectory data. Based on the grasping motion trajectory data, the robotic arm is controlled to perform part grasping processing to obtain robotic arm pose data and grasping status data; The robotic arm pose data, scene 3D point cloud information, robotic arm operating status, and grasping status data are integrated and processed to obtain an intelligent grasping control scheme.

2. The method according to claim 1, characterized in that, Based on the coordinate system transformation parameters, the color depth image data is input into a lightweight target detection network for part recognition processing to obtain the detection bounding boxes and graspability label data of each stacked part, including: Data augmentation and annotation segmentation are performed on images of stacked parts with different degrees of occlusion to obtain a parts stacking detection dataset; A lightweight target detection network is obtained by performing cross-stage feature fusion structure reconstruction, lightweight transformation of convolutional units, and attention mechanism embedding on the baseline target detection network. The part stack detection dataset is input into the lightweight object detection network for training and performance verification, resulting in a converged lightweight object detection network. Spatial alignment and input normalization are performed on the coordinate system transformation parameters and the color depth image data to obtain the image data to be detected. The image data to be detected is input into the lightweight object detection network that has been trained and converged for part identification and graspability evaluation, so as to obtain the detection bounding box and graspability label data of each stacked part.

3. The method according to claim 1, characterized in that, The process involves extracting depth point cloud data corresponding to the detected bounding box region, inputting this depth point cloud data into the 3D model of the part for six-degree-of-freedom pose estimation, and obtaining the spatial pose data of the target part, including: The detection bounding box and the color depth image are subjected to target region cropping and depth back projection processing to obtain the depth point cloud data of the region of interest corresponding to the target part; The three-dimensional model of the part is processed by neural implicit representation construction to obtain renderable part model data; Multiple initial pose sampling processes are performed on the six-DOF pose space to obtain a set of initial pose assumptions corresponding to different postures; The initial pose hypothesis set, the renderable part model data and the depth point cloud data of the region of interest are subjected to rendering comparison and iterative optimization processing to obtain multiple sets of optimized candidate pose data. The candidate pose data are subjected to hierarchical comprehensive scoring and sorting filtering to obtain the spatial pose data of the target part.

4. The method according to claim 1, characterized in that, The online crawling decision processing based on the pre-built offline crawling template library to obtain the optimal crawling pose data includes: The pre-built offline grasping template library and the spatial pose data are subjected to coordinate system mapping and transformation to obtain a set of candidate grasping poses in the robot arm base coordinate system; The candidate grasping pose set and the scene depth point cloud data are subjected to collision safety verification, grasping width adaptation verification and inverse kinematics reachability verification to obtain a valid candidate pose set that meets the constraints. Multi-objective comprehensive scoring processing is performed on the set of effective candidate poses to obtain comprehensive score data corresponding to each effective pose; The comprehensive scoring data and the set of effective candidate poses are sorted and filtered to obtain the optimal capture pose data.

5. The method according to claim 1, characterized in that, The dynamic constraints and scene collision constraints based on the robotic arm are used to perform trajectory planning and optimization on the optimal grasping pose data to obtain grasping motion trajectory data, including: The current joint state data of the robotic arm and the optimal grasping pose data are subjected to sampled collision-free motion planning to obtain joint space geometric path data. The joint space geometric path data and the robotic arm dynamics constraints are subjected to time-optimized parameterization to obtain the initial grasping trajectory data; The initial grasping trajectory data and scene collision constraints are subjected to multi-objective iterative optimization processing to obtain grasping motion trajectory data.

6. The method according to claim 1, characterized in that, The step of controlling the robotic arm to perform part grasping processing based on the grasping motion trajectory data, and obtaining robotic arm pose data and grasping state data, includes: The grasping motion trajectory data is processed by joint space control command conversion to obtain real-time motion control data of each joint of the robotic arm; The real-time motion control data and the robotic arm body are subjected to motion tracking and driving processing to obtain the real-time pose data of the robotic arm end effector. The gripping parameters of the robotic arm's end effector are adapted and the opening and closing control is processed to obtain the part gripping status data. The real-time pose data and the part gripping and holding status data are synchronized and integrated in time to obtain the robot arm pose data and gripping status data.

7. The method according to claim 1, characterized in that, The process involves acquiring color depth image data of the disordered stacked parts scene and robotic arm motion state data, performing multi-coordinate system calibration and time synchronization processing on the color depth image data and the robotic arm motion state data to obtain coordinate system transformation parameters, including: Intrinsic parameter calibration and lens distortion correction are performed on the depth camera to obtain the camera's internal imaging parameters; The robotic arm is modeled to obtain the coordinate transformation parameters of each link of the robotic arm; The robotic arm is controlled to move to multiple different spatial poses, and the calibration plate image data and the robotic arm end pose data under the corresponding poses are collected simultaneously to obtain calibration sample data. The calibration sample data is subjected to linear closed-form solution and nonlinear iterative optimization to obtain the rigid body transformation parameters between the camera coordinate system and the robot arm end coordinate system. The image acquisition timing sequence of the depth camera and the motion control timing sequence of the robotic arm are timestamped to obtain multi-system time synchronization parameters. The camera's internal imaging parameters, the coordinate system transformation parameters of each link, the calibration sample data, the rigid body transformation parameters, and the multi-system time synchronization parameters are integrated and processed to obtain the coordinate system transformation parameters.

8. A parts gripping device, characterized in that, include: The acquisition module is used to acquire color depth image data of a scene of disordered stacked parts and robotic arm motion state data, and to perform multi-coordinate system calibration and time synchronization processing on the color depth image data and the robotic arm motion state data to obtain coordinate system transformation parameters. The processing module is used to input the color depth image data into a lightweight target detection network for part recognition processing based on the coordinate system transformation parameters, so as to obtain the detection bounding box and graspability label data of each stacked part; The processing module is also used to extract the depth point cloud data of the region corresponding to the detection bounding box, input the depth point cloud data into the three-dimensional model of the part for six-degree-of-freedom pose estimation processing, and obtain the spatial pose data of the target part. The processing module is also used to perform online crawling decision processing on the spatial pose data based on a pre-built offline crawling template library to obtain the optimal crawling pose data. The processing module is also used to perform trajectory planning and optimization processing on the optimal grasping pose data based on the dynamic constraints of the robotic arm and the scene collision constraints, so as to obtain grasping motion trajectory data. The processing module is also used to control the robotic arm to perform part grasping processing based on the grasping motion trajectory data, so as to obtain robotic arm pose data and grasping state data. The integration module is used to integrate and process the robotic arm pose data, scene 3D point cloud information, robotic arm operating status and grasping status data to obtain an intelligent grasping control scheme.

9. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.