A hierarchical inference-based symmetric workpiece disorderly grabbing method and system

By using multimodal perception fusion and hierarchical topology reasoning model, the problem of physical hierarchy recognition of workpieces in disordered grasping of rotationally symmetric targets is solved, and stable grasping and efficient sorting in complex environments are achieved.

CN122391365APending Publication Date: 2026-07-14WUHAN UNIV OF TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN UNIV OF TECH
Filing Date
2026-05-06
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively acquire the true physical hierarchy and occlusion topology of workpieces when handling disordered grasping of rotationally symmetric targets. This leads to unstable grasping by robots in complex environments, which can easily cause material pile collapse and operation interruption.

Method used

By using multimodal perception fusion and full-view geometry reconstruction, a hierarchical topology reasoning model is constructed to obtain the complete physical boundary representation of the workpiece and its off-center load characteristic data. Top-level candidate workpieces are selected and their graspability is quantitatively evaluated. The baseline grasping pose is expanded into multiple equivalent candidate poses, and parallel feasibility verification is performed to generate collision-free grasping instructions.

Benefits of technology

It significantly improves the success rate of picking and operational safety in complex stacking environments, avoids material pile collapse and operation interruption, and enhances sorting flexibility and the stability of continuous operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122391365A_ABST
    Figure CN122391365A_ABST
Patent Text Reader

Abstract

The application provides a symmetric workpiece disorderly grabbing method and system based on hierarchical reasoning, relates to the technical field of industrial robots, and comprises the following steps: acquiring a multi-modal image of a stacking scene, and reconstructing a complete physical mask representation of each workpiece from the multi-modal image. Through multi-modal perception fusion and full-view geometric reconstruction processing, the application obtains complete physical boundary representation and unbalanced load feature data of each workpiece in a shielding state, and on this basis, a directed hierarchical topological model describing the physical pressing and stacking relationship between workpieces is constructed through spatial gradient analysis, a candidate workpiece in an absolute top layer in the current scene is screened out, and then, a grabbability quantitative evaluation is performed to determine a target to be grabbed, so that the robot can identify and eliminate the bottom shielding workpiece which is easy to cause the collapse of the material pile in advance based on the real physical support relationship before action execution, and interference and instability caused by blind grabbing are avoided from the root.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial robot technology, and in particular to a method and system for disordered grasping of symmetrical workpieces based on hierarchical reasoning. Background Technology

[0002] In automated assembly and material sorting operations of industrial robots, the ability to grasp disordered materials in confined environments such as densely stacked and narrow material frames is a core indicator determining whether a robot can handle complex and flexible tasks. This is especially true when handling rotationally symmetric standard parts, such as hexagonal head bolts. These scattered stacking scenarios are characterized by complex material spatial topology, severe surface metallic reflection, mutual occlusion of local features, and the inability of traditional two-dimensional visibility assumptions to meet actual three-dimensional stacking conditions. This places extremely high demands on the robot's environmental perception and motion decision-making. The robot must complete target recognition, occlusion relationship judgment, graspability assessment, pose calculation, and obstacle avoidance planning within milliseconds, avoiding gripper collisions, material pile collapses, and material slippage while ensuring the stability and efficiency of continuous operation.

[0003] Existing control technologies for unordered grasping of rotationally symmetric targets mainly fall into three core paths. The first is the 6D pose direct regression method based on end-to-end deep learning. This method relies on neural networks to extract geometric features of the object and output a single target pose. However, when dealing with rotationally symmetric targets, the inherent geometric rotational invariance of the object leads to manifold ambiguity in mathematical space. The pre-set unique supervision signal severely conflicts with the actual physically equivalent pose, easily resulting in model training non-convergence or random jumps in pose angles during inference, ultimately causing end-point docking failure. The second is the instance segmentation and heuristic grasping method based on visible contours. This method extracts the visible surface mask of the object using visual sensors to match grasping points. However, this method can only identify visible areas at the two-dimensional image level and cannot obtain the actual physical support properties of the material stack. The top-layer workpiece, completely exposed, and the semi-occluded workpiece, with its main structure pressed against the bottom layer, are highly similar in local visual features, but their physical graspability is vastly different. Forcibly extracting the bottom-layer workpiece will inevitably cause the material pile to collapse. Ultimately, although the robot can perceive the visual presence of the workpiece, it is difficult to accurately infer the true physical hierarchy and occlusion topology from the visual information, thus failing to effectively judge the safety of grasping and failing to fundamentally solve the instability problem caused by blind grasping. The third type is a rigid serial obstacle avoidance planning method based on a single pose output. It provides a fixed grasping posture through visual positioning and then hands it over to the robotic arm for inverse kinematics solution. However, the operable space inside the deep material frame is narrow. This unidirectional execution mechanism has inherent rigidity. Often, when the only grasping posture encounters joint limitations or frame wall collisions, the system can only passively report an error and stop, completely ignoring the equivalent posture redundancy that rotational symmetry can provide, and failing to avoid the risk of deadlock under complex constraints.

[0004] Overall, existing technologies typically only acquire shallow geometric information and design grasping strategies based on a single pose mapping model. These strategies are prone to failure when the actual physical stacking conditions do not align with the assumptions of unobstructed grasping. Such methods primarily rely on unidirectional visual recognition and fixed motion execution, making it difficult to assess grasping risks in advance based on the actual stacking state before the robot enters. Furthermore, they can only output a single target solution and lack the ability to adaptively expand the grasping pose space using geometric symmetry. Summary of the Invention

[0005] In view of this, the present invention proposes a method and system for disordered grasping of symmetrical workpieces based on hierarchical reasoning. Through multimodal perception fusion and full-view geometric reconstruction, the complete physical boundary representation and off-center loading characteristic data of each workpiece in the occlusion state are obtained. On this basis, a directed hierarchical topology model describing the physical stacking relationship between workpieces is constructed through spatial gradient analysis. Based on this, candidate workpieces at the absolute top level in the current scene are selected, and graspability is quantitatively evaluated to determine the target to be grasped. This allows the robot to identify and remove bottom-level occluded workpieces that are prone to causing the material pile to collapse before the action is executed, based on the real physical support relationship. This avoids interference and instability caused by blind grasping from the root, and greatly improves the success rate and operational safety of the first grasp in complex stacking environments.

[0006] The technical solution of this invention is implemented as follows: In a first aspect, the present invention provides a method for disordered grasping of symmetrical workpieces based on hierarchical reasoning, comprising: Acquire multimodal images of the stacked scene, reconstruct the complete physical mask representation of each workpiece from the multimodal images, and generate off-center load feature data accordingly. The off-center load feature data is used to describe the occlusion state of each workpiece. Based on the complete physical mask representation, a hierarchical topology reasoning model is constructed by analyzing the spatial gradient of the overlapping areas of different workpiece masks. The top-level candidate workpiece in the current scene is selected according to the hierarchical topology reasoning model, which is used to describe the physical overlapping relationship between workpieces. Based on the off-center load characteristic data, the graspability of the selected top-level candidate workpieces is quantitatively evaluated, and the target workpiece to be grasped is determined based on the evaluation results. Obtain the reference gripping pose of the target workpiece to be gripped, and expand the reference gripping pose into a pose set containing multiple equivalent candidate poses based on the inherent rotational symmetry characteristics of the workpiece. The pose set is mapped from the perception space to the robot execution space. In the execution space, the feasibility of each equivalent candidate pose in the set is checked in parallel, and a subset of globally safe poses that meet all constraints is selected. Select a target pose from the global safe pose subset and output motion control commands to drive the robotic arm to complete collision-free grasping.

[0007] In one embodiment, acquiring a multimodal image of the stacked scene, reconstructing a complete physical mask representation of each workpiece from the multimodal image, and generating off-center loading feature data accordingly includes the following steps: Simultaneously acquire color texture images and depth images of the stacked scene, and use the calibrated spatial mapping relationship to align and fuse the color texture images and depth images at the sub-pixel level to generate a multimodal input tensor; Guided by the high-frequency edge features in the color texture image, edge-preserving repair is performed on the cavity area in the depth image caused by the high reflectivity of the workpiece surface to obtain a repaired continuous depth map. The multimodal input tensor and the repaired continuous depth map are input into the pre-built full-view segmentation network. The full-view segmentation network injects the prior shape of the workpiece into implicit encoding and performs decoupling and cross-attention fusion of visible representation and occlusion representation in the latent space, and outputs the visible perception mask and complete physical mask representation of each workpiece simultaneously. Based on the visible perception mask and the complete physical mask, the geometric off-load vector, which represents the degree of spatial deviation between the centroid of the visible area of ​​the workpiece and the centroid of the complete physical area, is calculated. The geometric off-load vector and the occlusion rate determined based on the area ratio of the two types of masks are used together as the off-load feature data.

[0008] In one embodiment, the construction of a hierarchical topological reasoning model based on complete physical mask representation, by analyzing the spatial gradient of overlapping regions of different workpiece masks, includes the following steps: Extract the overlapping region of the complete physical mask representation of any two workpieces. When the overlapping region is not empty, estimate the normal vector of the corresponding spatial point cloud in the overlapping region. When the angle between the normal vectors of two workpieces or the average spatial depth difference does not meet the coplanarity determination condition, sampling is performed along the boundary of the overlapping area to calculate the spatial depth jump gradient in the preset neighborhood on both sides of the boundary. If the spatial depth gradient exceeds the preset physical thickness threshold, it is determined that there is a physical overlap relationship between the two workpieces, and this relationship is defined as a directed edge. The two workpieces connected by the directed edge are used as nodes to construct a directed acyclic graph describing the physical overlap relationship between the workpieces for the global scene, which serves as a hierarchical topological reasoning model.

[0009] In one embodiment, the step of quantitatively evaluating the graspability of the selected top-level candidate workpieces based on off-center load characteristic data, and determining the target workpiece to be grasped based on the evaluation results, includes the following steps: A multidimensional graspability score is constructed for each of the top-level candidate workpieces. The multidimensional graspability score is negatively correlated with the occlusion rate obtained from the off-center load feature data, negatively correlated with the modulus of the geometric off-center load vector obtained from the off-center load feature data, and positively correlated with the normalized height of the top-level candidate workpiece in the current frame depth direction. The multidimensional graspability scores of each of the top-level candidate workpieces are sorted in descending order, and the top-level candidate workpiece with the highest score is selected as the target workpiece to be grasped.

[0010] In one embodiment, obtaining the reference gripping pose of the target workpiece to be gripped, and expanding the reference gripping pose into a pose set containing multiple equivalent candidate poses based on the inherent rotational symmetry characteristics of the workpiece, includes the following steps: The pre-trained pose estimation network obtains any converged 6D reference pose of the target workpiece in the camera coordinate system as an anchor point. Determine the principal axis of symmetry of the target workpiece to be grasped, and calculate a series of equivalent rotation angles about the principal axis of symmetry based on at least one of the discrete rotational symmetry order of the principal axis of symmetry and the angle step size set for the continuous rotational symmetry body. By sequentially combining the anchor point with the rotation transformation containing each equivalent rotation angle, a set of candidate poses that are completely equivalent in physical grasping effect and contain the pre-grabbing proximity distance parameter are generated, forming a pose set.

[0011] In one embodiment, the feasibility verification of each equivalent candidate pose in the set is performed in parallel within the execution space, and a subset of globally safe poses that satisfy all constraints is selected. The following parallel verification operation is performed on each equivalent candidate pose: The equivalent candidate poses are solved by inverse kinematics to verify whether there is a joint configuration solution whose joint angles meet the hardware limit of the robotic arm within a preset safety margin. For the existing joint configuration solution, calculate the operability index based on the velocity transmission characteristics in the joint space, and verify whether it is in the flexible working range far away from the kinematic singularity. For equivalent candidate poses, the minimum distance between the robotic arm end-effector and the obstacles in the boundary model and the occupancy map is calculated when the robotic arm end-effector cuts into the pose, based on the digital twin boundary model of the material frame and the real-time three-dimensional occupancy map of the stacked workpieces in the material frame. The minimum distance is then verified to be greater than the preset safe expansion radius. The equivalent candidate pose that simultaneously satisfies the three constraints of having a legal joint configuration solution, being non-singular, and having no collision is determined as an element of the global safe pose subset.

[0012] In one embodiment, the method further includes exception handling when the global safe pose subset is empty, the exception handling including the following steps: Check if the global safe pose subset is empty; When it is determined that the global safe pose subset is empty, stop the output of motion commands for the current capture cycle and generate active perturbation commands; Execute the active disturbance command to control the robotic arm to use the end effector as a probe to perform a preset low-speed lateral tossing action along the path of the maximum gap area indicated in the hierarchical topology inference model, so as to change the rigid physical layout of the workpiece in the current stacking scene. After the toggle action is completed, the current decision queue is cleared, and the multimodal images of the stacked scene are reacquired to start a new grasping decision cycle.

[0013] Secondly, the present invention provides a hierarchical reasoning-based unordered grasping system for symmetric workpieces, used to implement the above-mentioned method, comprising: The image acquisition and representation generation module is used to acquire multimodal images of the stacked scene, reconstruct the complete physical mask representation of each workpiece from the multimodal images, and generate off-center feature data describing the occlusion state of each workpiece accordingly. The hierarchical topology reasoning module is used to construct a hierarchical topology reasoning model that describes the physical overlapping relationship between workpieces by analyzing the spatial gradient of the overlapping areas of different workpiece masks based on the complete physical mask representation, and to select the top-level candidate workpiece in the current scene based on the hierarchical topology reasoning model. The quantitative evaluation module is used to perform a quantitative evaluation of the graspability of the top-level candidate workpieces selected based on the off-center load characteristic data, and to determine the target workpiece to be grasped based on the evaluation results. The equivalent pose space mapping module is used to obtain the reference grasping pose of the target workpiece to be grasped. Based on the inherent rotational symmetry characteristics of the workpiece, the reference grasping pose is expanded into a pose set containing multiple equivalent candidate poses. The parallel constraint screening module is used to map the pose set from the perception space to the robot execution space, perform feasibility verification on each equivalent candidate pose in the set in parallel in the execution space, and screen out the global safe pose subset that satisfies all constraints. The global optimal instruction generation module is used to select a target pose from the global safe pose subset and output motion control instructions to drive the robotic arm to complete collision-free grasping.

[0014] Thirdly, the present invention provides an electronic device including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the above-described method.

[0015] Fourthly, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described method.

[0016] The hierarchical reasoning-based method and system for disordered grasping of symmetrical workpieces of the present invention has the following advantages over the prior art: 1. By processing multimodal perception fusion and full-view geometry reconstruction, the complete physical boundary representation and off-center load feature data of each workpiece in the occlusion state are obtained. Based on this, a directed hierarchical topology model describing the physical stacking relationship between workpieces is constructed through spatial gradient analysis. Based on this, the candidate workpieces at the absolute top level in the current scene are screened out. Then, the graspability is quantitatively evaluated to determine the target to be grasped. This enables the robot to identify and remove the bottom occluded workpieces that are very likely to cause the material pile to collapse before the action is executed, based on the real physical support relationship. This avoids interference and instability caused by blind grasping from the root, and greatly improves the success rate and operation safety of the first grasp in complex stacking environment. 2. By utilizing the rotational symmetry of the workpiece itself, the reference grasping pose obtained by the sensing end is expanded into a set of poses with equivalent physical effects, including multiple candidate directions and approach paths. This set of poses is then mapped to the robot's execution space for multi-constraint parallel feasibility verification. From this, a safe subset of poses that simultaneously satisfy the conditions of kinematic solvability, no singularity, and no collision is selected, thereby generating the optimal collision-free grasping motion command. This creatively transforms the pose ambiguity at the sensing end into a redundant decision space at the execution end, enabling the system to autonomously select feasible obstacle avoidance paths from multiple equivalent poses in confined environments such as deep and narrow material boxes. This effectively breaks through the bottleneck of operation interruption caused by the single pose in traditional systems, and significantly improves the sorting flexibility and clearing rate of scattered materials. 3. Anomaly self-healing is achieved through an active disturbance mechanism. This involves using a pre-built hierarchical topology reasoning model to identify the largest gap area within the material frame, controlling the robotic arm to perform a low-speed tossing motion along that area using the end effector, actively changing the rigid stacking layout of the workpieces, and re-triggering scene perception after the disturbance is completed to start a new decision cycle. This enables the system to autonomously break through deadlocks and recover from closed-loop conditions without human intervention, avoiding the production line interruption problems caused by passive shutdown and reliance on manual intervention in traditional systems during deadlocks. This ensures the operational stability and automation reliability of long-cycle continuous sorting operations. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the steps of the hierarchical reasoning-based unordered grasping method for symmetrical workpieces of the present invention. Figure 2 This is an overall flowchart of the hierarchical reasoning-based unordered grasping method for symmetrical workpieces of the present invention. Figure 3 The flowchart shows the steps of multimodal feature extraction and target full-view geometry reconstruction of the hierarchical reasoning-based symmetrical workpiece disordered grasping method of the present invention. Figure 4 This is a schematic diagram of the overlapping topological reasoning and graspability quantification process of the hierarchical reasoning-based symmetric workpiece disorder grasping method of the present invention. Figure 5 This is a schematic diagram of the equivalent pose space mapping modeling process for the hierarchical reasoning-based unordered grasping method for symmetric workpieces of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0020] like Figure 1-5 As shown, the hierarchical reasoning-based unordered workpiece grasping method of the present invention is used to solve the problem that the existing technology relies only on two-dimensional visible contours for grasping planning, which cannot penetrate occlusions to obtain complete physical information of the workpiece. This makes it difficult for the robot to accurately infer the real physical hierarchy and occlusion topology from visual information and to effectively predict the safety of grasping. For ease of understanding, this method is divided into steps S100-S600 for description.

[0021] Step S100: Obtain a multimodal image of the stacked scene, reconstruct the complete physical mask representation of each workpiece from the multimodal image, and generate off-center load feature data to describe the occlusion state of each workpiece.

[0022] In this step, a two-dimensional color texture image and a three-dimensional depth image containing the stacked workpieces to be processed are acquired simultaneously through an imaging sensing device. The two heterogeneous images are then precisely aligned and fused at the pixel level using a pre-calibrated spatial mapping relationship.

[0023] To eliminate the impact of voids and noise in the depth data caused by highly reflective materials on the workpiece surface on subsequent processing, high-frequency geometric edge features extracted from the color image are used as guiding constraints to repair and reconstruct the defective areas in the depth image while maintaining edge sharpness, resulting in a geometrically continuous repaired depth map.

[0024] Subsequently, the aligned and fused multimodal data is input into a full-view segmentation network model with shape prior injection capability. This network model uses the implicit encoding of the workpiece's prior shape as auxiliary information, and through internal feature decoupling and cross-modal attention fusion mechanisms, it completes and infers the occluded and invisible workpiece regions, thereby simultaneously outputting the visible perceptual mask and the complete physical mask representation of each workpiece.

[0025] Based on the two types of masks mentioned above, a geometric off-load vector is further calculated to characterize the degree of spatial deviation between the centroid of the visible area of ​​the workpiece and the centroid of the complete physical area. At the same time, the occlusion rate is determined based on the area ratio of the two types of masks. The geometric off-load vector and the occlusion rate are used together as off-load feature data to describe the occlusion state of each workpiece.

[0026] Step S200: Based on the complete physical mask representation, a hierarchical topology reasoning model is constructed by analyzing the spatial gradient of the overlapping areas of different workpiece masks, and the top-level candidate workpieces in the current scene are selected according to the hierarchical topology reasoning model.

[0027] In this step, based on the complete physical mask representation of each workpiece instance generated in step S100, the mask overlap region between workpieces is extracted pairwise. When an overlap is detected between the masks of any two workpieces, local geometric feature analysis is performed on the spatial point cloud data corresponding to the overlap region. Normal sampling is performed along the boundary of the overlap region, and the spatial depth jump amplitude within a preset neighborhood on both sides of the boundary is calculated. If the depth jump amplitude exceeds a preset physical thickness judgment threshold, it is determined that there is a real overlapping relationship between the two workpieces in physical space.

[0028] Each workpiece instance is abstracted as a node in a graph structure, and the stacking relationship obtained from the above determination is defined as a directed edge, with the direction of the directed edge pointing from the upper stacked workpiece to the lower stacked workpiece. Based on the irreversible nature of the physical stacking relationship under gravity constraints, a global topology graph with strict directed acyclic properties is constructed for the entire stacking scenario as a hierarchical topology reasoning model.

[0029] After the hierarchical topology reasoning model is constructed, by traversing the in-degree information of each node in the directed acyclic graph, all nodes that are not overlapped by any other workpieces, i.e., nodes with an in-degree of zero, are selected, and the workpiece instances corresponding to them are determined as the top-level candidate workpiece set in the current crawling cycle.

[0030] Step S300: Based on the off-center load characteristic data, perform a quantitative evaluation of the graspability of the selected top-level candidate workpieces, and determine the target workpiece to be grasped based on the evaluation results.

[0031] In this step, for each workpiece in the top-level candidate workpiece set selected in step S200, a multi-dimensional graspability score is constructed using the off-center load feature data generated in step S100. This scoring mechanism comprehensively considers the following physical factors: the lower the degree of workpiece occlusion, the smaller the off-center load amplitude between the visible area centroid and the complete physical centroid, and the closer the relative height of the workpiece in the current frame depth direction is to the top layer, the higher its graspability score.

[0032] All top-level candidate workpieces are ranked according to the above scores, and the workpiece with the highest score is determined as the target workpiece to be grasped in this grasping operation. This ensures at the decision-making level that the object with the most stable physical state and the lowest grasping risk is selected first in each grasping operation.

[0033] Step S400: Obtain the reference gripping pose of the target workpiece to be gripped, and based on the inherent rotational symmetry characteristics of the workpiece, expand the reference gripping pose into a pose set containing multiple equivalent candidate poses.

[0034] In this step, the initial pose estimation of the target workpiece to be grasped, determined in step S300, is first performed using a pre-trained pose estimation model to obtain a convergent and effective 6D reference pose of the workpiece in the camera coordinate system, which is then used as the anchor point for subsequent expansion.

[0035] Then, the direction of the main axis of symmetry of the target workpiece to be grasped is identified and determined. Based on the inherent rotational symmetry geometry of the workpiece itself, such as the six rotational symmetries of a hexagonal head bolt, the rotational angle sequence in which all physical grasping effects are completely equivalent when it rotates around the main axis of symmetry is calculated.

[0036] The anchor points are sequentially combined with the spatial rotation transformation corresponding to each equivalent rotation angle, and pre-grabbing parameters including safe approach distance are incorporated to generate a set of equivalent candidate poses that are completely equivalent in physical effect but differ in spatial cutting direction and approach path, thus forming the pose set for this grab.

[0037] Step S500: Map the pose set from the perception space to the robot execution space, perform feasibility verification on each equivalent candidate pose in the set in parallel within the execution space, and select a subset of globally safe poses that satisfy all constraints.

[0038] In this step, the pose set generated in step S400 in the camera perception space is first mapped to the robot base execution space using the offline calibrated hand-eye relationship transformation matrix.

[0039] Subsequently, for each equivalent candidate pose in the mapped pose set, the following multi-dimensional feasibility checks are performed in parallel: First, the pose is checked by inverse kinematics to determine whether there is a legal joint configuration that satisfies the hardware limit range of each joint of the robotic arm and leaves a safety margin; Second, the kinematic performance of the obtained joint configuration is checked to determine whether it is in a working range far from singular points and with good flexibility in all directions of motion; Third, the pose is checked by physical collision, and by combining the digital twin boundary model of the material frame with the real-time three-dimensional occupancy information of the stacked workpieces in the material frame, it is determined whether the robotic arm and the end effector collide or interfere with the surrounding environment when cutting in along the pose.

[0040] Based on the above verification results, all equivalent candidate poses that simultaneously satisfy the three core constraints of inverse kinematics solvability, non-singularity, and collision-free poses are retained to form a subset of globally safe poses.

[0041] Step S600: Select a target pose from the global safe pose subset and output motion control commands to drive the robotic arm to complete collision-free grasping.

[0042] In this step, the final target pose for this grasping operation is determined from the global safe pose subset selected in step S500 based on a preset optimal selection strategy (e.g., selecting the pose with the smallest joint spatial displacement and the highest execution efficiency). The spatial error between the current end-effector pose of the robotic arm and the target pose is calculated. Under the condition of satisfying the physical hard constraints on the position and velocity of each joint of the robotic arm, kinematic analysis and trajectory planning are performed on the spatial error. A sequence of joint-level motion control commands is generated and output in real time to drive the robotic arm to move smoothly to the target pose and complete the collision-free grasping action.

[0043] The primary objective of this invention is to address the 3D perception distortion caused by specular reflection from the surface of standard metal components, such as hexagonal head bolts, in scattered and stacked industrial environments, as well as the severe geometric truncation problem resulting from densely interwoven components. The system provides a high-fidelity physical basis for subsequent hierarchical inference by constructing a low-level multimodal data mapping and a closed-loop full-view topology reconstruction.

[0044] In some embodiments, the step S100 of acquiring a multimodal image of the stacked scene, reconstructing a complete physical mask representation of each workpiece from the multimodal image, and generating off-center load feature data accordingly can be specifically divided into steps S101-S104.

[0045] Step S101: Simultaneously acquire the color texture image and depth image of the stacked scene, and use the calibrated spatial mapping relationship to align and fuse the color texture image and depth image at the sub-pixel level to generate a multimodal input tensor.

[0046] By deploying a high-precision RGB-D camera at the robot's end effector or in the global field of view, and under the coordination of a hardware-triggered synchronization mechanism, the color texture stream and depth geometry stream of the stacked scene are simultaneously acquired. To eliminate the spatial parallax caused by the difference in physical installation positions between the color sensor and the depth sensor, this sub-step introduces a depth-color extrinsic affine transformation matrix obtained through offline calibration.

[0047] For any valid 3D spatial point in the depth image coordinate system, the extrinsic parameter matrix is ​​used to transform its rigid body to the color camera coordinate system, and then the intrinsic parameter matrix of the color camera is used to reproject it onto the pixel plane, thereby achieving sub-pixel-level precise registration between each valid measurement point in the depth image and the corresponding pixel in the color image, and finally generating a dense four-channel multimodal input tensor containing texture information and spatial topology information.

[0048] Specifically, the camera's intrinsic parameter matrix is ​​introduced. With offline calibration of the depth-color extrinsic affine transformation matrix For any valid 3D point in the depth map coordinate system Its reprojection to the corresponding pixel in the RGB pixel coordinate system The strict mathematical mapping relationship is as follows:

[0049]

[0050] By combining a millisecond-level hardware-triggered synchronization mechanism, the system achieves sub-pixel-level absolute alignment between texture features and 3D spatial topology at the underlying logic level, generating dense multimodal input tensors. .

[0051] Step S102: Guided by the high-frequency edge features in the color texture image, perform edge preservation repair on the void area in the depth image caused by the high reflectivity of the workpiece surface to obtain the repaired continuous depth map.

[0052] To address the issue of large-area numerical missing holes and discrete flying point noise in the depth map caused by the strong specular reflection effect on the surface of metal standard parts, this sub-step adopts an edge-preserving repair mechanism guided by the diffusion of high-frequency texture gradients of color images.

[0053] During the restoration process, an energy functional minimization objective is set to balance restoration smoothness and original data fidelity. The diffusion propagation coefficient is controlled by the color image edge gradient: in areas with large color image edge gradients, i.e., at the workpiece geometric boundaries, this coefficient decreases to suppress cross-edge smooth diffusion, thus preserving the sharp geometric edges of features such as bolt hexagonal heads; in areas with flat color image gradients within holes, this coefficient increases to promote adaptive propagation and filling of depth values. By iteratively solving this energy functional, a continuous and complete post-restoration depth map is obtained while maintaining the clarity of the target geometric edges.

[0054] Specifically, the system uses the high-frequency texture gradient of the RGB image. To guide features, the depth map Tensor voting repair is performed on the partial differential equations in the high-light void region. The energy functional minimization objective is set as follows:

[0055]

[0056] in, This is the original depth value. To ensure fidelity, The diffusion propagation coefficient is controlled by the edge gradient of the color image, and the fidelity weight λ ranges from 0.01 to 10.0. When λ is small, such as 0.01 to 1.0, the restoration process tends to fill smoothly, which is suitable for large-area hole regions; when λ is large, such as 1.0 to 10.0, the restoration result is more faithful to the original measurement value, which is suitable for removing sparse noise points. In typical industrial scenarios, a value of λ between 0.5 and 2.0 is recommended.

[0057] Step S103: Input the multimodal input tensor and the repaired continuous depth map into the pre-built full-view segmentation network. The full-view segmentation network injects the prior shape of the workpiece into implicit encoding and performs decoupling and cross-attention fusion of visible representation and occlusion representation in the latent space, and outputs the visible perception mask and complete physical mask representation of each workpiece simultaneously.

[0058] The multimodal input tensor x generated in sub-step S101 and the repaired continuous depth map obtained in sub-step S102 are jointly input into an offline-trained full-view segmentation network. This network is structurally equipped with a dual-blind branch decoder that combines feature decoupling and prior fusion. During the encoding phase, the network extracts high-dimensional visual features from the input data; during the decoding phase, the network forcibly decouples these high-dimensional visual features into two independent components: one is the visibility representation vector corresponding to the actual visible area in the image, and the other is the occlusion inference vector corresponding to the occluded invisible area. Simultaneously, the network injects implicit prior shape encodings of the workpiece type to be processed, which contain prior knowledge of the complete geometric shape of the workpiece.

[0059] Through a cross-attention mechanism, the network uses the visible representation vector as the query and the prior shape implicit encoding as the key to compute the feature response. This allows it to infer and complete the overlapping physical boundary within occluded regions lacking direct visual cues, based on prior shape constraints and contextual information from the visible portion, using a probability distribution. Finally, the network synchronously outputs the visible perceptual mask and the complete physical mask representation for each workpiece instance from two branches.

[0060] It should be noted that the full-view segmentation network employs a deep convolutional neural network with an encoder-decoder backbone architecture. In the specific implementation, the encoder part can utilize a backbone network pre-trained on large-scale data in the field of visual perception, such as a feature extraction architecture based on ResNet or Swin Transformer, to obtain robust low-level representation capabilities for general visual features. The decoder part is specifically designed for the full-view mask output requirements of this scheme, employing a double-blind branch structure.

[0061] The overall structure of the full-view segmentation network consists of the following three core functional layers: Encoder layer: The encoder receives the multimodal input tensor generated in sub-step S101 and the channel concatenation input of the repaired continuous depth map obtained in sub-step S102. It extracts multi-scale high-dimensional visual features through progressively downsampled convolution or self-attention operations. The feature map output by the encoder forms a fusion feature encoding in the deep latent space that combines texture information and spatial geometric information.

[0062] Feature Decoupling Layer: A feature decoupling layer is set between the encoder and decoder. This layer encodes the fused features output by the encoder through two parallel fully connected or convolutional projection heads, mapping them into a visible representation vector and an occlusion inference vector, respectively. The visible representation vector is responsible for encoding the geometric and texture attributes of the actual observable regions in the image, while the occlusion inference vector is responsible for encoding the latent features of the occluded regions based on context inference. The two vectors are mutually orthogonal in the latent space to achieve full decoupling of the features.

[0063] The dual-blind branch decoder layer comprises two parallel upsampling decoding paths: a visible branch and a full-view branch. The visible branch takes the visible representation vector as its primary input, recovering spatial resolution through progressive upsampling and skip connections, and outputting a visible perceptual mask. The full-view branch takes the fused features of the visible representation vector and the occlusion inference vector as its input, while also injecting the workpiece's prior shape implicit encoding. This prior shape implicit encoding is obtained by latent space compression of the workpiece's CAD model using a pre-trained shape autoencoder, carrying strong prior constraints on the workpiece's complete geometry. The full-view branch internally includes a cross-attention fusion module: it performs attention calculations using queries generated from the visible representation vector and key values ​​generated from the prior shape implicit encoding, filtering and activating prior shape features based on attention weights, and inferring and completing the workpiece's physical boundaries in invisible areas. Finally, the full-view branch outputs a complete physical mask representation.

[0064] In the initial training phase, a massive amount of synthetic image data covering different stacking configurations, occlusion levels, lighting conditions, and camera perspectives was generated in a virtual simulation environment using the workpiece CAD model and a physical rendering engine. For each frame of the synthetic image, accurate ground truth values ​​for the visible and full-view masks were pre-generated by the rendering engine. At this stage, the multimodal tensor of the synthetic image was used as input, and the ground truth values ​​for the visible and full-view masks were used as supervision signals. The network was fully pre-trained using binary cross-entropy loss or dice coefficient loss, enabling the network to initially learn the basic ability of full-view completion. A small amount of real multimodal image data from actual stacking scenarios was collected in industrial settings, and the corresponding ground truth values ​​for the visible and full-view masks were obtained through manual fine-tuning or semi-automatic annotation tools. The pre-trained network was then transferred and fine-tuned on this real dataset, while domain randomization and data augmentation strategies were introduced to reduce the perceptual difference between the synthetic and real domains. During fine-tuning, additional decoupling constraint loss was applied to the feature decoupling layer to strengthen the independence of visible and occlusion representations.

[0065] During the inference phase, the full-view segmentation network synchronously outputs the visible perception mask and the complete physical mask representation of each workpiece instance. Both masks are pixel-level precision binary or probabilistic maps, which provide the basis for the calculation of off-center feature data in the subsequent step S104 and the hierarchical topology inference in the step S200.

[0066] Specifically, the network in the deep hidden space Extracting high-dimensional visual features Then, it is forcibly decoupled into a visible representation vector. Occlusion inference vector Simultaneously, implicit prior coding of the CAD shape of hexagonal head bolts is injected into the network. The full-view decoder computes through a cross-attention mechanism. and The characteristic response, within a blind zone devoid of any visual cues, creates the illusion of overlapping physical boundaries in the form of a probability distribution. Finally, it synchronously outputs the visible perceptual mask of the workpiece. With full-view physical prediction mask .

[0067] Step S104: Based on the visible perception mask and the complete physical mask, calculate the geometric off-load vector that represents the degree of spatial deviation between the centroid of the visible area of ​​the workpiece and the centroid of the complete physical area. Use the geometric off-load vector and the occlusion rate determined based on the area ratio of the two masks as the off-load feature data.

[0068] After obtaining the visible sensing mask and the complete physical mask representations output from substep S103, this substep first defines indicator functions for the two types of masks respectively, and calculates the nonlinear spatial occlusion rate of the target instance based on the area ratio relationship between the complete physical mask and the visible sensing mask. This occlusion rate reflects the degree to which the workpiece is currently physically covered by other workpieces.

[0069] Secondly, the zeroth-order and first-order spatial moments of the visible perception mask are calculated separately to determine the apparent centroid of the visible region on the imaging plane. Simultaneously, the same spatial moment calculation is performed on the full-view physical mask to determine the complete physical centroid of the workpiece. The two-dimensional Euclidean distance between this apparent centroid and the physical centroid on the imaging plane is calculated as a geometric off-center load vector characterizing the risk of off-center loading during grasping under occlusion. Finally, the occlusion rate and the geometric off-center load vector are combined to output the off-center load feature data generated in step S100, providing quantitative physical basis for the hierarchical reasoning in step S200 and the graspability assessment in step S300.

[0070] Specifically, by defining a mask indicator function Accurate nonlinear spatial occlusion rate is calculated using double integrals. :

[0071] Secondly, the zeroth and first spatial moments of the image are introduced. Solve for the apparent centroid of the visible mask separately. Physical centroid of the full-view mask The system extracts the geometric offset vectors of the two components on the imaging plane.

[0072]

[0073] This feature not only provides a deep reference for the subsequent construction of hierarchical directed acyclic graphs, but also serves as the core physical boundary input for subsequent calculation of the risk of the off-center loading moment of the end gripper. It effectively bridges the semantic gap between data processing and state decision-making in traditional visual algorithms, and realizes cross-dimensional mapping from low-level pixel-level multimodal features to high-order physical and mechanical situation inference.

[0074] In some embodiments, step S200, which involves constructing a hierarchical topological reasoning model based on the complete physical mask representation by analyzing the spatial gradient of the overlapping regions of different workpiece masks, can be divided into sub-steps S201-S203.

[0075] Step S201: Extract the overlapping region of the complete physical mask representation of any two workpieces. When the overlapping region is not empty, estimate the normal vector of the corresponding spatial point cloud in the overlapping region.

[0076] After generating the complete physical mask representation of each workpiece instance in the scene in step S100, this sub-step begins to perform pairwise traversal detection of the spatial relationships between workpieces. For any two workpieces A and B, a pixel-by-pixel logical AND operation is performed on their complete physical mask representations on the image plane to extract the overlapping region of the two masks. If the overlapping region is empty, it indicates that the two workpieces do not have spatial intersection under the imaging viewpoint, and subsequent judgment is skipped; if the overlapping region is not empty, the 3D spatial point cloud of the overlapping region is reconstructed using the depth data of the corresponding position in the repaired continuous depth map obtained in step S100. Principal component analysis is performed on the point cloud data, and the eigenvector corresponding to the minimum eigenvalue of the point cloud covariance matrix is ​​used as the estimation result of the local surface normal vector of the overlapping region, and the normal vector directions are assigned to the local surfaces of workpieces A and B in this overlapping region.

[0077] In one specific embodiment, the system extracts the complete masks of any two instances from the global full-view segmentation result. and pixel intersection area .like If the data is not empty, then perform principal component analysis (PCA) on the corresponding depth point cloud within that region to calculate the local surface normal vector. and If the cosine value of the angle between two normal vectors The value is close to 1, and the average depth difference is less than the preset wall thickness threshold. If the two are found to be coplanar workpieces on the same plane, the overlapping determination is skipped; otherwise, the depth gradient fine analysis is entered.

[0078] Step S202: When the angle between the normal vectors of the two workpieces or the average spatial depth difference does not meet the coplanarity determination condition, sampling is performed along the boundary of the overlapping area, and the spatial depth jump gradient in the preset neighborhood on both sides of the boundary is calculated.

[0079] After obtaining the local surface normal vectors of workpieces A and B in the overlapping area in sub-step S201, this sub-step first performs preliminary screening for coplanarity. The cosine of the angle between the two normal vectors is calculated. If this cosine value is close to 1, meaning the two normal vectors are nearly parallel and the average spatial depth difference between the two workpieces in the overlapping area is less than a preset minimum wall thickness threshold, then it is determined that these two workpieces are likely coplanar workpieces belonging to the same planar layer, such as bolts laid side-by-side, without any overlapping relationship, and the subsequent fine-grained depth gradient analysis is skipped.

[0080] If the above coplanarity determination condition is not met, it indicates that there is a spatial misalignment between the two workpieces, and the overlapping direction needs to be further confirmed. This sub-step performs high-frequency spatial sampling along the boundary line segment of the visible sensing mask of the two workpieces in the overlapping area. Taking each sampling point as the center, a preset neighborhood of k pixels is extended to both sides along the normal direction of the boundary, and the jump amplitude of the depth value in the neighborhood is calculated.

[0081] Step S203: If the spatial depth jump gradient exceeds the preset physical thickness jump threshold, it is determined that there is a physical overlapping relationship between the two workpieces, and the relationship is defined as a directed edge. The two workpieces connected by the directed edge are used as nodes to construct a directed acyclic graph describing the physical overlapping relationship between the workpieces for the global scene, which serves as a hierarchical topological reasoning model.

[0082] After calculating the depth jump amplitude at each sampling point on the boundary of the overlapping area in sub-step S202, this sub-step compares the depth jump amplitude at each sampling point with a preset physical thickness jump threshold. This physical thickness jump threshold is set according to the actual physical thickness parameters of the workpiece to be grasped, and is used to distinguish between sudden depth changes caused by actual workpiece stacking and minor fluctuations caused by measurement noise.

[0083] If a depth jump exceeding the physical thickness jump threshold is detected on a continuous boundary segment, and the average depth values ​​of the two workpieces in the overlapping area are compared, confirming that one workpiece is significantly closer to the imaging sensor in the camera coordinate system and the other is significantly farther away, then it is determined that there is a strict physical stacking relationship between the two workpieces: the one with a smaller average depth value, i.e., closer to the camera, is the upper stacked workpiece, and the one with a larger average depth value, i.e. farther away from the camera, is the lower stacked workpiece.

[0084] After determining the overlapping relationships of all workpiece pairs, this sub-step constructs a global overlapping topology graph. Each identified workpiece instance in the scene is abstracted as a node in the graph structure, and each overlapping relationship obtained above is defined as a directed edge, with the direction of the directed edge pointing from the upper overlapping workpiece node to the lower stacked workpiece node. Due to the strict directional irreversibility of physical overlapping relationships under a gravitational field environment, the entire graph structure must mathematically be a directed acyclic graph, meaning there are no closed loops.

[0085] Finally, using this directed acyclic graph as the hierarchical topology reasoning model output from step S200, a complete description of the physical support and stacking topology between all workpieces in the current stacking scenario is obtained. In a specific example, high-frequency normal sampling is performed along the image boundary direction of the mask MvisA within the intersection region ΩAB. Using an improved adaptive Canny operator combined with the depth map matrix, the depth gradient jump ΔZ within the k-pixel neighborhood on both sides of the boundary is calculated. If ΔZ > δ exists on the continuous boundary segment... jump δ jump The threshold for the jump is determined by the physical thickness of the workpiece, and the average depth of region A is significantly smaller than that of region B, meaning it is closer to the sensor in the camera coordinate system. Therefore, in terms of physical constraints, it is strictly determined that instance A overlaps instance B.

[0086] It should be noted that δ jump The value of δ is determined based on the actual physical thickness of the workpiece to be gripped. In typical applications involving small standard parts such as hexagonal head bolts, this threshold is 0.3 to 0.8 times the actual physical thickness of the workpiece. For example, for a standard M8 hexagonal head bolt, the head thickness is approximately 5.5 mm, then δ jump The threshold can be set between 1.5 mm and 4.5 mm, with a recommended value of 0.5 times the workpiece thickness. This threshold can be preset after offline measurement of the workpiece's physical dimensions, or it can be automatically calibrated during actual operation based on the depth data from the first successful gripping.

[0087] Each identified workpiece instance within the material frame is abstracted as a graph node. The overlapping relationship determined in the above steps is defined as a directed edge. , representing a node Press on the node This allows for the construction of a global stacking topology network diagram for the entire material frame. Because physical overlap under gravity constraints is irreversible, this network strictly resembles a directed acyclic graph (DAG). The system can use a depth-first search algorithm to calculate the in-degree of this DAG, extracting all nodes with an in-degree of 0—that is, the absolute top-level targets not overlapped by any other objects—and using them as the safe candidate crawling pool for the current period. .

[0088] In some embodiments, step S300, which involves performing a quantitative evaluation of the graspability of the selected top-level candidate workpieces based on the off-center load characteristic data and determining the target workpiece to be grasped based on the evaluation results, can be divided into sub-steps S301-S302.

[0089] Step S301: Construct a multidimensional graspability score for each of the top-level candidate workpieces. The multidimensional graspability score is negatively correlated with the occlusion rate obtained from the off-center load feature data, negatively correlated with the modulus of the geometric off-center load vector obtained from the off-center load feature data, and positively correlated with the normalized height of the top-level candidate workpiece in the current frame depth direction.

[0090] This sub-step performs a multi-dimensional comprehensive quantitative evaluation of the crawlability of each workpiece instance in the top-level candidate workpiece set.

[0091] This evaluation mechanism considers factors from the following three physical dimensions: The first dimension is the occlusion rate. The occlusion rate index is extracted from the workpiece's off-center load characteristic data. This occlusion rate is determined by the ratio of the complete physical mask area to the visible perceived mask area, reflecting the degree to which the workpiece is physically overlapped and covered by surrounding workpieces. A lower occlusion rate indicates a larger visible exposed area and fewer physical constraints on the workpiece, resulting in less interference resistance from surrounding workpieces during grasping. Therefore, the graspability score is negatively correlated with this occlusion rate.

[0092] Secondly, there is the off-center load dimension. A geometric off-center load vector is extracted from the off-center load characteristic data of the workpiece, and its Euclidean modulus is calculated. This modulus reflects the degree of spatial deviation between the apparent centroid of the visible area of ​​the workpiece and its complete physical centroid. The greater the deviation, the greater the additional torque generated by the gripper due to the centroid offset when grasping the workpiece, and the higher the risk of the workpiece slipping or becoming unstable during lifting. Therefore, the gripability score is negatively correlated with the modulus of this geometric off-center load vector.

[0093] Thirdly, there is the relative height dimension. The spatial position information of the workpiece in the depth direction of the material frame is obtained and normalized to the height range from the bottom surface to the top opening of the material frame. The closer the workpiece is to the top opening of the material frame, the larger its normalized height value, the shallower the depth the robotic arm needs to extend into the material frame during gripping, and the lower the risk of collision with the frame wall and the material below. Therefore, the gripping capability score is positively correlated with this normalized height.

[0094] The factors from the three dimensions mentioned above are weighted and combined according to preset adaptive weights to generate a multi-dimensional graspability score for each top-level candidate artifact. The higher the score, the more suitable the artifact is as a priority grasping target for the current grasping cycle, considering both physical security and ease of execution.

[0095] In a specific example, targeting a secure candidate crawl pool In addition to considering the occlusion rate, relative height of the target, and physical off-center loading of the centroid, this invention further integrates these factors to construct a multi-dimensional graspability scoring function for each workpiece:

[0096] in, For the first The gripping capability score of each workpiece; the higher the score, the stronger the gripping safety and feasibility. The target occlusion rate reflects the degree to which the target is obscured by surrounding workpieces. This represents the normalized height of the workpiece in the depth direction of the material frame; a larger value indicates that it is closer to the top layer. The Euclidean off-center distance between the centroid of the two-dimensional visible mask and the centroid of the full-view three-dimensional reconstruction; These are adaptive weighting coefficients used to balance the contribution weights of each evaluation index to crawlability based on actual working conditions. The values ​​of all three coefficients are greater than zero and satisfy the normalization constraint α1 + α2 + α3 = 1. In practical applications, the typical ranges for the three coefficients are as follows: occlusion rate weight α1 is 0.3 to 0.5, off-center load weight α2 is 0.2 to 0.4, and height weight α3 is 0.2 to 0.4. These can be adjusted adaptively to suit different operational scenarios.

[0097] Step S302: Sort the multidimensional graspability scores of each of the top-level candidate workpieces in descending order, and select the top-level candidate workpiece with the highest score as the target workpiece to be grasped.

[0098] After sub-step S301 calculates the multi-dimensional graspability score for all workpieces in the top-level candidate workpiece set, this sub-step sorts all top-level candidate workpieces in descending order of their scores, forming a priority sequence for this grasping cycle. Workpieces ranked higher in this sequence represent those with more stable current physical states and a higher probability of successful grasping.

[0099] The system directly selects the workpiece with the highest score from the priority sequence, identifies it as the target workpiece to be grasped in step S300, and passes it to the subsequent step S400 for pose estimation and equivalent pose expansion processing. Through the above scoring and filtering mechanism, in each grasping decision, the system can prioritize the workpiece with the lowest occlusion degree, the lowest centroid off-center load risk, and the most favorable spatial position from multiple top-level candidate objects, thus maximizing the safety and success rate of the grasping operation from the decision logic level.

[0100] In some embodiments, the process of obtaining the reference gripping pose of the target workpiece in step S400, based on the inherent rotational symmetry characteristics of the workpiece, expands the reference gripping pose into a pose set containing multiple equivalent candidate poses, which can be specifically divided into sub-steps S401-S403.

[0101] Step S401: Obtain any converged 6D reference pose of the target workpiece in the camera coordinate system as an anchor point through a pre-trained pose estimation network.

[0102] After determining the target workpiece to be grasped in step S300, this sub-step inputs the visible sensing mask of the target workpiece and the corresponding spatial point cloud data into an offline pre-trained pose estimation network. During the training phase, this network does not employ the traditional globally unique solution constraint loss function. Instead, it introduces symmetry sensing loss or dense feature correspondence loss, allowing the network to freely converge to any valid solution in the pose multi-solution space caused by the workpiece's rotational symmetry. This avoids the network output oscillations and random angle jumps caused by forcibly fitting a unique solution.

[0103] It should be noted that the pose estimation network employs a deep neural network architecture specifically designed for 6D pose estimation tasks. In the specific implementation, a pose estimation network backbone based on dense feature correspondence or direct pose regression can be selected, such as an improved variant of the PoseCNN architecture based on convolutional neural networks, or an end-to-end pose estimation model based on Transformers. The key to network selection lies in ensuring that its loss function design matches the strategy of accepting symmetric multiple solutions in this scheme.

[0104] The overall structure of the pose estimation network consists of the following three functional layers: Feature extraction layer: Receives the region image block cropped by the visible perception mask of the target workpiece to be grasped and its corresponding spatial point cloud data as input. It obtains texture features and geometric features through two-dimensional and three-dimensional parallel feature extraction branches, and fuses the two in the feature channel dimension to form a comprehensive visual-geometric feature description of the target.

[0105] Pose regression layer: The fused integrated features are input to the translation regression head and the rotation regression head respectively. The translation regression head outputs the 3D translation vector of the target in the camera coordinate system. The rotation regression head outputs the rotation representation of the target in the camera coordinate system, which can be in parameterized form such as quaternions, rotation matrices, or axis angles.

[0106] Symmetry-aware output layer: A symmetry-aware output layer is set after the rotation regression head. This layer does not force the network to output a globally unique optimal rotation value, but allows the network to converge to any valid solution within the equivalent rotation subspace defined by the workpiece rotational symmetry. In specific implementation, this can be achieved by introducing a symmetry-aware loss function during the training phase: this loss function does not directly calculate the minimum distance between the network's predicted rotation value and a single ground truth value, but instead calculates the minimum distance between the network's predicted rotation value and all equivalent ground truth values ​​under the action of the workpiece symmetry group, and uses this minimum distance as the optimization objective.

[0107] For the specific training of this network, a synthetic training image covering different poses, occlusion levels, and lighting conditions can be generated in a virtual simulation environment using a workpiece CAD model and a physical rendering engine. For each frame, not only is a ground truth pose of the workpiece in the camera coordinate system recorded, but also, based on the known rotational symmetry characteristics of the workpiece, a set of all equivalent pose labels for that ground truth pose under the action of the symmetry group is pre-calculated. During training, for each set of input data, the pose distance loss is calculated for each label in the network's output predicted pose and the equivalent pose label set. This pose distance loss can be decomposed into the Euclidean distance loss of the translation component and the angular distance loss of the rotation component. Finally, the minimum pose distance between the predicted pose and all equivalent labels is taken as the loss value for that set of training samples. The design of this loss function ensures that the network will not generate contradictory gradient signals due to the rotational symmetry of the workpiece during training, thus avoiding loss oscillation and non-convergence problems. The loss of the translation component directly uses the Euclidean distance between the predicted translation vector and the ground truth translation vector. The loss for the rotation component is measured by the angular distance between the predicted rotation matrix and the ground truth rotation matrix. The network is trained iteratively using the aforementioned synthetic dataset and the symmetry-aware loss function until the loss converges. After training, the network is capable of stably outputting any valid pose solution when facing a rotationally symmetric target.

[0108] In the inference phase, the pose estimation network performs forward inference on the target workpiece to be grasped, outputting a convergent and valid 6D reference pose of the workpiece in the camera coordinate system. This reference pose includes rotation matrix components and translation vector components, serving as anchor points for the symmetry expansion in step S402. These anchor points are only used as mathematical benchmarks for subsequent expansion; their rotation components do not need to be globally uniquely optimal, but only need to belong to a valid element in the equivalent pose set corresponding to the symmetry group of the workpiece.

[0109] The network performs 6D pose inference on the target workpiece to be grasped, outputting a convergent effective pose in the camera coordinate system, containing a rotation matrix component and a translation vector component. This pose serves only as the mathematical benchmark for subsequent equivalent expansion, defined as the benchmark anchor pose, and does not need to strictly satisfy the optimality requirement of the rotation angle. This benchmark anchor pose is then passed to subsequent sub-steps for symmetric expansion processing.

[0110] In a specific example, by introducing symmetry-aware loss or dense feature correspondence, the network is allowed to... Output any convergent valid solution in the space, and define it as the target in the camera coordinate system. The reference anchor point pose :

[0111] in, As the reference rotation matrix, This is the translation vector. This anchor point serves only as a mathematical benchmark for subsequent redundant expansion and does not need to possess physical optimality for grasping.

[0112] Step S402: Determine the principal axis of symmetry of the target workpiece to be grasped, and calculate a series of equivalent rotation angles around the principal axis of symmetry based on at least one of the discrete rotational symmetry order of the principal axis of symmetry and the angle step size set for the continuous rotational symmetry body.

[0113] After obtaining the pose of the reference anchor point in sub-step S401, this sub-step analyzes the rotational symmetry geometry of the target workpiece to be grasped. First, based on the prior geometric information of this workpiece type, a local coordinate system with the workpiece itself as the reference is established, and the direction of its principal axis of symmetry is determined. For standard parts such as hexagonal head bolts, this principal axis of symmetry usually corresponds to the axial central axis of the bolt.

[0114] Then, the calculation method for the equivalent rotation angle is determined based on the symmetry type of the workpiece. For workpieces with discrete rotational symmetry characteristics, such as hexagonal head bolts which have six-fold rotational symmetry (discrete symmetry order six), a posture that is completely equivalent in physical gripping effect can be obtained by rotating 60 degrees around its principal symmetry axis. Based on this, a series of discrete equivalent rotation angles corresponding to the symmetry order are calculated and generated. For workpieces with continuous rotational symmetry characteristics, such as cylindrical pins, the angle step size is set according to the actual mechanical gripping tolerance of the end effector. The continuous 360-degree rotation range is discretized into a finite sequence of equivalent rotation angles at intervals of this step size.

[0115] In a specific example, a local coordinate system is established using the geometric principal symmetry axes of the workpiece. Let the axis of symmetry be... The axis, for discrete symmetry of order For the target group, this invention introduces Lie algebras. We perform dimensionality reduction analysis using spinor theory in [the context of the original text]. We define [the concept of] revolving around a local [location]. generator of the symmetry group of the axis By using the Rodriguez formula and the Lie group exponent mapping, the pose of a single anchor point can be reconstructed into a complete set of equivalent physical pose manifolds. :

[0116]

[0117] To improve the universality of this method, if the target to be grasped is a cylindrical pin or similar object... For workpieces with continuous symmetrical features, the continuous rotation angle will be... Discretize the step size based on the mechanical tolerance of the end gripper. A finite sequence.

[0118] Step S403: Combine the anchor point with the rotation transformation containing each equivalent rotation angle in sequence to generate a set of candidate poses that are completely equivalent in physical grasping effect and contain the pre-grabbing proximity distance parameter, thus forming a pose set.

[0119] After obtaining the reference anchor point pose in sub-step S401 and a series of equivalent rotation angles in sub-step S402, this sub-step combines the anchor point with the spatial rotation transformation corresponding to each equivalent rotation angle. For each equivalent rotation angle, a rotation transformation matrix of that angle about the principal axis of symmetry of the workpiece's local coordinate system is constructed. This rotation transformation matrix is ​​then combined with the rotation component in the reference anchor point pose to generate an equivalent grasping pose that is completely equivalent to the anchor point in terms of physical grasping effect, but with a different spatial orientation.

[0120] Meanwhile, considering that actual grasping operations not only require determining the final closed pose but also planning a safe approach trajectory to avoid collisions during the approach, this sub-step incorporates grasping configuration parameters when generating each equivalent grasping pose. Specifically, by combining the inherent grasping bias relationship of the end effector with preset safe pre-grasping distance parameters along the approach direction, a corresponding pre-grasping approach point is generated for each equivalent grasping pose. This parameterization ensures that each candidate pose in the generated pose set is not only legal in the static closed state but also physically feasible throughout its dynamic approach and approach path.

[0121] The poses with pre-captured parameters corresponding to all equivalent rotation angles are collected to form the equivalent candidate pose set output in this step, which is then passed to the subsequent step S500 for multi-constraint verification and filtering.

[0122] After generating the equivalent candidate pose set, this step also performs a physical feasibility-based attitude orientation pre-screening process on the set. Specifically, after mapping the pose set to the robot base coordinate system, the end effector approach direction vector of each candidate pose is extracted, and its spatial dot product with the normal vector of the bottom surface of the workpiece is calculated. If the dot product of the approach direction vector of a candidate pose with the normal vector of the bottom surface of the workpiece indicates that it attempts to penetrate the workpiece upward from the bottom of the workpiece for grasping, then the pose is determined to be absolutely invalid in physical topology and is directly removed from the pose set.

[0123] In a specific example, considering that actual physical grasping requires not only calculating the final closed pose but also planning a safe approach trajectory, grasping configuration parameters are further introduced in EPS. The intrinsic grasping bias matrix of the end effector is defined. And along the approach direction of the gripper tool center point (TCP), typically in the TCP coordinate system. Direction, set safe pre-grab distance Therefore, for For each pose, a corresponding pre-grab point mapping operator is generated. :

[0124]

[0125] in, The unit approximation direction vector. This parametric modeling ensures that the generated equivalent poses are not only statically valid, but also possess physically constrained semantics in the dynamic ingress path.

[0126] To transform the equivalent redundancy of perception into the robot's physical obstacle avoidance capability, an offline calibrated hand-eye extrinsic parameter matrix is ​​used. Map the equivalent set to the robot base coordinate system Next. Simultaneously, combining the aforementioned parameterized capture model, the complete set of equivalent TCP execution instructions required by the underlying controller is calculated. :

[0127]

[0128] Before outputting this set to the inverse kinematics solver, the system adds a pose pre-screening mechanism based on the task space inner product. Extraction gripper ingress vectors for each pose Calculate its normal vector to the bottom surface of the material frame. The dot product. If If the gripper attempts to penetrate and grasp the material from the bottom of the frame upwards, then this pose is physically invalid, and the system directly removes it from the set. (Threshold) The minimum allowable angle for determining whether the cutting direction of the equivalent candidate pose is physically valid is between 70 and 110 degrees.

[0129] Based on the aforementioned end-to-end isomorphic mapping and pre-screening mechanism, this invention effectively resolves the pose ambiguity induced by the geometric symmetry of the target at the visual perception end at the basic algorithm architecture level, and reconstructs it into a high-dimensional kinematic redundancy decision space for the robotic arm to perform active interference avoidance in a deep and narrow confined space.

[0130] In some embodiments, step S500 performs a parallel verification operation on each equivalent candidate pose in the pose set.

[0131] After generating the equivalent candidate pose set pre-screened by attitude direction in step S400, step S500 first uses the offline calibrated hand-eye relationship transformation matrix to uniformly map the pose set located in the camera perception space coordinate system to the robot base execution space coordinate system, obtaining the equivalent candidate pose set expressed in the robot base coordinate system. Subsequently, for each equivalent candidate pose in this set, the following three dimensions of feasibility verification operations are performed in parallel.

[0132] The first verification operation is to perform inverse kinematics on the equivalent candidate poses and verify whether there is a joint configuration solution that satisfies the hardware limit of the robotic arm within the preset safety margin.

[0133] The inverse kinematics solution is performed on the spatial description of the current equivalent candidate pose in the robot's base coordinate system. This solution process employs a nonlinear numerical iterative solver based on damped least squares or the Levenberg-Marquardt algorithm. The core criterion for verification is whether the solver can converge within a preset maximum number of iterations, and whether each joint angle in the obtained joint configuration solution strictly satisfies the boundary constraints of the robot arm's motor hardware and software limits. In actual verification, a safety dead zone margin is set inside the hardware limit boundaries. Only when all joint angles in the obtained joint configuration solution are within the safety range of the lower limit boundary plus margin and the upper limit boundary minus margin is the equivalent candidate pose deemed to have a valid joint configuration solution, and the inverse kinematics solvability verification passes.

[0134] In a specific example, for any candidate pose The system needs to map it from the three-dimensional Cartesian task space to the six-dimensional joint configuration space of the robotic arm. The underlying implementation of this invention employs a nonlinear inverse kinematics (IK) solver based on damped least squares (DLS) or Levenberg-Marquardt (LM) algorithms for numerical iteration. A joint limit safety discrimination functional is defined. If the IK solver converges within a preset number of iterations, and the obtained joint configuration set... Strictly meet the hardware and software limit boundaries of the motor, that is , For a safe dead zone margin, then If there is no solution or the limit is exceeded, then , The value of is determined based on a combination of the rated movement speed of the robotic arm joints and the response delay of the control system. In typical industrial applications, The value ranges from 0.5 degrees to 3.0 degrees.

[0135] The second verification operation is to calculate the operability index for the existing joint configuration solution based on its velocity transmission characteristics in the joint space, and verify whether it is in a flexible working range far away from the kinematic singularity.

[0136] Based on the valid joint configuration solution obtained in the first verification operation described above, the kinematic performance of the joint configuration is evaluated. The velocity transfer Jacobian matrix of the robotic arm from joint space to task space in the current state is calculated based on this joint configuration. According to the determinant of the product of this Jacobian matrix and its transpose, the operability index characterizing the flexibility of the robotic arm's current configuration in all directions is calculated.

[0137] A safety threshold for operability is set. If the calculated operability index is higher than this threshold, it indicates that the robotic arm's motion capability in all directions is well-balanced under this configuration, and it is far from the singular configuration region caused by the rank reduction of the Jacobian matrix. During the grasping and cutting process, it will not cause motor overload or control instability due to sudden changes in the angular velocity requirements of individual joints, and the kinematic singularity avoidance verification is passed. If the operability index is lower than or equal to this threshold, the configuration is identified as a singularity point and is discarded.

[0138] In a specific example, after confirming the legality of the joint limits, to prevent the robotic arm from getting stuck in singularities during the grasping and cutting process, leading to sudden changes in the angular velocity of individual joints and motor overload tripping, this invention introduces the Yoshikawa Jacobi maneuverability index in the initial screening stage. The current joint configuration is then calculated. Jacobian matrix under And construct singularity avoidance indicator function :

[0139]

[0140] like , The system's calibrated lower limit threshold for operability demonstrates that the robot possesses good isotropic motion flexibility in this posture, thus endowing it with... Otherwise, it is judged as a near-singular configuration and is eliminated. , The value of is related to the inherent kinematic parameters of the robotic arm. In practical applications, the operability index can be calculated for each configuration by offline traversal sampling of the robotic arm's joint space, and the value between the 5th and 15th percentiles of the distribution of all sampling results can be taken as . In typical six-axis industrial robot applications, it is recommended to Take a value in the range of 0.01 to 0.10.

[0141] The third verification operation: For the equivalent candidate pose, combined with the digital twin boundary model of the material frame and the real-time three-dimensional occupancy map of the stacked workpieces in the material frame, calculate the minimum distance between the end of the robotic arm and the obstacles in the boundary model and the occupancy map when it cuts in along the pose, and verify whether the minimum distance is greater than the preset safe expansion radius.

[0142] Physical collision detection is performed on the current equivalent candidate pose at two levels: static environment and dynamic stacking.

[0143] At the static environment level, a digital twin boundary model of the material frame, serving as the work container, is imported, along with hierarchical bounding box models of the robotic arm body and the end effector. The minimum Euclidean distance between the robotic arm end effector envelope and each wall of the material frame boundary model is calculated along the tangent direction corresponding to the equivalent candidate pose.

[0144] At the dynamic stacking level, the spatial point cloud of scattered materials inside the material frame that have not yet been grasped, acquired in real time by the depth camera, is transformed into a voxelized 3D occupancy map represented by a truncated symbolic distance field. Along the entry path of this equivalent candidate pose, a fast voxel query is performed on the sampling points on the outer surface of the end effector to obtain the penetration depth value of each sampling point in the occupancy map. The minimum penetration depth among all sampling points is used as the minimum distance between the gripper and the dynamically stacked materials.

[0145] Combining the above two levels, the smaller value between the minimum distance in the static environment and the minimum distance in the dynamic stacking is taken and compared with the preset safe expansion radius of the end effector. If the smaller value is greater than the safe expansion radius, it indicates that the robotic arm maintains sufficient safe clearance with the inner wall of the material frame and the stacked materials when cutting in along the equivalent candidate pose, and the multi-level physical collision verification is passed.

[0146] In a specific example, the CAD model of the material box is imported along with the hierarchical bounding box (BVH) of the robotic arm, such as an OBB tree. The separation axis theorem is used to calculate the gripper envelope along... Minimum Euclidean distance between the material frame and the inner wall during the orientation cut-in. Dynamic stacked point cloud verification: For scattered materials inside the material box that have not yet been grasped, the system introduces a truncated symbolic distance field in real time. The real-time environmental point cloud acquired by the depth camera is converted into a voxelized voxel mesh, allowing for rapid querying of the penetration depth from the sampling point on the end effector surface to the nearest obstacle point cloud. .

[0147] Construct a comprehensive collision-free discriminant function :like , If the gripper's safe expansion radius is defined, then this posture is absolutely safe in physical space. Otherwise, it is 0. The value is determined comprehensively based on the combined outer envelope dimensions of the end effector and the gripped workpiece, the repeatability of the robotic arm, and the spatial measurement error of the vision perception system. In typical industrial scenarios, The value ranges from 3 mm to 10 mm.

[0148] After performing the above three parallel verification operations on each equivalent candidate pose in the pose set, the equivalent candidate pose that simultaneously satisfies the three constraints of having a valid joint configuration solution, being in a non-singular flexible working range, and passing multi-level physical collision verification is determined as an element of the global safe pose subset. This global safe pose subset is the output result of step S500, which is passed to step S600 for the final target pose selection and motion command generation.

[0149] When the global safe pose subset is not empty, it indicates that in the equivalent pose redundancy space generated by utilizing the rotational symmetry characteristics of the workpiece, there exists at least one safe grasping pose that satisfies both kinematic solvability and collision risk. The system can then proceed to step S600 to complete the instruction calculation and output for this grasping operation.

[0150] In some embodiments, after determining the target pose from a global safe pose subset, the relative transformation between the current end-effector pose of the robotic arm and the target pose is calculated, and the relative transformation is converted into a spatial screw error vector using a Lie algebra-logarithmic mapping, which is then used as the tracking error input for subsequent trajectory planning.

[0151] From feasible subsets The optimal target pose for the current cycle is extracted based on the principle of minimum joint displacement. Calculate the current end effector pose of the robotic arm. and The relative transformation matrix between Using Lie's modern numbers The logarithmic mapping rigorously transforms this relative transformation into a six-dimensional spinor error vector containing three-dimensional linear velocity error and three-dimensional angular velocity error. :

[0152]

[0153] This spinor error provides the most stringent Cartesian space kinematic tracking gradient for the underlying servo system, avoiding the gimbal lock problem caused by Euler angles.

[0154] Based on the current joint configuration of the robotic arm Using initial values, a QP objective function integrating joint displacement penalty and soft constraint relaxation is constructed. To achieve minimal end-effector motion and strictly limit the peak energy consumption of the base motor, the following multi-objective optimization functional is defined:

[0155]

[0156] in, The sequence of joint angle increments to be solved; The heterogeneous penalty weight matrix is ​​positive definite and symmetric. It assigns a large weight coefficient to the base joint and shoulder joint with large inertia, forcing the system to prioritize the use of the end-effector lightweight wrist joint for attitude adjustment. This is the guiding vector for gradient descent; The vector consists of slack variables, coupled with a penalty weight matrix that has maximal eigenvalues. It is used to absorb nonlinear trajectory tracking errors under extreme conditions.

[0157] It should be noted that regarding the heterogeneous penalty weight matrix H, the diagonal weight coefficients for each joint are differentiated based on the magnitude of the inertia of the robotic arm links driven by the joint. Specifically, joints driving large-inertia base rotation, upper arm pitch, and shoulder movement are assigned larger penalty weights, such as weight coefficients ranging from 10.0 to 100.0, to suppress large-amplitude movements of these joints and reduce overall energy consumption; joints driving small-inertia wrist posture fine-tuning are assigned smaller penalty weights, such as weight coefficients ranging from 0.1 to 1.0, to encourage the system to prioritize scheduling these joints to complete posture adjustments, improving motion efficiency and smoothness. For W... δ The values ​​of the diagonal elements should be significantly larger than those of the diagonal elements of the heterogeneous penalty weight matrix H to ensure that the slack variables are only activated when there is a real physical constraint conflict, and remain close to zero under normal operating conditions. δ The recommended value for diagonal elements is 1×10. 3 Up to 1×10 6 The typical recommended value is 1×10. 4 .

[0158] Under the above optimization objective, the controller synchronously applies a strict set of physical constraints consisting of equality and inequality equations:

[0159]

[0160]

[0161] in, This is the basic Jacobian matrix under the current configuration; The underlying servo control cycle can be set to... To ensure The high-frequency response. By introducing slack variables. This method effectively resolves the unsolvable numerical optimization problem caused by feasible region conflicts under the dual hard constraints of multidimensional obstacle avoidance boundary and end-point trajectory tracking. Traditional QP solvers are prone to fatal flaws such as unsolvable crashes. The solver converges rapidly in milliseconds using the effective set method, and outputs smooth joint execution instructions in real time that are far from singularities, achieve absolute obstacle avoidance, and have optimal overall energy consumption. .

[0162] In some embodiments, when the global safe pose subset output in step S500 is empty, the method further includes an exception handling step, which can be divided into sub-steps S601-S604.

[0163] Step S601: Check if the global safe pose subset is empty.

[0164] After completing the multi-dimensional parallel verification of all poses in the equivalent candidate pose set in step S500, this sub-step summarizes and judges the verification results. It checks whether there are any elements in the global safe pose subset: if, after verification, the subset retains at least one equivalent candidate pose that satisfies all constraints, it is determined to be in a normal state, and the system enters the normal execution branch of step S600, selecting the optimal target pose from the safe pose subset and generating motion control commands.

[0165] If the verification result shows that the global safe pose subset is empty, that is, all equivalent candidate poses generated in step S400 and pre-screened are eliminated one by one due to the lack of inverse kinematics solution, neighboring singular configurations or physical collision interference, and none of them pass all the verifications, then it is determined that the current capture cycle has entered an abnormal deadlock state, triggering this abnormal handling branch.

[0166] Step S602: When it is determined that the global safe pose subset is empty, stop the motion command output of the current capture cycle and generate an active perturbation command.

[0167] After sub-step S601 detects that the global safe pose subset is empty and determines that none of the current equivalent candidate poses can be safely entered for execution, the system immediately suspends the current picking task.

[0168] Unlike traditional systems that issue fault alarms and passively shut down in this situation, awaiting manual intervention, this sub-step proactively generates an active perturbation command to change the physical state of the stacked scenario. This command is generated autonomously by the system based on currently perceived scenario information, without relying on external operator intervention. Its aim is to break the current rigid physical stack deadlock by introducing external mechanical energy, creating the possibility of re-establishing feasible grasping conditions.

[0169] Step S603: Execute the active disturbance command, control the robotic arm to use the end effector as a probe, and perform a preset low-speed lateral tossing action along the path of the maximum gap area indicated in the hierarchical topology inference model, so as to change the rigid physical layout of the workpiece in the current stacking scene.

[0170] After generating the active disturbance command in sub-step S602, this sub-step controls the robotic arm to execute the disturbance action. First, based on the hierarchical topology reasoning model constructed in step S200, which is a directed acyclic graph describing the physical stacking relationship between all workpieces in the current stacking scenario, the spatial structure information of the scenario is extracted from it to identify and locate the largest gap region inside the discharge frame that is currently not occupied by workpieces or has the lowest workpiece density. The identification of this gap region utilizes the spatial distribution information of the complete physical mask representation of each workpiece in the hierarchical topology reasoning model.

[0171] Subsequently, the system controls the robotic arm to close the end effector as a probe, planning a motion path from the current position to the maximum gap area. During path execution, the robotic arm performs a lateral tossing motion at a preset low speed. The speed of this tossing motion is strictly limited to a low speed range to ensure that the tossing process will not cause impact damage to the workpiece or the robotic arm itself, while the applied disturbance force is sufficient to break the rigid contact balance between the workpieces maintained by static friction. After being subjected to this lateral disturbance force, the workpiece group undergoes macroscopic displacement and attitude reconstruction. The workpieces that were originally in a mutually wedged and stuck state are redistributed and rearranged, thereby releasing the physical interference constraints that prevented all equivalent candidate poses from entering.

[0172] Step S604: After the toggle action is completed, clear the current decision queue and trigger the reacquisition of multimodal images of the stacked scene to start a new grasping decision cycle.

[0173] After the low-speed lateral tossing action executed in sub-step S603 is completed, this sub-step resets the system's decision and control state. First, it clears the decision queue that has been generated but not yet executed in the current grasping cycle, including the top-level candidate workpiece list filtered in step S200, the target workpiece information determined in step S300, the equivalent candidate pose set generated in step S400, and the intermediate results such as the global safe pose subset that has been determined to be empty and output in step S500, to avoid expired decision information interfering with the judgment of the new cycle.

[0174] Subsequently, the system automatically triggers the re-execution of step S100, synchronously acquiring the color texture image and depth image of the stacked scene again, and initiating a new grasping decision cycle. In the new cycle, the workpiece stacking state after perturbation reconstruction will be re-perceived. Step S200 will reconstruct the hierarchical topology reasoning model based on the new spatial layout, and steps S300 to S600 will sequentially re-execute the entire process of workpiece screening, pose expansion, verification, and instruction generation. Through the above closed-loop mechanism, the system has the ability to autonomously break through bottlenecks and self-recover in unattended long-cycle continuous operations, effectively avoiding production line downtime and manual intervention requirements caused by occasional extreme stacking deadlocks, and ensuring the continuous operational stability of automated sorting operations.

[0175] The present invention provides a hierarchical reasoning-based unordered grasping system for symmetric workpieces, which is used to implement the above-mentioned method. The system includes an image acquisition and representation generation module, a hierarchical topology reasoning module, a quantization evaluation module, an equivalent pose space mapping module, a parallel constraint screening module, and a global optimal instruction generation module.

[0176] The image acquisition and representation generation module is used to acquire multimodal images of the stacked scene, reconstruct the complete physical mask representation of each workpiece from the multimodal images, and generate off-center load feature data describing the occlusion state of each workpiece accordingly.

[0177] The hierarchical topology reasoning module is used to construct a hierarchical topology reasoning model that describes the physical overlapping relationship between workpieces by analyzing the spatial gradient of the overlapping areas of different workpiece masks based on the complete physical mask representation, and to select the top-level candidate workpiece in the current scene based on the hierarchical topology reasoning model.

[0178] The quantitative evaluation module is used to perform a quantitative evaluation of the grabbability of the selected top-level candidate workpieces based on the off-center load characteristic data, and to determine the target workpiece to be grabbed based on the evaluation results.

[0179] The equivalent pose space mapping module is used to obtain the reference grasping pose of the target workpiece to be grasped. Based on the inherent rotational symmetry characteristics of the workpiece, the reference grasping pose is expanded into a pose set containing multiple equivalent candidate poses.

[0180] The parallel constraint filtering module is used to map the pose set from the perception space to the robot execution space, and to perform feasibility verification on each equivalent candidate pose in the set in parallel within the execution space, and to filter out the global safe pose subset that satisfies all constraints.

[0181] The global optimal instruction generation module is used to select a target pose from the global safe pose subset and output motion control instructions to drive the robotic arm to complete collision-free grasping.

[0182] The electronic device of the present invention includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the above-described method.

[0183] The present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0184] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0185] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0186] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some service interfaces; indirect couplings or communication connections between apparatuses or units may be electrical or other forms.

[0187] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0188] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0189] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0190] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for unordered grasping of symmetrical workpieces based on hierarchical reasoning, characterized in that, include: Acquire multimodal images of the stacked scene, reconstruct the complete physical mask representation of each workpiece from the multimodal images, and generate off-center load feature data accordingly. The off-center load feature data is used to describe the occlusion state of each workpiece. Based on the complete physical mask representation, a hierarchical topology reasoning model is constructed by analyzing the spatial gradient of the overlapping areas of different workpiece masks. The top-level candidate workpiece in the current scene is selected according to the hierarchical topology reasoning model, which is used to describe the physical overlapping relationship between workpieces. Based on the off-center load characteristic data, the graspability of the selected top-level candidate workpieces is quantitatively evaluated, and the target workpiece to be grasped is determined based on the evaluation results. Obtain the reference gripping pose of the target workpiece to be gripped, and expand the reference gripping pose into a pose set containing multiple equivalent candidate poses based on the inherent rotational symmetry characteristics of the workpiece. The pose set is mapped from the perception space to the robot execution space. In the execution space, the feasibility of each equivalent candidate pose in the set is checked in parallel, and a subset of globally safe poses that meet all constraints is selected. Select a target pose from the global safe pose subset and output motion control commands to drive the robotic arm to complete collision-free grasping.

2. The method for disordered grasping of symmetrical workpieces based on hierarchical reasoning as described in claim 1, characterized in that, The process of acquiring multimodal images of the stacked scene, reconstructing the complete physical mask representation of each workpiece from the multimodal images, and generating off-center loading feature data accordingly includes the following steps: Simultaneously acquire color texture images and depth images of the stacked scene, and use the calibrated spatial mapping relationship to align and fuse the color texture images and depth images at the sub-pixel level to generate a multimodal input tensor; Guided by the high-frequency edge features in the color texture image, edge-preserving repair is performed on the cavity area in the depth image caused by the high reflectivity of the workpiece surface to obtain a repaired continuous depth map. The multimodal input tensor and the repaired continuous depth map are input into the pre-built full-view segmentation network. The full-view segmentation network injects the prior shape of the workpiece into implicit encoding and performs decoupling and cross-attention fusion of visible representation and occlusion representation in the latent space, and outputs the visible perception mask and complete physical mask representation of each workpiece simultaneously. Based on the visible perception mask and the complete physical mask, the geometric off-load vector, which represents the degree of spatial deviation between the centroid of the visible area of ​​the workpiece and the centroid of the complete physical area, is calculated. The geometric off-load vector and the occlusion rate determined based on the area ratio of the two types of masks are used together as the off-load feature data.

3. The method for disordered grasping of symmetrical workpieces based on hierarchical reasoning as described in claim 1, characterized in that, The method for constructing a hierarchical topological reasoning model based on complete physical mask representation, by analyzing the spatial gradient of the overlapping regions of different workpiece masks, includes the following steps: Extract the overlapping region of the complete physical mask representation of any two workpieces. When the overlapping region is not empty, estimate the normal vector of the corresponding spatial point cloud in the overlapping region. When the angle between the normal vectors of two workpieces or the average spatial depth difference does not meet the coplanarity determination condition, sampling is performed along the boundary of the overlapping area to calculate the spatial depth jump gradient in the preset neighborhood on both sides of the boundary. If the spatial depth gradient exceeds the preset physical thickness threshold, it is determined that there is a physical overlap relationship between the two workpieces, and this relationship is defined as a directed edge. The two workpieces connected by the directed edge are used as nodes to construct a directed acyclic graph describing the physical overlap relationship between the workpieces for the global scene, which serves as a hierarchical topological reasoning model.

4. The method for disordered grasping of symmetrical workpieces based on hierarchical reasoning as described in claim 1, characterized in that, The process of quantitatively evaluating the graspability of the top-level candidate workpieces selected based on off-center load characteristic data, and determining the target workpiece to be grasped based on the evaluation results, includes the following steps: A multidimensional graspability score is constructed for each of the top-level candidate workpieces. The multidimensional graspability score is negatively correlated with the occlusion rate obtained from the off-center load feature data, negatively correlated with the modulus of the geometric off-center load vector obtained from the off-center load feature data, and positively correlated with the normalized height of the top-level candidate workpiece in the current frame depth direction. The multidimensional graspability scores of each of the top-level candidate workpieces are sorted in descending order, and the top-level candidate workpiece with the highest score is selected as the target workpiece to be grasped.

5. The method for disordered grasping of symmetrical workpieces based on hierarchical reasoning as described in claim 1, characterized in that, The process of obtaining the reference gripping pose of the target workpiece and expanding the reference gripping pose into a pose set containing multiple equivalent candidate poses based on the inherent rotational symmetry characteristics of the workpiece includes the following steps: The pre-trained pose estimation network obtains any converged 6D reference pose of the target workpiece in the camera coordinate system as an anchor point. Determine the principal axis of symmetry of the target workpiece to be grasped, and calculate a series of equivalent rotation angles about the principal axis of symmetry based on at least one of the discrete rotational symmetry order of the principal axis of symmetry and the angle step size set for the continuous rotational symmetry body. By sequentially combining the anchor point with the rotation transformation containing each equivalent rotation angle, a set of candidate poses that are completely equivalent in physical grasping effect and contain the pre-grabbing proximity distance parameter are generated, forming a pose set.

6. The method for disordered grasping of symmetrical workpieces based on hierarchical reasoning as described in claim 1, characterized in that, The feasibility verification of each equivalent candidate pose in the set is performed in parallel within the execution space, and a subset of globally safe poses that satisfy all constraints is selected. Specifically, the following parallel verification operation is performed on each equivalent candidate pose: The equivalent candidate poses are solved by inverse kinematics to verify whether there is a joint configuration solution whose joint angles meet the hardware limit of the robotic arm within a preset safety margin. For the existing joint configuration solution, calculate the operability index based on the velocity transmission characteristics in the joint space, and verify whether it is in the flexible working range far away from the kinematic singularity. For equivalent candidate poses, the minimum distance between the robotic arm end-effector and the obstacles in the boundary model and the occupancy map is calculated when the robotic arm end-effector cuts into the pose, based on the digital twin boundary model of the material frame and the real-time three-dimensional occupancy map of the stacked workpieces in the material frame. The minimum distance is then verified to be greater than the preset safe expansion radius. The equivalent candidate pose that simultaneously satisfies the three constraints of having a legal joint configuration solution, being non-singular, and having no collision is determined as an element of the global safe pose subset.

7. The method for disordered grasping of symmetrical workpieces based on hierarchical reasoning as described in claim 1, characterized in that, It also includes exception handling when the global safe pose subset is empty, and the exception handling includes the following steps: Check if the global safe pose subset is empty; When it is determined that the global safe pose subset is empty, stop the output of motion commands for the current capture cycle and generate active perturbation commands; Execute the active disturbance command to control the robotic arm to use the end effector as a probe to perform a preset low-speed lateral tossing action along the path of the maximum gap area indicated in the hierarchical topology inference model, so as to change the rigid physical layout of the workpiece in the current stacking scene. After the toggle action is completed, the current decision queue is cleared, and the multimodal images of the stacked scene are reacquired to start a new grasping decision cycle.

8. A symmetric workpiece unordered grasping system based on hierarchical reasoning, characterized in that, To implement the method of any one of claims 1-7, comprising: The image acquisition and representation generation module is used to acquire multimodal images of the stacked scene, reconstruct the complete physical mask representation of each workpiece from the multimodal images, and generate off-center feature data describing the occlusion state of each workpiece accordingly. The hierarchical topology reasoning module is used to construct a hierarchical topology reasoning model that describes the physical overlapping relationship between workpieces by analyzing the spatial gradient of the overlapping areas of different workpiece masks based on the complete physical mask representation, and to select the top-level candidate workpiece in the current scene based on the hierarchical topology reasoning model. The quantitative evaluation module is used to perform a quantitative evaluation of the graspability of the top-level candidate workpieces selected based on the off-center load characteristic data, and to determine the target workpiece to be grasped based on the evaluation results. The equivalent pose space mapping module is used to obtain the reference grasping pose of the target workpiece to be grasped. Based on the inherent rotational symmetry characteristics of the workpiece, the reference grasping pose is expanded into a pose set containing multiple equivalent candidate poses. The parallel constraint screening module is used to map the pose set from the perception space to the robot execution space, perform feasibility verification on each equivalent candidate pose in the set in parallel in the execution space, and screen out the global safe pose subset that satisfies all constraints. The global optimal instruction generation module is used to select a target pose from the global safe pose subset and output motion control instructions to drive the robotic arm to complete collision-free grasping.

9. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1-7.