Robot double-arm grabbing method for object with double-side handle structure

By using visual sensors and local point cloud analysis, combined with constraint-guided insertion slot search, high-precision grasping and positioning of objects with dual handles and dual-arm collaborative insertion grasping are achieved. This solves the problems of grasping stability and real-time performance in existing technologies and improves the executability and applicability of the method.

CN121870805APending Publication Date: 2026-04-17CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING UNIV
Filing Date
2026-03-17
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing robot grasping methods struggle to match the force distribution with the structural design when handling objects with dual-handle structures. This results in insufficient positioning accuracy, poor grasping feasibility, and high computational overhead for functional grasping methods, making it difficult to meet real-time requirements.

Method used

The robot uses a vision sensor to identify the approximate three-dimensional position of the two handles. Through geometric analysis of local point clouds and constraint-guided insertion slot search, it generates insertion slot poses that meet preset geometric constraints, thereby enabling the robot's two arms to perform cooperative insertion gripping.

Benefits of technology

It improves the stability and reliability of grasping objects with dual handles, reduces the dependence on complete 3D models, and enhances the engineering deployability and real-time performance of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121870805A_ABST
    Figure CN121870805A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of robot control, and discloses a robot double-arm grabbing method for an object with a bilateral handle structure, which comprises the following steps: identifying a target object based on a visual sensor of a robot and detecting rough three-dimensional space positions of bilateral handles of the target object, generating a pre-grabbing pose according to the rough three-dimensional space positions, and grabbing the pre-grabbing pose according to the rough three-dimensional space positions; the two arms of the robot are controlled to move to the pre-grabbing posture; acquiring a local point cloud of the side surface of a target object under the pre-grabbing pose, approximately modeling a robot finger into a cuboid, and executing constraint-guided insertion slot search in a search space defined by the local point cloud so as to find a pose of the cuboid meeting a preset geometric constraint condition as a feasible insertion slot pose; and according to the searched insertion groove poses, robot fingers are controlled to execute insertion motion in the preset direction of the cuboid so as to complete grabbing. The grabbing stability and reliability can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot control technology, and more specifically to a method for a robot to grasp objects with dual handles on both sides. Background Technology

[0002] With the continuous development of robotics technology, its applications in industrial manufacturing, warehousing and logistics, household services, and public services are constantly expanding. Grasping, as a fundamental capability for robots to perform handling, assembly, and collaborative tasks, is a key link in achieving autonomous operation. In recent years, grasping technology based on visual perception and intelligent algorithms has gradually matured, enabling robots to identify target objects in complex environments and complete basic grasping and movement operations. However, when facing large or structurally unique objects, existing grasping technologies still have certain limitations in terms of stability and adaptability.

[0003] Objects with double-sided handles are widely used in daily life, industrial warehousing, and logistics handling scenarios, such as storage boxes, laundry baskets, toolboxes, and various containers. These objects are typically used for storage and transportation, and are frequently handled and moved during use, making them high-frequency objects in actual operating environments. Their structural design generally employs double-sided handles to allow humans to insert their hands into the handles to form symmetrical support for handling, thereby achieving stable force distribution and balanced load-bearing. These objects are structurally well-suited for "internal support" handling methods, rather than simply relying on external gripping. When robots grasp these objects, if external surface gripping or overall lateral clamping is still used, it is often difficult to create a force distribution mode that matches their structural design, especially when the object is large or heavily loaded, which can easily lead to slippage, posture deviation, or even detachment. At the same time, double-sided handles are usually cavities with limited internal space, often close to the size of the robot's end effector, thus placing higher demands on the positioning accuracy and insertion posture accuracy of the handle area than conventional external surface gripping. Therefore, how to effectively identify the handle position and achieve high-precision internally supported dual-arm gripping has become a key technical problem in the operation of this type of object.

[0004] To address the robot grasping problem, existing technologies have been researched from the perspectives of different end effector morphologies and task objectives. In gripper grasping, GraspNet-1 Billion, trained on large-scale data, generates candidate grasping postures, enabling automatic grasping of conventional objects in complex scenes. In dexterous hand grasping, DexGraspAnything, by constructing a multi-finger contact model and combining it with physical constraints, achieves multi-finger grasping of objects of different shapes, improving grasping stability. In functional grasping, SemGrasp associates grasping representations with semantic information, enabling grasping results to respond to task instructions or functional requirements, achieving grasping generation oriented towards functional goals. In dual-arm collaborative grasping, COMBO-Grasp achieves dual-arm collaborative grasping based on a constraint-guided learning framework; BimanGrasp-DDPM generates dual-dexterous hand collaborative grasping postures through a diffusion model. These methods improve the robot's autonomous grasping capabilities from different angles.

[0005] However, the aforementioned methods still have significant shortcomings when handling objects with dual-handle structures. First, most existing grasping methods rely primarily on contact and clamping between the object's external surface. Their grasping mechanism depends on the force exerted by external friction, failing to fully consider the structural characteristics of dual-handle objects, which achieve stable transport through internal support. This makes it difficult to meet the special requirements of such objects under high loads. Second, some dual-arm grasping methods are mainly trained and validated in simulation environments, typically using a complete 3D model as input. In real-world scenarios, only single-view or local observation information is often available, limiting the practical deployment of these methods. Third, some research on functional grasping relies on large-scale pre-trained models or generative networks for semantic reasoning and pose synthesis. The reasoning process is computationally intensive and time-consuming, hindering real-time operation and exhibiting problems such as low interpretability and severe illusions. Finally, some existing methods focus on generating the final grasping pose without adequately considering accessibility and collision risks during the insertion process, potentially leading to unsuccessful execution of the generated grasping result in actual operation. Summary of the Invention

[0006] To address the aforementioned shortcomings in existing technologies, this invention provides a robotic dual-arm grasping method for objects with dual-handle structures. This method solves the technical problems of existing robotic grasping methods when handling objects with dual-handle structures, such as mismatch between force distribution and structural design, insufficient accuracy in handle positioning, poor grasping feasibility under real-world observation conditions, and high computational overhead and difficulty in meeting real-time requirements for some functional grasping methods. The invention achieves high-precision grasping and positioning of the handle area of ​​objects with dual-handle structures and collaborative dual-arm insertion-type grasping, improving the stability and reliability of grasping in heavy-duty handling scenarios. Furthermore, by using geometric analysis methods on local point clouds, the method reduces reliance on the complete 3D model of the object and reasoning from highly complex models, enhancing its engineering deployability and practical application value.

[0007] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows: A method for robotic dual-arm grasping of objects with dual-handle structures includes the following steps: The robot uses a visual sensor to identify the target object and detect the approximate three-dimensional spatial position of its two handles. Based on the approximate three-dimensional spatial position, a pre-grasping pose is generated, and the robot's two arms are controlled to move to the pre-grasping pose. In the pre-grasping pose, a local point cloud of the side of the target object is acquired. The robot finger is approximated as a cuboid, and a constraint-guided insertion slot search is performed in the search space defined by the local point cloud to find the pose of the cuboid that satisfies the preset geometric constraints, which is taken as a feasible insertion slot pose. The preset geometric constraints include at least an internal cavity constraint and an upper support continuity constraint. The internal cavity constraint means that the space occupied by the cuboid does not contain any points in the local point cloud. The upper support continuity constraint means that above the top of the cuboid and on the side facing the object, there are multiple sub-segments along the horizontal division, and there are point cloud points with a distance less than a first threshold from the corresponding sub-segments. Based on the found insertion slot pose, the robot finger is controlled to perform an insertion motion along the preset direction of the cuboid to complete the grasping.

[0008] Preferably, the step of identifying the target object and detecting the approximate three-dimensional spatial position of its two handles includes: Based on the open vocabulary target detection model, the target object is detected according to the first text prompt, and the candidate handle region is detected according to the second text prompt; Candidate handle regions that overlap with the target object region above the overlap threshold and are located on both sides of the center of the target object in the horizontal direction are selected as the approximate positions of the dual handles.

[0009] Preferably, if the open-vocabulary target detection fails to detect both handles, the method further includes: Principal component analysis is performed on the point cloud of the target object to obtain the principal axis of its directed bounding box; Based on the center and longest axis direction of the oriented bounding box, the approximate positions of the dual handles are inferred.

[0010] Preferably, before performing the constraint-guided insertion slot search, the method further includes: The local point cloud is voxelized to generate a three-dimensional voxel-occupied grid. Based on the three-dimensional voxel occupancy grid, a three-dimensional integral volume is constructed to accelerate the query of spatial region occupancy status.

[0011] Preferably, the preset geometric constraints further include position validity constraints; The position validity constraint means that the spatial position of the cuboid must be located in the middle region of the local point cloud bounding box and avoid its bottom region.

[0012] Preferably, the preset geometric constraints further include forward insertion space constraints; The forward insertion space constraint means that a space of a predetermined length is reserved in front of the cuboid along its insertion direction, which does not contain points from the local point cloud.

[0013] Preferably, the constraint-guided insertion slot search includes sampling the rotational orientation of the cuboid; For each sampled rotational posture, the local point cloud is transformed to the local coordinate system corresponding to that rotational posture to determine the geometric constraints.

[0014] Preferably, the constraint-guided insertion slot search includes: Prioritize searching for feasible insertion slot poses in a zero-rotation attitude; If there is no feasible solution under the zero-rotation attitude, then rotation sampling and searching are performed within a small angle range centered on the zero-rotation attitude.

[0015] Preferably, the constraint-guided insertion slot search employs a batch parallel evaluation strategy, including: Based on the voxelization process, a set of all candidate insertion slot center positions is generated; The geometric constraints described above are constructed as vectorized decision functions, and these vectorized decision functions are applied in batches to the current set of candidate center positions in order of increasing computational complexity to gradually reduce the feasible candidate set.

[0016] Preferably, controlling the robot finger to perform the insertion motion includes: The robot's arms move synchronously, so that the fingers of both hands are inserted into the insertion slots corresponding to the handles on both sides of the target object, until a preset depth is reached and then a grasping action is performed.

[0017] The present invention has the following beneficial effects: 1) This invention addresses the internal support-type force characteristics of objects with double-handle structures by proposing a dual-arm cooperative insertion gripping strategy. This strategy matches the gripping method with the object's structural design, enabling the formation of a stable load-bearing structure during the handling of large, heavy-duty objects and significantly improving gripping stability and reliability. 2) This invention improves the positioning accuracy and insertion posture accuracy of the handle area by combining phased visual perception with fine grasping decision-making. It can achieve fine grasping even with limited space inside the handle cavity, and enhances the adaptability to high-precision handle grasping tasks. 3) This invention can complete the grasping decision without relying on a complete 3D model, reducing the dependence on high-precision global reconstruction data and improving the executability and applicability of the method in real-world environments; 4) By constructing a structured crawling decision process, this invention reduces the dependence on high-complexity model reasoning and large-scale training data, reduces computational overhead, and improves the real-time performance, interpretability, and engineering deployability of the method, thus having high practical application value. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the framework of the present invention; Figure 2 This is a schematic diagram of a robot dual-arm grasping method for an object with a dual-handle structure according to the present invention. Figure 3 Flowchart for coarse localization of the handle region using object detection and segmentation models; Figure 4 This is a schematic diagram of the various constraints; Figure 5 Example image of the generated result. Detailed Implementation

[0019] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0020] This invention proposes a robotic dual-arm insertion-type grasping method for objects with dual-handle structures. This method combines visual perception with a structured decision-making mechanism to achieve high-precision localization and stable grasping of the grasping area. The overall scheme employs a two-stage processing framework: initial localization of the handle area is achieved based on long-distance scene perception; then, fine analysis and grasping pose determination are performed on the handle area under close-range observation conditions, thereby achieving stable dual-arm collaborative insertion-type grasping. This method improves grasping accuracy and execution reliability without relying on a complete 3D model or large-scale training data. The framework structure of this invention is shown in the attached figure. Figure 1 As shown, it mainly includes the following two key stages and steps: ① Coarse localization of the handle area: This stage uses the depth camera on the robot's head to identify and segment the target object, and further detects the area where the handles are located on both sides, calculates the approximate three-dimensional position of the handles, and moves the robot's two arms to the corresponding pre-grasping position, providing a reasonable observation angle and initial pose conditions for subsequent fine grasping.

[0021] ② Precise localization of the handle area and generation of grasping pose: Based on the close-range depth observation information of the robot's hand, this stage performs a detailed analysis of the handle area and introduces a constraint-guided slot search algorithm (CG-SS). By constructing a feasible grasping area search mechanism under spatial geometric constraints, an insertion-type grasping pose that meets multiple requirements such as the cavity of the grasping area and the continuity of the supporting surface is generated. The dual-arm grasping action is then executed to achieve stable lifting of the object with both handles.

[0022] The relevant concepts and principles will be introduced below.

[0023] In the field of robotics, a relatively systematic technological framework has been developed for the problem of autonomous object grasping. Among these, RGB-D depth imaging, open-vocabulary visual perception, 3D point cloud processing, and voxel-based spatial representation are key foundational technologies in current research and engineering practice.

[0024] In visual perception, RGB-D imaging technology has been widely used in robot environment modeling. An RGB-D camera simultaneously outputs a color image and pixel-by-pixel depth information. Let the color image be represented as... ,in and These represent the image height and width, respectively; the depth map is represented as... ,in Represents pixel coordinates The depth value at that location. The camera intrinsic parameter matrix is: in For focal length parameters, The principal point coordinates are used. The pixel coordinates and depth values ​​can be back-projected into 3D space using the camera intrinsic parameter matrix. For any pixel... Its corresponding three-dimensional point It can be represented as: in This represents the backprojection operator. Backprojecting all pixels within an image region constructs a 3D point cloud set. in This represents the effective depth pixel count. Point clouds, as a discrete 3D geometric representation, are an important data format for robot spatial understanding and grasping analysis.

[0025] In recent years, open-vocabulary object detection techniques have been extensively studied. These methods are typically based on vision-language pre-trained models, capable of detecting semantic regions in images based on textual cues. Grounding DINO is a representative method of this type. The detection results are usually represented as a set of two-dimensional bounding boxes: in Indicates the first Candidate regions, To determine the number of detections. To further obtain pixel-level fine boundaries, an instance segmentation model is typically used to generate a mask set: in For the first A set of pixels for each instance. SAM (Segment Anything Model) is a representative method of this type of technique. By performing 3D backprojection on the pixels in the masked region, a semantically constrained subset of the point cloud can be obtained: This semantically guided target detection mechanism enables robots to locate specific targets in complex scenes.

[0026] In 3D geometric analysis, point clouds typically require pose and principal orientation estimation. Principal Component Analysis (PCA) is a commonly used method for geometric feature extraction. Let the mean of the point cloud be: The covariance matrix is ​​then: Perform eigenvalue decomposition on the covariance matrix: in For eigenvalues, This corresponds to the feature vector. The direction corresponding to the largest eigenvalue represents the principal axis direction of the point cloud distribution. Based on this principal axis, an Oriented Bounding Box (OBB) can be constructed, with its center denoted as . The axis direction is a unit vector. The side lengths are respectively OBBs play an important role in spatial geometric alignment.

[0027] In spatial representation and accelerated computation, voxelization is a common discretization technique. Let's consider a three-dimensional spatial region... Divide it into sections with side length of For a cube unit (voxel), the voxel index is: in These are discrete quantities in three directions. Define the occupancy function: but This forms a three-dimensional occupancy grid. To improve the efficiency of region queries, a three-dimensional integral volume can be further constructed: Using an integral volume, the number of points within an arbitrary axis-aligned cuboid region can be calculated in constant time, thus quickly determining whether the space is empty.

[0028] Reference Figure 2 The present invention provides a method for a robot to grasp an object with a double-handled structure using a dual-arm gripper, comprising the following steps S1 to S3: S1. Based on the robot's visual sensor, identify the target object and detect the rough three-dimensional spatial position of its two handles. Generate a pre-grasping pose based on the rough three-dimensional spatial position and control the robot's two arms to move to the pre-grasping pose. In an optional embodiment of the present invention, step S1, identifying the target object and detecting the approximate three-dimensional spatial position of its two handles, includes: Based on the open vocabulary object detection model, target objects are detected according to the first text prompt, and candidate handle regions are detected according to the second text prompt; Candidate handle regions that overlap with the target object region above the overlap threshold and are located on both sides of the target object's center in the horizontal direction are selected as the approximate positions of the two handles.

[0029] If the open-vocabulary target detection fails to detect both handles, then it also includes: Principal component analysis is performed on the point cloud of the target object to obtain the principal axes of its directed bounding box; Based on the center and longest axis direction of the oriented bounding box, the approximate positions of the handles on both sides are inferred.

[0030] Step S1 aims to estimate the approximate three-dimensional spatial position of the handles on both sides of the target object, guiding the hands to a symmetrical pre-insertion pose. Given an RGB image... and its aligned depth map First, the Grounding DINO method, combined with text prompts, is used to detect the target object, resulting in a two-dimensional object bounding box. Subsequently, the SAM model is used to process the detection results and generate the corresponding object segmentation mask. Based on camera intrinsic parameters, pixels within the masked area are back-projected to construct a 3D point cloud of the target object. in, This represents the back projection operator.

[0031] Define the object's bounding box The center is Its width and height are respectively and To obtain candidate handle regions, the text prompt "handle" is used again for object detection via the Grounding DINO model, resulting in a set of candidate handle bounding boxes. When candidate bounding boxes bounding box of the target object Only when the overlap at the pixel level is sufficiently large is it considered to belong to the target object. Where the threshold .set up The two-dimensional center is The corresponding candidate handle position Defined as to The centroid of the 3D point cloud obtained by back-projecting the depth values ​​within the region.

[0032] If exactly two candidate regions remain after overlapping filtering, it is further determined whether they correspond to the left and right handles of the object. Specifically, the centers of the two handles must satisfy the condition that they are located on either side of the center of the object: Ensure that one handle is located to the left of the object's center and the other to the right. Furthermore, each handle must be sufficiently offset horizontally from the object's center to avoid being located in the central area. in The width of the object's bounding box. When all of the above conditions are met, As a rough spatial position of the handle. This process is as follows: Figure 3 As shown.

[0033] When the detection fails (i.e., the above conditions are not met), this step will utilize point clouds. The geometric structure is used to infer the handle position. Specifically, principal component analysis is performed on the point cloud to estimate its principal axis orientation and construct a directed minimum bounding box (OBB). Considering that containers with handles are typically placed in an approximately horizontal orientation, an axis is only accepted as a valid direction if the azimuth deviation between the longest axis of the OBB and the horizontal axis of the table is less than a given threshold, ensuring the physical plausibility of the geometric inference. Let the center of the OBB be... The unit vector along the longest effective axis is The corresponding length is The positions of the two handles deduced by this method are defined as follows: In obtaining Then (regardless of the method used to obtain it), it can be extended at both ends along the direction connecting the two handles. This generates a symmetrical pre-insertion position and positions the wrist camera toward the midpoint of the line connecting them, providing a suitable observation angle for the subsequent precise positioning stage.

[0034] S2. Under the pre-grasping pose, acquire the local point cloud of the side of the target object, approximate the robot finger as a cuboid, and perform constraint-guided slot search in the search space defined by the local point cloud to find the pose of the cuboid that satisfies the preset geometric constraints, which is taken as a feasible slot pose; wherein, the preset geometric constraints include at least internal cavity constraints and upper support continuity constraints; the internal cavity constraint means that the space occupied by the cuboid does not contain points in the local point cloud; the upper support continuity constraint means that above the top of the cuboid and on the side facing the object, there are point cloud points above the multiple sub-segments divided laterally, and the distance between the points and the corresponding sub-segments is less than a first threshold. In an optional embodiment of the invention, after the hand reaches the pre-grasping position in step S2, the remaining task is to determine a feasible finger insertion cavity. This step approximates each group of fingers (index, middle, ring, and little fingers) as a flattened cuboid, and formalizes the problem as: finding a directed cuboid in the side-view local point cloud that satisfies geometric constraints, where the cuboid simultaneously represents the finger's size, insertion direction, and insertion depth. This formalized problem definition is based on the five-finger dexterous hand structure, but is not limited to it; it can also be applied to other robotic end effectors with similar structures, such as a two-finger gripper (where one finger is considered a cuboid structure).

[0035] With the hand roughly aligned with one of the handles, a side-view RGB-D observation is acquired using a hand camera, and a compact 2D bounding box corresponding to the object is obtained on the visible side surface of the object using Grounding DINO. This bounding box is then combined with a depth map, and points whose depth values ​​deviate significantly from the handle depth are removed, resulting in a filtered point cloud. The point cloud primarily contains information about the visible side surface of the object, while effectively suppressing interference from background structures and the handle on the other side. During the hand's movement to the pre-grabbing position, the wrist is controlled to a specific posture, ensuring the image plane remains vertical and the optical axis is aligned with the side of the container. The resulting side-view object coordinate system is aligned with the camera coordinate system. The axis points to the right of the image. The axis points to the bottom of the image. The axis points from the camera to the object along the line of sight.

[0036] In an optional embodiment of the invention, step S2 further includes, before performing the constraint-guided insertion slot search: The local point cloud is voxelized to generate a three-dimensional voxel-occupied grid. Based on the three-dimensional voxel occupancy grid, a three-dimensional integral volume is constructed to accelerate the query of spatial region occupancy status.

[0037] This embodiment demonstrates the feasibility of efficiently detecting collision-free finger insertion during insertion, using side-view point clouds. Voxelization is performed, dividing the space covered by the point cloud into regular three-dimensional grid units (voxels), and modeling the space occupancy using voxels as the smallest unit, so that the point cloud data is converted into a structured raster representation, thereby reducing the computational complexity of subsequent geometric queries and collision detection.

[0038] set up Represents side-view point clouds The spatial bounding volume, that is, the smallest axis-aligned 3D region that can completely cover the point cloud. According to fixed voxel size By performing uniform discretization, a resolution of [resolution value] can be obtained. A three-dimensional voxel mesh, in which They represent in , and The number of voxels in a direction. Each voxel is represented by its index. A unique identifier, corresponding to a side length of The coordinates of the smallest corner point of a cube are defined as follows: in Represents the bounding volume of space The minimum coordinates in three-dimensional space. This is achieved by processing point clouds. All points in the array are mapped to their corresponding voxel indices to construct a binary occupancy raster. This is used to indicate whether each voxel contains at least one observation point. The occupancy grid geometrically characterizes the occupancy distribution of the object's surface and the space around it.

[0039] Based on this, further from occupying the grid Pre-calculated three-dimensional integral volume This allows the number of voxels within an arbitrary axis-aligned cuboid region to be calculated in constant time using a finite number of table lookups and addition / subtraction operations. Using the integral volume representation, it is possible to efficiently determine whether a candidate insertion region spatially overlaps with the observed point cloud, thus providing computational efficiency assurance for subsequent large-scale geometric feasibility searches. For simplicity, the voxelized representation is uniformly denoted as... ,in Used to store the boundary information of voxel meshes in three-dimensional space.

[0040] The potential finger insertion cavity is geometrically modeled as a regular cuboid with dimensional parameters... These correspond to the lateral opening width, longitudinal thickness, and insertion depth required for finger insertion, respectively. This cuboid approximates the volume occupied by a set of fingers in space, thus transforming the insertion feasibility problem into a geometric constraint determination problem in three-dimensional space.

[0041] In parametric representation, the pose of the cuboid is determined by its center position. and rotation matrix Uniquely determined. Among them, This indicates the spatial center position of the cuboid in the camera coordinate system. This describes the orientation of the local coordinate system of the cuboid relative to the camera coordinate system, used to characterize the small-angle tilt or rotation that the finger may have. Let represent a special orthogonal group of three-dimensional rotation matrices, satisfying the constraints of orthogonality and determinant of 1.

[0042] In the local coordinate system of a cuboid, assuming it is aligned with the coordinate axes, the set of points inside the cuboid can be defined as follows: in, Let represent any point in the local coordinate system. The above inequality constraints respectively limit the range of values ​​of this point in the three coordinate axes, thus defining a coordinate system centered at the origin with side lengths _____. , and An axis-aligned cuboid.

[0043] After mapping this local representation to the camera coordinate system, the corresponding cuboid region can be represented as: in, This represents a point in the camera coordinate system. The mapping process begins with a rotation matrix. Points in the local coordinate system Rotate to the camera coordinate system orientation, then use a translation vector. Move it to the corresponding spatial location. Thus, This precisely describes the volume range occupied by the finger in three-dimensional space during insertion, given its position and orientation. All subsequent judgments regarding geometric constraints are based on the occupancy relationship of this cuboid model in voxel space.

[0044] In an optional embodiment of the present invention, the preset geometric constraints set in step S2 further include position validity constraints; The position validity constraint means that the spatial position of the cuboid must be located in the middle region of the local point cloud bounding box and avoid its bottom region.

[0045] The preset geometric constraints set in step S2 also include forward insertion space constraints; The forward insertion space constraint means that a space of a predetermined length is reserved in front of the cuboid along its insertion direction, which does not contain points from the local point cloud.

[0046] This embodiment ensures that the finger follows the local coordinate system of the cuboid. Insertion along the axial direction is geometrically feasible and functionally meaningful; the parameters corresponding to the candidate insertion slots are relevant. The following geometric constraints must be satisfied simultaneously.

[0047] The first constraint is a position validity constraint, which requires that candidate cuboids should avoid the bottom region and the extreme edge regions on the left and right sides of the object. The bottom region of the object usually contacts the supporting plane (such as a tabletop), which would physically obstruct finger insertion and is therefore considered an infeasible region. The extreme edge regions on the left and right sides of the object often have missing or incomplete point clouds due to limited viewing angles, resulting in lower reliability of their geometric information. Furthermore, in practical designs, handles are usually located in the middle region of the object's side rather than the outermost edge; therefore, excluding these extreme regions helps reduce the number of erroneous candidates and improve search efficiency.

[0048] set up as well as These represent side-view point clouds. exist shaft and Minimum and maximum coordinate values ​​along the axis. Define height threshold: Used to exclude the lowest point in the point cloud. A proportional area is defined to avoid candidate insertion slots being located near the bottom of the object. Simultaneously, the inner lateral boundary is defined. To exclude the left and right sides each occupy Extreme regions of proportion.

[0049] Let the candidate cuboid be along and The axis alignment bounding intervals in the axial direction are respectively and Then the position validity constraint can be formalized as: The above inequalities together ensure that the candidate cuboid is located in the middle region of the object's side surface, away from the bottom support plane and the left and right edges.

[0050] The second constraint is the internal cavity constraint, which requires that the entire candidate cuboid must not contain any observed point cloud data, that is, the space occupied by the cuboid should be completely empty: This constraint ensures that there is no direct collision with the object's surface or other structures within the target area where the finger is inserted. In the voxelized discrete representation, this condition is achieved by detecting candidate cuboids. Is all the voxels covered implemented as empty voxels? If these voxels occupy a grid... If the sum of the occupancy values ​​is zero, the cuboid is considered empty. Using a three-dimensional integral volume representation, this judgment can be completed in constant time, significantly improving search efficiency.

[0051] The third constraint is the continuity constraint of the upper support. This constraint requires that the candidate insertion slot must be located directly below a continuous support surface; otherwise, after insertion, the finger will bend upwards and enter the cavity space, failing to form effective contact with the inner surface of the handle, thus causing gripping failure. Let... This represents the horizontal point cloud region at the top of the candidate cuboid, facing the camera side. This point cloud region is then uniformly divided along the horizontal direction (i.e., the x-axis direction) into... Each sub-segment, denoted as the first... Each section is From the observed point cloud, located in Any point above is denoted as And located in the first The points above each segment are denoted as Set thresholds separately. and It is used to control the tolerance range of close contact determination and surface continuity determination.

[0052] One of the necessary conditions for accepting a candidate insertion slot is that... There exists at least one sufficiently close point in the point cloud above, satisfying the following: This condition ensures that once the finger is fully inserted, its base immediately contacts the inner surface of the handle, rather than remaining suspended in the air.

[0053] In addition, the continuity of the upper supporting surface in the lateral direction must also be met. Specifically, only when Each sub-segment The condition is considered true only if there is at least one sufficiently close point in the point cloud above: This constraint ensures that the support surface above the insertion slot remains continuous throughout the entire lateral range, preventing breaks or hollow structures, thereby providing stable and uniform force support for the finger.

[0054] The fourth constraint is the forward insertion space constraint, used to ensure sufficient forward collision-free space when the finger moves along the expected insertion direction. Let... Let represent the depth position of the candidate cuboid's front face in the camera coordinate system. Then, a length of is reserved before this front face along the insertion direction. The spatial region is defined as: This area must not contain any data from side-view point clouds. This constraint ensures that the finger does not collide with the outer surface of the container or other structures in the path as it approaches and inserts into the handle cavity, thus guaranteeing the feasibility of the insertion action.

[0055] The overall schematic diagram of the above four geometric constraints is attached. Figure 4 As shown.

[0056] In an optional embodiment of the present invention, step S2 performs a constraint-guided insertion slot search, including sampling the rotational orientation of the cuboid. For each sampled rotational posture, the local point cloud is transformed to the local coordinate system corresponding to that rotational posture to determine the geometric constraints.

[0057] Step S2 performs constraint-guided insertion slot search, including: Prioritize searching for feasible insertion slot poses in a zero-rotation attitude; If there is no feasible solution under the zero-rotation attitude, then rotation sampling and searching are performed within a small angle range centered on the zero-rotation attitude.

[0058] Step S2 performs constraint-guided slot search using a batch parallel evaluation strategy, including: Based on the voxelization process, a set of all candidate insertion slot center positions is generated; The geometric constraints described above are constructed as vectorized decision functions, and these vectorized decision functions are applied in batches to the current set of candidate center positions in order of increasing computational complexity to gradually reduce the feasible candidate set.

[0059] In real-world scenarios, due to slight tilting of the object's orientation, the handle's geometry not being strictly aligned with the camera coordinate axes, and unavoidable pose errors when the robotic arm reaches the pre-grasping pose, the space occupied by the finger during the insertion action is not necessarily strictly parallel to the coordinate axes in the camera coordinate system. If only axis-aligned cuboid models are considered, geometrically feasible but slightly tilted valid insertion slots may be missed. Therefore, this embodiment explicitly introduces rotational degrees of freedom during the search process to improve the insertion slot search's adaptability to the uncertainties of real-world scenarios.

[0060] Based on the above considerations, this step applies a rotational perturbation to the cuboid within a limited range. Rotation matrix The search is restricted to the vicinity of the attitude aligned with the camera coordinate system, allowing only small-angle yaws on selected rotation axes. Rotation is performed in steps along the roll, yaw, and pitch directions (rotation in a specific direction can also be cancelled for certain cases). In the interval Discrete sampling within the range can generate a finite set of rotational candidates. This constrained rotation search can cover reasonable deflections caused by attitude errors and handle tilt, while avoiding the computational overhead of high-dimensional continuous rotation spaces.

[0061] For each rotation candidate The system first transforms the side-view point cloud into the corresponding local coordinate system, reconstructs the voxel mesh and 3D integral volume in this coordinate system, and models the insertion slot candidates as axis-aligned cuboids. Then, it performs constraint condition checks. If a candidate passes all constraints under this rotational posture, its corresponding cuboid parameters are processed using a rotation matrix. Mapping back to the camera coordinate system yields the final insertion slot.

[0062] Based on the above constraints and rotation settings, a constraint-guided insertion slot search algorithm CG-SS can be constructed to search for reasonable insertion hand poses in the point cloud space. The feasibility determination of each candidate cuboid depends only on the local geometric information in the discrete voxel mesh and is independent of other candidates, making it a typical parallel computation problem. Therefore, unlike testing each candidate one by one, this method performs batch evaluation of all candidates under each rotation: a complete candidate center set is constructed from the voxel mesh, and each geometric constraint is used as a vectorized decision function acting on this set for batch parallel computation. Each constraint is evaluated on the current candidate set, generating a new subset of candidates that satisfy the constraint. Subsequently, the next constraint is only filtered on this updated set. Since different constraints have different computational costs, this embodiment prioritizes applying constraints with lower computational costs and stronger filtering capabilities to eliminate a large number of candidate positions that do not meet the conditions as early as possible, thereby reducing the evaluation scale of subsequent complex constraints. Constraints are applied sequentially in this order, and the candidate set is gradually shrunk until the final feasible set is obtained. If multiple feasible solutions exist, the final slot is selected according to deterministic rules (e.g., selecting the position closest to the coarse positioning handle estimate). A non-zero rotation search is only performed when no feasible solution exists in the zero rotation case.

[0063] S3. Based on the searched insertion slot position, control the robot finger to perform insertion movement along the preset direction of the cuboid to complete the grasping.

[0064] In an optional embodiment of the present invention, step S3, controlling the robot finger to perform an insertion motion, includes: The robot's arms move synchronously, so that the fingers of both hands are inserted into the insertion slots corresponding to the handles on both sides of the target object, until a preset depth is reached and then a grasping action is performed.

[0065] Figure 5Examples of generation using the CG-SS algorithm on different objects are given. Once the optimal insertion slot is determined... The hand will move along the local coordinate system of the cuboid. The machine performs a linear insertion motion along the axis until it reaches the preset insertion depth; then, the finger joints rotate and close, forming a stable contact with the inner surface of the handle, thereby completing the grasping action.

[0066] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0067] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0068] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0069] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

[0070] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.

Claims

1. A method for robot dual-arm grasping of objects with dual-handle structures, characterized in that, Includes the following steps: The robot uses a vision sensor to identify the target object and detect the approximate three-dimensional spatial position of its two handles. Based on the approximate three-dimensional spatial position, a pre-grasping pose is generated, and the robot's two arms are controlled to move to the pre-grasping pose. In the pre-grasping pose, a local point cloud of the side of the target object is acquired. The robot finger is approximated as a cuboid, and a constraint-guided insertion slot search is performed in the search space defined by the local point cloud to find the pose of the cuboid that satisfies the preset geometric constraints, which is taken as a feasible insertion slot pose. The preset geometric constraints include at least an internal cavity constraint and an upper support continuity constraint. The internal cavity constraint means that the space occupied by the cuboid does not contain any points in the local point cloud. The upper support continuity constraint means that above the top of the cuboid and on the side facing the object, there are multiple sub-segments along the horizontal division, and there are point cloud points with a distance less than a first threshold from the corresponding sub-segments. Based on the found insertion slot pose, the robot finger is controlled to perform an insertion motion along the preset direction of the cuboid to complete the grasping.

2. The method for robot dual-arm grasping of an object with a dual-handle structure according to claim 1, characterized in that, The process of identifying the target object and detecting the approximate three-dimensional spatial position of its two handles includes: Based on the open vocabulary target detection model, the target object is detected according to the first text prompt, and the candidate handle region is detected according to the second text prompt; Candidate handle regions that overlap with the target object region above the overlap threshold and are located on both sides of the center of the target object in the horizontal direction are selected as the approximate positions of the dual handles.

3. A method for robot dual-arm grasping of an object with a dual-handle structure according to claim 2, characterized in that, If the open-vocabulary target detection fails to detect both handles, then it also includes: Principal component analysis is performed on the point cloud of the target object to obtain the principal axis of its directed bounding box; Based on the center and longest axis direction of the oriented bounding box, the approximate positions of the dual handles are inferred.

4. A method for robot dual-arm grasping of an object with a dual-handle structure according to claim 1, characterized in that, Prior to performing the constraint-guided slot search, the following is also included: The local point cloud is voxelized to generate a three-dimensional voxel-occupied grid. Based on the three-dimensional voxel occupancy grid, a three-dimensional integral volume is constructed to accelerate the query of spatial region occupancy status.

5. A method for robot dual-arm grasping of an object with a dual-handle structure according to claim 1, characterized in that, The preset geometric constraints also include position validity constraints; The position validity constraint means that the spatial position of the cuboid must be located in the middle region of the local point cloud bounding box and avoid its bottom region.

6. A method for robot dual-arm grasping of an object with a dual-handle structure according to claim 1, characterized in that, The preset geometric constraints also include forward insertion space constraints; The forward insertion space constraint means that a space of a predetermined length is reserved in front of the cuboid along its insertion direction, which does not contain points from the local point cloud.

7. A method for robot dual-arm grasping of an object with a dual-handle structure according to claim 1, characterized in that, The constraint-guided insertion slot search includes sampling the rotational orientation of the cuboid; For each sampled rotational posture, the local point cloud is transformed to the local coordinate system corresponding to that rotational posture to determine the geometric constraints.

8. A method for robot dual-arm grasping of an object with a dual-handle structure according to claim 1, characterized in that, The constraint-guided slot search includes: Prioritize searching for feasible insertion slot poses in a zero-rotation attitude; If there is no feasible solution under the zero-rotation attitude, then rotation sampling and searching are performed within a small angle range centered on the zero-rotation attitude.

9. A method for robot dual-arm grasping of an object with a dual-handle structure according to claim 4, characterized in that, The constraint-guided slot search employs a batch parallel evaluation strategy, including: Based on the voxelization process, a set of all candidate insertion slot center positions is generated; The geometric constraints described above are constructed as vectorized decision functions, and these vectorized decision functions are applied in batches to the current set of candidate center positions in order of increasing computational complexity to gradually reduce the feasible candidate set.

10. A method for robot dual-arm grasping of an object with a dual-handle structure according to claim 1, characterized in that, The control of the robot's finger to perform the insertion movement includes: The robot's arms move synchronously, so that the fingers of both hands are inserted into the insertion slots corresponding to the handles on both sides of the target object, until a preset depth is reached and then a grasping action is performed.