Material grabbing method, device, equipment and program product

By combining top and wrist cameras and utilizing visual models and multimodal data processing, the motion sequence is adjusted in real time, solving the problem of grasping failure in disordered stacked material scenarios and achieving efficient dynamic environment adaptation and improved grasping success rate.

CN121515249APending Publication Date: 2026-02-13杭州普联系统技术有限公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511947552.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing technologies are prone to failure when grasping materials in disordered stacked material scenarios, making it difficult to adapt to dynamically changing scenarios. Furthermore, they are highly dependent on 3D point cloud matching, leading to matching failures and grasping interference.

Method used

The method combines a top camera and a wrist camera. The top camera captures the first image and uses the first visual model to determine the target grasping area. The wrist camera captures the second image and uses the second visual model to determine the action sequence and grasping pose. This avoids the use of 3D point cloud matching and uses multimodal data and a transformer encoder for feature capture and denoising, and adjusts the action sequence in real time.

Benefits of technology

It improves the reliability and success rate of grasping in disordered stacking scenarios, can adapt to dynamically changing scenarios, reduces the performance requirements of depth cameras, reduces grasping failures, and increases the probability of grasping success.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121515249A_ABST
    Figure CN121515249A_ABST
Patent Text Reader

Abstract

The invention relates to the field of industrial automation, in particular to a material grabbing method, device and equipment and a program product. The method comprises the following steps: collecting a first image comprising a material through a top camera; determining a target capture area of the first image through a preset first visual model; the mechanical arm is controlled to move to the target grabbing area, and a second image including the target grabbing area is collected through the wrist camera; determining an action sequence and a grabbing pose of the mechanical arm for grabbing the material in the second image through a preset second visual model; and the materials are grabbed according to the action sequence and the grabbing poses. According to the method, the reliability in a disordered stacking scene can be improved, even if the materials move in the grabbing process, the output action sequence can be adjusted in time through timely updated observation of the wrist camera, the grabbing success probability can be improved, and the method can effectively adapt to grabbing operation in a dynamic change scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of industrial automation, and in particular to methods, apparatus, equipment and program products for grasping materials. Background Technology

[0002] In industrial automation, using robots to perform tasks such as sorting and grasping is a common operation. When performing grasping tasks, the disordered stacking of materials creates a diverse and complex state space, making it difficult to complete using traditional pre-coding grasping methods. Introducing 3D vision information into grasping methods is considered an effective way to solve this problem and has been widely researched and practiced in recent years. By using 3D point cloud matching to obtain the target grasping pose, the grasping process is completed.

[0003] However, due to the stacking of materials, point cloud information is easily incomplete, and point cloud matching is prone to failure. If only local point clouds are used to select targets and determine grasping poses, grasping may fail due to interference from surrounding materials in disordered stacking scenarios. This makes it difficult to adapt to dynamically changing scenarios. Summary of the Invention

[0004] In view of this, embodiments of this application provide methods, apparatus, devices and program products to solve the problem that material grabbing in the prior art is prone to failure due to material stacking and is difficult to adapt to dynamically changing scenarios.

[0005] A first aspect of this application provides a method for grasping materials. The grasping device for grasping materials includes a robotic arm, a top camera, and a wrist camera. The top camera is located directly above the material to be grasped, and the wrist camera is located at the end of the robotic arm. The method includes: A first image including the material is captured by the top camera; The target grasping area of ​​the first image is determined by a preset first visual model. The robotic arm is controlled to move to the target grasping area, and a second image including the target grasping area is acquired through the wrist camera; The robotic arm's action sequence and grasping posture for grasping materials in the second image are determined by a preset second visual model. The material is grasped according to the action sequence and the grasping pose.

[0006] In conjunction with the first aspect, in the first possible implementation of the first aspect, the first visual model is a first visual language model; The target grasping region of the first image is determined by a preset first visual model, including: The first image and natural language instructions are input into the first visual language model, and a two-dimensional heat map is output through the first visual language model. The target mask is determined based on the pixels in the heat map whose heat values ​​are greater than a preset heat threshold; Based on the internal and external parameters of the top camera and the depth image, the target point cloud corresponding to the target mask is determined, and the target grasping area is determined based on the target point cloud.

[0007] In conjunction with the first aspect, in a second possible implementation of the first aspect, the action sequence and grasping pose of the robotic arm grasping the material in the second image are determined by a preset second visual model, including: The multimodal data is encoded using a preset encoding model to extract multimodal features. The multimodal data includes the second image, natural language commands, robotic arm status, and point cloud data acquired by the wrist camera. The transformer encoder captures the intrinsic correlations between features of different modes and outputs observed features. The observed features, the noisy action sequence, and the noisy target grasping pose are input into the diffusion transformer to obtain the denoised action sequence and the denoised grasping pose.

[0008] In conjunction with the first aspect, in a third possible implementation of the first aspect, the action sequence and grasping pose of the robotic arm grasping the material in the second image are determined by a preset second visual model, including: The multimodal data is encoded using a preset encoding model to extract multimodal features. The multimodal data includes the second image, natural language commands, robotic arm status, and point cloud data acquired by the wrist camera. The transformer encoder captures the intrinsic correlations between features of different modes and outputs observed features. The observed features and the noisy action sequence are input into the diffusion transformer to obtain the denoised action sequence; The observed features are input into a preset pose prediction network to obtain the grasping pose.

[0009] In conjunction with the first aspect, in a fourth possible implementation of the first aspect, before determining the target grasping region of the first image through a preset first visual model, the method further includes: Collect training data, which includes a third image captured by the top camera, a fourth image captured by the wrist camera, the end-effector pose, and the robotic arm motion commands during the implementation of the standard grasping task. Receive annotation instructions to annotate target data in the training data, wherein the target data includes a target mask, a target grasping pose, and a target action sequence; The prediction loss is determined based on the prediction data of the first visual model and the second visual model and the target data. The parameters of the first visual model and the second visual model are adjusted based on the prediction loss until a predetermined training termination condition is met.

[0010] In conjunction with the fourth possible implementation of the first aspect, in the fifth possible implementation of the first aspect, the prediction loss includes occlusion loss. Determining the prediction loss based on the prediction data from the first visual model and the second visual model and the target data includes: Obtain the 3D point cloud region corresponding to the labeled target mask, and calculate the minimum envelope sphere of the 3D point cloud region; For each time step in the training process, with the current wrist camera position as the vertex, a tangent is drawn to the minimum envelope sphere, and the cone is determined by the line connecting the tangent point and the vertex; The three-dimensional point cloud located inside the cone and outside the smallest envelope sphere is projected onto the bottom circle of the cone; The points projected onto the bottom circle are clustered according to density, and the sum of the convex hull areas of each cluster is used as the occlusion area. The shading loss is determined based on the ratio of the shading area to the area of ​​the bottom circle.

[0011] In conjunction with the fifth possible implementation of the first aspect, in the sixth possible implementation of the first aspect, determining the prediction loss based on the prediction data of the first visual model and the second visual model and the target data includes: The pose prediction loss is determined based on the predicted grasping pose determined by the second visual model and the calibrated target grasping pose. The action prediction loss is determined based on the predicted action sequence determined by the second visual model and the calibrated target action sequence. The prediction loss of the second visual model is determined based on the occlusion loss, the pose prediction loss, and the action prediction loss.

[0012] A second aspect of this application provides a material gripping device. The gripping device for gripping materials includes a robotic arm, a top camera, and a wrist camera. The top camera is located directly above the material to be gripped, and the wrist camera is located at the wrist of the robotic arm. The device includes: The first image acquisition unit is used to acquire a first image including the material through the top camera; The target grasping region determination unit is used to determine the target grasping region of the first image through a preset first visual model. The second image acquisition unit is used to control the robotic arm to move to the target grasping area and acquire a second image including the target grasping area through the wrist camera; The motion pose determination unit is used to determine the motion sequence and grasping pose of the robotic arm grasping the material in the second image through a preset second visual model; A gripping control unit is used to perform gripping operations on the material according to the action sequence and the gripping pose.

[0013] A third aspect of this application provides a material gripping device, the gripping comprising: robotic arm; A top camera, positioned above the working area of ​​the robotic arm, is used to capture a first image of the overall field of view; A wrist camera is positioned at the end of the robotic arm at a predetermined tilt angle to acquire a second image of a local field of view near the gripper. A control unit configured to perform steps of the method described in any of the first aspects.

[0014] A fourth aspect of this application provides a computer program product that, when run on a computer, causes the computer to execute the methods described in the first aspect or its various implementations.

[0015] A fifth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method as described in any of the first aspects.

[0016] A sixth aspect of this application provides a chip for implementing the methods in the various implementations of the first aspect described above. Specifically, the chip includes a processor for calling and running a computer program from a memory, causing a device equipped with the chip to perform the methods as described in the first aspect or its various implementations.

[0017] The beneficial effects of this application embodiment compared with the prior art are as follows: This application embodiment acquires a first image through a top camera, determines the target grasping area in the first image based on a first visual model, acquires a second image through a wrist camera set at the end of the robotic arm, and determines the action sequence and grasping pose of the robotic arm to grasp the material through a second visual model. By grasping the material through this action sequence and grasping pose, it is possible to avoid using 3D point cloud matching to obtain the grasping pose, which is beneficial to improving the reliability in disordered stacking scenarios. Even if the material moves during the grasping process, the timely observation through the wrist camera can adjust the output action sequence in time, which is beneficial to improving the grasping success rate and can effectively adapt to grasping operations in dynamically changing scenarios. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of a material gripping device provided in an embodiment of this application; Figure 2 This is a schematic diagram illustrating the implementation process of a material grasping method provided in an embodiment of this application; Figure 3 This is a schematic diagram of a target grasping region prediction process provided in an embodiment of this application; Figure 4 This is a flowchart illustrating the process of adding pose prediction to an action denoising network, as provided in an embodiment of this application. Figure 5 This is a structural diagram of an implementation of a novel network-based prediction and grasping pose, provided in an embodiment of this application. Figure 6 This is a schematic diagram illustrating the implementation process of a model data fine-tuning method provided in an embodiment of this application; Figure 7 This is a schematic diagram of a material gripping device provided in an embodiment of this application; Figure 8 This is a schematic diagram of a material gripping device provided in an embodiment of this application. Detailed Implementation

[0020] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0021] To illustrate the technical solution described in this application, specific embodiments are provided below.

[0022] In the field of industrial automation, using robots to perform grasping tasks such as sorting is a common operation. However, when faced with disordered stacked materials, traditional pre-coded grasping methods often struggle to cope effectively due to the high diversity and complexity of the state space. In recent years, introducing 3D vision information has been seen as an effective way to solve such problems and has been widely explored in research and practice. A typical approach is to use 3D point cloud matching technology to obtain the target grasping pose, thereby guiding the robot to complete the grasping process.

[0023] Currently, 3D point cloud matching-based grasping methods face the following challenges when processing disordered stacked materials: Incomplete point cloud information leads to matching failure: Due to mutual occlusion between stacked materials, the acquired 3D point cloud often has a large number of missing parts, which seriously affects the reliability of point cloud matching and makes it difficult to apply effectively in most practical scenarios.

[0024] Limitations of local perception lead to grasping interference: Existing methods typically rely solely on local point cloud information for target recognition and grasping pose planning, lacking a physical understanding of the overall stacked structure. In scenarios where materials are tightly stacked, even if an apparent "best match" target is found, actual grasping is still easily affected by interference from surrounding materials or bins, leading to failure.

[0025] Insufficient adaptability to dynamic scenes: These methods mostly rely on single-frame RGB-D images for pose estimation, which makes it difficult to cope with dynamic changes in materials in the scene and unable to achieve continuous and stable grasping operations.

[0026] High dependence on point cloud quality and sensor performance: The accuracy of the matching results depends heavily on the integrity and precision of the input point cloud, thus placing high demands on the imaging quality and anti-interference capabilities of the depth camera, which limits its applicability in complex industrial environments.

[0027] To address the aforementioned problems, this application proposes a method for grasping materials. Figure 1This is a schematic diagram of a material gripping device according to an embodiment of this application. The gripping device includes a robotic arm 1, a top camera 2, and a wrist camera 3. The robotic arm 1 has multi-degree-of-freedom motion capabilities and is equipped with a gripper at its end, which can be used to perform material gripping actions. The top camera 2 is located directly above the material bin and is used to acquire a first image of the entire stacked area from a top-down angle, i.e., an RGB-D image. The wrist camera 3 is located at the end of the robotic arm 1 and moves in coordination with the gripper. By being tilted at a certain angle (e.g., 45 degrees), it is used to observe the target material and its surrounding environment at close range. When grasping materials, the first image including the material is first captured by the top camera 2. The target grasping area is determined by the preset first visual model. The robotic arm 1 is then controlled to move to the target grasping area. During the movement, the second image is captured by the wrist camera 3. The action sequence and grasping pose of the robotic arm 1 in the second image are determined by the second visual model. Based on the determined action sequence and grasping pose, the use of 3D point cloud matching to obtain the grasping pose can be avoided, which is beneficial to improving the reliability in disordered stacking scenarios. Even if the material moves during the grasping process, the timely observation through the wrist camera can adjust the output action sequence in time, which is beneficial to improving the grasping success rate and can effectively adapt to grasping operations in dynamically changing scenarios.

[0028] Figure 2 This is a schematic diagram illustrating the implementation process of a material grasping method proposed in an embodiment of this application. The method is based on... Figure 1 The method for gripping the material shown is described in detail below: In S201, a first image including the material is acquired by the top camera.

[0029] In this embodiment of the application, when the first image is captured by the top camera, in order to improve the reliability of the first image, the robotic arm can be controlled to move out of the field of view of the top camera, thereby reducing the obstruction of materials and improving the integrity of materials in the first image.

[0030] The first image may include depth information and RGB information, which can determine the distance between each pixel in the first image and the top camera.

[0031] In S202, the target grasping area of ​​the first image is determined by a preset first visual model.

[0032] The first visual model in this embodiment can be used to extract the target grasping region from the first image. The first visual model can be a language-visual model. For example, the first visual model can be BridgeVLA (Bridge Vision-Language-Action Model), or it can be other visual-language models, including RGB-based VLM heatmap prediction models, etc.

[0033] BridgeVLA was originally a Visual-Language-Action (VLA) strategy for receiving 3D point cloud input. Typically, the 3D point cloud is projected onto three fixed planes to generate three 2D RGB images. Then, a grasping heatmap encoding is generated on the RGB images using a 2D visual model. Finally, the grasping heatmap is projected back onto the 3D point cloud, and the optimal grasping prediction is obtained using an action head mapping.

[0034] However, in disordered stacking scenarios, mutual occlusion of materials leads to incomplete point clouds and significant loss of RGB image information in the projected data, making it difficult for the original BridgeVLA to reach its theoretical performance limit. In actual testing, it was found that the 2D vision model of the intermediate layer of BridgeVLA can output a reliable grasping region from a complete RGB image. To fully utilize the performance of the BridgeVLA pre-trained model, this embodiment eliminates the original motion head and grasping pose prediction, instead using only the 2D vision model to output a feasible grasping region based on the RGB image in the first image captured by the top camera. Specifically, the prediction process for the target grasping region is as follows: Figure 3 As shown, it includes: In S301, the first image and natural language instructions are input into the first visual language model, and a two-dimensional heat map is output through the first visual language model.

[0035] The first image captured in real-time by a top camera, along with pre-defined natural language instructions, is input into the 2D visual model of the first visual language model, BridgeVLA. Internally, the model uses a visual encoder (such as CNN (Convolutional Neural Network) or ViT (Vision Transformer)) to extract deep features of the image, while using a text encoder (such as BERT (Bidirectional Encoder Representations from Transformers)) to extract semantic features of the instructions. Then, a cross-modal attention mechanism fuses the visual and linguistic features, enabling the model to understand "what regions in the current image match the textual description 'graspable'." In one possible implementation, the first visual language model, BridgeVLA, can use PaliGemma (a large visual language model) as its backbone. This backbone includes a SigLIP (Sigmoid language-image pre-trained model) visual encoder, a word segmenter, and a Gemma Transformer backbone. Its working principle is as follows: PaliGemma takes a two-dimensional image and prefix text (i.e., natural language instructions) as input. For the image token (lexicon) and the prefix text token, a bidirectional attention mechanism is used to fuse them. The model outputs a suffix text token. In the application of BridgeVLA, these tokens directly correspond to the heat values ​​of each pixel in the image. Subsequently, image reconstruction and upsampling operations are performed based on the token positions to generate a heat map with the same resolution as the first input image.

[0036] The fused multimodal features can be upsampled using a decoder (such as a deconvolutional network) to generate a heatmap with the same resolution as the input first image. For example, for a 640x480 RGB image, the model will output a 640x480 heatmap. On the heatmap, material surface areas exposed at the top of the stack, with clear outlines and sufficient surrounding space will be highlighted.

[0037] In S302, a target mask is determined based on the pixels in the heat map whose heat values ​​are greater than a preset heat threshold.

[0038] For each pixel in the generated heatmap, if its heat value is greater than or equal to the preset heat threshold, then mark the pixel position as 1 (foreground) in the target mask; otherwise, mark it as 0 (background).

[0039] After marking is complete, further processing can be performed, including morphological opening operations (erosion followed by dilation) to remove small noise points and smooth the boundaries of the mask area.

[0040] In S303, the target point cloud corresponding to the target mask is determined based on the internal and external parameters of the top camera and the depth image, and the target grasping area is determined based on the target point cloud.

[0041] We can iterate through all pixels with a value of 1 in the target mask to obtain their coordinates (u, v) in the image and the corresponding depth value d in the depth image. Using camera intrinsics, the 2D pixel coordinates (u, v) and depth value d are transformed into 3D points in the camera coordinate system. Using camera extrinsic parameters (including rotation matrices and translation vectors), the points in the camera coordinate system are transformed into 3D points in the world coordinate system (usually aligned with the robot arm's base coordinate system). The set of all transformed 3D points generates the target point cloud. The 3D space occupied by this target point cloud is the finally determined target grasping area.

[0042] In S203, the robotic arm is controlled to move to the target grasping area, and a second image including the target grasping area is acquired by the wrist camera.

[0043] The coordinates of the center point of the target grasping area can be calculated. A collision-free trajectory is planned around this center point, and the robotic arm is controlled to move along this trajectory so that the end effector of the robotic arm (such as the gripper) is suspended directly above the center point.

[0044] After the robotic arm hovers directly above the center point, a close-up RGB-D image is captured as a second image using a wrist camera mounted at a predetermined angle to the horizontal plane at the end of the robotic arm. This predetermined angle, for example, can be a 45-degree angle, to capture the relative positional relationship between the gripper and the target material in real time.

[0045] In S204, the action sequence and grasping posture of the robotic arm grasping the material in the second image are determined by a preset second vision model.

[0046] The second vision model can be a second language-visual model. For example, the second vision model can be an FP3 model, or it can be a model that combines a diffusion model with a transformer, etc.

[0047] FP3 is a 3D VLA strategy. The FP3 model has a high performance ceiling. However, in unordered stacking scenarios, the FP3 model suffers significant performance degradation, mainly due to the following issues: 1. When multiple identical materials exist in the field of view at the same time, the FP3 model shows signs of frequent switching in target selection, resulting in poor consistency of the model's output actions before and after. 2. The model's target selection is not always optimal. When moving towards a non-optimal target, the wrist camera's field of view can easily be obstructed by other materials.

[0048] 3. In complex stacked scenarios, the point clouds of materials stick together. It is difficult to segment individual materials and determine their poses using only sparse point clouds, resulting in poor model output of motion poses.

[0049] To address the above issues, we modified and optimized FP3, including: 1. Remove the side camera point cloud from the input.

[0050] When grasping materials from a bin, the bin and robotic arm can severely obstruct the view of the material being grasped. In this situation, the side camera point cloud not only fails to provide any effective visual information, but the redundant network parameters also degrade model performance. Therefore, this embodiment removes the side camera point cloud from the model input and disables the network layers related to side camera point cloud processing.

[0051] 2. Add a second image captured by the wrist camera to the input.

[0052] A second image captured by a wrist camera is added to the input of the FP3 model and encoded using DINO V2 (second-generation unlabeled distillation model). This image is then fed into the Transformer Encoder structure along with other modal codes for alignment. The inclusion of the RGB image in the second image compensates for the deficiencies of the sparse point cloud representation, enabling the model to adjust the gripper pose in real time according to the target pose, resulting in a more reasonable output action sequence.

[0053] 3. Add prediction of target grasping pose to the output.

[0054] The main reason for point cloud failure and material interference is that FP3 relies solely on the matching degree between point cloud features and natural language to select potential targets, lacking a spatiotemporal assessment of whether a potential target is graspable. This application's embodiments can add target grasping pose prediction to the output, explicitly outputting potential targets to the model. This allows the model to optimize its actions based on constraints such as visibility constraints when outputting actions moving towards the target, avoiding trajectories with visual obstruction. Furthermore, it can standardize target selection during model training, using manually selected targets as guidance, prompting the model to learn more effective target selection methods.

[0055] In the embodiments of this application, when determining the action sequence and grasping pose, pose prediction can be added to the action denoising network, or a new network layer can be used for pose prediction.

[0056] The flowchart for adding pose prediction to an action denoising network can be shown as follows: Figure 4 As shown, multimodal data can be encoded using a pre-defined encoding model to extract multimodal features. This multimodal data can include a second image captured by a wrist camera, natural language commands, robotic arm status, and point cloud data acquired by the wrist camera. The second image can be encoded using DINO V2; the point cloud data acquired by the wrist camera can be encoded using Uni3D ViT (a large-scale 3D vision model); the robotic arm status can be encoded using MLP (Multilayer Perceptron); and the natural language commands can be encoded using CLIP (Contrastive Language-Image Pre-trained Model). The encoded multimodal features can be captured by a Transformer Encoder to capture the intrinsic relationships between different modal features and output the observed features. By inputting the observed features, a noisy action sequence, and a noisy target grasping pose into a diffusion transformer, the denoised action sequence and the denoised grasping pose can be obtained.

[0057] By adding dimensions at the end of the original model's action denoising network to denoise the grasping pose, the action sequence can interact with the newly added grasping pose thanks to the attention mechanism in the denoising network, resulting in a high consistency between the action sequence and the grasping pose output by the model.

[0058] like Figure 5 The diagram illustrates a novel network-based structure for predicting grasping poses, as proposed in this application. For multimodal data, a pre-defined encoding model can be used to encode and extract multimodal features. This multimodal data can include a second image captured by a wrist camera, natural language commands, the state of the robotic arm, and point cloud data acquired by the wrist camera. The second image can be encoded using DINO V2; the point cloud data acquired by the wrist camera can be encoded using Uni3D ViT (3D Vision Basic Large Model); the robotic arm state can be encoded using MLP (Multilayer Perceptron); and the natural language commands can be encoded using CLIP (Contrastive Language-Image Pre-trained Model). The encoded multimodal features can be captured by a Transformer Encoder, which captures the intrinsic relationships between different modal features and outputs observed features. The observed features, a noisy action sequence, and a noisy target grasping pose are input into a diffusion transformer to predict a denoised action sequence. The observed features are then input into a prediction pose prediction network to obtain the grasping pose.

[0059] After obtaining the observation encoding using the Transformer Encoder, this method feeds the observation features into two different networks. In the action prediction part, an action prediction network is used to predict the action sequence; in the pose prediction part, an additional pose prediction network is used, taking the observation features as input, to predict the output pose sequence. In this structure, action prediction and pose prediction are performed in parallel, thus preserving the original model's action planning capabilities to the greatest extent possible.

[0060] Therefore, the two pose prediction methods mentioned above can be flexibly selected according to actual needs. Figure 4 The pose prediction method for motion-pose collaborative noise reduction, as shown, emphasizes the strong correlation between motion and pose, and is suitable for scenarios where the material grasping path needs to strictly match the pose. Figure 5 The parallel independent prediction method shown places greater emphasis on the independence of pose decision-making, making it easier to adjust quickly in dynamic environments.

[0061] In S205, the material is grasped according to the action sequence and the grasping pose.

[0062] The motion sequence output by the second vision model (containing a series of robotic arm end-effector poses and gripper opening and closing commands) is converted into low-level joint motor control commands to drive the robotic arm to perform a continuous "descend-grasp-lift" motion. The gripping pose serves as the target pose at the moment the gripper closes in the motion sequence.

[0063] During execution, two indicators can be monitored in real time: whether all gripper commands in the action sequence are in a "closed" state, and whether there is a significant upward movement at the end of the robotic arm. If both conditions are met, the current gripping action is considered complete.

[0064] After the grasping is complete, the latest second image from the wrist camera can be input into a pre-trained visual-language model (VLM). The VLM determines whether the grasping was successful (e.g., whether the material was stably gripped). If successful, the robotic arm is controlled to transport the material to the placement area; if it fails, the robotic arm is controlled to exit the top camera's field of view and the grasping operation is repeated.

[0065] The closed-loop process of this solution enables the system to adapt to dynamic environments. Even if the material moves during the grasping process, the system can re-perceive and replan its strategy using a second image, without human intervention, effectively increasing the success rate of grasping.

[0066] In the embodiments of this application, the first and second visual models, such as BridgeVLA and FP3, can be fine-tuned by collecting human teaching data based on pre-training. The implementation process can be as follows: Figure 6 As shown, it includes: In S601, training data is collected.

[0067] The training data includes the third image captured by the top camera, the fourth image captured by the wrist camera, the end-effector pose, and the robotic arm motion commands during the implementation of the standard grasping task.

[0068] Fine-tuning data for the first and second vision models (BridgeVLA and FP3) can be acquired simultaneously within a single topology. Data acquisition can be achieved remotely. During acquisition, a robotic arm model can be set up for operator control, and the operations on the robotic arm model can be mapped to the actual robotic arm for action execution. Data can be saved at a certain frame rate (e.g., 10fps) during the acquisition process. This data includes: (1) RGB-D images from the top camera, i.e., the third image acquired by the top camera; (2) RGB-D images from the wrist camera, i.e., the fourth image acquired by the wrist camera; (3) the end-effector pose of the robotic arm; and (4) the robotic arm action commands. To efficiently improve the model's grasping ability, only the "descend-grasp-lift" segment can be retained from the acquired data, and the placement and reset segments can be excluded from the training data.

[0069] In S602, a labeling instruction is received to label the target data in the training data.

[0070] The target data includes the target mask, the target grab pose, and the target action sequence.

[0071] The grasping task can be broken down into five segments: "descend - grasp - lift - place - reset," with only the "descend - grasp - lift" segment retained for training. Therefore, timestamps need to be labeled between the segments for differentiation.

[0072] Labeling can be done in a "pre-labeling" manner: First, the starting moment of the downward movement can be selected based on whether the change in the robotic arm's pose between adjacent time steps in the collected training data exceeds a threshold. Then, starting from the starting moment of the downward movement, the starting moment of the lifting movement is determined based on the closing of the gripper and the rising of the robotic arm position in the training data. The lifting start moment is then delayed by a predetermined number of time steps, such as 10 time steps, as the lifting end moment, making the model's output after grasping more convenient for judging the completion of subsequent tasks. The data between the starting moment of the downward movement and the lifting end moment is used as training data.

[0073] The target mask can be labeled in the first frame image of the top camera and the wrist camera, based on the actual material being captured.

[0074] The target grasping pose, as one of the targets predicted by the model, needs to be labeled in the training data. Labeling can be done automatically. The timing of the grasp can be determined by whether the gripper is closed and whether there is a sudden change in the gripper width. Then, the actual pose of the robotic arm's end effector gripper at the moment of grasping is used as the target grasping pose.

[0075] The target action sequence is the sequence of actions of the robotic arm in the process of grasping materials, which is found in the training data.

[0076] In S603, a prediction loss is determined based on the prediction data of the first visual model and the second visual model and the target data, and the parameters of the first visual model and the second visual model are adjusted based on the prediction loss until a predetermined training termination condition is met.

[0077] During model fine-tuning, the first-vision model, such as BridgeVLA, aims to predict the grasping pose. It fine-tunes all network layer weights, retaining only the encoded heatmap results from intermediate layers during deployment. The second-vision model, such as FP3, aims to predict both the grasping pose and the action. It can use LoRa (Low-Rank Adaptation) to fine-tune the weights of all linear layers in the original model and fine-tune the weights of the newly added grasping pose prediction network. During fine-tuning, the training loss of the second-vision model, such as FP3, can be expressed as: loss = loss pose +loss action +loss occluded Where, loss pose Loss for pose prediction action Loss for action prediction occluded To cover up the damage.

[0078] That is, the pose prediction loss can be determined based on the predicted grasping pose determined by the second vision model and the calibrated target grasping pose, and the action prediction loss can be determined based on the predicted action sequence determined by the second vision model and the calibrated target action sequence.

[0079] When determining the occlusion loss, the 3D point cloud region corresponding to the labeled target mask (the mask determined by the shape of the material) can be obtained, and the minimum envelope sphere of this 3D point cloud region can be calculated. For each time step in the training process, with the current wrist camera position as the vertex, a tangent line is drawn to the minimum envelope sphere, and a cone is determined by the line connecting the tangent point and the vertex. The 3D point cloud located inside the cone but outside the minimum envelope sphere is projected onto the base circle of the cone. The points projected onto the base circle are clustered according to density, and the sum of the convex hull areas formed by each cluster is taken as the occlusion area. Based on the shading area and the area of ​​the base circle The ratio determines the shading loss. , can be represented as:

[0080] By labeling the target mask of the capture target in the training data and calculating the corresponding 3D point cloud to obtain the labeled region, the minimum envelope sphere of the labeled region can be solved using the stochastic incremental method or the second-order cone programming method. For each time step, with the wrist camera position as the vertex, tangents are drawn to the minimum envelope sphere, and the line segments connecting all tangent points to the vertex form a cone. For all points in the 3D point cloud that are inside the cone but not inside the envelope sphere, they are projected onto the base circle of the cone, starting from the wrist camera position. Finally, a density clustering algorithm can be used to cluster the projected points according to density, forming convex hulls for each cluster as occlusion regions.

[0081] In disordered stacking scenarios, occlusion point clouds may come from different materials. For example, two objects may occlude a portion of the target material on the left and right sides, while the middle area remains unoccluded. If only a convex hull formed by the entire occlusion point cloud is used to represent the occlusion area, then most of the unoccluded area in the middle will also be considered as occlusion, leading to unreasonable calculation results. Therefore, by adding clustering operations, the accuracy of the calculation results can be improved.

[0082] In addition, by using the envelope sphere and the base circle of the cone for calculation, the envelope sphere only needs to be calculated and generated at the initial position. Subsequent camera position movements will not affect the position and size of the envelope sphere, thus effectively improving calculation efficiency, especially the calculation efficiency of occlusion loss of irregular materials.

[0083] In summary, the material grasping method in this application embodiment can effectively avoid using 3D point cloud matching to obtain the grasping pose, thereby making it more versatile and reliable in disordered stacking scenarios.

[0084] This method can independently complete continuous grasping tasks using images captured by a top camera and a wrist camera. No human intervention is required during the task.

[0085] The first and second vision models, through continuous observation input and action sequence output, enable this grasping method to have strong adaptability to dynamic environments. Even if the material or material box moves during the grasping process, timely updated observations allow the second vision model to adjust the output action sequence in a timely manner.

[0086] In addition, this method does not require high-performance depth cameras. Even if the point cloud image is poor at the initial position due to the distance, the noise of the point cloud will gradually decrease as the camera approaches the target.

[0087] Furthermore, this method, through modifications to BridgeVLA, makes the localization of the grasping area more focused and reliable. Performance tests were conducted on the FP3 model before and after the modification in the same scenario of disordered stacked materials. The grasping success rate of the model before modification was approximately 40%, while the grasping success rate after modification reached over 85%. This method significantly improves the success rate of grasping tasks by modifying the FP3 model.

[0088] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0089] Figure 7 A material gripping device provided in this application embodiment includes a robotic arm, a top camera, and a wrist camera. The top camera is located directly above the material to be gripped, and the wrist camera is located at the wrist of the robotic arm. The device includes: The first image acquisition unit 701 is used to acquire a first image including the material through the top camera.

[0090] The target grasping region determination unit 702 is used to determine the target grasping region of the first image through a preset first visual model.

[0091] The second image acquisition unit 703 is used to control the robotic arm to move to the target grasping area and acquire a second image including the target grasping area through the wrist camera.

[0092] The motion pose determination unit 704 is used to determine the motion sequence and grasping pose of the robotic arm grasping the material in the second image through a preset second visual model.

[0093] The gripping control unit 705 is used to perform gripping operations on the material according to the action sequence and the gripping pose.

[0094] Figure 7 The material gripping device shown is, with Figure 2 The corresponding material grabbing method is shown.

[0095] Figure 8 This is a schematic diagram of a material gripping device provided in an embodiment of this application. Figure 8As shown, the material gripping device in this embodiment includes: a robotic arm 81, a top camera 82, a wrist camera 83, and a control unit 84. The top camera 82 is positioned above the working area of ​​the robotic arm to capture a first image of the overall field of view. The wrist camera 83 is positioned at a predetermined tilt angle at the end of the robotic arm to acquire a second image of a local field of view near the gripper. The control unit executes the material gripping method described above. The control unit 84 includes a processor 840, a memory 841, and a computer program 842 stored in the memory 841 and executable on the processor 840, such as a material gripping program. When the processor 840 executes the computer program 842, it implements the steps in the various material gripping method embodiments described above. Alternatively, when the processor 840 executes the computer program 842, it implements the functions of each module / unit in the various device embodiments described above.

[0096] For example, the computer program 842 may be divided into one or more modules / units, which are stored in the memory 841 and executed by the processor 840 to complete this application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 842 in the material gripping device.

[0097] The material gripping device may include, but is not limited to, a processor 840 and a memory 841. Those skilled in the art will understand that... Figure 8 This is merely an example of a material gripping device and does not constitute a limitation on the material gripping device. It may include more or fewer components than shown, or combine certain components, or different components. For example, the material gripping device may also include input / output devices, network access devices, buses, etc.

[0098] The processor 840 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0099] The memory 841 can be an internal storage unit of the material gripping device, such as a hard drive or memory of the material gripping device. The memory 841 can also be an external storage device of the material gripping device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the material gripping device. Furthermore, the memory 841 can include both internal and external storage units of the material gripping device. The memory 841 is used to store the computer program and other programs and data required by the material gripping device. The memory 841 can also be used to temporarily store data that has been output or will be output.

[0100] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0101] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0102] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0103] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0104] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0105] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0106] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by hardware related to computer program instructions. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0107] In addition, this application also provides a computer program product that, when run on a computer, causes the computer to execute the methods in the above-described implementations.

[0108] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for grasping materials, characterized in that, The gripping device for grasping materials includes a robotic arm, a top camera, and a wrist camera, wherein the top camera is located directly above the material to be grasped, and the wrist camera is located at the end of the robotic arm. The method includes: A first image including the material is captured by the top camera; The target grasping area of ​​the first image is determined by a preset first visual model. The robotic arm is controlled to move to the target grasping area, and a second image including the target grasping area is acquired through the wrist camera; The robotic arm's action sequence and grasping posture for grasping materials in the second image are determined by a preset second visual model. The material is grasped according to the action sequence and the grasping pose.

2. The method according to claim 1, characterized in that, The first visual model is a first visual language model; The target grasping region of the first image is determined by a preset first visual model, including: The first image and natural language instructions are input into the first visual language model, and a two-dimensional heat map is output through the first visual language model. The target mask is determined based on the pixels in the heat map whose heat values ​​are greater than a preset heat threshold; Based on the internal and external parameters of the top camera and the depth image, the target point cloud corresponding to the target mask is determined, and the target grasping area is determined based on the target point cloud.

3. The method according to claim 1, characterized in that, The robotic arm determines the action sequence and grasping pose for grasping materials in the second image using a preset second visual model, including: The multimodal data is encoded using a preset encoding model to extract multimodal features. The multimodal data includes the second image, natural language commands, robotic arm status, and point cloud data acquired by the wrist camera. The transformer encoder captures the intrinsic correlations between features of different modes and outputs observed features. The observed features, the noisy action sequence, and the noisy target grasping pose are input into the diffusion transformer to obtain the denoised action sequence and the denoised grasping pose.

4. The method according to claim 1, characterized in that, The robotic arm determines the action sequence and grasping pose for grasping materials in the second image using a preset second visual model, including: The multimodal data is encoded using a preset encoding model to extract multimodal features. The multimodal data includes the second image, natural language commands, robotic arm status, and point cloud data acquired by the wrist camera. The transformer encoder captures the intrinsic correlations between features of different modes and outputs observed features. The observed features and the noisy action sequence are input into the diffusion transformer to obtain the denoised action sequence; The observed features are input into a preset pose prediction network to obtain the grasping pose.

5. The method according to claim 1, characterized in that, Before determining the target grasping region of the first image using a preset first visual model, the method further includes: Collect training data, which includes a third image captured by the top camera, a fourth image captured by the wrist camera, the end-effector pose, and the robotic arm motion commands during the implementation of the standard grasping task. Receive annotation instructions to annotate target data in the training data, wherein the target data includes a target mask, a target grasping pose, and a target action sequence; The prediction loss is determined based on the prediction data of the first visual model and the second visual model and the target data. The parameters of the first visual model and the second visual model are adjusted based on the prediction loss until a predetermined training termination condition is met.

6. The method according to claim 5, characterized in that, The predicted loss includes occlusion loss. Determining the prediction loss based on the prediction data from the first visual model and the second visual model and the target data includes: Obtain the 3D point cloud region corresponding to the labeled target mask, and calculate the minimum envelope sphere of the 3D point cloud region; For each time step in the training process, with the current wrist camera position as the vertex, a tangent is drawn to the minimum envelope sphere, and the cone is determined by the line connecting the tangent point and the vertex; The three-dimensional point cloud located inside the cone and outside the smallest envelope sphere is projected onto the bottom circle of the cone; The points projected onto the bottom circle are clustered according to density, and the sum of the convex hull areas of each cluster is used as the occlusion area. The shading loss is determined based on the ratio of the shading area to the area of ​​the bottom circle.

7. The method according to claim 6, characterized in that, Determining the prediction loss based on the prediction data from the first visual model and the second visual model and the target data includes: The pose prediction loss is determined based on the predicted grasping pose determined by the second visual model and the calibrated target grasping pose. The action prediction loss is determined based on the predicted action sequence determined by the second visual model and the calibrated target action sequence. The prediction loss of the second visual model is determined based on the occlusion loss, the pose prediction loss, and the action prediction loss.

8. A material gripping device, characterized in that, A gripping device for grasping materials includes a robotic arm, a top camera, and a wrist camera. The top camera is located directly above the material to be grasped, and the wrist camera is located at the wrist of the robotic arm. The device includes: The first image acquisition unit is used to acquire a first image including the material through the top camera; The target grasping region determination unit is used to determine the target grasping region of the first image through a preset first visual model. The second image acquisition unit is used to control the robotic arm to move to the target grasping area and acquire a second image including the target grasping area through the wrist camera; The motion pose determination unit is used to determine the motion sequence and grasping pose of the robotic arm grasping the material in the second image through a preset second visual model; A gripping control unit is used to perform gripping operations on the material according to the action sequence and the gripping pose.

9. A material gripping device, characterized in that, The crawling includes: robotic arm; A top camera, positioned above the working area of ​​the robotic arm, is used to capture the first image of the global field of view; A wrist camera is positioned at the end of the robotic arm at a predetermined tilt angle to acquire a second image of a local field of view near the gripper. A control unit configured to perform the steps of the method according to any one of claims 1 to 7.

10. A computer program product comprising computer program instructions, characterized in that, When the computer program is run, the method as described in any one of claims 1-7 is performed.

Citation Information

Cited By

  • VLA large model evolution method oriented to industrial manufacturing scene

    CN121809606A

  • Point cloud generation method and device based on multi-modal information, equipment and storage medium

    CN122089967A