Multi-model flexible six-axis robot vision grabbing system

Through a multi-model flexible six-axis robot visual capture system, combining quality data and depth point cloud data, the optimal crawling posture is generated, which solves the problem of unstable crawling of the visual capture system in complex scenarios, and achieves high-precision and stable object capture.

CN120480912AActive Publication Date: 2025-08-15GUANGDONG SHENGHUI TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510774630.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-08-15
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

In the prior art, when the visual grasping system has missing texture or uneven light, the edge accuracy of the segmentation mask is significantly reduced, resulting in contour recognition errors, which in turn causes deviations in position judgment of the target object, affecting the grasping stability, especially for non-uniform mass objects to cause slip or overturning.

Method used

A multi-model flexible six-axis robot visual capture system is adopted to obtain quality data, two-dimensional density maps and depth images through the data acquisition module. The encoder-decoder Transformer architecture is used to extract contour feature data, combine the depth point cloud data to generate mass distribution data, and filter out the optimal grab posture through the NMS algorithm. Finally, the six-axis robot control module performs the grab operation.

Benefits of technology

Accurately extracting object contour and spatial position information in complex scenarios, improving the robot's grasping accuracy, ensuring the stability and success rate of grabbing, and adapting to the grasping needs of non-uniform mass objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120480912A_ABST
    Figure CN120480912A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of industrial automation, and particularly relates to a multi-model flexible six-axis robot vision grabbing system which comprises a data acquisition module used for acquiring quality data, a two-dimensional density map, an RGB image and a depth image of a target object; the target detection module is used for acquiring depth point cloud data; the weight compensation module is used for fusing the depth point cloud data and the quality data of the target object to generate quality distribution data; the pose calculation module is used for generating a plurality of groups of grabbing pose samples on the basis of the mass distribution data and screening out an optimal grabbing pose from the grabbing pose samples on the basis of an NMS algorithm; and the optimal grabbing posture is converted into a joint movement instruction of a six-axis robot of a specific model, the six-axis robot is controlled to execute grabbing operation based on the joint movement instruction, the optimal grabbing posture is generated through mass distribution data fused by the point cloud data and the mass data, and the precision of the grabbing action executed by the robot is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of industrial automation, and in particular relates to a multi-machine flexible six-axis robot visual grasping system. Background Art

[0002] In automated production, intelligent robots are used to automatically grasp, move objects, or operate tools according to a set control program. Robots typically use a visual grasping system to identify the surface features of the target object, thereby determining the desired grasping position for the robot's mechanical arm, enabling precise grasping of the target object.

[0003] Existing solutions rely on visual information (such as RGB-D images) to plan the grasping posture, but lack the actual mass distribution characteristics of the object to be grasped, such as the center of gravity and moment of inertia. For objects with non-uniform mass (such as liquid containers and irregularly shaped workpieces), torque imbalance during grasping can easily cause slippage or overturning. When the recognition system cannot determine the weight of the object, the control system port cannot adaptively adjust the gripping force or suction force by controlling the end effector, resulting in overpressure damage to the gripper or insufficient gripping force.

[0004] Moreover, when the target object has missing texture or uneven lighting, the edge accuracy of the segmentation mask decreases significantly, resulting in contour recognition errors, which in turn causes the recognition system to deviate from its position judgment of the target object. The posture deviation will further affect the trajectory planning of the robot arm and reduce the success rate of grasping. Summary of the Invention

[0005] In order to solve the above-mentioned problems existing in the prior art, the present invention provides a multi-model flexible six-axis robot visual grasping system to solve the problem that when the target object has texture loss or uneven lighting, the edge accuracy of the segmentation mask is significantly reduced, resulting in contour recognition errors, and then causing the recognition system to deviate from the position judgment of the target object and thus cause unstable grasping.

[0006] The purpose of the present invention can be achieved by the following technical solution: A multi-machine flexible six-axis robot visual grasping system, comprising:

[0007] A data acquisition module is used to obtain the mass data, two-dimensional density map, RGB image and depth image of the target object;

[0008] An object detection module is configured to synthesize a point cloud image of the target object based on the RGB image and the depth image, extract contour feature data from the RGB image using a Transformer architecture with an encoder-decoder structure, and apply the contour feature data to the depth image to extract depth point cloud data of the target object;

[0009] The weight compensation module is used to generate mass distribution data based on the depth point cloud data and mass data of the target object;

[0010] The posture calculation module is used to generate several groups of grasping posture samples based on the mass distribution data, and select the optimal grasping posture from the several grasping posture samples based on the NMS algorithm;

[0011] The six-axis robot control module converts the optimal grasping posture into joint motion instructions of a specific model of the six-axis robot and controls the six-axis robot to perform a grasping operation based on the joint motion instructions.

[0012] Preferably, in the weight compensation module, the depth point cloud data and the quality data are fused using the following steps:

[0013] S1: Based on the Radial Basis Function (RBF) interpolation algorithm, the discrete sampling points of the two-dimensional density map are mapped to the three-dimensional point cloud surface through the distance-weighted kernel function to generate a continuous mass density field;

[0014] S2: Optimizing the mass density field using a variable density SIMP algorithm to minimize the inhomogeneity of mass distribution;

[0015] S3: Obtain the ambient temperature factor and perform temperature compensation correction on the coordinate points of the optimized mass density field point by point;

[0016] S4: performing coordinate fusion on the corrected mass density field and the point cloud coordinates corresponding to the depth point cloud data to generate mass distribution data;

[0017] Use Open3D algorithms to visualize the mass distribution of the target object and compare it with the mass data to verify the accuracy.

[0018] Preferably, step S1 includes the following steps:

[0019] The pixel coordinates of the two-dimensional density map are projected onto the three-dimensional point cloud surface through perspective transformation (PT) to generate the corresponding sampling points c i ;

[0020] Construct an interpolation system to calculate the mass density value f(x) for each 3D point cloud coordinate. The calculation formula of f(x) is:

[0021]

[0022] Among them, x is the coordinate of the target object in the three-dimensional data, c i is the three-dimensional projection coordinate corresponding to the two-dimensional density sampling point, is the radial basis function, p(x) is a low-order polynomial, and the constant term is taken according to the experimental data, w iis the weight coefficient, which is solved using the least squares method.

[0023] Preferably, using the NMS algorithm to screen the optimal grasping posture includes:

[0024] Input: all candidate grasping postures and their corresponding scores;

[0025] Sort all candidate postures in descending order of score and select the grasping posture with the highest score and put it into the optimal solution set;

[0026] Traverse the remaining candidate poses:

[0027] Calculate the similarity between the current candidate pose and all existing poses in the optimal solution set;

[0028] If the similarity between the current candidate pose and any pose in the optimal solution set exceeds a preset threshold, the current candidate pose is suppressed or eliminated;

[0029] If the similarity between the current candidate posture and all postures in the optimal solution set is lower than the threshold, it will be added to the optimal solution set;

[0030] Repeat the traversal process until all candidate poses are processed;

[0031] Output the optimal grasping posture.

[0032] Preferably, the data acquisition module further includes inputting a preset generative adversarial network into an occluded contour image of the occluded portion of the target object, and outputting a complete contour of the target object after completion. The determination of whether the target object is occluded includes:

[0033] The network model is back-projected onto the target object’s images at each viewing angle, and the projection residual is calculated. If the projection residual is greater than a preset threshold, it is determined that the target object is occluded at that viewing angle.

[0034] Preferably, the Transformer architecture includes:

[0035] Encoder: Uses Vision Transformer (ViT) as the backbone network to split the RGB image into fixed-size tiles, linearly embed them, add positional encoding, and then process them through a multi-layer Transformer encoder block;

[0036] Decoder: Receives the encoder output, progressively upsamples, and fuses features from different levels. The Transformer decoder block uses a cross-attention mechanism to output a contour feature map with the same resolution as the input RGB image.

[0037] Preferably, the Transformer architecture includes enhancing edge gradients of the contour feature map by a Laplacian operator.

[0038] Preferably, the data acquisition module includes:

[0039] Use load cells to obtain mass data;

[0040] Use an industrial CT scanner to obtain a two-dimensional density map;

[0041] Acquire RGB images using an RGB-D camera

[0042] Use stereo vision technology to obtain depth images;

[0043] And use GPIO signal lines to connect the trigger interface of the industrial CT scanner and RGB-D camera.

[0044] The multi-machine flexible six-axis robot visual grasping system according to claim 1 is characterized in that: the data acquisition module further includes aligning the pixels in the acquired RGB image and depth image, including:

[0045] Use the hand-eye calibration method to obtain the rotation matrix R and translation vector t of the RGB-D camera and depth sensor;

[0046] Map the RGB pixel coordinate A to the depth camera's coordinate system B through external parameters:

[0047] B=R·A+t.

[0048] Preferably, the six-axis robot includes end effectors of different models, and the end effectors of different models are used to grasp different target objects.

[0049] The beneficial effects of the present invention are:

[0050] This method constructs a 3D point cloud by fusing RGB images, depth images, and 2D density maps, and implements cross-modal feature alignment using an encoder-decoder Transformer architecture. This allows for accurate extraction of object contours and spatial position information even in complex scenes, such as those with varying lighting or missing textures. Furthermore, mass distribution data is used to understand the mass distribution of the target object, thereby improving the robot's grasping accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] To facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.

[0052] Figure 1 This is a system structure diagram of the present invention. DETAILED DESCRIPTION

[0053] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the specific implementation methods, structures, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.

[0054] See also Figure 1 This embodiment provides a multi-machine flexible six-axis robot visual grasping system, including:

[0055] The data acquisition module is used to obtain the mass data, two-dimensional density map, RGB image and depth image of the target object, including:

[0056] Use a load cell (such as an electronic scale integrated under a conveyor belt or on a workbench) to measure the mass of the target object as it enters the gripping station;

[0057] Obtain a two-dimensional density map. Use specific imaging techniques (such as industrial CT scanners, X-ray imaging, terahertz imaging, acoustic imaging, or density estimation algorithms based on multi-view reconstruction) to obtain the density distribution information of the target object on a two-dimensional plane (usually a plane perpendicular to the grasping direction). This generates a two-dimensional array or image DensityMap(x, y), whose values represent the local relative density.

[0058] Use an RGB-D camera to capture an RGB image of the target object;

[0059] Use a depth sensor (such as a structured light camera, binocular stereo vision camera, or ToF camera) to synchronously capture a depth image aligned with the RGB image, where each pixel value represents the distance from the point to the camera (Z coordinate).

[0060] Use GPIO signal lines to connect the industrial CT scanner, RGB-D camera, and the trigger interface of the sensor used to obtain the depth image of the target object, so that the data acquisition can be aligned and synchronized in time, which facilitates the subsequent fusion of multiple images of the target object.

[0061] Since the data acquisition module uses different cameras to capture and extract RGB images and depth images respectively, the target detection module synthesizes the point cloud map of the target object based on the acquired RGB images and depth images. In order to facilitate the next step of stereo matching in the target detection module, the data acquisition module also includes pixel alignment in the acquired RGB images and depth images, including:

[0062] Use the hand-eye calibration method to obtain the rotation matrix R and translation vector t of the RGB-D camera and depth sensor;

[0063] Map the RGB pixel coordinate A to the depth camera's coordinate system B through external parameters, and calculate the coordinate system B using the following formula:

[0064] B=R·A+t.

[0065] The target detection module synthesizes a point cloud map of the target object based on the RGB image and the depth image, uses the Transformer architecture of the encoder-decoder structure to extract the contour feature data of the RGB image, and applies the contour feature data to the depth image to extract the depth point cloud data corresponding to the target object;

[0066] Point cloud synthesis: Using the camera's intrinsic parameters, each valid depth pixel is back-projected into 3D space to generate the initial point cloud coordinates that contain all objects in the scene (including the background);

[0067] RGB contour feature extraction: Input the RGB_image into a Transformer network based on an encoder-decoder structure (such as SETR, SegFormer);

[0068] Encoder: Typically uses the Vision Transformer (ViT) or its variants as the backbone network. The RGB image is split into fixed-size tiles, linearly embedded, positional encoding is added, and then processed through a multi-layer Transformer encoder block. The encoder outputs a feature map or feature sequence containing rich contextual information.

[0069] Decoder: Receives the output of the encoder, gradually upsamples and fuses features from different levels. The Transformer decoder block (or a lighter CNN decoder) uses self-attention / cross-attention mechanisms to focus on the boundary information of the object. Finally, it outputs a contour feature map (ContourFeatMap) or segmentation mask (SegMask) with the same resolution as the input image. The contour feature map may contain high-dimensional feature vectors that highlight the precise boundary and shape information of the object. The contour feature provides object location information that is more robust and more focused on the boundary than the original RGB;

[0070] Deep point cloud data extraction: Using the contour feature map or segmentation mask obtained in step 2 as a guide, the point cloud data is processed, including:

[0071] Based on segmentation mask: If the decoder output is a segmentation mask SegMask (binary image or multi-class image, the target object area is 1), the segmentation mask is directly used as a mask to filter out the point cloud data with the corresponding mask value of the target category, and obtain the depth point cloud data of the target object;

[0072] Based on contour feature map: If the output is a high-dimensional ContourFeatMap, its information needs to be "fused" or "guided" deep point cloud segmentation, including:

[0073] Upsample the ContourFeatMap to the same resolution as the depth image pixels;

[0074] Concatenate or weightedly fuse the channel dimension of ContourFeatMap (including contour information) with the depth image pixels (or their derived features) at the feature level;

[0075] The fused features are fed into a lightweight segmentation head (e.g., a 1x1 convolutional layer) to predict the probability of each pixel belonging to the target object.

[0076] Get the segmentation mask SegMask by thresholding or using argmax according to the predicted probability;

[0077] Use SegMask to extract the depth point cloud data of the target object from the point cloud data;

[0078] Output depth point cloud data.

[0079] The weight compensation module generates mass distribution data based on the fusion of the target object's depth point cloud data and mass data, and extracts the center of gravity coordinates of the target object from the mass distribution data;

[0080] The fusion of depth point cloud data and quality data adopts the following method:

[0081] S1: Based on the Radial Basis Function (RBF) interpolation algorithm, the discrete sampling points of the 2D density map are mapped to the 3D point cloud surface through the distance-weighted kernel function to generate a continuous mass density field. Specifically, it includes:

[0082] The pixel coordinates of the two-dimensional density map are projected onto the three-dimensional point cloud surface through perspective transformation (PT) to generate the corresponding sampling points c i ;

[0083] Construct an interpolation system to calculate the mass density value f(x) for each 3D point cloud coordinate. The calculation formula of f(x) is:

[0084]

[0085] Among them, x is the coordinate of the target object in the three-dimensional data, c i is the three-dimensional projection coordinate corresponding to the two-dimensional density sampling point, is the radial basis function, p(x) is a low-order polynomial, and the constant term is taken according to the experimental data, w i , which is solved using the least squares method.

[0086] S2: Optimizing the mass density field using a variable density SIMP algorithm to minimize the inhomogeneity of mass distribution;

[0087] S3: Obtain the ambient temperature factor and perform temperature compensation correction on the coordinate points of the optimized mass density field point by point;

[0088] S4: Fusing the corrected mass density field with the point cloud coordinates corresponding to the depth point cloud data to generate mass distribution data;

[0089] Use Open3D algorithms to visualize the mass distribution of the target object and compare it with the mass data to verify the accuracy.

[0090] The posture calculation module generates and selects the most stable and feasible optimal grasping posture of the six-axis robot based on the mass distribution data of the target object, including:

[0091] The mass distribution data is used as input to a pre-trained neural network, and the point cloud is sampled as the grasping position. Feature extraction is performed on a fixed-size image centered at the grasping position. At each grasping position, the grasping angle is sampled from 0° to 170° in steps of 10°. This means that angle estimation is treated as an 18-category classification problem. The NMS algorithm is used to select the grasping position and angle corresponding to the highest score as the optimal grasping posture.

[0092] The neural network here contains convolutional layers and fully connected layers, and is generated by self-supervised learning training using historical grasping position and posture data of different robot models;

[0093] The NMS algorithm is used to remove highly overlapping (similar) candidate grasps and retain the optimal grasping postures with the highest scores and sufficient differences from each other. The execution process includes:

[0094] Input: all candidate grasping postures and their corresponding scores;

[0095] Sort all candidate postures in descending order of score and select the grasping posture with the highest score and put it into the optimal solution set;

[0096] Traverse the remaining candidate poses:

[0097] Calculate the similarity between the current candidate posture and all existing postures in the optimal solution set (such as whether the Euclidean distance of the grasping point, the angle of the direction vector, etc. is less than the threshold);

[0098] If the similarity between the current candidate pose and any pose in the optimal solution set exceeds a preset threshold (i.e., too similar), the current candidate pose is suppressed (eliminated);

[0099] If the similarity between the current candidate pose and all poses in the optimal solution set is lower than the threshold (i.e., sufficiently different), it will be added to the optimal solution set;

[0100] The traversal process is repeated until all candidate postures are processed, and the first posture in the optimal solution set (i.e., the one with the highest score and not suppressed) is selected as the optimal grasping posture.

[0101] Output: The calculated optimal robot end-effector grasping pose.

[0102] The six-axis robot control module converts the optimal grasping posture into joint motion instructions for a specific six-axis robot model, and controls the six-axis robot to complete the grasping action safely and accurately based on the joint motion instructions. Its execution process includes:

[0103] Based on the optimal grasping posture, the robot is transformed into the base coordinate system of a specific model to obtain the target joint position of the robot;

[0104] Planning the robot's joint motion trajectory: planning the terminal straight line or arc path in Cartesian space, and then converting it into the joint trajectory of the robot from the current joint position to the target joint position;

[0105] Discretize the planned motion trajectory into a series of target joint angles at time points, and package these target joint angles (and velocities, accelerations, etc.) into joint motion instructions (such as MoveJ, ServoJ, etc.) that can be understood by a specific robot controller;

[0106] Send instructions to the robot controller of a specific target model through the robot control interface;

[0107] A robot of a specific target model executes the instructions, and force sensor feedback and position feedback are set on the robot's end effector to detect the gripper status to confirm successful grasping.

[0108] In order to improve the accuracy of the target object's contour recognition by the target detection module, in one embodiment, the self-attention layer of the Transformer encoder uses a deformable attention mechanism, specifically including:

[0109] A lightweight CNN branch is inserted before the encoder self-attention layer, consisting of three 3×3 convolutional layers (with 643216 channels, respectively) and a ReLU function. Its input is the current query feature map (size H×W×C) and it outputs a spatial offset Δp (size H×W×2N, where N is the number of sampling points). The relative offset is learned through a residual structure to avoid training oscillations caused by direct regression of absolute coordinates.

[0110] The Transformer architecture cross-modally fuses contour feature data with the geometric feature data of the depth image. It constructs a three-layer feature pyramid, performs Laplacian enhancement on the features of each layer, and then upsamples them to a uniform size. The contour feature data and the geometric feature data of the depth image are weightedly fused through the channel attention mechanism (SE-Block), and generated using a weighted generation network. The weighted generation network contains global average pooling and two fully connected layers, and enhances the edge contour feature gradient through the Laplacian operator. Specifically, by using a 3×3 Laplacian convolution kernel to perform a convolution operation on the Transformerd output feature map, the positive gradient response is retained to strengthen the significant edges, thereby improving the target object.

[0111] The data acquisition module also includes inputting a preset generative adversarial network into the occluded contour image of the occluded target object portion, and outputting a complete contour of the target object after completion. The determination of whether the target object is occluded includes:

[0112] The network model is back-projected onto the target object’s images at each viewing angle, and the projection residual is calculated. If the projection residual is greater than a preset threshold, it is determined that the target object is occluded at that viewing angle.

[0113] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as above in terms of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can, without departing from the scope of the technical solution of the present invention, make some changes or modifications to equivalent embodiments using the technical contents disclosed above. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A multi-machine flexible six-axis robot visual grasping system, characterized by: include: A data acquisition module is used to obtain the mass data, two-dimensional density map, RGB image and depth image of the target object; An object detection module is configured to synthesize a point cloud image of the target object based on the RGB image and the depth image, extract contour feature data from the RGB image using a Transformer architecture with an encoder-decoder structure, and apply the contour feature data to the depth image to extract depth point cloud data of the target object; The weight compensation module is used to generate mass distribution data based on the depth point cloud data and mass data of the target object; The posture calculation module is used to generate several groups of grasping posture samples based on the mass distribution data, and select the optimal grasping posture from the several grasping posture samples based on the NMS algorithm; The six-axis robot control module converts the optimal grasping posture into joint motion instructions of a specific model of the six-axis robot and controls the six-axis robot to perform a grasping operation based on the joint motion instructions.

2. The multi-machine flexible six-axis robot visual grasping system according to claim 1, characterized in that: In the weight compensation module, the depth point cloud data and the quality data are fused using the following steps: S1: Based on the Radial Basis Function (RBF) interpolation algorithm, the discrete sampling points of the two-dimensional density map are mapped to the three-dimensional point cloud surface through the distance-weighted kernel function to generate a continuous mass density field; S2: Optimizing the mass density field using a variable density SIMP algorithm to minimize the inhomogeneity of mass distribution; S3: Obtain the ambient temperature factor and perform temperature compensation correction on the coordinate points of the optimized mass density field point by point; S4: performing coordinate fusion on the corrected mass density field and the point cloud coordinates corresponding to the depth point cloud data to generate mass distribution data; Use Open3D algorithms to visualize the mass distribution of the target object and compare it with the mass data to verify the accuracy.

3. The multi-machine flexible six-axis robot visual grasping system according to claim 2, characterized in that: Step S1 includes the following steps: The pixel coordinates of the two-dimensional density map are projected onto the three-dimensional point cloud surface through perspective transformation (PT) to generate the corresponding sampling points c i ; Construct an interpolation system to calculate the mass density value f(x) for each 3D point cloud coordinate. The calculation formula of f(x) is: Among them, x is the coordinate of the target object in the three-dimensional data, c i is the three-dimensional projection coordinate corresponding to the two-dimensional density sampling point, is the radial basis function, p(x) is a low-order polynomial, and the constant term is taken according to the experimental data, w i is the weight coefficient, which is solved using the least squares method.

4. The multi-machine flexible six-axis robot visual grasping system according to claim 1, characterized in that: Using the NMS algorithm to select the optimal grasping posture includes: Input: all candidate grasping postures and their corresponding scores; Sort all candidate postures in descending order of score and select the grasping posture with the highest score and put it into the optimal solution set; Traverse the remaining candidate poses: Calculate the similarity between the current candidate pose and all existing poses in the optimal solution set; If the similarity between the current candidate pose and any pose in the optimal solution set exceeds a preset threshold, the current candidate pose is suppressed or eliminated; If the similarity between the current candidate posture and all postures in the optimal solution set is lower than the threshold, it will be added to the optimal solution set; Repeat the traversal process until all candidate poses are processed; Output the optimal grasping posture.

5. The multi-machine flexible six-axis robot visual grasping system according to claim 1, characterized in that: The data acquisition module further includes inputting a preset generative adversarial network into the occluded contour image of the occluded target object portion, and outputting a complete contour of the target object after completion. The determination of whether the target object is occluded includes: The network model is back-projected onto the target object’s images at each viewing angle, and the projection residual is calculated. If the projection residual is greater than a preset threshold, it is determined that the target object is occluded at that viewing angle.

6. The multi-machine flexible six-axis robot visual grasping system according to claim 1, characterized in that: The Transformer architecture includes: Encoder: Uses Vision Transformer (ViT) as the backbone network to split the RGB image into fixed-size tiles, linearly embed them, add positional encoding, and then process them through a multi-layer Transformer encoder block; Decoder: Receives the output of the encoder, gradually upsamples and fuses features from different levels. The Transformer decoder block uses a cross-attention mechanism to output a contour feature map with the same resolution as the input RGB image.

7. The multi-machine flexible six-axis robot visual grasping system according to claim 6, characterized in that: The Transformer architecture includes enhancing the edge gradients of the contour feature map through the Laplacian operator.

8. The multi-machine flexible six-axis robot visual grasping system according to claim 1, characterized in that: The data acquisition module includes: Use load cells to obtain mass data; Use an industrial CT scanner to obtain a two-dimensional density map; Acquire RGB images using an RGB-D camera Use stereo vision technology to obtain depth images; And use GPIO signal lines to connect the trigger interface of the industrial CT scanner and RGB-D camera.

9. The multi-machine flexible six-axis robot visual grasping system according to claim 1, characterized in that: The data acquisition module further includes aligning pixels in the acquired RGB image and depth image, including: Use the hand-eye calibration method to obtain the rotation matrix R and translation vector t of the RGB-D camera and depth sensor; The RGB pixel coordinate A is mapped to the coordinate system B of the depth camera through external parameters: B = R·A+t.

10. The multi-machine flexible six-axis robot visual grasping system according to claim 1, characterized in that: The six-axis robot includes end effectors of different models, and the end effectors of different models are respectively used to grasp different target objects.

Citation Information

Patent Citations

  • Protective edge point cloud hole repair method based on two-dimensional projection

    CN107610061A

  • Intelligent robot grabbing method based on three-dimensional vision

    CN111906782A

  • Deep learning-based six-degree-of-freedom mechanical arm grabbing pose detection method

    CN116277014A

  • Grasp generation using a variational autoencoder

    US20200361083A1

  • Image recognition method and apparatus, storage medium and electronic device

    WO2025060766A1