Mechanical arm autonomous grabbing method and system based on image instance segmentation
Through the autonomous capture method of robotic arm based on image instance segmentation, multi-angle image acquisition and depth point cloud data processing are used to generate the optimal capture position, which solves the problem of insufficient generalization and stability of traditional methods in the kitchen environment, and achieves efficient and reliable tool capture.
Patent Information
- Application Number
- CN202510294656.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-01
AI Technical Summary
Traditional robotic arm visual grasping methods are insufficient in generalization and stability in kitchen environments, making them difficult to adapt to non-standardized appliances and complex backgrounds. The existing deep learning solutions have high computational load and limited response speed, which cannot meet real-time operation requirements.
The robotic arm autonomous grasping method based on image instance segmentation is adopted, and the optimal grasping position is generated through multi-angle image acquisition, depth point cloud data preprocessing and principal component analysis, and combined with visual sensors and robotic arm coordinated optimization, the autonomous grasping of the target object is achieved.
It improves the environmental adaptability and operational reliability of the grab, reduces the computational complexity, improves the crawling success rate and response speed, and is suitable for diversified tool grabbing in unstructured kitchen environments.
Smart Images

Figure CN120235935A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and robotic arm grasping, and particularly relates to a robotic arm autonomous grasping method and system based on image instance segmentation. Background Art
[0002] As an important application scenario of the integration of artificial intelligence and robotic technology, the core link of an unmanned intelligent kitchen lies in realizing the autonomous grasping operation of kitchen utensils and food ingredients. In this scenario, the robotic arm system needs to accurately identify and manipulate objects with various shapes such as pots, pans, bowls, spoons, and condiment containers, which poses stringent requirements on the visual perception system. The current mainstream RGBD depth camera has become the core sensor for environmental perception because it can synchronously obtain the color information and three-dimensional spatial data of objects. However, how to construct an efficient and reliable grasping pose estimation system based on depth camera data remains the key technical bottleneck restricting the improvement of the intelligent level of unmanned kitchens.
[0003] Traditional visual grasping algorithms mainly rely on three-dimensional point cloud matching technology. By pre-building a high-precision three-dimensional model of the target object and performing feature matching between the real-time collected point cloud data and the model library, the grasping pose is then solved. Such methods show good positioning accuracy in the grasping scenarios of standardized workpieces, but their technical limitations are significantly magnified in the kitchen scenario: on the one hand, the utensils in the kitchen environment have diverse shapes and non-standard features, and it is neither realistic nor economical to pre-establish a three-dimensional model library of all possible objects; on the other hand, tableware is often accompanied by complex situations such as oil stains and partial occlusion during use, resulting in a decline in the quality of point cloud data and seriously affecting the reliability of model matching. More critically, this method cannot adapt to the dynamic introduction of new kitchen utensils, severely restricting the environmental adaptability of the system.
[0004] In recent years, the emerging deep learning grasping algorithms have provided new ideas for breaking through the limitations of traditional methods. End-to-end models represented by GraspNet directly map the scene point cloud to the grasping pose and obtain generalization ability through a large amount of data training. Although such methods show advantages in grasping unknown objects, they also expose multiple defects in the actual kitchen scenario application: First, the network model is sensitive to point cloud noise and is prone to pose estimation deviation in complex situations such as tableware stacking and reflective surfaces, resulting in grasping action failures; second, the existing algorithms lack the ability to identify objects at a fine-grained level and are difficult to adapt differential grasping strategies for specific types of tableware (such as fragile glassware or frying pans that require special clamping); third, the full-scene processing of point cloud data leads to a sharp increase in computational load, and the single inference time is generally long, making it difficult to meet the real-time operation requirements in a dynamic environment. These technical shortcomings make the grasping success rate of existing deep learning solutions in real kitchen scenarios still not meet the engineering application requirements, severely restricting the practical process of the system.
[0005] The existing technical system has not yet achieved an effective balance between generalization and stability. Although traditional methods can ensure the grasping accuracy of specific objects, they sacrifice environmental adaptability. Deep learning solutions have expanded the range of target objects, but have inherent problems such as insufficient reliability and limited response speed. Therefore, there is an urgent need to construct a new visual grasping algorithm framework to improve the real-time response speed and operation robustness of the system while maintaining the generalization ability, and ultimately achieve the intelligent and safe grasping of multi-category objects in an unmanned kitchen environment. Summary of the Invention
[0006] The purpose of the present invention is to solve the limitations such as poor generalization and stability in the visual grasping process of traditional robotic arms. Therefore, a robotic arm autonomous grasping method and system based on image instance segmentation are proposed. The present invention has the characteristics of small computational complexity and insensitivity to environmental changes. The robotic arm can autonomously move to the best position to observe the target object, and then grasp the target object placed at any position in space, and the grasping accuracy is also improved.
[0007] The present invention adopts the following technical solutions to achieve the purpose:
[0008] A robotic arm autonomous grasping method based on image instance segmentation, including:
[0009] Collect multi-angle images of the target object through the visual sensor of the robotic arm. According to the object mask information output by the target detection and segmentation model, adjust the robotic arm to the observation position of the target object, obtain the depth point cloud data containing the complete three-dimensional information of the target object, and perform preprocessing;
[0010] Perform principal component analysis on the depth point cloud data of the target object, extract the principal direction vector representing the spatial distribution of the target object, and construct the principal direction coordinate system of the target object;
[0011] Input the depth point cloud data of the target object into the grasping pose detection model to generate a set of candidate grasping poses. By calculating the spatial consistency parameters between each candidate grasping pose in the set and the principal direction coordinate system of the target object, filter out the optimal grasping pose that meets the preset direction constraints;
[0012] Convert the optimal grasping pose to the coordinate data in the base coordinate system of the robotic arm, and drive the robotic arm to perform the grasping action to complete the autonomous grasping operation of the target object.
[0013] Specifically, the method includes the following steps:
[0014] S1. The robotic arm moves to the initial observation point, takes a picture of the target object through the visual sensor, and obtains the first RGB image and the first depth map of the target object;
[0015] S2. Input the first RGB image into the target detection and segmentation model to perform category recognition and segmentation of the target object. According to the object mask information obtained by segmentation, combine it with the first depth map projection to obtain the three-dimensional position data of the target object. After the robotic arm moves to the observation position corresponding to the three-dimensional position data, take another picture of the target object to obtain the second RGB image and the second depth map of the target object.
[0016] S3. Input the second RGB image into the target detection and segmentation model. According to the object mask information obtained by segmentation, combine it with the second depth map to generate the depth point cloud data of the target object, and perform denoising processing on the depth point cloud data.
[0017] S4. Input the depth point cloud data into the grasping pose detection model to generate multiple candidate grasping poses. From the multiple candidate grasping poses, select the candidate grasping poses with the preset score ranking quantity to form a candidate grasping pose set.
[0018] S5. Perform principal component analysis on the depth point cloud data, extract the first three main variance directions of the depth point cloud data, generate three mutually orthogonal representative vectors, and form a rotation matrix describing the main direction characteristics of the target object. From the candidate grasping pose set, select the candidate grasping pose closest to the rotation matrix as the optimal grasping pose. The optimal grasping pose is the target position data corresponding to the target object.
[0019] S6. Send the optimal grasping pose to the robotic arm. Based on the hand-eye calibration result, the corresponding target position data is converted into the coordinate data in the base coordinate system of the robotic arm. The robotic arm adjusts its pose and moves based on the coordinate data to complete the autonomous grasping operation of the target object.
[0020] Preferably, in steps S1 and S2, when obtaining the first RGB image and the first depth map of the target object, and when obtaining the second RGB image and the second depth map of the target object, the time synchronization calibration module is used to control the timing of the RGB image acquisition unit and the depth map acquisition unit of the visual sensor, so that the acquisition timestamps of the RGB image and the depth map are aligned or the deviation is controlled within the preset threshold range, and image data fusion processing is performed based on the synchronized timestamps to eliminate the distortion of the point cloud spatial coordinates caused by timing asynchrony.
[0021] Further, in steps S2 and S3, the YOLOv8-seg model is used as the target detection and segmentation model to perform category recognition and segmentation of the target object based on the YOLOv8-seg model.
[0022] When the first RGB image is input, according to the corresponding center point coordinates in the object mask information obtained by segmentation, combined with the first depth map projection, the three-dimensional position data of the target object is obtained. Subsequently, the robotic arm moves to the central position directly in front of the target object represented by the three-dimensional position data as the observation position;
[0023] When the second RGB image is input, according to all the coordinates in the object mask information obtained by segmentation, all the coordinates are used for projection and combined with the second depth map to generate the depth point cloud data of the target object.
[0024] Specifically, in step S2, the center point coordinates of the object mask information are combined with the first depth map projection to obtain the three-dimensional position data of the target object, which specifically includes the following calculation formula:
[0025] Z = D(u, v)
[0026]
[0027] In the formula, X, Y, and Z respectively represent the corresponding axis coordinates of the three-dimensional coordinates of the target object; f x 、f y 、c x and c y are all camera internal parameters, where f x and f y are the focal length parameters of the camera in the x-axis and y-axis directions, c x and c y are the horizontal and vertical coordinates of the camera optical center in the pixel coordinate system; u and v are the center point coordinates of the object mask information, and D(u, v) represents the depth value corresponding to the center point coordinates.
[0028] Specifically, when the robotic arm moves to the observation position corresponding to the three-dimensional position data, the observation position is determined based on the coordinates (X, Y, 0) in the camera coordinate system, that is, when moving to Z = 0, no further movement is made in the depth direction represented by the coordinate value Z.
[0029] Preferably, in step S3, the conditional Euclidean clustering algorithm is used to denoise the depth point cloud data of the target object, specifically: performing a clustering operation on the depth point cloud data, and selecting the cluster with the largest number of point clouds from the clustering results as the depth point cloud data of the target object after denoising.
[0030] Furthermore, in step S4, the grasping pose detection model is a GPD (Grasp Pose Detection) algorithm model improved by prior constraints; the prior constraints are prior distance constraints based on the preset direction of the target object. After presetting the distance threshold between the grasping end of the robotic arm and the target object, the prior distance constraints are divided into grasping of distant objects with a grasping distance greater than or equal to the distance threshold and grasping of close objects with a distance less than the distance threshold.
[0031] Specifically, under the prior distance constraint for grasping distant objects, the proximity between the z-axis direction of the candidate grasping poses generated by the improved GPD algorithm model and the camera depth direction meets the preset first constraint requirement; under the prior distance constraint for grasping close objects, the proximity between the z-axis direction of the candidate grasping poses generated by the improved GPD algorithm model and the camera depth direction meets the preset second constraint requirement.
[0032] The present invention also provides a robotic arm autonomous grasping system based on image instance segmentation, including:
[0033] A multi-view image acquisition module configured to collect image data of a target object from multiple angles through a vision sensor at the end of the robotic arm, and adjust the observation pose of the robotic arm based on the object mask information output by the object detection and segmentation model to obtain depth point cloud data containing the complete three-dimensional information of the target object;
[0034] A point cloud preprocessing and main direction extraction module configured to preprocess the depth point cloud data and extract the main direction vector representing the spatial distribution of the target object through principal component analysis to construct a target object coordinate system based on the main direction;
[0035] A grasping pose generation and optimization module, including:
[0036] A grasping pose prediction unit configured to input the preprocessed depth point cloud data into a grasping pose detection model to generate multiple candidate grasping poses;
[0037] A direction consistency screening unit configured to calculate the spatial consistency parameters of each candidate grasping pose with the target object coordinate system and screen the optimal grasping pose that meets the preset direction constraint;
[0038] A motion planning and execution module configured to convert the optimal grasping pose to the robotic arm base coordinate system, generate the motion trajectory of the robotic arm joints, and further drive the robotic arm to perform a grasping action.
[0039] In summary, due to the adoption of the present technical solution, the beneficial effects of the present invention are as follows:
[0040] Through the collaborative optimization of the robotic arm's autonomous positioning and visual perception, the present invention significantly improves the environmental adaptability and operation reliability of target object grasping. The visual sensor mounted on the robotic arm can dynamically adjust the observation pose to ensure that the target object can be in the best observation perspective, effectively overcoming the problems of point cloud missing or image occlusion caused by the randomness of the object's spatial pose, maintaining high precision and integrity in the object segmentation process, and greatly reducing the risk of recognition errors caused by complex background interference. By integrating the grasping pose detection model with the principal component analysis (PCA) direction constraint mechanism of the point cloud, on the basis of making full use of the generalization ability of the deep neural network, prior geometric feature knowledge is introduced to verify the physical rationality of the grasping pose, which not only expands the adaptation range of the algorithm to unknown objects but also ensures the stability of the grasping pose through multi-dimensional constraints, effectively solving the imbalance problem between the robustness and generalization of traditional single methods.
[0041] Compared with the existing grasping algorithms based on full-scene point cloud processing, the present invention adopts a target-oriented data processing method, which only needs to perform pose calculation on the point cloud of the segmented target object, avoiding the computational load of redundant environmental data. This focused processing strategy greatly reduces the computational complexity of the algorithm, enables faster real-time inference speed on typical embedded platforms, and at the same time, the grasping success rate can be increased to more than 95% because the interference of environmental noise on pose estimation is eliminated. In addition, the present invention shows strong robustness to environmental variables such as light changes, cluttered backgrounds, and object surface reflections, and can operate stably without relying on prior configurations of specific scenarios, especially suitable for the grasping requirements of diverse utensils placed unstructured in the kitchen environment, and achieving a closed-loop grasping operation with high precision and low latency under limited computational resources. Brief Description of the Drawings
[0042] Figure 1 It is a schematic diagram of the overall steps of the method of the present invention;
[0043] Figure 2 It is a detailed example flowchart of the method of the present invention. Detailed Embodiments
[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. The processes of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0045] Accordingly, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0046] Embodiment 1
[0047] A robotic arm autonomous grasping method based on image instance segmentation, and the core points of this method are as follows:
[0048] Collect multi-angle images of the target object through the visual sensor of the robotic arm. According to the object mask information output by the target detection and segmentation model, adjust the robotic arm to the observation position of the target object, obtain the depth point cloud data containing the complete three-dimensional information of the target object, and perform preprocessing.
[0049] Perform principal component analysis on the depth point cloud data of the target object, extract the principal direction vector characterizing the spatial distribution of the target object, and construct the principal direction coordinate system of the target object.
[0050] Input the depth point cloud data of the target object into the grasping pose detection model to generate a set of candidate grasping poses. By calculating the spatial consistency parameters between each candidate grasping pose in the set and the principal direction coordinate system of the target object, select the optimal grasping pose that meets the preset direction constraints.
[0051] Convert the optimal grasping pose to the coordinate data in the base coordinate system of the robotic arm, and drive the robotic arm to perform the grasping action to complete the autonomous grasping operation of the target object.
[0052] Based on the above method points, this embodiment introduces the specific situation during the method execution through relatively detailed step diagrams; refer to Figure 1 for the diagram, and the general steps of this method are as follows:
[0053] S1. The robotic arm moves to the initial observation point, and takes a picture of the target object through the visual sensor to obtain the first RGB image and the first depth map of the target object.
[0054] S2. Input the first RGB image into the target detection and segmentation model to perform category recognition and segmentation of the target object. According to the object mask information obtained by the segmentation, combine it with the first depth map to project and obtain the three-dimensional position data of the target object. After the robotic arm moves to the observation position corresponding to this three-dimensional position data, take a picture of the target object again to obtain the second RGB image and the second depth map of the target object.
[0055] S3. Input the second RGB image into the target detection and segmentation model. According to the object mask information obtained by segmentation, combine it with the second depth map to generate the depth point cloud data of the target object, and perform denoising processing on the depth point cloud data;
[0056] S4. Input the depth point cloud data into the grasping pose detection model to generate multiple candidate grasping poses. From the multiple candidate grasping poses, select the candidate grasping poses with the preset score ranking quantity to form a candidate grasping pose set;
[0057] S5. Perform principal component analysis on the depth point cloud data, extract the first three main variance directions of the depth point cloud data, generate three mutually orthogonal representative vectors, and form a rotation matrix describing the main direction characteristics of the target object; from the candidate grasping pose set, select the candidate grasping pose closest to this rotation matrix as the optimal grasping pose; the optimal grasping pose is the target position data corresponding to the target object;
[0058] S6. Send the optimal grasping pose to the robotic arm. Based on the hand-eye calibration result, the corresponding target position data is converted into the coordinate data in the base coordinate system of the robotic arm; the robotic arm adjusts its pose and moves based on this coordinate data to complete the autonomous grasping operation of the target object.
[0059] This embodiment will further introduce the details of each of the above steps. For the execution process of the involved method, reference can also be made to the detailed illustration in Figure 2 the detailed illustration.
[0060] As a preference of this embodiment, in steps S1 and S2, when obtaining the first RGB image and the first depth map of the target object, and when obtaining the second RGB image and the second depth map of the target object, the time synchronization calibration module is used to control the timing of the RGB image acquisition unit and the depth map acquisition unit of the visual sensor, so that the acquisition timestamps of the RGB image and the depth map are aligned or the deviation is controlled within a preset threshold range, and image data fusion processing is performed based on the synchronized timestamps to eliminate the distortion of the point cloud spatial coordinates caused by timing asynchrony; the misalignment or large deviation of the acquisition timestamps will lead to errors in the subsequent depth point cloud data, and further lead to errors in the calculation of the grasping pose of the robotic arm. Therefore, the problem of the basic accuracy of the data is solved by aligning the timestamps.
[0061] In steps S2 and S3 of this embodiment, the target detection and segmentation model uses the YOLOv8-seg model to perform class recognition and segmentation of target objects based on the YOLOv8-seg model. In the YOLOv8-seg model, it realizes the efficient segmentation of multi-scale targets by fusing the improved SPPF (Spatial Pyramid Pooling) module and the dynamic convolution layer. The model adopts a decoupled head structure to decouple the target detection and instance segmentation tasks, and respectively outputs the target bounding box and the pixel-level mask through a two-branch network. Aiming at problems such as target stacking and occlusion in the robotic arm grasping scenario, the model can introduce an attention mechanism to strengthen edge feature extraction and combine a weighted bidirectional feature pyramid network (BiFPN) to achieve multi-scale feature fusion, significantly improving the segmentation accuracy and instance discrimination of small target objects in complex backgrounds.
[0062] In this embodiment, the methods for obtaining the three-dimensional position data and depth point cloud data of the target object in steps S2 and S3 are similar. The difference is that in step S3, all coordinates in the object mask information are used, while in step S2, only the corresponding center point coordinates in the object mask information are used. Specifically: when inputting the first RGB image in step S2, according to the corresponding center point coordinates in the segmented object mask information, combined with the first depth map projection, the three-dimensional position data of the target object is obtained. Subsequently, the robotic arm moves to the central position directly in front of the target object represented by the three-dimensional position data as the observation position; when inputting the second RGB image in step S3, according to all coordinates in the segmented object mask information, all coordinates are used for projection and combined with the second depth map to generate the depth point cloud data of the target object.
[0063] In step S2 of this embodiment, the center point coordinates of the object mask information are combined with the first depth map projection to obtain the three-dimensional position data of the target object, which specifically includes the following calculation formula:
[0064] Z = D(u, v)
[0065]
[0066] In the formula, X, Y, and Z respectively represent the corresponding axis coordinates of the three-dimensional coordinates of the target object; f x , f y , c x and c y are all camera internal parameters, where f x and f y are the focal length parameters of the camera in the x-axis and y-axis directions, c x and c y are the horizontal and vertical coordinates of the camera optical center in the pixel coordinate system; u and v are the center point coordinates of the object mask information, and D(u, v) represents the depth value corresponding to the center point coordinates.
[0067] Moreover, for the robotic arm to subsequently move to the central position directly in front of the target object represented by the three-dimensional position data in step S2 as the observation position, this observation position does not refer to directly using the corresponding axis coordinates (X, Y, Z) of the target object's three-dimensional coordinates, but is determined based on the coordinates (X, Y, 0) in the camera coordinate system. That is, when the robotic arm moves to Z = 0, it no longer moves in the depth direction represented by the coordinate value Z.
[0068] As a preference of this embodiment, in step S3, a conditional Euclidean clustering algorithm is used to denoise the depth point cloud data of the target object. Point cloud denoising is a key guarantee for the robotic arm to achieve autonomous grasping with good results. This is because the point cloud of the target object after segmentation and projection will contain a large number of noise points, which will significantly affect the calculation accuracy of the final grasping posture. The conditional Euclidean clustering algorithm can dynamically adjust the clustering parameters to adapt to the noise interference in complex scenarios. It first divides the original point cloud data in space and constructs a dynamic neighborhood search range in combination with local geometric features (such as curvature, normal vector, etc.). By setting an adaptive threshold, it effectively distinguishes the surface points of the target object from the background noise points and excludes the outlier points caused by sensor errors or environmental interference.
[0069] When this embodiment performs clustering operations on the depth point cloud data, the cluster with the largest number of point clouds is selected from the clustering results as the depth point cloud data of the target object after denoising. Specifically, the algorithm will preferentially retain the point cloud clusters with continuous spatial distribution and high density, and screen out the main cluster representing the true geometric features of the target object by iteratively optimizing the clustering center and boundary conditions. This processing method significantly suppresses the interference of abnormal points introduced by occlusion, reflection, or motion blur while retaining the detailed features of the object. The point cloud data after denoising has higher spatial consistency, providing a stable and reliable three-dimensional structure information basis for the accurate calculation of the subsequent grasping posture.
[0070] In step S4 of this embodiment, the grasping pose detection model is a GPD algorithm model improved by prior constraints. For the traditional GPD algorithm model, its core process includes three main steps: (1) generating a large number of candidate grasps; (2) using a CNN network to classify the candidate grasps into valid grasping postures and invalid grasping postures; (3) clustering the feasible grasps with similar geometric characteristics. This embodiment makes targeted improvements on this basis, that is, by prior distance constraints based on the preset direction of the target object, after presetting the distance threshold between the grasping end of the robotic arm and the target object, the prior distance constraints are divided into grasping of distant objects with a grasping distance greater than or equal to the distance threshold and grasping of close objects with a distance less than the distance threshold.
[0071] Specifically, under the prior distance constraint for grasping distant objects, the proximity between the z-axis direction of the candidate grasping poses generated by the improved GPD algorithm model and the camera depth direction meets the preset first constraint requirement, which requires that the z-axis direction of the grasping pose must be close to the camera depth direction. This can be achieved by making the cosine value of the z-axis direction of the candidate grasping pose and the z-axis direction of the camera greater than 0.98.
[0072] Under the prior distance constraint for grasping nearby objects, the proximity between the z-axis direction of the candidate grasping poses generated by the improved GPD algorithm model and the camera depth direction meets the preset second constraint requirement, which only requires that the z-axis direction of the grasping pose be as close as possible to the camera depth direction, without being as strict as for distant objects. This can be achieved by making the cosine value of the z-axis direction of the candidate grasping pose and the z-axis direction of the camera greater than 0.75.
[0073] This improvement significantly optimizes the screening mechanism of candidate grasping poses by introducing constraint conditions based on the directionality of the grasping pose. For distant targets, the angle between the z-axis of the grasping pose and the camera depth direction is strictly restricted to ensure that the grasping direction is highly aligned with the observation perspective. For nearby targets, the constraint is appropriately relaxed to enhance the flexibility of pose generation while ensuring grasping stability. This method effectively eliminates two types of invalid candidates through prior distance constraints: one is the illegal pose where the grasping pose penetrates obstacles (such as the container wall or support surface); the other is the pose where the z-axis direction deviates severely from the observation perspective, resulting in insufficient reachability of the robotic arm's end effector. This method improves the efficiency of eliminating invalid grasping proposals, reduces the computational time consumption in the pose generation stage, thereby enhancing the robustness and accuracy of the algorithm, and further effectively improving the real-time performance and reliability of the robotic arm.
[0074] In this embodiment, after the improved GPD algorithm model generates multiple candidate grasping poses, the scores of these candidate grasping poses are uniformly normalized, and the top 20 candidate grasping poses in terms of scores are selected to form the set for subsequent determination of the optimal grasping pose, thereby realizing the final autonomous grasping function of the robotic arm.
[0075] Embodiment 2
[0076] Based on the method process in Embodiment 1, corresponding to it, this embodiment provides a robotic arm autonomous grasping system based on image instance segmentation, including:
[0077] A multi-view image acquisition module, configured to collect image data of the target object from multiple angles through the vision sensor at the end of the robotic arm, and adjust the observation pose of the robotic arm based on the object mask information output by the object detection and segmentation model to obtain depth point cloud data containing the complete three-dimensional information of the target object;
[0078] The point cloud preprocessing and principal direction extraction module is configured to preprocess the depth point cloud data, extract the principal direction vector characterizing the spatial distribution of the target object through principal component analysis, and construct a target object coordinate system based on the principal direction;
[0079] The grasping pose generation and optimization module includes:
[0080] The grasping pose prediction unit is configured to input the preprocessed depth point cloud data into a grasping pose detection model to generate multiple candidate grasping poses;
[0081] The direction consistency screening unit is configured to calculate the spatial consistency parameters between each candidate grasping pose and the target object coordinate system, and screen the optimal grasping pose that meets the preset direction constraints;
[0082] The motion planning and execution module is configured to convert the optimal grasping pose to the base coordinate system of the robotic arm, generate the motion trajectory of the robotic arm joints, and then drive the robotic arm to perform the grasping action.
[0083] Each module and unit in the above system of this embodiment can be implemented in the form of a computer system. Such computer systems all include a memory, a processor, and computer programs / instructions stored on the memory. When the processor executes the computer programs / instructions, the relevant steps of the robotic arm autonomous grasping method based on image instance segmentation can be realized.
[0084] In this embodiment, these computer programs / instructions can also be stored in any form of computer-readable storage medium that can guide a computer or other programmable data processing device to work in a specific manner, so that the computer programs / instructions stored in the computer-readable storage medium generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in one or more steps of the robotic arm autonomous grasping method.
Claims
1. A robotic arm autonomous grasping method based on image instance segmentation, characterized in that: include: The robot arm's visual sensor collects multi-angle images of the target object, adjusts the robot arm to the target object's observation position based on the object mask information output by the target detection and segmentation model, obtains the deep point cloud data containing the target object's complete three-dimensional information, and performs preprocessing; Perform principal component analysis on the depth point cloud data of the target object, extract the main direction vector representing the spatial distribution of the target object, and construct the main direction coordinate system of the target object; The depth point cloud data of the target object is input into the grasping posture detection model to generate a set of candidate grasping postures. The optimal grasping posture that meets the preset direction constraints is screened out by calculating the spatial consistency parameters of each candidate grasping posture in the set and the main direction coordinate system of the target object. The optimal grasping posture is converted into coordinate data in the robot base coordinate system, and the robot is driven to perform the grasping action to complete the autonomous grasping operation of the target object.
2. The robotic arm autonomous grasping method according to claim 1, characterized in that: The method comprises the following steps: S1, the robot arm moves to the initial observation point, shoots the target object through the visual sensor, and obtains the first RGB image and the first depth map of the target object; S2, inputting the first RGB image into the target detection and segmentation model, performing category recognition and segmentation of the target object, and obtaining three-dimensional position data of the target object based on the object mask information obtained by segmentation and combining the first depth map projection; After the robot arm moves to the observation position corresponding to the three-dimensional position data, it photographs the target object again to obtain a second RGB image and a second depth map of the target object; S3, inputting the second RGB image into the target detection and segmentation model, generating depth point cloud data of the target object according to the object mask information obtained by segmentation and combining the second depth map, and performing denoising on the depth point cloud data; S4, inputting the depth point cloud data into a grasping posture detection model to generate multiple candidate grasping postures, and selecting a preset number of candidate grasping postures with a ranking of scores from the multiple candidate grasping postures to form a candidate grasping posture set; S5. Perform principal component analysis on the depth point cloud data, extract the first three main variance directions of the depth point cloud data, generate three mutually orthogonal representative vectors, and form a rotation matrix that describes the main direction characteristics of the target object; from the candidate grasping posture set, select the candidate grasping posture closest to the rotation matrix as the optimal grasping posture; the optimal grasping posture is the target position data corresponding to the target object; S6. The optimal grasping posture is sent to the robotic arm. Based on the hand-eye calibration result, the corresponding target position data is converted into coordinate data in the robotic arm base coordinate system; the robotic arm adjusts and moves its posture based on the coordinate data to complete the autonomous grasping operation of the target object.
3. The robot arm autonomous grasping method according to claim 2, characterized in that: In step S1 and step S2, when acquiring the first RGB image and the first depth map of the target object, and when acquiring the second RGB image and the second depth map of the target object, the time synchronization calibration module is used to perform timing control on the RGB image acquisition unit and the depth map acquisition unit of the visual sensor, so that the acquisition timestamps of the RGB image and the depth map are aligned or the deviation is controlled within a preset threshold range, and image data fusion processing is performed based on the synchronization timestamp to eliminate the distortion of the point cloud space coordinates caused by timing asynchrony.
4. The robot arm autonomous grasping method according to claim 2, characterized in that: In step S2 and step S3, the target detection and segmentation model adopts the YOLOv8-seg model, and the category recognition and segmentation of the target object are performed based on the YOLOv8-seg model; When the first RGB image is input, the three-dimensional position data of the target object is obtained according to the coordinates of the center point corresponding to the object mask information obtained by segmentation, combined with the projection of the first depth map, and the robot arm then moves to the center position in front of the target object represented by the three-dimensional position data as the observation position; When the second RGB image is input, all coordinates in the object mask information obtained by segmentation are used as projections and then combined with the second depth map to generate depth point cloud data of the target object.
5. The robot arm autonomous grasping method according to claim 4, characterized in that: In step S2, the coordinates of the center point of the object mask information are combined with the first depth map projection to obtain the three-dimensional position data of the target object, which specifically includes the following calculation formula: Z=D(u,v) Where X, Y, and Z represent the corresponding axis coordinates of the three-dimensional coordinates of the target object; x 、f y 、c x and c y All are camera internal parameters, among which f x and f y is the focal length parameter of the camera in the x-axis and y-axis directions, c x and c y are the horizontal and vertical coordinates of the camera optical center in the pixel coordinate system; u and v are the center point coordinates of the object mask information, and D(u,v) represents the depth value corresponding to the center point coordinates.
6. The robot arm autonomous grasping method according to claim 5, characterized in that: When the robot moves to the observation position corresponding to the three-dimensional position data, the observation position is determined based on the coordinates (X, Y, 0) in the camera coordinate system, that is, when it moves to Z=0, it no longer moves in the depth direction represented by the coordinate value Z.
7. The robot arm autonomous grasping method according to claim 2, characterized in that: In step S3, the conditional Euclidean clustering algorithm is used to denoise the depth point cloud data of the target object. Specifically, a clustering operation is performed on the depth point cloud data, and the cluster with the largest number of point clouds is selected from the clustering results as the depth point cloud data of the target object after denoising.
8. The robot arm autonomous grasping method according to claim 2, characterized in that: In step S4, the grasping posture detection model is a GPD algorithm model improved by a priori constraints; the priori constraints are a priori distance constraints based on a preset direction of the target object. After presetting the distance threshold between the grasping end of the robotic arm and the target object, the priori distance constraint is divided into grasping long-distance objects with a grasping distance greater than or equal to the distance threshold and grasping close-distance objects with a grasping distance less than the distance threshold.
9. The robot arm autonomous grasping method according to claim 8, characterized in that: Under the prior distance constraint of long-distance object grasping, the proximity between the z-axis direction of the candidate grasping pose generated by the improved GPD algorithm model and the camera depth direction meets the preset first constraint requirement; Under the prior distance constraint of close-range object grasping, the degree of proximity between the z-axis direction of the candidate grasping pose generated by the improved GPD algorithm model and the camera depth direction meets the preset second constraint requirement.
10. A robotic arm autonomous grasping system based on image instance segmentation, characterized in that: include: A multi-view image acquisition module is configured to collect image data of the target object from multiple angles through the visual sensor at the end of the robotic arm, and adjust the observation posture of the robotic arm based on the object mask information output by the target detection and segmentation model to obtain depth point cloud data containing complete three-dimensional information of the target object; A point cloud preprocessing and main direction extraction module is configured to preprocess the depth point cloud data, extract the main direction vector representing the spatial distribution of the target object through principal component analysis, and construct a target object coordinate system based on the main direction; Grasping posture generation and optimization module, including: A grasping posture prediction unit is configured to input the preprocessed depth point cloud data into a grasping posture detection model to generate a plurality of candidate grasping postures; A direction consistency screening unit is configured to calculate the spatial consistency parameters of each candidate grasping posture and the target object coordinate system, and screen the optimal grasping posture that meets the preset direction constraint; The motion planning and execution module is configured to convert the optimal grasping posture into the robot base coordinate system, generate the motion trajectory of the robot joint, and then drive the robot to perform the grasping action.
Citation Information
Patent Citations
Object grabbing method applied to intelligent equipment, intelligent equipment and storage medium
CN114037753A
Rotary correction photographing method and device for environmental test and medium
CN116668814A
Mechanical arm control method, device and equipment based on machine vision and storage medium
CN116690568A
Object recognition and pose estimation method for grabbing robot
CN117036470A
2D image and 3D point cloud combined mechanical arm target grabbing detection system and method
CN117437216A
Cited By
Collaborative robot target identification and grabbing attitude planning system
CN120620227A
Mechanical arm grabbing method based on self-adaptive updating of stimulation control
CN121105001A
Space robot sensing system and method based on active vision strategy
CN121200028A
Control method for boron neutron capture therapy system beam limiting device maintenance robot
CN121245779A
Navel orange grabbing pose estimation method based on multi-feature segmentation and visual hedgehog algorithm
CN121353348A