Vision-based data collection method for robot grasping training

By using Intel RealSense camera and visual tracking algorithm in the robot capture training data acquisition, we collect and synthesize the viewing image and hand posture information during the capture process of the operator, which solves the problem that traditional methods are difficult to collect complete data, and achieves low-cost and low-cost data acquisition and high generalization capabilities.

CN114782774BActive Publication Date: 2025-05-09ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210409528.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-19
Publication Date
2025-05-09
Estimated Expiration
2042-04-19

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively collect complete robots to capture training data, especially in dexterity grab tasks. Traditional methods are difficult to record hand posture information, and data collection costs are high and are not suitable for large-scale acquisition.

Method used

By building a simple data acquisition operation platform, using Intel RealSense binocular stereoscopic depth camera to collect first-person and third-person perspective image sequences of the operator during the capture process, combining frame-to-frame visual tracking method and palm tracking algorithm, the motion trajectory and hand joint posture information are estimated, and combined into a robot to capture training data set.

Benefits of technology

It realizes low-cost and low-cost robotic acquisition training data collection, complete data, suitable for research on various bionic clever grasping algorithms, and does not require complex training or precision equipment, reducing data acquisition costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114782774B_ABST
    Figure CN114782774B_ABST
Patent Text Reader

Abstract

A vision-based robot grasping training data collection method includes the following steps: Step 1: Collect the operator's first-person perspective and third-person perspective image sequences that change over time during the complete grasping process. Step 2: Estimate the grasping motion trajectory from the collected first-person perspective image sequence. Step 3: Estimate the hand joint posture information from the collected third-person perspective image sequence. Step 4: Combine the dual-perspective image group that changes over time with the corresponding grasping motion trajectory and hand joint posture information into robot grasping training data. The robot grasping training data collection method proposed in the present invention relies only on visual information, has a low collection cost, and the collected robot grasping training data is complete and has strong generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a vision-based robot grasping training data acquisition method. Background Art

[0002] With the development of artificial intelligence research, intelligent robots have gradually appeared in different fields and positions of human society, and intelligent robot research has also received more and more attention. Among the basic capabilities of intelligent robots, grasping ability is the most basic and important. From industrial robots that complete heavy picking and placing tasks to household robots that help the elderly with daily grasping tasks. Giving robots grasping capabilities is a long-term goal of the robotics discipline.

[0003] In the design of intelligent robot grasping function, the end effector is an important component. Different from simple end effectors, multi-jointed finger grippers are more compatible with objects designed and built for humans in the real world. Dexterous end effectors can perform more delicate and functional operations, and can pick up objects in a more reasonable posture, ready to use or place the object. Such characteristics make the research on dexterous grasping have greater potential and application prospects.

[0004] At present, the dexterous grasping algorithms are mainly driven by visual information. Due to the complexity of visual information and dexterous grasping functions, it is difficult to build models with traditional methods, and the grasping ability of novel objects is almost zero. Therefore, most researchers design dexterous grasping algorithms based on machine learning methods. How to obtain effective training data is the key to the research of dexterous grasping algorithms. The training data needs to retain the complete visual, trajectory and posture information of the grasp, and reduce the cost of data collection as much as possible.

[0005] To solve the above problems, there are methods to grasp the hand posture information through classification, and classify the complex hand posture into a certain number of types. However, the classification method still has certain limitations and cannot record the hand information completely. In addition, the data is highly unique and is not suitable for generalization to other grasping tasks. There are also methods that use various sensors such as data gloves to record hand joint information, but this type of data collection method is expensive and requires professional equipment and operators, which is not suitable for large-scale collection. Summary of the invention

[0006] The present invention overcomes the above-mentioned problems of the prior art and proposes a method for collecting robot grasping training data which relies only on visual information and has low collection cost. The robot grasping training data collected by this method is complete and has strong generalization ability.

[0007] The present invention first builds a simple data acquisition operation platform, and uses a wristband-type Intel RealSense binocular stereo depth camera and an Intel Real Sense binocular stereo depth camera directly above the operation platform to collect the operator's first-person and third-person perspective image sequences that change over time during the grasping process. Then, the frame-to-frame visual tracking method is used to estimate the motion trajectory that changes over time during the grasping process from the first-person perspective image. The palm tracking algorithm is then used to estimate the hand joint position that changes over time during the grasping process from the third-person perspective image, and the palm constraint is used to obtain the joint bending angle information. The above information is combined to form a robot grasping training data set.

[0008] The technical solution adopted by the present invention to solve the problems of the prior art is: a robot grasping data acquisition method based on vision, characterized in that it comprises the following steps:

[0009] Step 1: Collect the operator's first-person perspective and third-person perspective image sequences that change over time during the complete grasping process;

[0010] Step 2: Estimate the grasping motion trajectory from the collected first-person perspective image sequence;

[0011] Step 3: Estimate the hand joint posture information from the collected third-person perspective image sequence;

[0012] Step 4: Combine the time-varying dual-view image group with the corresponding grasping motion trajectory and hand joint posture information into robot grasping training data.

[0013] The step 1 specifically includes:

[0014] Step 1-1: First, use hard materials to build a cubic frame as the image collection space, and surround each side of the cube with green cloth to shield the collection environment;

[0015] Step 1-2: Fix one depth camera directly above the operator and fix the other depth camera on the operator's wrist using a camera wrist strap. Place the object to be grasped at the center of the bottom surface of the operating space. Complete the preparation of the grasping environment;

[0016] Step 1-3: After starting the recording, the operator stops recording after completing a complete grasping process. The two cameras respectively collect image sequences of different perspectives of the grasping process that change over time.

[0017] The step 2 specifically includes:

[0018] Step 2-1: First, use the SURF algorithm to detect the feature points of each captured first-person perspective image. The pixels in the image can be represented as I(x, y), where x is the horizontal coordinate of the pixel and y is the vertical coordinate of the pixel. The size of the box filter template is N×N, and the corresponding scale is σ=1.2×9 / N. The Hessian matrix of the pixel is defined as follows:

[0019]

[0020] Where L xx (x,σ) is the second-order derivative of the Gaussian filter L xy (x,σ),L yy The definition of (x,σ) is similar. Find the scale extreme value of the Hessian matrix, and the extreme point is the candidate feature point. Then interpolation operation is performed in scale space and image space to obtain stable feature points;

[0021] Step 2-2: Use the FLANN algorithm to obtain the point matching mapping set of feature points of two adjacent frames. First, for the SIFT key point of the previous frame, find the two key points in the next frame with the closest Euclidean distance to the point, divide the closer distance by the next closest distance, and when the ratio is lower than the threshold of 0.4, the key point is successfully matched;

[0022] Step 2-3: Use the least squares algorithm to calculate the rigid transformation between the matching key points and obtain the rigid transformation matrix between the two frames of images;

[0023] Step 2-4: Repeat steps 2 and 3 to obtain a rigid transformation matrix group between all images captured once;

[0024] Step 2-5: Multiply the rigid transformation matrix between the palm and the depth camera with the rigid transformation matrix group obtained in step 4 to obtain the palm motion trajectory that changes over time during the grasping process.

[0025] The step 3 specifically includes:

[0026] Step 3-1: First, use the MediaPipe Hands detection method based on machine learning to obtain the palm information in each captured third-person perspective image;

[0027] Step 3-2: extracting the three-dimensional coordinate information of 21 palm joints from the obtained palm information;

[0028] Step 3-3: Use the constraint relationship between the joints of the palm to obtain the bending angle information of each joint

[0029]

[0030] Among them, Vec(n) represents the normal vector of the palm plane, Vec(a,b) represents the vector from joint point a to joint point b, A x represents the bending angle of the joint point, and x represents the joint point number;

[0031] Step 3-4: Repeat steps 2 and 3 to generate the palm posture information that changes over time during a complete grasp.

[0032] The advantages of the present invention are: the present invention collects images from two perspectives during the grasping process, and can basically restore the motion trajectory during the grasping process, as well as complete information such as the spatial position and bending angle of each joint of the hand through the images, which is suitable for various bionic dexterous grasping algorithm research, and can also be extended to the research of ordinary parallel two-finger grasping algorithms, with complete data retention and strong generalization ability. In addition, the present invention is completely based on visual acquisition of human grasping demonstration data, and the operator does not need to undergo complex training or wear sophisticated motion capture equipment, which reduces the cost of data acquisition. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a flow chart of the implementation scheme of the present invention;

[0034] Figure 2 It is a schematic diagram of the image acquisition process of the present invention;

[0035] Figure 3 It is a schematic diagram of palm joint labeling defined in the present invention. DETAILED DESCRIPTION

[0036] The following is a further detailed description of the present invention in conjunction with the accompanying drawings:

[0037] The vision-based robot grasping training data acquisition method of the present invention is specifically implemented as follows:

[0038] Step 1: Collect the operator's first-person perspective and third-person perspective image sequences that change over time during the complete grasping process, such as Figure 1 As shown;

[0039] Step 1-1: First, use hard materials to build a cubic frame as the image collection space, and surround each side of the cube with green cloth to shield the collection environment;

[0040] Step 1-2: Fix one depth camera directly above the operator and fix the other depth camera on the operator's wrist using a camera wrist strap. Place the object to be grasped at the center of the bottom surface of the operating space. Complete the preparation of the grasping environment;

[0041] Step 1-3: After starting the recording, the operator stops recording after completing a complete grasping process. The two cameras respectively collect image sequences of different perspectives of the grasping process that change over time.

[0042] Step 2: Estimate the grasping motion trajectory from the collected first-person perspective image sequence;

[0043] Step 2-1: First, use the SURF algorithm to detect the feature points of each captured first-person perspective image. Find the scale extreme value of the Hessian matrix of each image pixel. The extreme point is the candidate feature point, and then perform interpolation operations in the scale space and image space to obtain stable feature points;

[0044] Step 2-2: Use the FLANN algorithm to obtain the point matching mapping set of feature points of two adjacent frames. First, for the SIFT key point of the previous frame, find the two key points in the next frame with the closest Euclidean distance to the point, divide the closer distance by the next closest distance, and when the ratio is lower than the threshold of 0.4, the key point is successfully matched;

[0045] Step 2-3: Use the least squares algorithm to calculate the rigid transformation between the matching key points and obtain the rigid transformation matrix between the two frames of images;

[0046] Step 2-4: Repeat steps 2 and 3 to obtain a rigid transformation matrix group between all images captured once;

[0047] Step 2-5: Multiply the rigid transformation matrix between the palm and the depth camera with the rigid transformation matrix group obtained in step 4 to obtain the palm motion trajectory that changes over time during the grasping process.

[0048] Step 3: Estimate the hand joint posture information from the collected third-person perspective image sequence;

[0049] Step 3-1: First, use the MediaPipe Hands detection method based on machine learning to obtain the palm information in each captured third-person perspective image;

[0050] Step 3-2: extracting the three-dimensional coordinate information of 21 palm joints from the obtained palm information;

[0051] Step 3-3: Use the constraint relationship between the joints of the palm to obtain the bending angle information of each joint. Vec(a,b) represents the vector from point a to point b. Cross product Vec(0,17) and Vec(0,5) to obtain the normal vector of the palm plane. The bending angles of joints 1, 5, 9, 13, and 17 are the angles formed by the vector from joint 0 to the palm normal vector. The bending angles of joints 2, 3, 6, 7, 10, 11, 14, 15, 18, and 19 are the angles between the vectors formed by the joint point and the two adjacent joints;

[0052] Step 3-4: Repeat steps 2 and 3 to generate the palm posture information that changes over time during a complete grasp.

[0053] Step 4: Combine the time-varying dual-view image group with the corresponding grasping motion trajectory and hand joint posture information into robot grasping training data.

[0054] It should be emphasized that the contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept, and the protection scope of the present invention should not be regarded as limited to the specific forms described in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.

Claims

1. A vision-based robot grabbing data collection method, characterized in that: The following steps are included: Step 1: Collect the operator's first-person perspective and third-person perspective image sequences that change over time during the complete grasping process; Step 2: Estimate the grasping motion trajectory from the collected first-person perspective image sequence; Step 3: Estimate the hand joint posture information from the collected third-person perspective image sequence; specifically including: 3.1): First, the MediaPipe Hands detection method based on machine learning is used to obtain the palm information in each captured third-person perspective image; 3.2): Extract the three-dimensional coordinate information of 21 palm joints from the obtained palm information; 3.3): Use the constraint relationship between the joints of the palm to obtain the bending angle information of each joint; Among them, Vec(n) represents the normal vector of the palm plane, Vec(a,b) represents the vector from joint point a to joint point b, A x represents the bending angle of the joint point, and x represents the joint point number; 3.4): Repeat steps 3.2 and 3.3 to generate the palm posture information that changes over time during a complete grasp; Step 4: Combine the time-varying dual-view image group with the corresponding grasping motion trajectory and hand joint posture information into robot grasping training data.

2. The method for collecting robot data based on vision according to claim 1, characterized in that: Step 1 specifically includes: Step 1-1: First, use hard materials to build a cubic frame as the image collection space, and surround each side of the cube with green cloth to shield the collection environment; Step 1-2: Fix one depth camera directly above the operator and fix another depth camera on the operator's wrist using a camera wrist strap; place the object to be grasped at the center of the bottom surface of the operating space to complete the preparation of the grasping environment; Step 1-3: After starting recording, the operator stops recording after completing a complete grasping process; the two cameras respectively collect image sequences of different perspectives of the grasping process that change over time.

3. The method for collecting robot data based on vision according to claim 1, characterized in that: The step 2 specifically includes: 2.1): First, use the SURF algorithm to detect the feature points of each captured first-person perspective image. The pixels in the image are represented by I(x, y), where x is the horizontal coordinate of the pixel and y is the vertical coordinate of the pixel. The size of the box filter template is N×N, and the corresponding scale is σ=1.2×9 / N. The Hessian matrix of the pixel is defined as follows: Where L xx (x,σ) is the second-order derivative of the Gaussian filter L xy (x,σ),L yy The definition of (x,σ) is similar; find the scale extreme value of the Hessian matrix, the extreme point is the candidate feature point, and then perform interpolation operations in the scale space and image space to obtain stable feature points; 2.2): Use the FLANN algorithm to obtain the point matching mapping set of feature points of two adjacent frames of images; first, for the SIFT key point of the previous frame image, find the two key points in the next frame image that are closest to the point in Euclidean distance, divide the closer distance by the next closest distance, and when the ratio is lower than the threshold of 0.4, the key point matching is successful; 2.3): Use the least squares algorithm to calculate the rigid transformation between the matching key points and obtain the rigid transformation matrix between the two frames of images; 2.4): Repeat steps 2.2 and 2.3 to obtain a rigid transformation matrix group between all images captured once; 2.5): Multiply the rigid transformation matrix between the palm and the depth camera with the rigid transformation matrix group obtained in step 2.4 to obtain the palm motion trajectory that changes over time during the grasping process.

Citation Information

Patent Citations

  • Control method, system and device of mechanical arm vision positioning sorting grabbing

    CN111515945A

  • Robot disordered grabbing method and system based on machine vision and storage medium

    CN112070818A