Systems and methods for improving 3D eye-hand coordination accuracy of a robotic system
By dividing the movement of the robotic arm into small steps and utilizing a 3D machine vision system and neural network technology, the challenge of achieving sub-millimeter eye-hand coordination accuracy for robotic systems in 3D space was addressed, enabling high-precision object positioning.
Patent Information
- Application Number
- CN202210637151.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-05-23
- Filing Date
- 2022-06-07
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2042-06-07
AI Technical Summary
Existing robotic systems face challenges in achieving submillimeter eye-hand coordination accuracy in 3D space, mainly due to positioning errors of the robotic arm and camera, 3D visual measurement errors, and errors in the calibration target, which lead to overall system errors and limit the operational accuracy of the robotic system.
The movement of the robot arm is divided into multiple small steps, and the displacement of each small step is less than or equal to the predetermined maximum displacement value. A 3D machine vision system is used to capture images and determine the posture of the object in each small step. A segmentation neural network and a structured light projector are combined to generate a 3D point cloud. The transformation matrix is used to transform the posture from the camera coordinate system to the robot base coordinate system, and the motion plan is adjusted through small steps until the object reaches the target posture.
It effectively reduces the impact of errors in the transformation matrix on positioning, achieves sub-millimeter positioning accuracy, and meets the requirements of automated assembly of consumer electronic products.
Smart Images

Figure CN115446847B_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 208,816, filed June 9, 2021, inventors Sabarish Kuduwa Sivanath and Zheng Xu, titled “SYSTEM AND METHOD FOR CORRECTING AND COMPENSATING ERRORS OF 3D EYE-TO-HAND COORDINATION,” attorney docket number EBOT21-1001PSP, and U.S. Provisional Patent Application No. 63 / 209,933, filed June 11, 2021, inventors Zheng Xu, Sabarish Kuduwa Sivanath, and Ming Du Kang, titled “SYSTEM AND METHOD FOR IMPROVING ACCURACY OF 3D EYE-TO-HAND COORDINATION OF A ROBOTIC SYSTEM,” attorney docket number EBOT21-1002PSP, which disclosures are incorporated by reference herein in their entireties for all purposes.
[0003] The present disclosure is related to U.S. Application No. 17 / 751,228 (attorney docket number EBOT21-1001NP), filed May 23, 2022, inventors Sabarish Kuduwa Sivanath and Zheng Xu, titled “SYSTEM AND METHOD FOR ERROR CORRECTION AND COMPENSATION FOR 3D EYE-TO-HAND COORDINATION,” the disclosure of which is incorporated by reference herein in its entirety for all purposes. TECHNICAL FIELD
[0004] The present disclosure relates generally to computer vision systems for robotic applications. In particular, the present invention relates to a system and method for improving the accuracy of 3D eye-to-hand coordination of a robotic system. BACKGROUND
[0005] Robots have been widely adopted and developed in modern industrial factories, representing a particularly important element in the production flow. The requirements for higher flexibility and fast reconfigurability have driven the advancement of robotics technology. The position accuracy and repeatability of industrial robots are essential properties required to achieve automation of flexible manufacturing tasks. The position accuracy and repeatability of robots can vary greatly within the robot workspace, and vision-guided robot systems have been introduced to improve the flexibility and accuracy of robots. A great deal of work has been done to improve the accuracy of machine vision systems in terms of the end-effector of a robot, so-called eye-hand coordination. Achieving a high level of accuracy in eye-hand coordination is a challenging task, especially in three-dimensional (3D) space. Positioning or movement errors from the robot arm and end-effector, measurement errors of 3D vision, and errors contained in the calibration target all contribute to overall system errors, limiting the operational accuracy of the robot system. Achieving sub-millimeter accuracy throughout the workspace of a 6-axis robot can be challenging. SUMMARY
[0006] One embodiment can provide a robot system. The system can include a machine vision module, a robot arm including an end-effector, and a robot controller configured to control movement of the robot arm to move a component held by the end-effector from an initial pose to a target pose. In controlling the movement of the robot arm, the robot controller can be configured to move the component in a plurality of steps. The displacement of the component in each step is less than or equal to a predetermined maximum displacement value.
[0007] In a variation of this embodiment, the machine vision module can be configured to determine a current pose of the component after each step.
[0008] In a further variation, the robot controller can be configured to determine a next step based on the current pose of the component and the target pose.
[0009] In a further variation, the machine vision module can include a plurality of cameras and one or more structured light projectors, and the cameras can be configured to capture images of a workspace of the robot arm under illumination of the structured light projectors.
[0010] In a further variation, in determining the current pose of the component, the machine vision module can be configured to generate a three-dimensional (3D) point cloud based on the captured images.
[0011] In a further variation, in determining the current pose of the end-effector, the machine vision module can be configured to compare surflet pairs associated with the 3D point cloud and surflet pairs associated with a computer-aided design (CAD) model of the component.
[0012] In yet further variants, the robotic system can further include a coordinate transformation module configured to transform the pose determined by the machine vision module from a camera-centric coordinate system to a robot-centric coordinate system.
[0013] In yet further variants, the coordinate transformation module can be further configured to determine the transformation matrix based on a predetermined number of measured poses of the calibration target.
[0014] In variants of this embodiment, the predetermined maximum displacement value is determined based on a required level of positioning accuracy of the robotic system.
[0015] One embodiment can provide a method for controlling movement of a robotic arm including an end effector. The method can include determining a target pose of a component held by the end effector, and controlling movement of the robotic arm by a robot controller to move the component from an initial pose to the target pose in a plurality of steps. A displacement of the component in each step is less than or equal to a predetermined maximum displacement value. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 An exemplary robotic system is illustrated in accordance with one embodiment.
[0017] Figure 2 An exemplary trajectory of an object-on-hand is illustrated in accordance with one embodiment.
[0018] Figure 3 A flowchart is presented that illustrates an exemplary process for calibrating a robotic system and obtaining a transformation matrix in accordance with one embodiment.
[0019] Figure 4 A flowchart is presented that illustrates an exemplary operation process of a robotic system in accordance with one embodiment.
[0020] Figure 5 A flowchart is presented that illustrates an exemplary process for determining a pose of a component in accordance with one embodiment.
[0021] Figure 6 A block diagram of an exemplary robotic system is shown in accordance with one embodiment.
[0022] Figure 7 An exemplary computer system that facilitates small-step movement in a robotic system is illustrated in accordance with one embodiment.
[0023] In the drawings, like reference numerals refer to like elements throughout the various drawings. DETAILED DESCRIPTION
[0024] The following description presents embodiments so that any person skilled in the art can make and use the embodiments, and the following description provides the context in which the particular applications and their requirements exist. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of the disclosure. Thus, the present disclosure is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0025] SUMMARY
[0026] The embodiments described herein address the technical problem of improving the accuracy of eye-hand coordination of a robotic system. More specifically, to improve the positioning accuracy of an object on a robot arm, a large movement can be divided into a series of small steps, each guided by a 3D machine vision system. In each small step, the 3D vision system captures an image of the work scene and determines the pose of the object on the hand with the help of a segmentation neural network. The robot controller can adjust the motion plan according to the determined pose until the object on the hand reaches the destination pose (e.g., a pose that matches a mounting location. In some embodiments, determining the pose of the object on the hand can involve comparing measured surflets of the object to surflets of a known computer-aided design (CAD) model of the object.
[0027] Eye-hand coordination errors
[0028] Figure 1 An example robotic system is illustrated in accordance with one embodiment. The robotic system 100 can include a robot arm 102 and a 3D machine vision module 104. In some embodiments, the robot arm 102 can include a base 106, a plurality of joints (e.g., joints 108 and 110), and a gripper 112. The combination of the plurality of joints can enable the robot arm 102 to have a wide range of movement and have six degrees of freedom (6DoF). Figure 1 A Cartesian coordinate system (e.g., X-Y-Z) is shown in which the robot controller uses to control the pose of the robot arm 102. This coordinate system is referred to as the robot base coordinate system. In the example shown in FIG. 1, the origin of the robot base coordinate system is located at the robot base 106, so this coordinate system is also referred to as a robot-centric coordinate system. Figure 1 In the example shown in FIG. 1, the origin of the robot base coordinate system is located at the robot base 106, so this coordinate system is also referred to as a robot-centric coordinate system.
[0029] Figure 1The 3D machine vision module 104 (which can include multiple cameras) is also shown configured to capture images of the robot arm 102, including images of a calibration target 114 held by the gripper 112. The calibration target 114 typically includes a predefined pattern, such as an array of dots as shown in the enlarged view of the target 114. Capturing images of the calibration target 114 allows the 3D machine vision module 104 to determine the exact location of the gripper 112. Figure 1 A Cartesian coordinate system used by the 3D machine vision system 104 to track the pose of the robot arm 102 is also shown. This coordinate system is referred to as the camera coordinate system or camera-centric coordinate system. In the example shown in FIG. 1, the origin of the camera coordinate system is at one of the cameras. Figure 1 In the example shown in FIG. 1, the origin of the camera coordinate system is at one of the cameras.
[0030] Robot eye-hand coordination refers to transforming coordinates from the camera coordinate system to the robot base coordinate system so that machine vision can be used to guide movement of the robot arm. The transformation between coordinate systems can be expressed as:
[0031]
[0032] where b H c is the transformation matrix, is a vector in the robot base space (i.e., it is expressed using coordinates in the robot base coordinate system), is a vector in the camera space (i.e., it is expressed using coordinates in the camera coordinate system). Equation (1) can be expanded by expressing each vector using its X, Y, and Z components to obtain:
[0033]
[0034] where X c , Y c , Z c are coordinates in the camera space; X r , Y r , Z r are coordinates in the robot base space; R ij are rotation coefficients, i = 1, 2, 3 and j = 1, 2, 3; and T x , Y y , T z are translation coefficients.
[0035] The transformation matrix can be obtained by performing an eye-hand calibration process. During the calibration process, a user can safely install the robot arm and the cameras of the 3D machine vision system, and then place a calibration target (e.g., the calibration target 114) in the field of view of the cameras. The 3D machine vision module 104 can then capture images of the calibration target 114 from different viewpoints, and use the images to determine the transformation matrix H. Figure 1The target 114 shown in the middle) is attached to the end effector g of a robot arm. The robot arm can move the end effector g to a plurality of planned poses within the camera field of view (FOV). The robot controller records the pose of the end effector g relative to the robot base (i.e., relative to the origin of the robot base coordinate frame) as b H g And the 3D machine vision system records the pose of the calibration target relative to the camera (i.e., relative to the origin of the camera coordinate frame) as c H t The poses in the robot base space and the camera space satisfy the following equation:
[0036]
[0037] where i and j correspond to the poses, g(i) H b and g(j) H b is the pose of the robot base relative to the end effector g (where g(i) H b = [ b H g(i) ] -1 And g(j) H b = [ b H g(j) ] -1 ; c H t(i) and c H t(j) is the pose of the calibration target relative to the origin in the camera space, and b H c is the camera pose relative to the origin of the robot base space, which is actually the transformation matrix from the camera space to the robot base space. In other words, knowing b H c , one can convert the camera observed pose of the target to the pose of the end effector g controlled by the robot controller. Equation (3) can be rearranged to obtain:
[0038]
[0039] Various numerical methods have been developed to solve equation (4) in order to derive the transformation matrix b H c). It has been proven that at least three poses (or two pairs of poses) are required to solve equation (4). Linear least squares techniques or singular vector decomposition (SVD) can be used to derive the transformation matrix. Lie theory can also be used to derive the transformation matrix by minimizing a distance metric on the Euclidean group. More specifically, least squares fitting can be introduced to obtain a solution for the transformation matrix using the canonical coordinates of the Lie group. Additional methods can include using quaternions and non-linear least squares to improve the robustness of the solution, using Kronecker products and vectorization to improve robustness for small rotation angles, and using SVD to implement dual quaternions and simultaneous solution for rotation and translation to improve the accuracy of the transformation matrix.
[0040] While the above methods have been proven to improve the accuracy of the transformation matrix, errors can still exist due to the non-linearity of kinematics and the inherent nature of numerical calculations. Furthermore, input data from the robot controller and camera can also contain errors, which can lead to unavoidable errors in the transformation matrix. For example, in current robotic systems, the error in the rotation coefficient ij is on the order of 10 -3 above. Errors in the transformation matrix lead to positioning / pose errors of the robot.
[0041] The eye-hand coordination error can be determined by the following equation:
[0042]
[0043] where is the error in the robot base space, is the measurement error in the camera space, and b H c is the error contained in the transformation matrix. Equation (5) can be expanded as follows:
[0044]
[0045] where [△X r △Y r △Z r ] T is the position error in the robot base space, [△X c △Y c △Z c 0] T is the measurement error in the camera space, and [X c Y c Z c 1] T is the actual position of the object in the camera space.
[0046] From Equation (6), we can see that the positioning error of the object in the robot base space (i.e., [△X r △Y r △Z r ] T ) and the object's displacement in camera space (i.e., [X c Y c Z c 1] T ) or the distance between the object and the origin in camera space. Therefore, it is inevitable that the eye-hand coordination error increases with the distance between the object on the hand and the camera, that is, the distance from the eye (camera) to the hand (object on the hand) is the main factor of the error.
[0047] When the object on the hand moves from R1 and R2 in the robot base space, the object on the hand moves from C1 and C2 in the camera space (or within the FOV of the 3D vision system). The displacement of the object in the robot base space is expressed as And the displacement of the object in camera space is expressed as in:
[0048]
[0049] And (7.2)
[0050]
[0051] Accordingly, the positioning error of the object can be determined by the following equation:
[0052]
[0053] As can be seen from equation (8), using small steps, the positioning error can be obtained by the transformation matrix (i.e., b H c ), the change of the transformation matrix (i.e., Δ[ b H c ]), the displacement of the object in camera space (i.e., )and changes (i.e., ) is determined. Note that the absolute distance from the camera to the object is eliminated. Equation (8) can be expanded as follows:
[0054]
[0055] In practical applications, the change in the rotation coefficient (i.e., ΔR ij ) can be in 2e -3 within the range, and R ij Can be in e-2 within the range of 5mm and 10mm, and [△x c y c z c 1] T within the range of 5mm and 10mm, and [△x c △y c △z c 0] T can be about 50pm, then the positioning error [△x r △y r △z r ] T can be controlled within 100pm, which is sufficient to meet the requirements of automated assembly of consumer electronic products.
[0056] Small step robot movements
[0057] In some embodiments, to improve the positioning accuracy of the object, the movement of the object can be divided into a plurality of small steps, in each of which the object moves within a predetermined small range (e.g., the 5-10mm range described above). Note that the object can be an end effector of a robot arm, or an object held by the end effector, and it can be referred to as a hand-on object. At each small step, after the object moves, the 3D machine vision system can determine the actual position of the object and adjust the motion plan for the next small step until the object reaches the destination pose.
[0058] Figure 2 Fig. 3 illustrates an exemplary trajectory of a hand-on object according to one embodiment. Note that for simplicity of illustration, only the displacement of the hand-on object is shown, without showing the 3D poses. In Figure 2 , the object will be moved from a start position 202 to an end position 204 by a robot arm. The distance between the start position 202 and the end position 204 is greater than a predetermined maximum displacement. Note that the value of the maximum displacement can depend on the desired positioning accuracy. The higher the desired accuracy, the lower the value of the maximum displacement. In some embodiments, sub-millimeter level of positioning accuracy is required, and the value of the maximum displacement can be between 5mm and 10mm.
[0059] Figure 2 As shown, instead of moving the object in one step following the direct path 200, the robot arm moves the object in a plurality of small steps, e.g., steps 206 and 208. The displacement of the object in each step can be equal to or less than the predetermined maximum displacement. The small steps can ensure that the positioning amount caused by the error in the transformation matrix can be kept small, and the actual path taken by the object will not deviate much from the direct path 200. For illustration purposes, Figure 2 the path deviation in is exaggerated.
[0060] The way the robot arm uses a series of small steps to move the object to the destination pose and adjusts the path based on visual feedback at each step can be similar to the way a human performs a movement that requires accurate manual positioning (e.g., threading a needle). Just like a human relies on their eyes to determine the current pose of the needle thread to adjust the movement of the hand, a robot relies on its eyes (e.g., a 3D machine vision system) to determine the actual pose of the object in each small step. Note that there can be multiple objects within the workspace, and the machine vision system can need to determine the pose of multiple objects in order to guide the robot arm to move the objects from one pose to another.
[0061] In some embodiments, the 3D machine vision system can include two vision modules, each including two cameras and one structured light projector. One of the two vision modules can be mounted directly above the workspace, while the other vision module can be mounted at an angle. A detailed description of the 3D machine vision system can be found in PCT Application No. PCT / US2020 / 043105, filed July 22, 2020, inventors Ming Du Kang, Kai C Yung, Wing Tsui, and Zheng Xu, titled “SYSTEM AND METHOD FOR 3D POSE MEASUREMENT WITH HIGH PRECISION AND REAL-TIME OBJECT TRACING,” the disclosure of which is incorporated by reference herein in its entirety.
[0062] To identify the objects in the workspace, the 3D machine vision system can control all four cameras to capture one image of the workspace and send the image to an instance segmentation neural network to better understand the scene. In some embodiments, the instance segmentation neural network can generate a semantic map that classifies each pixel as belonging to the background or an object and an instance center for each object. Based on the semantic map and the instance centers, the 3D machine vision system can identify the objects in the scene and generate a mask for each object.
[0063] Subsequently, the 3D machine vision system can control the structured light projector to project a pattern onto the objects and control the cameras to capture images of the objects under the illumination of the structured light. The captured images can be used to construct a 3D point cloud of the environment. More specifically, constructing the 3D point cloud can include generating a decoded map based on the projected image of the structured light projector, associating camera pixels with projector pixels, and triangulating 3D points based on camera and projector intrinsic matrices, relative positions between the camera and the projector, and camera-projector pixel associations.
[0064] When there are multiple objects in the workspace (or specifically in the FOV of the 3D machine vision system), based on the masks generated by the instance segmentation neural network, the 3D machine vision system can isolate the object of interest and generate a 3D point cloud for the object.
[0065] In some embodiments, surflet-based template matching techniques can be used to estimate or determine the 3D pose of an object in the FOV of the machine vision system. A surflet refers to a directed point on the surface of a 3D object. Each surflet can be described as a pair (p, n), where p is a position vector and n is a surface normal. The surflet-pair relationship can be considered as a generalization of curvature. A surflet pair can be expressed using vectors:
[0066]
[0067] where, is a 3D position vector of a surface point, is a vector normal to the surface, and is a distance vector from to .
[0068] Surflet-based template matching techniques are based on a known CAD model of the object. In some embodiments, a plurality of surflet pairs can be extracted from the CAD model of the object, and the extracted surflet pairs (as expressed using equation (10)) can be stored in a 4D hash map. During operation of the robotic arm, at each small step, the 3D machine vision system can generate a 3D point cloud of the object and compute surflet pairs based on the 3D point cloud. The surflet pairs of the object are associated with the surflet pairs of the CAD model using the pre-computed hash map, and the 3D pose of the object can be estimated accordingly.
[0069] Before a robotic system is used to complete an assembly task, the robotic system needs to be calibrated and a transformation matrix needs to be derived. Note that even though the derived transformation matrix is likely to contain errors, these errors will only minimally affect the positioning accuracy of the object at each small step. Figure 3A flowchart is presented that illustrates an exemplary process for calibrating a robotic system and obtaining a transformation matrix according to one embodiment. During operation, a 3D machine vision system is installed (operation 302). The 3D machine vision system may include multiple cameras and a structured light projector. The 3D machine vision system may be mounted and secured above the end actuator of a robotic arm being calibrated, with the lenses of the cameras and the structured light projector facing the end actuator of the robotic arm. A robotic operator may mount and secure a 6-axis robotic arm so that the end actuator of the robotic arm can freely move to all possible poses within the FOV and depth of field (DOV) of the 3D machine vision system (operation 304).
[0070] For calibration purposes, a calibration target (e.g. Figure 1 The target 114 shown in FIG. 1 can be attached to the end actuator (operation 306). The predetermined pattern on the calibration target can facilitate the 3D machine vision system to determine the pose of the end actuator (which includes not only the position but also the tilt angle). The surface area of the calibration target is smaller than the FOV of the 3D machine vision system.
[0071] The controller of the robot arm can generate a plurality of predetermined poses in the robot base space (operation 308) and sequentially move the end actuator to these poses (operation 310). At each pose, the 3D machine vision system can capture an image of the calibration target and determine the pose of the calibration target in the camera space (operation 312). A transformation matrix can then be derived based on the pose generated in the robot base space and the pose determined by the machine vision in the camera space (operation 314). Various techniques can be used to determine the transformation matrix. For example, equation (4) can be solved based on the predetermined poses in the robot base space and the camera space using various techniques, including but not limited to: linear least squares or SVD techniques, techniques based on Lie theory, techniques based on quaternions and nonlinear minimization or dual quaternions, techniques based on Kronecker products and vectorization, etc.
[0072] After calibration, the robotic system can be used to complete assembly tasks such as picking up components in a workspace, adjusting the component's pose, and installing the component at the installation site. In some embodiments, the robotic arm moves in small steps, and in each small step, a 3D machine vision system is used to measure the current pose of the end effector or component, and the next pose is calculated based on the measured pose and the target pose.
[0073] Figure 4A flowchart is presented that illustrates an example operational process of a robotic system according to one embodiment. The robotic system can include a robotic arm and a 3D machine vision system. During operation, an operator can install a gripper on the robotic arm and calibrate its TCP (operation 402). The robotic controller can move the gripper to the vicinity of a component to be assembled in the workspace under the guidance of the 3D machine vision system (operation 404). At this stage, the 3D machine vision system can generate low resolution images (e.g., using a camera with a large FOV) to guide the movement of the gripper. At this stage, the positioning accuracy of the gripper is not very important. For example, the gripper can be moved from its original location to the vicinity of the component to be assembled in one large step.
[0074] The 3D machine vision system can determine the pose of the component in the camera space (operation 406). Note that there can be multiple components within the workspace, and determining the pose of the component to be assembled can involve the steps of identifying the component and generating a 3D point cloud for the component.
[0075] Figure 5 A flowchart is presented that illustrates an example process for determining the pose of a component according to one embodiment. During operation, the 3D machine vision system can capture one or more images of the workspace (operation 502). The images can be input to an instance segmentation machine learning model (e.g., a deep learning neural network) that outputs a semantic map of the scene and instance centers of the components (operation 504). Masks of the objects can be generated based on the semantic map and the instance centers (operation 506).
[0076] The 3D machine vision system can also capture images of the scene under the illumination of structured light (operation 508) and generate a 3D point cloud for the component based on the captured images and the masks of the components (operation 510). Note that generating the 3D point cloud can include the steps of generating a decoding map based on the images captured under the illumination of structured light and triangulating each 3D point based on the intrinsic matrices of the camera and the projector, the relative position between the camera and the projector, and the camera-projector pixel correspondence.
[0077] The 3D machine vision system can also compute surflet pairs from the 3D point cloud (operation 512) and compare the computed surflet pairs of the 3D point cloud with surflet pairs of a 3D CAD model of the component (operation 514). The 3D pose of the component can then be estimated based on the comparison result (operation 516). Note that the estimated pose is in the camera space.
[0078] Back to Figure 4After determining / estimating the pose of the part, the robotic system can then use the transformation matrix derived during calibration to convert the part pose from camera space to robot base space (operation 408). In this example, it is assumed that the TCP pose of the gripper should be aligned with the part to facilitate the gripper picking up the part. Thus, the converted part pose can be the target pose of the gripper TCP. Based on the current pose of the gripper (which is in robot base space and known by the robot controller) and the target pose, the robot controller can determine a small step for moving the gripper (operation 410). Determining the small step can include computing intermediate poses to move the gripper towards the target pose. For example, the intermediate poses can be on a straight path between the initial pose and the target pose. The amount of displacement of the gripper in this small step can be equal to or less than a predetermined maximum displacement. The robot controller can generate a motion command based on the determined small step and send the motion command to the robot arm (operation 412). The gripper moves accordingly to the intermediate pose (operation 414). After the movement of the gripper stops, the 3D machine vision system estimates the pose of the gripper (operation 416) and determines whether the gripper has reached its target pose (operation 418). Note that the process shown in Figure 5
[0079] If the gripper has moved to its target pose, the gripper can grasp the part (operation 420). Otherwise, the robot controller can determine a next small step for moving the gripper (operation 410). Note that determining the next small step can include computing intermediate poses of the part.
[0080] After the gripper securely grasps the part, the robot controller can move the gripper with the part to the vicinity of the installation site of the part under the guidance of the 3D machine vision system (422). As in operation 404, the 3D machine vision system can operate at low resolution and the gripper with the part can move in large steps. The 3D machine vision system can determine the pose of the installation site in camera space and convert such pose to robot base space (operation 424). For example, if the grasped part is to mate with another part, the 3D machine vision system can determine the pose of the other part. Based on the pose of the installation site, the robotic system can determine a target installation pose of the part held by the gripper (operation 426). The robot controller can then move the part to its target installation pose using a plurality of small steps (operation 428). In each small step, the 3D machine vision system can determine the actual pose of the part in camera space and convert the pose to robot base space. The robot controller can then determine a small step to be taken to move the part towards its target pose. Once the part held by the gripper reaches the target installation pose, the robot controller can control the gripper to install and secure the part (operation 430).
[0081] In the example shown in FIG. 3, the components are mounted to stationary positions. In practice, it is also possible that both components to be assembled are held by two robotic arms. The two robotic arms can each move their end effectors in small steps while determining the pose of both components in each step until the two components can be aligned for assembly. In this case, both robotic arms are moved, similar to how human hands move in order to align the end of a thread with the hole of a needle when threading a needle. Figure 4
[0082] To further improve the positioning accuracy, in some embodiments, the error in the transformation matrix can also be compensated for at each small step. In some embodiments, a trained machine learning model (e.g., a neural network) can be used to generate an error matrix at any given location / pose in the 3D workspace, and this error matrix can be used to correlate the camera-indicated pose of the component with the controller-desired pose. In other words, given the camera-indicated pose (or target pose) determined by the 3D machine vision system, the system can compensate for the error in the transformation matrix by having the robotic controller generate commands for the controller-desired pose. A detailed description of compensating for the error in the transformation matrix can be found in co-pending U.S. Application No. 17 / 751,228 (Attorney Docket No. EBOT21-1001 NP) filed on May 23, 2022, entitled “SYSTEM AND METHOD FOR ERROR CORRECTION AND COMPENSATION FOR 3D EYE-TO-HAND COORDINATION” by inventors Sabarish Kuduwa Sivanath and Zheng Xu, the disclosure of which is incorporated by reference herein in its entirety for all purposes.
[0083] Figure 6 A block diagram of an exemplary robotic system according to one embodiment is illustrated. The robotic system 600 can include a 3D machine vision module 602, a six-axis robotic arm 604, a robotic control module 606, a coordinate transformation module 608, an instance segmentation machine learning model 610, a point cloud generation module 612, a template matching module 614, and a 3D pose estimation module 616.
[0084] The 3D machine vision module 602 can use 3D machine vision techniques (e.g., capturing images under structured light illumination, constructing 3D point clouds, etc.) to determine the 3D pose of the objects (including the two components to be assembled and the gripper) within the FOV and DOV of the camera. In some embodiments, the 3D machine vision module 602 can include multiple cameras with different FOVs and DOVs and one or more structured light projectors.
[0085] The six-axis robotic arm 604 can have multiple joints and 6DoF. The end effector of the six-axis robotic arm 604 can be free to move in the FOV and DOV of the cameras of the 3D machine vision module 602. In some embodiments, the robotic arm 604 can include multiple segments, with adjacent segments coupled to each other via a rotational joint. Each rotational joint can include a servo motor that can continuously rotate in a particular plane. The combination of multiple rotational joints can enable the robotic arm 604 to have a wide range of movement with 6DoF.
[0086] The robot control module 606 controls the movement of the robotic arm 604. The robot control module 606 can generate a motion plan, which can include a series of motion commands that can be sent to each individual motor in the robotic arm 604 to facilitate movement of the gripper to complete a particular assembly task, such as picking up a component, moving the component to a desired installation site, and installing the component. In some embodiments, the robot control module 606 can be configured to limit each movement of the gripper to small steps, such that the displacement of each small step is equal to or less than a predetermined maximum displacement value. The maximum displacement value can be determined based on a desired level of positioning accuracy. Higher positioning accuracy means that the maximum displacement value for each small step is smaller.
[0087] The coordinate transformation module 608 can be responsible for transforming the pose of the gripper or a component from camera space to robot base space. The coordinate transformation module 608 can maintain a transformation matrix and use the transformation matrix to transform the pose seen by the 3D machine vision module 602 in camera space to a pose in robot base space. The transformation matrix can be obtained through a calibration process that measures multiple poses of a calibration target.
[0088] The instance segmentation machine learning model 610 applies machine learning techniques to generate semantic maps and instance centers for a captured image that includes multiple components. A mask for each object can be generated based on the output of the instance segmentation machine learning model 610. The point cloud generation module 612 can be configured to generate a 3D point cloud for a component to be assembled. The template matching module 614 can be configured to compare surflet pairs of the 3D point cloud to surflet pairs of a 3D CAD model using template matching techniques. The 3D pose estimation module 616 can be configured to estimate a 3D pose of a component to be assembled based on the output of the template matching module 614.
[0089] Figure 7An exemplary computer system that facilitates small step movement in a robotic system is illustrated in accordance with one embodiment. The computer system 700 includes a processor 702, a memory 704, and a storage device 706. In addition, the computer system 700 can be coupled to peripheral input / output (I / O) user devices 710, such as a display device 712, a keyboard 714, and a pointing device 716. The storage device 706 can store an operating system 720, a small step movement control system 722, and data 740.
[0090] The small step movement control system 722 can include instructions that, when executed by the computer system 700, can cause the computer system 700 or the processor 702 to perform the methods and / or processes described in this disclosure. In particular, the small step movement control system 722 can include instructions for controlling a 3D machine vision module to measure the actual pose of a gripper (machine vision control module 724), instructions for controlling movement of a robotic arm so as to place the gripper in a particular pose (robotic control module 726), instructions for transforming the pose from camera space to robot base space (coordinate transformation module 728), instructions for executing an instance segmentation machine learning model to generate a mask of a part to be assembled in a captured image of the workspace (instance segmentation model execution module 730), instructions for generating a 3D point cloud for the part to be assembled (point cloud generation module 732), instructions for applying a template matching technique to compare the 3D point cloud and a surflet pair of a CAD model (template matching module 734), and instructions for estimating a 3D pose of the part to be assembled (3D pose estimation module 736). The data 740 can include part CAD models 742.
[0091] In general, embodiments of the present invention can provide systems and methods for detecting and compensating for pose errors in a robotic system in real time. The system can use machine learning techniques (e.g., train a neural network) to predict an error matrix that can transform a camera seen pose (i.e., an indicated pose) to a controller controlled pose (i.e., a desired pose). Thus, to align a gripper with a part in a camera view, the system can first obtain the camera seen pose of the part, and then use the trained neural network to predict the error matrix. By multiplying the camera seen pose with the error matrix, the system can obtain the controller controlled pose. The robot controller can then use the controller controlled pose to move the gripper to the desired pose.
[0092] The methods and processes described in the DETAILED DESCRIPTION section can be implemented as code and / or data, which can be stored in a computer-readable storage medium as described above. When the computer system reads and executes the code and / or data stored on the computer-readable storage medium, the computer system performs the methods and processes embodied as data structures and code and stored within the computer-readable storage medium.
[0093] Furthermore, the above methods and processes may be included in hardware modules or devices. Hardware modules or devices may include, but are not limited to, application-specific integrated circuit (ASIC) chips, field-programmable gate arrays (FPGAs), dedicated or shared processors that execute a specific software module or piece of code at a specific time, and other programmable logic devices now known or later developed. When the hardware modules or devices are activated, they execute the methods and processes contained therein.
[0094] The foregoing descriptions of the embodiments of the present invention have been presented for purposes of illustration and description only. They are not intended to be exhaustive or to limit the invention to the forms disclosed. Therefore, many modifications and variations will be apparent to those skilled in the art. Furthermore, the foregoing disclosure is not intended to limit the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. A robotic system, comprising: Machine vision module; a robotic arm including an end effector; as well as a robotic controller configured to control movement of the robotic arm to move a component held by the end effector from an initial pose to a target pose; wherein, in controlling the movement of the robotic arm, the robotic controller is configured to move the component in a plurality of steps, and Wherein, in each step, the machine vision module is configured as follows: Capture an image of the robotic arm, generating a three-dimensional 3D point cloud based on the captured images of the end-actuator and the part, and determining a current pose of the component by comparing surflet pairs associated with the 3D point cloud with surflet pairs associated with a computer-aided design (CAD) model of the component; and The displacement of the component in each step is less than or equal to a predetermined maximum displacement value. 2 . The robotic system of claim 1 , wherein the robotic controller is configured to determine a next step based on a current pose and a target pose of the component.
3. The robotic system of claim 1 , wherein the machine vision module comprises a plurality of cameras and one or more structured light projectors, and wherein the cameras are configured to capture images of a workspace of the robotic arm under illumination by the structured light projectors. 4 . The robot system of claim 1 , wherein the robot system further comprises a coordinate transformation module, wherein the coordinate transformation module is configured to transform the current pose from a camera-centered coordinate system to a robot-centered coordinate system. 5 . The robotic system of claim 4 , wherein the coordinate transformation module is further configured to determine a transformation matrix based on a predetermined number of measured poses of the calibration target.
6. The robotic system of claim 1, wherein the predetermined maximum displacement value is determined based on a level of positioning accuracy required by the robotic system.
7. A computer-implemented method for controlling movement of a robotic arm including an end effector, the method comprising: determining a target pose of the part held by the end effector; as well as controlling, by a robot controller, movement of the robot arm to move the component from an initial pose to a target pose in a plurality of steps, wherein a displacement of the component in each step is less than or equal to a predetermined maximum displacement value; wherein the controller controls the movement of the robot arm including determining the current posture of the component by the machine vision module after each step; and Determining the current pose of the component includes generating, by a machine vision module, a three-dimensional 3D point cloud based on the captured image and comparing surflet pairs associated with the 3D point cloud with surflet pairs associated with a computer-aided design (CAD) model of the component.
8. The method of claim 7, further comprising: The next step is determined by the robot controller based on the current pose and the target pose of the part.
9. The method of claim 7, further comprising capturing images of the workspace of the robotic arm under illumination by one or more structured light projectors.
10. The method of claim 7, further comprising transforming the pose determined by the machine vision module from a camera-centric coordinate system to a robot-centric coordinate system.
11. The method of claim 10, further comprising determining a transformation matrix based on a predetermined number of poses of a calibration target measured by the machine vision module.
12. The method of claim 7, wherein the predetermined maximum displacement value is determined based on a desired level of positioning accuracy.
Citation Information
Patent Citations
Aloneness estimation device
US20200043105A1
System and method for error correction and compensation for 3D eye-hand coordination
CN115519536A
Intelligent robots
US20190184570A1
Method and apparatus of non-contact tool center point calibration for a mechanical arm, and a mechanical arm system with said calibration function
US20200198145A1
System and method for automatic hand-eye calibration of vision system for robot motion
US20200238525A1