System and method for error correction and compensation for 3D eye-hand coordination

CN115519536BActive Publication Date: 2026-08-28EBOTS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210638283.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-05-23
Filing Date
2022-06-07
Publication Date
2026-08-28
Estimated Expiration
2042-06-07

AI Technical Summary

Technical Problem

来自机器人臂和末端致动器的定位/移动误差、3D视觉的测量误差以及校准目标中包含的误差都会导致整体系统误差,从而限制机器人系统的操作准确度

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115519536B_ABST
    Figure CN115519536B_ABST
Patent Text Reader

Abstract

The present disclosure relates to systems and methods for error correction and compensation for 3D eye-hand coordination. One embodiment can provide a robotic system. The system can include a machine vision module, a robotic arm including an end effector, a robot controller configured to control movement of the robotic arm, and an error compensation module configured to compensate for an error in a pose of the robotic arm by determining a controller-desired pose corresponding to a camera-indicated pose of the end effector as observed by the machine vision module, such that when the robot controller controls movement of the robotic arm based on the controller-desired pose, the end effector achieves the camera-indicated pose as observed by the machine vision module. The error compensation module can include a machine learning model configured to output an error matrix that relates the camera-indicated pose to the controller-desired pose.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 208,816, filed June 9, 2021, entitled “SYSTEM AND METHOD FOR CORRECTING AND COMPENSATING ERRORS OF 3D EYE-TO-HAND COORDINATION”, Attorney’s File No. EBOT21-1001PSP, and U.S. Provisional Patent Application No. 63 / 209,933, filed June 11, 2021, entitled “SYSTEM AND METHOD FOR IMPROVING ACCURACY OF 3D EYE-TO-HAND COORDINATION OF A ROBOTICSYSTEM”, Attorney’s File No. EBOT21-1002PSP, the disclosures of which are incorporated herein by reference in their entirety for all purposes. Technical Field

[0003] This disclosure generally relates to computer vision systems for robotic applications. In particular, the present invention relates to a system and method for error correction and compensation of 3D eye-hand coordination in a robotic system by training a neural network to derive an error matrix. Background Technology

[0004] Robots have been widely adopted and developed in modern industrial factories, representing a particularly important element in production processes. The demand for greater flexibility and rapid reconfigurability has driven advancements in robotics. Positional accuracy and repeatability of industrial robots are fundamental attributes required for automating flexible manufacturing tasks. Robot positional accuracy and repeatability can vary significantly within the robot's workspace, leading to the introduction of vision-guided robotic systems to improve robot flexibility and accuracy. Extensive work has been done to improve the accuracy of machine vision systems with respect to robot end effectors, a process known as eye-hand coordination. Achieving highly accurate eye-hand coordination is a challenging task, especially in three-dimensional (3D) space. Positioning / movement errors from the robot arm and end effector, measurement errors from 3D vision, and errors inherent in calibration targets all contribute to overall system errors, limiting the operational accuracy of the robot system. For 6-axis robots, achieving sub-millimeter accuracy across their entire workspace can be particularly challenging. Summary of the Invention

[0005] One embodiment may provide a robotic system. The system may include a machine vision module, a robotic arm including an end effector, a robot controller configured to control movement of the robotic arm, and an error compensation module configured to compensate for pose errors of the robotic arm by determining a pose desired by the controller corresponding to a camera-indicated pose of the end effector, such that when the robot controller controls movement of the robotic arm based on the desired pose, the end effector achieves the pose as observed by the machine vision module. The error compensation module may include a machine learning model configured to output an error matrix that correlates the camera-indicated pose with the controller-desired pose.

[0006] In a variant of this embodiment, the machine learning model may include a neural network.

[0007] In other variations, the neural network may include an embedding layer and a processing layer, and each of the embedding and processing layers may include a multilayer perceptron.

[0008] In other variations, the embedding layer can be configured to embed separate translation and rotation components of the pose.

[0009] In other variants, the embedding layer can use the rectified linear unit (ReLU) as the activation function, and the processing layer can use the leaky ReLU as the activation function.

[0010] In other variations, the system may also include a model training module configured to train a neural network by collecting training samples. During neural network training, the model training module is configured to: cause the robot controller to generate pose samples desired by the controller; control the movement of the robot arm based on the desired pose samples; determine the actual pose of the end effector using a machine vision module; and calculate an error matrix based on the desired pose samples and the actual pose.

[0011] In other variations, the model training module can be configured to train a neural network until the error matrix generated by the machine learning model reaches a predetermined level of accuracy.

[0012] In a variant of this embodiment, the system may further include a coordinate transformation module configured to transform the pose determined by the machine vision module from a camera-centric coordinate system to a robot-centric coordinate system.

[0013] In other variations, the coordinate transformation module can also be configured to determine the transformation matrix based on a predetermined number of measurement attitudes of the calibration target.

[0014] In other variations, the coordinate transformation module can also be configured to associate the orientation of the part held by the end effector with the corresponding orientation of the end effector.

[0015] One embodiment may provide a computer-implemented method. The method may include: a machine vision module determining a camera-indicated pose of an end effector of a robotic arm for performing an assembly task; a robot controller determining a controller-desired pose corresponding to the camera-indicated pose of the end effector, including applying a machine learning model to obtain an error matrix that correlates the camera-indicated pose with the controller-desired pose; and controlling the movement of the robotic arm based on the controller-desired pose, thereby facilitating the end effector to achieve the camera-indicated pose in order to complete the assembly task.

[0016] One embodiment may provide a computer-implemented method. The method may include modeling attitude errors associated with an end effector of a robotic arm using a neural network; training the neural network using multiple training samples, the training samples including camera-indicated attitudes of the end effector and corresponding error matrices that associate the camera-indicated attitudes of the end effector with attitudes desired by the controller; and applying the trained neural network to compensate for attitude errors during robotic arm operation. Attached Figure Description

[0017] Figure 1 An exemplary robot system according to one embodiment is illustrated.

[0018] Figure 2 An exemplary pose error detection neural network according to one embodiment is illustrated.

[0019] Figure 3 A flowchart is presented illustrating an exemplary process for calibrating a robot system and obtaining a transformation matrix according to one embodiment.

[0020] Figure 4 A flowchart is presented illustrating an exemplary process for training a neural network for pose error detection according to one embodiment.

[0021] Figure 5A A flowchart is presented illustrating an exemplary operational process of a robot system according to one embodiment.

[0022] Figure 5B The illustration depicts a scenario where a flexible cable is picked up by an end effector of a robotic arm according to one embodiment.

[0023] Figure 6 A block diagram of an exemplary robot system according to one embodiment is illustrated.

[0024] Figure 7An exemplary computer system is illustrated to facilitate error detection and compensation in a robotic system according to one embodiment.

[0025] In each figure, the same reference numerals refer to the same figure elements. Detailed Implementation

[0026] The following description is presented to enable any person skilled in the art to make and use the embodiments, and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of this disclosure. Therefore, the invention is not limited to the illustrated embodiments, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0027] Overview

[0028] The embodiments described herein address the technical problem of correcting pose errors in robotic systems. More specifically, a machine learning model (e.g., a neural network) can be trained to learn an error matrix that characterizes the pose error of the robotic arm at each location within its workspace. Training the machine learning model may include supervised training processing. During training, the robotic arm's gripper can be moved and positioned into a predetermined pose, and the actual pose of the gripper can be measured using a 3D machine vision system. The error matrix can be derived based on the difference between the predetermined pose and the measured pose. Once sufficiently trained, the machine learning model can infer the error matrix corresponding to any pose within the workspace. During operation, possible pose errors of the gripper can be compensated for in real time based on the inferred error matrix.

[0029] Error matrix

[0030] Figure 1 An exemplary robot system according to one embodiment is illustrated. The robot system 100 may include a robotic arm 102 and a 3D machine vision module 104. In some embodiments, the robotic arm 102 may include a base 106, multiple joints (e.g., joints 108 and 110), and a gripper 112. The combination of multiple joints enables the robotic arm 102 to have a wide range of motion and six degrees of freedom (6DoF). Figure 1 The Cartesian coordinate system (e.g., XYZ) used by the robot controller to control the attitude of robot arm 102 is shown. This coordinate system is referred to as the robot's fundamental coordinate system. Figure 1 In the example shown, the origin of the robot's base coordinate system is located at the robot's base 106, so this coordinate system is also called the robot-centric coordinate system.

[0031] Figure 1 Also shown is a 3D machine vision module 104 (which may include multiple cameras) configured to capture images of the robotic arm 102, including images of a calibration target 114 held by the gripper 112. The calibration target 114 typically includes a predefined pattern, such as the dot array shown in the magnified view of the target 114. Capturing an image of the calibration target 114 allows the 3D machine vision module 104 to determine the exact location of the gripper 112. Figure 1 The Cartesian coordinate system used by the 3D machine vision system 104 to track the pose of the robotic arm 102 is also shown. This coordinate system is referred to as the camera coordinate system or the camera-centric coordinate system. Figure 1 In the example shown, the origin of the camera coordinate system is located at one of the cameras.

[0032] Robot eye-hand coordination refers to transforming the coordinate system from the camera coordinate system to the robot's base coordinate system, enabling machine vision to guide the robot arm's movement. The transformation between coordinate systems can be expressed as:

[0033]

[0034] in b H c It is a transformation matrix. It is a vector in the robot's base space (i.e., it is represented using coordinates in the robot's base coordinate system). It is a vector in camera space (i.e., it is represented using coordinates in the camera coordinate system). Equation (1) can be expanded by expressing each vector using its X, Y, and Z components to obtain:

[0035]

[0036] Where X c Y c Z c These are the coordinates in camera space; X r Y r Z r These are the coordinates in the robot's basic space; R ij These are the rotation coefficients, i = 1, 2, 3 and j = 1, 2, 3; and T x T y T z It is the translation coefficient.

[0037] The transformation matrix can be obtained by performing eye-hand calibration. During the calibration process, the user can safely mount the robot arm and the camera of the 3D machine vision system, and then calibrate the target (e.g., Figure 1The target (114) shown is attached to the end effector g of the robot arm. The robot arm can move the end effector g to multiple planned poses within the camera's field of view (FOV). The robot controller records the pose of the end effector g relative to the robot base (i.e., relative to the origin of the robot's base coordinate system) as... b H g Furthermore, the 3D machine vision system records the pose of the calibration target relative to the camera (i.e., relative to the origin of the camera coordinate system) as... c H t The robot's pose in its base space and camera space satisfies the following equations:

[0038] g(i) H b b H c c H t(i) = g(j) H b b H c c H t(j) (3)

[0039] Where i and j correspond to the pose. g(i) H b and g(j) H b It is the attitude of the robot's base relative to the end effector g (where g is the attitude of the robot's base relative to the end effector g). g(i) H b =[ b H g(i) ] -1 and g(j) H b =[ b H g(j) ] -1 ; c H t(i) and c H t(j) It calibrates the target's attitude relative to the origin in camera space, and b H c It represents the camera pose relative to the origin of the robot's base space; it is actually the transformation matrix from the camera space to the robot's base space. In other words, knowing... b H c This allows the camera's observation posture of the target to be converted into the posture controlled by the robot controller of the end effector g. Equation (3) can be rearranged to obtain:

[0040]

[0041] Various numerical methods have been developed to solve equation (4) in order to derive the transformation matrix (b H c It has been proven that solving equation (4) requires at least three poses (or two pairs of poses). Linear least squares techniques or singular vector decomposition (SVD) can be used to derive the transformation matrix. Lie theory can also be used to derive the transformation matrix by minimizing distance metrics on the Euclidean group. More specifically, least squares fitting can be introduced to obtain a solution to the transformation matrix using the canonical coordinates of the Lie group. Additional methods can include using quaternions and nonlinear minimization to improve the robustness of the solution, using Kronecker products and vectorization to improve robustness in small rotation angles, and using SVD to achieve simultaneous solutions of double quaternions and rotations and translations to improve the accuracy of the transformation matrix.

[0042] While the methods described above have been shown to improve the accuracy of the transformation matrix, errors may still exist due to the nonlinearity of kinematics and the inherent nature of numerical computation. Furthermore, input data from the robot controller and camera may also contain errors, which can lead to unavoidable errors in the transformation matrix. For example, in current robot systems, the rotation coefficient ΔR... ij The error is within 10 -3 The above explains that errors in the transformation matrix can lead to positioning / attitude errors in the robot.

[0043] To improve localization / pose accuracy, the ability to correct errors in the transformation matrix in real time is desired. Various methods have been explored to correct localization and pose errors, including machine learning-based approaches. For example, one method trains a neural network to achieve eye-hand coordination instead of the transformation matrix, and another similar approach applies a neural network to eye-hand coordination instead of the transformation matrix, and to eye-joint coordination instead of inverse kinematics. However, these methods may still result in localization accuracy within a few millimeters. Another approach is to build a special neural network to predict localization errors and compensate for errors along a prescribed end effector path. While this approach can reduce localization errors to less than one milliliter after compensation, it does not address the problems associated with pose errors. In general, existing robotic systems cannot meet the accuracy and repeatability requirements for manufacturing consumer electronics (e.g., smartphones, tablets, wearables, etc.). Assembling consumer electronics typically involves handling many small (e.g., millimeter or smaller) parts in a confined space and requires robot localization / pose accuracy in the sub-millimeter range or even higher (sometimes as low as 10) millimeters. -3 mm).

[0044] To reduce the localization / pose error of a robot throughout its workspace in real time, the concept of an error matrix can be introduced. The error matrix indicates the difference between the controller-desired pose of the robot's end effector (i.e., the pose programmed by the robot controller) and the actual pose of the end effector (which can be captured by a camera and transformed from camera space to the robot's base space using a transformation matrix), and it can change as the end effector's position changes within the workspace. In some embodiments, the error matrix can be expressed as a transformation from the indicated pose to the desired pose in the robot's base space:

[0045]

[0046] Where H td H is the desired pose of the controller at the tool center position (TCP) in the robot's basic space (or simply the desired pose). ti This is the actual pose transformed from camera space to robot base space using a transformation matrix, and is called the camera pointing pose (or simply pointing pose). It is the error matrix, which is the position vector. The function. In one example, the robot controller can send commands to move the end effector to the desired TCP pose H. td However, due to errors in the robot system (e.g., drive errors in joints and end effectors), when the controller instructs the robot arm to achieve this pose, the resulting pose is often inconsistent with H. td Different. The actual pose of the end effector, measured by the 3D machine vision module and transformed from camera space to the robot's base space, can be the indicated pose H. ti Therefore, given the indicated pose (i.e., the pose known to the camera), if the error matrix... Given that the controller can be used to instruct the robot arm to move the end effector to the indicated pose, thereby achieving the desired pose of eye (camera) and hand (robot controller) coordination.

[0047] Real-time error detection and compensation

[0048] Although it can be deduced However, given that the robot can have six degrees of freedom (6DoF) (meaning that TCP pose can include at least six components) and the error matrix is ​​a nonlinear function of position, this task is computationally intensive. TCP pose can be expressed as [x,y,z,r...]. x ,r y ,r z ], where [x,y,z] are translation components, and [r x ,r y ,r z[ ] represents the rotational components of the attitude (e.g., roll, pitch, and yaw). Furthermore, the nonlinear nature of the error means that the error matrix may have infinite dimensions. To reduce the derivation of the error matrix... The required computation can be mitigated in some embodiments of this application using machine learning techniques, where a trained machine learning model (e.g., a neural network) can be used to learn the error matrix. Once the error matrix is ​​learned, the system can calculate the indicated TCP pose to achieve the desired TCP pose. The robot controller can then send appropriate pose commands to the robot arm.

[0049] In some embodiments, the error detection machine learning model may include a neural network (e.g., a deep learning neural network). The input to the model may be an indication pose H. ti And the output of the model can be an error matrix. In other words, given an indicated pose, the model can predict the error, and then the desired pose can be calculated using equation (5). The controller can use the desired pose to control the robot's movement. The neural network can be constructed to include an embedding layer (which can be used to map discrete variables (e.g., TCP pose) to continuous vectors) and a processing layer. In some embodiments, to reduce embedding complexity and improve efficiency, translation components (i.e., [x,y,z]) and rotation components (i.e., [r]) are used. x ,r y ,r z ]) can be embedded individually (e.g., using two parallel embedding layers).

[0050] Figure 2An exemplary pose error detection neural network according to one embodiment is illustrated. The neural network 200 may include embedding layers 202 and 204, a cascade module 206, and a processing layer 208. Embedding layer 202 may be used to embed the rotational component of the pose, and embedding layer 204 may be used to embed the translational component of the pose. Note that, depending on the application, the pose input to the neural network 200 may be a desired pose or an indication pose. For example, if the application wants to calculate the desired pose based on the indication pose, then the indication pose is used as input. On the other hand, if the application wants to determine the indication pose for a given desired pose, then the desired pose will be used as input. In both cases, the definition of the error matrix may differ. However, the error matrix indicates the transformation between the desired pose and the indication pose. Cascade module 206 may be used to cascade the embedding of translational and rotational components to obtain the pose embedding. In some embodiments, each embedding layer (layer 202 or 204) may be implemented using a multilayer perceptron (MLP), which may include multiple inner layers (e.g., an input layer, hidden layers, and an output layer). In a further embodiment, each embedding layer may use a rectified linear unit (ReLU) as the activation function at each node. Besides ReLU, other types of activation functions, such as nonlinear activation functions, may also be used in the embedding layers.

[0051] Cascaded embeddings can be sent to processing layer 208, which learns the mapping between pose and error matrix. In some embodiments, processing layer 208 can also be implemented using an MLP, and at each node of processing layer 208, leaked ReLU can be used as the activation function. In addition to leaked ReLU, processing layer 208 can also use other types of activation functions, such as nonlinear activation functions.

[0052] Before training the pose error detection neural network, the system needs to be calibrated and the transformation matrix needs to be derived. Even if the derived transformation matrix may contain errors, these errors are taken into account and corrected by the error matrix learned by the neural network. Figure 3 A flowchart is presented illustrating an exemplary process for calibrating a robot system and obtaining a transformation matrix according to one embodiment. During operation, a 3D machine vision system is installed (operation 302). The 3D machine vision system may include multiple cameras and a structured light projector. The 3D machine vision system may be mounted and secured above the end effector of the robot arm being calibrated, with the camera lenses and the structured light projector facing the end effector of the robot arm. The robot operator may mount and secure the 6-axis robot arm such that the end effector of the robot arm can freely move to all possible poses within the FOV and DOV of the 3D machine vision system (operation 304).

[0053] For calibration purposes, the calibration target (e.g., Figure 1 The target 114 shown can be attached to an end effector (operation 306). A predetermined pattern on the calibration target can facilitate the 3D machine vision system in determining the attitude of the end effector (which includes not only location but also tilt angle). The surface area of ​​the calibration target is smaller than the FOV of the 3D machine vision system.

[0054] The robot arm controller can generate multiple predetermined poses in the robot's base space (operation 308) and sequentially move the end effector to these poses (operation 310). In each pose, the 3D machine vision system can capture an image of the calibration target and determine the pose of the calibration target in the camera space (operation 312). The transformation matrix can then be derived based on the poses generated in the robot's base space and the poses determined by machine vision in the camera space (operation 314). Various techniques can be used to determine the transformation matrix. For example, equation (4) can be solved using various techniques based on the predetermined poses in the robot's base space and the camera space, including but not limited to: linear least squares or SVD techniques, Lie theory-based techniques, quaternion-based and nonlinear minimization or dual quaternion-based techniques, Kronecker product-based and vectorization techniques, etc.

[0055] Figure 4 A flowchart is presented illustrating an exemplary process for training a neural network for posture error detection according to one embodiment. In some embodiments, training the neural network may include supervised training processes. During training, a robot operator may replace the calibration target with a gripper and calibrate the gripper's tool center point (TCP) (operation 402). The robot arm's controller may generate a desired posture within the robot arm's workspace or within the FOV and DOV of the 3D machine vision system (operation 404). The desired posture may be generated randomly or following a predetermined path. In some embodiments, the workspace may be divided into a 3D mesh with a predetermined number of cells, and the controller may generate an indication posture for each cell. This ensures that training covers a sufficient portion of the workspace. The robot controller may then move the gripper based on the desired posture (operation 406). For example, the robot controller may generate motion commands based on the desired posture and send these commands to various motors on the robot arm to adjust the gripper's TCP posture. In an alternative embodiment, instead of having the controller generate a random posture, the end effector may move to a random posture, and the desired posture of the gripper's TCP may be determined by reading the values ​​of the robot arm's encoder.

[0056] After the gripper stops moving, the 3D machine vision module can measure the gripper's TCP pose (operation 408). Due to the high accuracy of 3D machine vision, the measured pose can be considered the actual pose of the gripper. In other words, any measurement error from the 3D machine vision module can be considered negligible and can be ignored. Note that the measurement output of the 3D machine vision module can be in camera space. The measured pose in camera space can then be transformed into the measured pose in the robot's base space using a previously determined transformation matrix to obtain the indicated pose (operation 410). Based on the measured pose in the robot's base space and the desired pose (which is also in the robot's base space), the current location can be calculated. error matrix (Operation 412). For example, the error matrix can be calculated as follows: Where H ti It indicates the attitude and H td This is the desired pose. The system can record H. ti and As training samples (operation 414), determine whether the predetermined number of training samples has been collected (operation 416). If so, then the collected samples, including those from different locations... If yes, the pose error detection neural network can be used (operation 418); if no, then the controller can generate additional poses (operation 404). In one embodiment, the system can also collect multiple pose samples at a single location.

[0057] In some embodiments, training of the neural network can be stopped when a sufficient portion of the workspace (e.g., 50%) has been covered. For example, if the workspace is divided into a 3D mesh of 1000 cells and more than 50% of the cells have been randomly selected for training (i.e., the robotic arm has moved to these cells and collected training samples), then training can be stopped. In an alternative embodiment, training of the neural network can be stopped after the neural network can predict / detect errors with an accuracy higher than a predetermined threshold level. In this case, after the initial training in operation 418, the controller can generate a test desired pose (operation 420) and move the gripper's TCP according to the test desired pose (operation 422). The 3D machine vision module measures the pose of the gripper TCP in camera space (operation 424). The measured pose in camera space can be transformed to the robot's base space using a transformation matrix to obtain a test indication pose (operation 426). The neural network can predict / infer the error matrix corresponding to the test indication pose (operation 428). Furthermore, the system can use the test desired pose and the test indication pose to calculate the error matrix (operation 430). The predicted error matrix and the calculated error matrix can be compared to determine if the difference is less than a predetermined threshold (operation 432). If yes, then training is complete. If not, then additional training samples are collected by returning to operation 404. The threshold can be varied depending on the localization accuracy required for robot operation.

[0058] Once the posture error detection neural network is sufficiently trained, the robotic system can operate with real-time error correction capabilities. For example, for any indicated posture in the workspace of the robotic arm, the system can use the neural network to infer / predict the corresponding error matrix, and then determine the desired posture of the gripper by multiplying the inferred error matrix by the indicated posture. In one example, the gripper's indicated posture can be obtained by using a 3D machine vision module to measure the posture of the part to be assembled in the workspace. Therefore, by using the desired posture H... td The robot controller generates commands that allow the gripper to move to H. ti Align with the component to facilitate gripping the component.

[0059] Figure 5A A flowchart is presented illustrating exemplary operational processes of a robotic system according to one embodiment. The robotic system may include a robotic arm and a 3D machine vision system. During operation, an operator may mount a gripper on the robotic arm and calibrate its TCP (operation 502). The robot controller may move the gripper to the vicinity of the part to be assembled in the workspace, guided by the 3D machine vision system (operation 504). At this stage, the 3D machine vision system may generate low-resolution images (e.g., using a camera with a large FOC) to guide the movement of the gripper.

[0060] The 3D machine vision system can determine the part's pose in camera space (operation 506), and then use a transformation matrix to transform the part's pose from camera space to the robot's base space (operation 508). In this example, it is assumed that the gripper's TCP pose should be aligned with the part to facilitate gripper pickup. Therefore, the transformed part pose can be the indication pose of the gripper's TCP. The pose error detection and compensation system can then use a neural network to infer the error matrix of the indication pose (operation 510). Based on the indication pose and the error matrix, the pose error detection and compensation system can determine the desired pose of the gripper so that the gripper can successfully grasp the part in the desired pose (operation 512). For example, the desired pose can be calculated by multiplying the indication pose by the predicted error matrix. The robot controller can generate motion commands based on the desired pose and send the motion commands to the robot arm (operation 514). The gripper moves accordingly to grasp the part (operation 516).

[0061] When the gripper firmly grasps the part, the robot controller can move the gripper carrying the part to the vicinity of the part's mounting location under the guidance of the 3D machine vision system (518). As in operation 504, in operation 518, the 3D machine vision system can operate at low resolution. The 3D machine vision system can determine the pose of the mounting location in camera space (operation 520). For example, if the gripped part is to mate with another part, the 3D machine vision system can determine the pose of that other part. A pose error detection and compensation system can similarly determine the desired pose of the mounting location (operation 522). For example, the 3D machine vision system can measure the pose of the mounting location in camera space, and the pose error detection and compensation system can convert the measured pose in camera space into a measured pose in the robot's base space, and then apply an error matrix to obtain the desired pose of the mounting location.

[0062] Given the desired posture, the robot controller can generate motion commands (operation 524) and send the motion commands to the robot arm to move the gripper to align with the mounting location (operation 526). The gripper can then mount and secure the component at the mounting location (operation 528).

[0063] The system can accurately infer / predict the error matrix of any possible posture of the gripper within the workspace, thereby significantly improving the robot's operational accuracy and reducing the amount of time required to adjust robot movements. Furthermore, the posture error detection neural network can be continuously trained by collecting additional samples to improve its accuracy or recalibrate. For example, after each robot movement, the 3D machine vision system can measure and record the gripper's actual posture, and such measurements can be used to generate additional training samples. In some embodiments, training processes can be performed periodically or as needed (e.g., Figure 4 (The processing is shown in the diagram). For example, if a pose error exceeding a predetermined threshold is detected during normal operation of the robot system, operation can be stopped and the neural network can be retrained.

[0064] In some situations, robotic arms need to pick up and install flexible components. For example, a robotic arm may pick up RF cables, align cable connectors with receptacles, and insert cable connectors into receptacles. Because cables are flexible, the relative position of the end effector / gripper changes each time the robotic arm's end effector / gripper grasps the cable. Furthermore, the curvature of the cable may change in mid-air, making it difficult to align the cable connector with the receptacle, even with error compensation. Simply controlling the attitude of the end effector may not be sufficient to complete the task of installing or connecting cables.

[0065] Figure 5B The illustration depicts a scenario where a flexible cable, according to one embodiment, is picked up by the end effector of a robotic arm. Figure 5B In this process, the end effector 532 of the robotic arm 530 picks up a flexible cable 534. The goal of the end effector 532 is to move the cable 534 to a desired location so that the connector 536 at the end of the cable 534 can mate with the corresponding connector. Due to the flexible nature of the cable 534, simply controlling the orientation of the end effector 532 may not guarantee the success of this cable connection task. However, associating the orientation of the connector 536 with the orientation of the end effector 532, and determining the desired orientation of the connector 536, can facilitate determining the desired orientation of the end effector 532.

[0066] In some embodiments, the system may use an additional transformation matrix to extend the TCP from the tip of the end effector to the center of the connector, such that the controller's desired orientation is referenced to the center of the RF connector. This additional transformation matrix may be referred to as the component transformation matrix T. c It transforms / associates the orientation of a component to the orientation of the end effector holding that component (both orientations have been transformed into the robot's base space). More specifically, given an end effector H... e The attitude of the component, H c The following equation can be used to calculate:

[0067] H c =T c ×H e (6)

[0068] The component transformation matrix can be determined in real time. During the operation of the robotic arm, the end effector and components (i.e., Figure 5BThe poses of the end effector 532 and connector 536 shown can be determined using a 3D machine vision system, and the component transformation matrix can then be calculated using the following equation:

[0069] T c =H c ×H e -1 (7)

[0070] By extending TCP to components, the controller's desired pose can be calculated as follows:

[0071]

[0072] Where H td It is the desired attitude of the end effector controller, and H ci This refers to the camera-indicated pose of the component. In other words, once the camera determines the target pose of the component, the system can calculate the pose desired by the controller, which can be used to generate motion commands for moving the end effector, allowing the component to be moved to its target pose. In some embodiments, to ensure accuracy, the system can repeatedly (e.g., at short intervals) measure and calculate the component transformation matrix, so that even if the component can move relative to the end effector, changes in this relative pose can be captured. For example, the system can calculate T every 300ms. c The recent T c It will be used to calculate the desired pose of the controller.

[0073] Figure 6 The diagram illustrates a block diagram of an exemplary robot system according to one embodiment. The robot system 600 may include a 3D machine vision module 602, a six-axis robot arm 604, a robot control module 606, a coordinate transformation module 608, a posture error detection machine learning model 610, a model training module 612, and an error compensation module 614.

[0074] The 3D machine vision module 602 can use 3D machine vision techniques (e.g., capturing images under structured light illumination, constructing 3D point clouds, etc.) to determine the 3D pose of an object (including two parts to be assembled and a gripper) within the FOV and DOV of the camera. In some embodiments, the 3D machine vision module 602 may include multiple cameras with different FOVs and DOVs and one or more structured light projectors.

[0075] The six-axis robotic arm 604 may have multiple joints and a 6DoF (6th-degree-of-field). The end effector of the six-axis robotic arm 604 can move freely within the FOV and DOV of the camera of the 3D machine vision module 602. In some embodiments, the robotic arm 604 may include multiple segments, wherein adjacent segments are coupled to each other via rotary joints. Each rotary joint may include a servo motor capable of continuous rotation within a specific plane. The combination of multiple rotary joints allows the robotic arm 604 to have a wide range of motion with a 6DoF.

[0076] The robot control module 606 controls the movement of the robot arm 604. The robot control module 606 can generate a motion plan, which may include a series of motion commands that can be sent to each individual motor in the robot arm 604 to facilitate the movement of the gripper to perform specific assembly tasks, such as picking up parts, moving parts to a desired installation location, and installing parts. Due to errors included in the system (e.g., encoder errors at each motor), when the robot control module 606 instructs the gripper to move to one posture, the gripper may end up moving to a slightly different posture. Such positioning errors can be compensated for.

[0077] The coordinate transformation module 608 is responsible for transforming the gripper's pose from camera space to robot base space. The coordinate transformation module 608 can maintain a transformation matrix and use this matrix to transform or correlate the pose observed by the 3D machine vision module 602 in camera space to the pose in robot base space. The transformation matrix can be obtained through calibration processing of multiple poses of the calibration target. Errors contained in the transformation matrix can be accounted for and compensated for using an error matrix. In further embodiments, the coordinate transformation module 608 can also maintain a component transformation matrix that correlates the pose of a component held by an end effector (e.g., the end of a flexible cable) with the pose of the end effector.

[0078] The pose error detection machine learning model 610 applies machine learning techniques to learn the error matrix of all poses in the workspace of the robot arm 604. In some embodiments, the pose error detection machine learning model 610 may include a neural network that takes the pose indicated / seen by the 3D machine vision module 602 as input and outputs an error matrix that can be used to calculate the desired pose of the robot controller to achieve the pose seen / indicated by the camera. The neural network may include an embedding layer and a processing layer, both of which are implemented using an MLP. The embedding of the rotational and translational components of the pose can be done separately, and the embedding results are concatenated before being sent to the processing layer. The activation functions used in the embedding layer include ReLU, while leaky ReLU is used as the activation function in the processing layer. The model training module 612 trains the neural network, for example, through supervised training. More specifically, the model training module 612 collects training samples by instructing the robot control module 606 to generate poses and then calculates the error matrix of these poses.

[0079] Error compensation module 614 can compensate for posture errors. To this end, for a desired posture, error compensation module 614 can obtain the corresponding error matrix by applying posture error detection machine learning model 610. Error compensation module 614 can compensate for posture errors by calculating the posture desired by the controller to achieve the actual posture or the posture seen / indicated by the camera. Error compensation module 614 can send the posture desired by the controller to robot control module 606 to allow it to generate appropriate motion commands to move the gripper to the desired posture.

[0080] Figure 7 An exemplary computer system for facilitating error detection and compensation in a robotic system, according to one embodiment, is illustrated. The computer system 700 includes a processor 702, a memory 704, and a storage device 706. Furthermore, the computer system 700 may be coupled to a peripheral input / output (I / O) user device 710, such as a display device 712, a keyboard 714, and a pointing device 716. The storage device 706 may store an operating system 720, an error detection and compensation system 722, and data 740.

[0081] The error detection and compensation system 722 may include instructions that, when executed by the computer system 700, cause the computer system 700 or processor 702 to perform the methods and / or processes described in this disclosure. Specifically, the error detection and compensation system 722 may include instructions for controlling a 3D machine vision module to measure the actual posture of the gripper (machine vision control module 724), commands for controlling the movement of the robot arm to place the gripper in a specific posture (robot control module 726), instructions for transforming the posture from camera space to robot base space (coordinate transformation module 728), instructions for training a posture error detection machine learning model (model training module 730), instructions for executing the machine learning model during robot arm operation to infer an error matrix associated with the posture (model execution module 732), and instructions for error compensation based on the inferred error matrix (error compensation module 734). Data 740 may include collected training samples 742.

[0082] Generally, embodiments of the present invention can provide systems and methods for real-time detection and compensation of posture errors in a robotic system. The system can use machine learning techniques (e.g., training a neural network) to predict an error matrix that transforms the posture seen by the camera (i.e., the indicated posture) into a posture controlled by the controller (i.e., the desired posture). Therefore, to align the gripper with a part in the camera view, the system can first obtain the camera-seen posture of the part and then use a trained neural network to predict the error matrix. By multiplying the camera-seen posture by the error matrix, the system can obtain the controller-controlled posture. The robot controller can then use the controller-controlled posture to move the gripper to the desired posture.

[0083] The methods and processes described in the Detailed Description section can be implemented as code and / or data, which can be stored in a computer-readable storage medium as described above. When a computer system reads and executes the code and / or data stored on the computer-readable storage medium, the computer system executes the methods and processes implemented as data structures and code and stored in the computer-readable storage medium.

[0084] Furthermore, the methods and processes described above may be incorporated into hardware modules or devices. These hardware modules or devices may include, but are not limited to, application-specific integrated circuit (ASIC) chips, field-programmable gate arrays (FPGAs), dedicated or shared processors that execute specific software modules or pieces of code at specific times, and other programmable logic devices now known or developed hereafter. When a hardware module or device is activated, it executes the methods and processes contained therein.

[0085] The foregoing description of embodiments of the invention is presented for illustrative and descriptive purposes only. It is not intended to be exhaustive or to limit the invention to the forms disclosed. Therefore, many modifications and variations will be apparent to those skilled in the art. Furthermore, the foregoing disclosure is not intended to limit the invention. The scope of the invention is defined by the appended claims.

Claims

1. A robot system, the system comprising: A machine vision module, including a camera configured to capture images of a work scene; Robotic arms including end effectors; The robot controller is configured to control the movement of the robot arm; as well as The error compensation module is configured to compensate for the robot arm's posture error by determining the controller's desired posture corresponding to the posture indicated by the camera of the end effector, such that when the robot controller controls the movement of the robot arm based on the controller's desired posture, the end effector achieves the posture indicated by the camera as observed by the machine vision module. The pose indicated by the camera of the end effector is derived from an image of the work scene; The desired posture of the robot controller is achieved by programming the robot controller; and The error compensation module includes a machine learning model configured to output an error matrix that correlates the pose indicated by the camera with the pose desired by the controller. The error matrix can be expressed as a transformation from the pose indicated by the camera to the pose desired by the controller in the robot's base space: in It is the desired pose of the controller in the robot's base space. It is the pose indicated by the camera when transitioning to the robot's base space, and It is the error matrix.

2. The robot system of claim 1, wherein the machine learning model comprises a neural network.

3. The robot system of claim 2, wherein the neural network comprises an embedding layer and a processing layer, and wherein each of the embedding layer and the processing layer comprises a multilayer perceptron.

4. The robot system of claim 3, wherein the embedding layer is configured to embed separate translational and rotational components of the pose.

5. The robot system of claim 3, wherein the embedding layer uses a rectified linear unit as the activation function, and wherein the processing layer uses a leaky rectified linear unit as the activation function.

6. The robot system of claim 2, further comprising a model training module configured to train a neural network by collecting training samples, wherein, when training the neural network, the model training module is configured to: Enables the robot controller to generate pose samples desired by the controller; The robot arm's movement is controlled based on the pose samples expected by the controller. The actual orientation of the end effector is determined using a machine vision module; and The error matrix is ​​calculated based on the controller's expected attitude samples and the actual attitude.

7. The robot system of claim 6, wherein the model training module is configured to train a neural network until the error matrix generated by the machine learning model reaches a predetermined accuracy level.

8. The robot system of claim 1 further includes a coordinate transformation module configured to transform the pose determined by the machine vision module from a camera-centric coordinate system to a robot-centric coordinate system.

9. The robot system of claim 8, wherein the coordinate transformation module is further configured to determine the transformation matrix based on a predetermined number of measured postures of the calibration target.

10. The robot system of claim 8, wherein the coordinate transformation module is further configured to associate the orientation of the component held by the end effector with the corresponding orientation of the end effector.

11. A computer-implemented method, the method comprising: The pose indicated by the camera of the end effector of the robotic arm used to complete the assembly task is determined by a machine vision module, wherein the machine vision module includes a camera configured to capture images of the work scene; Determine the desired pose of the controller corresponding to the pose indicated by the camera of the end effector, including applying a machine learning model to obtain an error matrix that associates the pose indicated by the camera with the desired pose of the controller. as well as The robot arm moves according to the robot controller’s desired posture, thereby enabling the end effector to achieve the posture indicated by the camera in order to complete the assembly task. The pose indicated by the camera of the end effector is derived from an image of the work scene; The desired posture of the robot controller is achieved by programming the robot controller. The error matrix can be expressed as a transformation from the pose indicated by the camera to the pose desired by the controller in the robot's base space: in It is the desired pose of the controller in the robot's base space. It is the pose indicated by the camera when transitioning to the robot's base space, and It is the error matrix.

12. The method of claim 11, wherein the machine learning model comprises a neural network, wherein the neural network comprises an embedding layer and a processing layer, and wherein each of the embedding layer and the processing layer comprises a multilayer perceptron.

13. The method of claim 12, wherein applying the machine learning model includes embedding translational and rotational components of the pose separately by the embedding layer.

14. The method of claim 12, wherein applying the machine learning model further comprises: Implement rectified linear units as activation functions at the embedding layer, and A leaky rectifier linear unit is implemented as the activation function at the processing layer.

15. The method of claim 12, further comprising training the neural network by collecting training samples, wherein collecting the training samples includes: Enables the robot controller to generate pose samples desired by the controller; The robot arm's movement is controlled based on the pose samples expected by the controller. The actual orientation of the end effector is determined using a machine vision module; and The error matrix is ​​calculated based on the controller's expected attitude samples and the actual attitude.

16. The method of claim 11, further comprising determining a transformation matrix for transforming the pose determined by the machine vision module from a camera-centric coordinate system to a robot-centric coordinate system.

17. The method of claim 11, further comprising determining a component transformation matrix for transforming the attitude of the component held by the end effector to a corresponding attitude of the end effector.

18. A computer-implemented method, the method comprising: The attitude error associated with the end effector of the robotic arm is modeled using a neural network; The neural network is trained using multiple training samples, where the corresponding training samples include the camera-indicated pose of the end effector and the corresponding error matrix that associates the camera-indicated pose of the end effector with the pose desired by the controller. as well as A trained neural network is used to compensate for posture errors during the operation of the robotic arm; The camera-indicated pose of the end effector is derived from an image of the working scene captured by the camera; The desired posture of the robot arm is achieved by programming the robot controller, wherein the robot controller is configured to control the movement of the robot arm; The error matrix can be expressed as a transformation from the pose indicated by the camera to the pose desired by the controller in the robot's base space: in It is the desired pose of the controller in the robot's base space. It is the pose indicated by the camera when transitioning to the robot's base space, and It is the error matrix.

19. The computer-implemented method of claim 18, wherein the neural network comprises an embedding layer and a processing layer, and wherein each of the embedding layer and the processing layer comprises a multilayer perceptron.

20. The computer-implemented method of claim 19, wherein modeling the attitude error includes embedding the translation and rotation components of the attitude separately by the embedding layer.

21. The computer-implemented method of claim 19, wherein modeling the attitude error comprises: A rectified linear unit is implemented as the activation function at the embedded layer, and A leakage rectification linear unit is implemented as the activation function at the processing layer.

22. The computer-implemented method of claim 18, further comprising collecting training samples, wherein collecting the corresponding training samples includes: Enables the robot controller to generate pose samples desired by the controller; The robot arm's movement is controlled based on the pose samples expected by the controller. The actual orientation of the end effector is determined using a machine vision module; and The error matrix is ​​calculated based on the controller's expected attitude samples and the actual attitude.