Learning visual pose estimation of robotic operations
By collecting training images on the robot arm and training neural networks, combining visual servo and flexibility control, the problem of insufficient accuracy of robot visual pose estimation in strict tolerance assembly is solved, and an efficient and economical assembly solution is achieved.
Patent Information
- Application Number
- CN202510053467.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-19
- Filing Date
- 2025-01-14
- Publication Date
- 2025-08-19
AI Technical Summary
The prior art is difficult to implement robot visual pose estimation in strict tolerance assembly operations, resulting in insufficient assembly accuracy and traditional methods are time-consuming and expensive.
By using a camera mounted on the robot arm to collect training images at different locations and lighting conditions, the neural network is trained to minimize posture differences and combined with visual servo control and compatibility control to achieve precise placement of the workpiece.
Improves robot assembly accuracy, reduces manual adjustment time, reduces cost, and maintains efficient and stable in complex environments.
Smart Images

Figure CN120503166A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for robot skill learning, and in particular to a method for robot visual posture estimation applied to high-precision placement tasks, wherein images from a camera mounted on a robot arm at different positions and lighting conditions are used to train a neural network to infer the posture of a workpiece relative to a target posture, and then the trained neural network is used to perform visual servo control on the robot performing the task. Background Art
[0002] The use of industrial robots to repetitively perform a wide variety of manufacturing and assembly operations is well known. However, some types of tight-tolerance assembly operations, such as installing a pin into a hole or inserting one component into another, remain challenging for robots to perform. These types of operations are typically performed manually because robots struggle to detect and correct the complex misalignments that can arise during tight-tolerance assembly tasks. That is, due to slight deviations in component pose caused by gripping and fixturing uncertainties, the robot cannot simply move the component to its nominal installation position but must instead "feel around" to properly align and assemble one component into another.
[0003] In order to make assembly tasks robust to these unavoidable placement uncertainties, robotic systems typically utilize force controllers (also known as compliance control or admittance control), in which force and torque feedback is used to provide the motion commands required to complete the assembly operation. The traditional way to set up and adjust force controllers for robotic assembly tasks is through manual adjustment, in which an operator programs the actual robotic system for the assembly task, runs the program, and carefully adjusts the force control parameters in a trial-and-error manner. However, using physical testing to adjust and set these force control functions is time-consuming and expensive because manual trial and error must be performed. Parameter adjustments on a real physical test system can also be dangerous because the robot is not compliant and unexpected forceful contact between components can therefore damage the robot, the components, or surrounding fixtures or structures.
[0004] Visual servo control systems are also known, which use visual images of the operating environment to guide robot motion. Visual servo control can be used to guide the robot until part-to-part contact is achieved, at which point force control takes effect. However, in some types of assembly and other operations, visual servo systems have difficulty identifying geometric features that can be used for pose detection and correction. Traditional methods often resort to manual teaching of the visual servo system for feature recognition, or require special visual markers to enable more robust recognition of robot position and orientation.
[0005] In view of the above, improved methods for robot visual pose estimation are needed, especially in tight tolerance applications. Summary of the Invention
[0006] The following disclosure describes methods and systems for robotic skill learning using visual pose estimation. A robotic arm that performs a task such as an assembly or workpiece placement operation has a camera mounted thereon. The camera provides training images of the operating scene from various positions and under various lighting conditions. For each image, the relative pose of the tool center point with respect to the target pose is recorded. A neural network is trained using the images in a supervised learning process to minimize the difference between the inferred pose and the relative pose. Once trained, the neural network is used to calculate the relative target position used in visual servo control of the robot. The robot may also employ a force controller for final placement of the workpiece once contact is made with the mating workpiece.
[0007] Additional features of the present disclosure will become apparent from the following description and appended claims, taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 is an illustration of a robotic assembly operation performed on a tight tolerance part illustrating sources of part placement uncertainty that create challenges for the robotic assembly operation;
[0009] Figure 2 is an illustration of a component undergoing robotic assembly, wherein the component needs to be aligned in such a way that the robot performs a hole search in a plane perpendicular to the axis of insertion;
[0010] Figure 3 is a block diagram illustration of a system configured for robotic assembly operations using a compliant controller (i.e., force or admittance control) as is known in the art;
[0011] Figure 4 is an illustration of a system for learning visual pose estimation in robotic manipulation, including data collection for training a pose estimation neural network, according to an embodiment of the present disclosure;
[0012] Figure 5 is configured according to an embodiment of the present disclosure to use Figure 4 A block diagram of a system for learning visual pose estimation in a neural network using image and relative pose data;
[0013] Figure 6 is a schematic diagram of an embodiment of the present disclosure Figure 5 A block diagram illustrating the structure of the visual pose estimation neural network;
[0014] Figure 7is a block diagram illustration of a system configured for robotic placement of a workpiece using a trained visual pose estimation neural network in a visual servo controller according to an embodiment of the present disclosure;
[0015] Figure 8 is a block diagram illustrating a system according to an embodiment of the present invention, wherein the system is configured to use Figure 7 Robotic placement of workpieces using a visual servo controller for preliminary placement and a compliance controller for final placement;
[0016] Figure 9 is a block diagram illustration of a system configured for robotic placement of a workpiece using integrated compliance control and visual servo control with a visual pose estimation neural network according to an embodiment of the present disclosure; and
[0017] Figure 10 1 is a flowchart of a method for learning visual pose estimation for robotic manipulation according to an embodiment of the present disclosure, including offline training of a neural network using images at different positions and lighting conditions, and online visual servoing using the trained neural network. DETAILED DESCRIPTION
[0018] The following discussion of embodiments of the present disclosure relating to systems and methods for learning visual pose estimation in robotic manipulation is merely exemplary in nature and is in no way intended to limit the disclosed technology or its applications or uses.
[0019] The use of industrial robots for a variety of manufacturing and assembly operations is well known. The present disclosure is directed to overcoming challenges encountered in many robotic operations, such as component assembly, where vision placement is employed and a considerable degree of precision is required in placing the workpiece.
[0020] Figure 1 is an illustration of a robotic assembly operation performed on a tight tolerance component, illustrating several sources of component placement uncertainty that challenge robotic assembly operations. A robot 100 having a gripper 102 grasps a first component 110 to be assembled with a second component 120. In this example, the first component 110 is a pin component, while the second component 120 is a hole structure. The pin component 110 is to be inserted into a hole in the hole structure 120. The tolerances of the components in a pin-in-hole assembly are typically very tight so that the assembly can operate without excessive looseness after assembly. Some pin-in-hole assemblies have two coaxial pins on one component, or two pins with parallel axes on one component, which must be inserted simultaneously into two holes in another component, making the assembly operation more difficult. Many other types of mating component assemblies, such as electrical connectors, complex planar shapes, etc., exhibit similar tight tolerances.
[0021] Assembly operations of this type are typically performed manually because robots struggle to detect and correct the complex misalignments that can arise in tight-tolerance assembly tasks. That is, due to slight deviations in component pose, the robot cannot simply move the components to their nominal installation positions, but must instead "feel" the alignment and assembly of one component to another. There are many possible sources of error and uncertainty in component pose. First, the precise position and orientation (collectively, "pose") of the pin component 110, as grasped in the fixture 102, can vary by very small amounts from the expected pose. Similarly, the exact pose of the hole component 120 in its fixture may also differ from the expected pose. In systems that use a camera 130 to provide images of the workspace scene for position recognition, perception errors can also contribute to relative component placement uncertainty. Furthermore, calibration errors in the placement of the robot 100 and the fixtures holding the component 120 in the workspace, as well as minor variations in the robot's joint positions, can further contribute to component placement uncertainty. These factors combine to make it impossible for the robot 100, controlled by the controller 140, to simply pick up the pin component 110 and insert it into the hole structure 120 in a single motion.
[0022] Figure 2 is an illustration of a part being robotically assembled where the part needs to be aligned in such a way that the robot performs a hole search in a plane perpendicular to the axis of insertion. Figure 1 The same method is shown for grasping part 210, which must be inserted into hole 220. Distance 230, which is exaggerated for visual effect, represents the uncertainty in the lateral position of part 210 relative to hole 220. To find the correct alignment of part 210 with hole 220, the robot may need to perform a hole search, in which gripper 202 moves part 210 back and forth in a zigzag pattern 240 in a plane perpendicular to the axis of part 210. In a system with camera image input, distance 230 can be minimized through vision control, but the robot still cannot simply insert part 210 into hole 220 in a single movement.
[0023] Figure 1 and 2 An example of an assembly operation is shown, where tight tolerances require high precision in workpiece placement. One known technique that has been developed for these types of robotic assembly operations is the use of force or compliance controllers, as described below. Other types of robotic operations, such as placing an in-process workpiece into a tool fixture or placing a finished component into a form-fitting compartment of a container, similarly require varying degrees of workpiece placement accuracy. In some of these types of operations, compliance control is inadequate, and accurate visual workpiece placement is preferred.
[0024] Figure 3 is a block diagram illustration of a system 300 configured for robotic assembly operations using a compliance controller (ie, force or admittance control) as is known in the art. In the physical world, the robotic controller 310 interacts with objects such as Figure 1 The controller 310 provides joint motion commands to the robot 100 and receives status feedback from the robot 100, as is known in the art and discussed below. Figure 1 As shown, the robot 100 has a gripper 102 that grasps a first component 110 , and the controller 310 provides commands to assemble the first component 110 with (or into) a second component 120 .
[0025] Block 320 represents the controller 310 and the robot 100 in block diagram form. The controller 310 is configured as a compliance controller, the functionality of which is discussed below. Block 330 provides a nominal target position for the first part 110. The nominal target position can be predetermined and constant for a particular robot work cell, or the nominal target position can be provided by a vision system based on an observed position of the second part 120, which can be moving on a conveyor, for example. For the purposes of this discussion, it is assumed that the position of the second part 120 in the robot work cell is known, and the nominal target position from block 320 defines the position of the first part 110 for mounting it into the second part 120. The nominal target position of the first part 110 can then be converted into fixture coordinates, which can then be converted into robot joint positions using inverse kinematics in a known manner.
[0026] Summing node 340 is included after block 330. Although Figure 3 Node 340 does not have a second input, but a second input may be added in some embodiments. Block 350 defines motion limits for the robot 100. The motion limits ensure that the robot 100 does not take excessively large motion steps during the control process, where excessive motion could create a dangerous situation or result in forceful contact between the robot 100 and / or parts 110 / 120 and each other or other objects in the work cell. If the difference between the target position and the current position is greater than the motion limit, the motion limit takes precedence and limits the step size.
[0027] Block 360 includes an admittance control function that interacts with a robot in block 370 that performs an assembly task in block 380. Blocks 370 and 380 represent the physical actions of the robot 100 as it installs the first component 110 into the second component 120. The robot in block 370 provides state feedback to the admittance control function in block 360 on line 372. The state feedback provided on line 372 includes the robot's joint states (position / velocity) and contact forces and torques. Alternatively, the position and velocity state data can be provided in Cartesian coordinates, which can be easily converted to joint coordinates, or vice versa, via the aforementioned transformation calculations. Force and torque sensors (not shown) are required for operation of the robot 100 to measure the contact forces between the components 110 and 120. The force and torque sensors can be positioned between the robot 100 and the fixture 102, between the fixture 102 and the first component 110, or between the second component 120 and its "ground" (fixture). Contact forces and torques can also be measured from robot joint torque sensors or estimated from other signals such as motor currents.
[0028] As is known in the art and only briefly discussed here, the admittance control function in block 360 operates in the following manner. Impedance control (or admittance control) is a method of dynamic control involving force and position. It is often used in applications where a manipulator interacts with its environment and the force-position relationship is of interest. Mechanical impedance is the ratio of force output to motion input. Controlling the impedance of a mechanism means controlling the resistance to external motion imposed by the environment. Mechanical admittance is the inverse of impedance and defines the motion produced by a force input. The theory behind the impedance / admittance control method is to treat the environment as an admittance and the manipulator as an impedance.
[0029] Using the target position from node 340 and the motion constraints from block 350, the admittance control function in block 360 calculates the target velocity (in the six degrees of freedom) to move the workpiece from its current position to the target position (or step size where the motion is constrained). The admittance control function then calculates the command velocity by adjusting the target velocity with a force compensation term using an equation such as: Where V is the command velocity vector (this equation applies to translational motion), V d is the target velocity vector, is the inverse of the admittance gain matrix, and F is the contact force vector measured from a force sensor mounted to the robot or workpiece. In this example, the vectors each include three translational degrees of freedom. A similar equation is used to calculate the rotational command velocity ω using contact torque feedback.
[0030] Then, the command velocity calculated as described above is converted to the command joint velocity with respect to all robot joints by multiplying the inverse of the Jacobian matrix J by the transpose of the command velocity vector as follows: After calculating the command joint velocity, a low-pass filter may also be provided to ensure smoothness and feasibility of the command velocity. The calculated command joint velocity is provided to the robot, which moves and measures the new contact forces. The target position is again compared to the current position, and the velocity calculation is repeated. While attempting to reach the target position from node 340, the admittance control function in block 360 repeatedly provides motion commands to the robot in block 370 using force feedback and robot state feedback on line 372.
[0031] Elements 330-360 are programmed as algorithms in the controller 310. The interaction of the controller 310 with the robot 100 occurs in the real world via cables (or wirelessly), and this interaction is represented in box 320 by the forward arrow (motion command) and feedback line 372 (joint state and contact forces) from box 360 to box 370.
[0032] As mentioned previously, conventional compliance control techniques are not effective for all types of assembly tasks. For example, in tight tolerance assembly operations, and in Figure 2 In the event that the initial position in block 370 deviates significantly from the target position, the robot in block 360 may spend a long time in the feedback loop with the admittance control function in block 360 and may ultimately fail to complete the assembly task, including the possibility of component damage in the process. This situation can be somewhat mitigated by fine-tuning the impedance / admittance control parameters, but this is only effective in some cases and only for specific workpiece assembly operations.
[0033] Another robotic control technology that has been developed for assembly and other precision motion applications is visual servoing. In visual servoing, rather than controlling the robot to move a workpiece to a nominal target position at a predetermined, constant location, the desired workpiece movement is dynamically adjusted based on images of the workpiece and the operating environment. As the robot moves the workpiece, new images are provided, and the adjusted target position is refined in a feedback control loop.
[0034] Similar to compliance control, visual servoing has been shown to be effective in some robotics applications. However, visual servoing also has drawbacks, including occlusions in the image, depth inaccuracies when using 2D cameras, processing speed when using 3D cameras, and others, and visual servoing technology is therefore often unable to meet the challenges of robotic assembly and other precision motion applications.
[0035] The techniques of the present disclosure have been developed to address the shortcomings of existing visual servoing and compliance control techniques and are discussed in detail below.
[0036] Figure 41 is an illustration of a system for learning visual pose estimation in robotic manipulation, including data collection for training a pose estimation neural network, according to an embodiment of the present disclosure. Robot 400 is similar to robot 100 discussed previously. Robot 400 includes a multi-jointed robot arm having a gripper 402 at the end of the arm. The gripper 402 grasps a workpiece 410 to be installed in or assembled with a workpiece 420. The gripper 402 and the workpiece 410 are positioned in a manner similar to the embodiment of the present disclosure. Figure 4 The is shown in multiple positions for reasons explained below.
[0037] The robot 400 communicates with a controller 440, which controls the movement of the robot 400 and also receives data from the robot 400. A camera 450 is mounted on an outer robot arm so that the camera 450 is fixed in orientation relative to the fixture 402. The camera 450 has a field of view that preferably includes the workpiece 410 and the mating workpiece 420 in most images as the robot 400 moves. The camera 450 is preferably a two-dimensional (2D) camera with color (i.e., "RGB") imaging capabilities. 2D cameras are preferred because 2D camera images are processed much faster than three-dimensional (3D) camera data, and the technology of the present disclosure eliminates the need for 3D camera data. The camera 450 provides its images to the controller 440, or to a separate computer (not shown) that is also in communication with the controller 440.
[0038] A plurality of lights (460, 462, 464) are arranged in the workspace of the robot. The lights 460-464 serve as a controllable source of illumination in the workspace where the robot is performing operations. Each of the lights 460-464 can be any type of light fixture deemed suitable for illuminating the workspace, including wide-dispersion ambient lights (e.g., overhead fluorescent lights), floodlights, etc. The lights 460-464 can simply be available light fixtures that are already installed in the building (e.g., a factory or warehouse) where the robot's workspace is located. The purpose and operation of the lights 460-464 are discussed further below.
[0039] According to the technology disclosed in this disclosure, Figure 4The system is used to collect data for training a pose estimation neural network, which will be used for visual servo robot control after training. The data used for training is in the form of multiple images from camera 450, with the robot pose recorded for each image. The recorded robot pose is the relative pose in six degrees of freedom of the tool center point relative to the target pose (when the workpiece 410 is assembled with the workpiece 420). The images are taken by camera 450, the robot 400 is configured so that the fixture 402 and the workpiece 410 are in multiple positions near the target pose, and the lights 460-464 provide a variety of lighting conditions. Each image and its corresponding relative pose are recorded by controller 440 or a separate computer, and this data is used for neural network training. Control of the robot position, lighting conditions and image capture can be manual or automatic.
[0040] By training with multiple images depicting scenes with various relative poses and various lighting conditions, the pose estimation neural network becomes robust to variations in the part grasp orientation and starting position, robust to camera calibration parameters, and robust to changes in lighting conditions and shadows. The details of the neural network training method and the subsequent use of the trained neural network for visual servoing robot control are discussed below.
[0041] Figure 5 is a configuration according to an embodiment of the present disclosure for using Figure 4 Block diagram illustration of a system 500 for learning visual pose estimation in a neural network using image and relative pose data of the system. Figure 5 The neural network training system can be used in the controller 440 or above about Figure 4 The execution of the program from the stand-alone computer discussed above, or on another computer that has access to the input data, is performed. Figure 4 input of the system.
[0042] Box 510 contains multiple images of the workpiece mounting scene from camera 450. As described above, the images in box 510 are captured from various positions near the target pose under various lighting conditions. The positions captured in the images can form a grid pattern or some other geometric or defined pattern around the target pose, in which case the movement of the robot 400 and the triggering of each image by the camera 450 can be defined in the control program. Alternatively, the movement of the robot 400 and the capture of the images can be manually controlled by an operator (e.g., using a teach pendant). In any case, the position variation covers a range of offset distances in at least two lateral directions (and optionally vertical) and offset angles around all three axes. For example, the position deviation range can be defined as + / - 60 mm in two lateral directions (e.g., X and Y), no deviation in the vertical (Z) direction, + / - 5 degrees around the pitch and roll axes, and + / - 10 degrees around the yaw axis (the axis of the gripper). The range can be selected to suit a particular application. While the robot remains in a given pose, the lighting conditions can change, and / or the lighting conditions can change from one pose to the next. The lighting condition changes may include any combination of turning on, off, and / or dimming individual lights in the lights 460-464. An image is captured at each unique combination of robot pose and lighting condition. All images are provided in block 510.
[0043] Box 520 contains the relative pose corresponding to each image in box 510, recorded by the robot controller 440. The relative pose defines the robot tool center point position relative to the target workpiece position and, as described above, is captured for each image (each unique combination of robot pose and lighting conditions). The relative pose provided for each image is a vector defining all six degrees of freedom; for example, x / y / z offset distances and yaw / pitch / roll offset angles. Each image in box 510 has its own uniquely identified relative pose vector in box 520.
[0044] Box 530 is the visual pose estimation neural network. The structure of the neural network 530 will be referred to below. Figure 6 Block 540 is the inferred pose for each image output from neural network 530. For each image, the relative pose from block 520 is compared to the inferred pose from block 540, and the difference is calculated as a numerical value at summing node 550. The difference value can be a weighted sum of the difference values (relative pose minus inferred pose) in six dimensions (x, y, z, yaw, pitch, roll).
[0045] For each image, the difference value from summing node 550 is used in a cost function to train neural network 530 in a supervised learning process. This is illustrated by dashed line 560 passing back through neural network 530. This supervised learning is an automatic process in which a large difference on line 560 (between the relative pose and the inferred pose) tells neural network 530 that its pose estimate for that image is not very good, and conversely, a small difference on line 560 tells neural network 530 that its pose estimate for that image is good. With a sufficient number of training images, neural network 530 learns which parameter settings provide the most accurate estimate of the relative pose of the workpiece relative to the mounted target pose.
[0046] Figure 6 is a schematic diagram illustrating an embodiment of the present disclosure Figure 5 A block diagram illustrating the structure of the visual pose estimation neural network 530. Figure 6 In the preferred embodiment shown, visual pose estimation neural network 530 includes portion 532 comprising layers from a pre-trained neural network and portion 534 comprising linear projection layers customized to match task data from the robot operation.
[0047] Section 532 includes several layers of neural networks that have been developed and pre-trained to be effective in feature extraction from images. In these neural networks, the structure (number of layers and nodes, node connectivity, etc.) and preliminary values of parameters (weights) are pre-trained for feature extraction effectiveness. Such pre-trained neural network packages are available from various commercial and public domain sources. By using the portions of the pre-trained neural network in section 532, the training time of the visual pose estimation neural network 530 is significantly reduced compared to starting with a "blank slate" neural network architecture.
[0048] The linear projection layer in portion 534 performs linear matrix multiplication to project the higher dimensional discrete vector (from portion 532) into a lower dimensional continuous vector. The exact matrix multiplication parameters are learned through a training process.
[0049] exist Figure 6 In the example above, one of the images (510A) from block 510 is provided to a visual pose estimation neural network 530. Using the current parameters in the feature extraction layer portion 532 and the linear projection layer 534, the visual pose estimation neural network 530 outputs an inferred pose 540A corresponding to the image 510A. This is exactly the same as before. Figure 5As shown. Then, supervised learning training is performed on the neural network 530 using a cost function based on the difference between the relative pose (corresponding to image 510A) and the inferred pose 540A. In block 510, an incremental training process is performed for each image. In a preferred embodiment, only the parameters of the neural network 530 are modified during training (in the feature extraction layer portion 532 and the linear projection layer 534); the structure is not modified during training.
[0050] Figure 4-6 The structure of the visual pose estimation neural network 530 described in and above and the corresponding training method have proven to be advantageous for several reasons. First, the training process is easy and is performed using Figure 4 The data collected is automatically processed. No manual calibration of the neural network 530 is required in terms of structure or parameters, and no physical features need to be identified or selected in the image. Furthermore, no calibration of the camera 450 is required, as any camera calibration or alignment inaccuracies are automatically compensated for during the neural network training process.
[0051] Furthermore, neural network training and execution are fast. In actual experiments, training of neural network 530 has been shown to converge to highly accurate pose estimates in just a few minutes on readily available computing equipment. This training is performed using a number of training images (in block 510) in the low hundreds. Once trained, execution of neural network 530 (in inference mode) is also very fast, on the order of a few milliseconds. This execution in inference mode is fast enough for production robotic operations, where neural network 530 receives camera images and infers relative poses in visual servo robotic control for workpiece placement.
[0052] The visual pose estimation neural network 530 also provides accurate (at sub-millimeter levels) and robust results to changes in lighting conditions. This accuracy and robustness are necessary for high-precision placement tasks in real-world environments where shadows and poor / variable lighting conditions are a reality.
[0053] The simple and straightforward data collection, fast and automatic training, and the accuracy and robustness of pose estimates obtained from the trained networks make visual pose estimation neural networks and the related training methods discussed above highly effective in visual servoing robotic control applications.
[0054] Figure 7 is a block diagram illustration of a system 700 configured for robotic placement of a workpiece using a trained visual pose estimation neural network in a visual servo controller according to an embodiment of the present disclosure. Block 710 is configured as a robotic visual servo controller. Figure 4 ) robot 400 is shown in box 710. Figure 7The robot and camera in the visual servo control system do not have to be Figure 4 The same single device for data collection in the data collection system ( Figure 4 ) and production systems ( Figure 7 The visual servo controller block 710 may be executed on the controller 440 or a similar robotic controller.
[0055] Camera 450 (mounted on the robot arm, as previously discussed) provides an image of the workpiece placement scene to block 710. Specifically, camera 450 provides the image to previously trained visual pose estimation neural network 530. Upon receiving the image, trained neural network 530, operating in inference mode, outputs an estimated relative target pose as described above. Thus, an estimated target pose of the workpiece relative to the robot's current position is provided on line 720 for processing. The target pose on line 720 defines the position to which the robot 400 needs to move in order to properly place the workpiece.
[0056] The target pose (relative to the current position) can be provided to an optional motion constraint module 730, which ensures that the robot 400 does not take excessively large motion steps during the control process. The target position (if motion constrained, if appropriate) is provided to a position planner module 740. The position planner module 740 calculates the robot motion required to move the workpiece from the current position to the target position. For example, the module 740 can calculate a spline function designed to move the workpiece from its current (6-degree-of-freedom) position and orientation to the target position and orientation provided by the neural network 530. Based on the spline curve trajectory, the position planner module 740 can calculate the corresponding robot joint motion using inverse kinematics calculations in a known manner.
[0057] As indicated by the forward arrow, the position planner module 740 provides robot motion commands to the robot 400. The robot provides state data (e.g., joint positions and velocities) back to the position planner module 740 on feedback line 742. The feedback control loop from the position planner module 740 to the robot 400 and back on line 742 can continue for several steps in real time as the robot moves along the calculated path (e.g., the spline curve trajectory described above). There is also a feedback loop on line 750 back to the pose estimation neural network 530. This is the essence of a visual servo control system that moves the robot toward a target a certain distance, then acquires another image and updates the control commands based on the new image. In Figure 7In the case of a system, the new image is processed by the pose estimation neural network 530, the new relative target pose is estimated by the neural network 530, and the new target is provided to the position planner module 740, which calculates new robot motion commands. This process continues until the robot reaches the target position and the neural network 530 indicates that the target position is equal to the current position.
[0058] If the workpiece placement or assembly operation is not particularly difficult, Figure 7 The visual servo controller system can complete the operation, release the workpiece and return to the starting position to pick up a new workpiece. However, in very tight tolerance applications, Figure 7 The visual servo control method may need to be integrated with or followed by the force controller operation in order to complete the workpiece placement or assembly. The combination of visual servo control and compliance control is shown in the following two figures and discussed below.
[0059] Figure 8 is a block diagram illustration of a system 800 configured to use Figure 7 Robotic placement of workpieces is performed using a visual servo controller for preliminary placement and a compliance controller for final placement. Figure 8 The system can be run on controller 440 or a similar controller. The visual servoing controller in block 710 is Figure 8 shown in the upper part of, and exactly as above with respect to Figure 7 Operationally, block 710 receives images from camera 450 and visual pose estimation neural network 530 estimates a relative target pose that is used by the position planner to control the robot. Figure 8 In the system, visual servo control is used to perform preliminary workpiece placement in the first step (indicated at ①).
[0060] After the preliminary workpiece placement, the final workpiece placement (e.g., installation or assembly) is performed in a second overall step (as indicated at ②). The final workpiece placement is performed using the compliance controller shown in block 820. The compliance controller in block 820 receives the target position from the neural network 530 on line 810. This transfer of the target position from the visual servo controller block 710 to the compliance controller block 820 is a one-time transfer that indicates what relative position movement is required to complete the workpiece placement. Afterwards, the visual servo controller block 710 is no longer active.
[0061] In the compliance controller block 820, the relative target position is received in block 830. From that point, the compliance controller block 820 operates as before for Figure 3320 of . That is, the target position is motion constrained if necessary (block 840) and is provided to the admittance control block 850, which communicates with the robot 400, provides motion commands to the robot 400 and receives state feedback (including forces and torques and robot motion) on line 852. The assembly task is shown in block 860, and the interaction of the robot 400 with the physical environment (the workpiece contacts the mating component in a manner that produces contact forces and torques) is depicted by the double-headed arrows. The compliance controller block 820 operates in a feedback loop (admittance control block 850 and robot 400) until the workpiece is placed in the final target position. Figure 8 In the system, there is no feedback loop to box 830 to determine the new target position.
[0062] Figure 8 The two-level control shown in is one embodiment of a system that combines visual servo control (using a visual pose estimation neural network) with compliance control for precision workpiece placement applications. Figure 9 Another embodiment is shown where visual servo control and compliance control are integrated in nested loops.
[0063] Figure 9 is a block diagram illustration of a system 900 configured for robotic placement of a workpiece using integrated compliance control and visual servo control with a visual pose estimation neural network, according to an embodiment of the present disclosure. Figure 9 The system may be run on controller 440 or a similar controller. An integrated visual servoing and compliance controller is shown in block 910. Camera 450 provides images of the workpiece placement scene to a trained visual pose estimation neural network 530, which operates as previously described to estimate relative target position.
[0064] The target position from the neural network 530 is motion-constrained, if necessary (block 920), and is provided to the admittance control block 930, which communicates with the robot 400, providing motion commands to the robot 400 and receiving state feedback (including forces and torques as well as robot motion) on line 932. The assembly task is shown in block 940, and the interaction of the robot 400 with the physical environment (the workpiece contacts the mating component in a manner that generates contact forces and torques) is depicted by the double-headed arrows as previously described. The admittance control block 930 and the robot 400 operate in an inner feedback loop for several cycles, after which the robot position state is provided to the pose estimation neural network 530 on an outer feedback loop line 950. Using the new image from the camera 450 and the actual relative pose on line 950, the neural network 530 estimates a new relative target pose, which is provided to the admittance control block 930.
[0065] Figure 9The system provides integrated control that leverages the benefits of visual servo control using a pose estimation neural network 530, which is particularly effective for large motions as the workpiece moves from an initial position toward a target position, and compliance control, which is particularly effective for precise seating motions required during tight tolerance placement of the workpiece, such as when assembled with a mating component.
[0066] Whether used in a robot controller employing pure visual servo control or incorporated into a robot controller that also employs compliance control, the visual pose estimation neural network 530 provides the engine for workpiece placement. The visual pose estimation neural network 530 is easily trained using the techniques discussed previously, is cost-effective, accurate, and fast using a 2D camera, and is robust to varying and suboptimal lighting conditions due to the variations included in the training image set.
[0067] Figure 10 1000 is a flowchart of a method for learning visual pose estimation for robotic manipulation according to an embodiment of the present disclosure, including offline training of a neural network using images in different positions and lighting conditions, and online visual servoing using the trained neural network.
[0068] At block 1002, a robot / camera system is provided for data collection. Figure 4 The system shown in , has a robot 400 and a controller 440, and an arrangement for mounting a first workpiece 410 into a second workpiece 420. This is all arranged in a work cell with lights 460 / 462 / 464 that can be controlled to change the lighting conditions.
[0069] At block 1004, training images are collected using the robot / camera system. As previously described, the robot 400 moves the gripper 402 with the workpiece 410 to various positions near the target position and changes the lighting conditions while capturing many images. For each image captured, the relative pose of the workpiece with respect to the target pose is also recorded by the controller 440. The images and poses are stored in a file or database 1006. The database 1006 may reside on the controller 440, or alternatively may reside on a separate computer ( Figure 4 For example, a separate computer may store the training data (database 1006) and may also be used in the neural network training process.
[0070] At block 1008, a pre-configured pose estimation neural network is provided. Figure 5 and 6 The neural network 530 is preferably configured as Figure 6The design shown and discussed previously is pre-configured. At block 1010, a pose estimation neural network is trained in a supervised learning process using the image and relative pose data from database 1006. Figure 5 As discussed, training involves the neural network inferring relative poses with respect to images, and the difference between the recorded relative pose and the inferred relative pose is used in a cost function for training the parameters of the neural network, where large differences are penalized and small differences are rewarded.
[0071] The above process is repeated until a maximum count is reached, or until the neural network converges to a consistently accurate inferred pose that satisfies a certain threshold of difference between the recorded relative pose and the inferred relative pose. At decision diamond 1012, when the neural network performance has not converged to a satisfactory level, the process loops back to block 1010 for the next image. From decision diamond 1012, when the neural network performance has converged to a satisfactory level, the process moves to block 1014.
[0072] At block 1014, the trained pose estimation neural network is used in inference mode for visual servo control of the robot. This may include Figure 7-9 Any control system architecture shown in , where the trained visual pose estimation neural network 530 is used for a pure visual servo control system ( Figure 7 ), or visual servo control of the neural network 530 for preliminary placement, followed by compliance control for final placement ( Figure 8 ), or neural network 530 for integrated visual servo / compliance control system ( Figure 9 ). In all of these system architectures, the trained neural network 530 provides an inferred or calculated value of the relative object pose from the most recent camera image. Figure 10 The method steps may be performed on the robot controller 440, and optionally on other robot controllers and separate computers as described above. For example, neural network training may be performed on a computer other than the robot controller, and the trained neural network then provided and used in a control module running on the robot controller.
[0073] The methods and systems disclosed herein enable simple, automatic training of visual pose estimation neural networks, which are then fast and accurate enough to be used for real-time visual servoing robotic control in conjunction with compliance control, where appropriate. The trained neural networks are robust to changes in lighting conditions resulting from the training data image set, and the system employs cost-effective 2D cameras in both training and inference modes. In these ways, the disclosed methods and systems offer significant improvements over conventional image-based robotic control systems.
[0074] Throughout the foregoing discussion, various computers and controllers are described and implied. It should be understood that the software applications and modules of these computers and controllers are executed on one or more computing devices having processors and memory modules that are configured to learn visual pose estimation in robotic operations. In particular, this includes the processor in the robot controller 440 and the processors for Figure 4 Optional separate computer for data collection in Figure 5 The computer of the training process, as well as the robot controller and optionally other computers, which perform the functions of visual servoing and compliance controller, including Figure 7-9 Operation of the visual pose estimation neural network in inference mode as depicted in [1].
[0075] The foregoing discussion discloses and describes only exemplary embodiments of the present disclosure. Those skilled in the art will readily recognize from such discussion and from the accompanying drawings and claims that various changes, modifications and variations may be made therein without departing from the spirit and scope of the present disclosure as defined in the appended claims.
Claims
1. A learning visual pose estimation robot system, the system comprising: a robot having a gripper configured to perform an operation on a workpiece; a camera coupled to an outer arm of the robot proximate to the gripper, the camera providing an image of a workpiece manipulation scene; one or more lights that illuminate the robot's workspace; as well as at least one computing device in communication with the robot and the camera, the at least one computing device configured with a neural network, wherein the neural network is trained for visual pose estimation using a plurality of images having various workpiece positions and various lighting conditions along with an actual relative pose for each image, and after training, the neural network operates in an inference mode to perform visual pose estimation for use in visual servo control of the robot performing the operation.
2. The system according to claim 1, wherein: The workpiece operation scene in each image includes at least a portion of the workpiece in the fixture of the robot and a placement target area, and the actual relative pose with respect to each image defines the relative position of the workpiece relative to the target position determined based on the robot joint position.
3. The system according to claim 1, wherein: The at least one computing device is configured with a supervised learning algorithm that trains the neural network for visual pose estimation by computing, for each of the plurality of images, a difference between an inferred pose from the neural network with respect to the image and the actual relative pose with respect to the image, and applying a cost function that rewards small differences and penalizes large differences.
4. The system according to claim 1, wherein: The various lighting conditions are achieved by turning on, turning off, and / or dimming the one or more lights individually or collectively.
5. The system according to claim 1, wherein The neural network has a structure including a plurality of layers pre-configured for image feature extraction and a linear projection layer that receives outputs from the plurality of layers and provides an inferred pose.
6. The system according to claim 5, wherein: Parameter values of the neural network are modified during training to improve the accuracy of the inferred pose, without modifying the structure.
7. The system according to claim 1, wherein: In the visual servo control of the robot, the neural network receives camera images and calculates an inferred relative pose, and a position planning module calculates robot joint motions required to move the workpiece to a target position based on the inferred relative pose.
8. The system according to claim 1, wherein: The visual servo control is used in conjunction with compliance control of the robot performing the operation.
9. The system according to claim 8, wherein: The visual servo control is used to perform preliminary placement of the workpiece and the compliance control is subsequently used to perform final placement of the workpiece, or during placement of the workpiece, the visual servo control operates in an outer feedback control loop and the compliance control operates in an inner feedback control loop.
10. The system according to claim 1, wherein: The camera is a two-dimensional (2D) camera.
11. The system according to claim 1, wherein: The operation is moving the workpiece to a destination location, or assembling the workpiece with or into a second workpiece.
12. The system according to claim 1, wherein: The at least one computing device is a robot controller that controls movement of the robot and receives joint state data from the robot, and also performs training of the neural network.
13. The system of claim 1, wherein: The at least one computing device includes a computer in communication with a robot controller, wherein the computer receives images from the camera and joint position data from the robot controller and performs training of the neural network, and the robot controller performs the visual servo control of the robot using the trained neural network.
14. A learning visual pose estimation robot system, the system comprising: a robot having a gripper configured to perform an operation on a workpiece; a camera coupled to an outer arm of the robot proximate to the gripper, the camera providing an image of a workpiece manipulation scene; one or more lights that illuminate the robot's workspace; and at least one computing device in communication with the robot and the camera, wherein the at least one computing device is configured with a supervised learning algorithm that trains a neural network for visual pose estimation using a plurality of images having various workpiece positions and various lighting conditions, together with an actual relative pose for each image, wherein the supervised learning algorithm computes, for each image of the plurality of images, a difference between an inferred pose from the neural network for the image and the actual relative pose for the image, and trains parameters of the neural network using a cost function by rewarding small differences and penalizing large differences, And after training, the at least one computing device operates the neural network in an inference mode to perform visual pose estimation used in visual servo control of the robot performing the operation, wherein the neural network receives camera images and calculates an inferred relative pose, and a position planning module calculates the robot joint motion required to move the workpiece to a target position based on the inferred relative pose.
15. A method for learning visual pose estimation in robotic manipulation, the method comprising: providing a robot having a gripper configured to perform an operation on a workpiece, one or more lights to illuminate a workspace of the robot, and a camera coupled to an outer arm of the robot proximate the gripper, wherein the camera is configured to provide an image of a workpiece operation scene; collecting training data by a computing device in communication with the robot and the camera, wherein the training data includes a plurality of images having various workpiece positions and various lighting conditions and an actual relative pose for each image; training a neural network for visual pose estimation using the training data; as well as The neural network is run in an inference mode to perform visual pose estimation for use in visual servo control of the robot performing the operation.
16. The method according to claim 15, wherein The workpiece operation scene in each image includes at least a portion of the workpiece in the fixture of the robot and a placement target area, and the actual relative pose with respect to each image defines the relative position of the workpiece relative to the target position determined based on the robot joint position.
17. The method according to claim 15, wherein: The computing device is configured with a supervised learning algorithm that trains the neural network for visual pose estimation by computing, for each of the plurality of images, a difference between an inferred pose from the neural network with respect to the image and the actual relative pose with respect to the image, and applying a cost function that rewards small differences and penalizes large differences.
18. The method according to claim 15, wherein The various lighting conditions are achieved by turning on, turning off, and / or dimming the one or more lights individually or collectively.
19. The method according to claim 15, wherein The neural network has a structure including a plurality of layers pre-configured for image feature extraction and a linear projection layer that receives outputs from the plurality of layers and provides an inferred pose, and wherein parameter values of the neural network are modified during training to improve the accuracy of the inferred pose.
20. The method according to claim 15, wherein In the visual servo control of the robot, the neural network receives camera images and calculates an inferred relative pose, and a position planning module calculates robot joint motions required to move the workpiece to a target position based on the inferred relative pose.
21. The method according to claim 15, wherein The visual servo control is used in conjunction with compliance control of the robot performing the operation.
22. The method according to claim 21, wherein The visual servo control is used to perform preliminary placement of the workpiece and the compliance control is subsequently used to perform final placement of the workpiece, or during placement of the workpiece, the visual servo control operates in an outer feedback control loop and the compliance control operates in an inner feedback control loop.
23. The method according to claim 15, wherein The operation is moving the workpiece to a destination location, or assembling the workpiece with or into a second workpiece.