Learning of visual sensation pose estimation for robot work
A neural network trained with images from a robot-mounted camera addresses positioning uncertainties in industrial robots, enabling precise assembly by combining visual and force controllers for improved safety and efficiency in tight-tolerance tasks.
Patent Information
- Application Number
- JP2025024045
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-19
- Filing Date
- 2025-02-18
- Publication Date
- 2025-08-29
AI Technical Summary
Industrial robots face challenges in performing tight-tolerance assembly tasks due to positioning uncertainties, requiring manual tuning of force controllers and conventional visual servoing systems struggle with feature recognition, making these tasks time-consuming and potentially dangerous.
A method using a robot-mounted camera to collect images under varying conditions for training a neural network to estimate the pose of a workpiece, enabling visual servoing with a combination of visual and force controllers for precise assembly.
The method provides robust and accurate visual pose estimation, reducing manual tuning time, enhancing safety, and improving precision in tight-tolerance assembly tasks.
Smart Images

Figure 2025126909000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to methods for learning robotic skills, and more particularly to a method for visual pose estimation of a robot applicable to high-precision positioning tasks, in which images from a robot arm-mounted camera at different positions and lighting conditions are used to train a neural network to infer the pose of a workpiece relative to a target pose, and the trained neural network is then used for visual servoing of the robot to perform the task. [Background technology]
[0002] The use of industrial robots to perform a wide range of repetitive manufacturing and assembly tasks is well-known. However, tight-tolerance types of assembly tasks, such as placing a peg into a hole or inserting one part into another, remain difficult for robots to perform. These types of tasks are often performed manually due to the difficulty for robots to detect and correct the complex misalignments that can occur in tight-tolerance assembly tasks. That is, due in part to slight deviations caused by uncertainties in both gripping and fixturing, a robot cannot simply move a part to its nominal placement position but rather must "feel around" to properly align and mate one part with another. Summary of the Invention [Problem to be solved by the invention]
[0003] To make assembly tasks robust to these unavoidable positioning uncertainties, robotic systems typically utilize force controllers (also known as compliance or admittance control) that use force and torque feedback to provide the motion commands necessary to complete the assembly task. The traditional method for setting up and tuning force controllers for robotic assembly tasks is manual tuning, in which a human operator programs the actual robot system for the assembly task, runs the program, and carefully adjusts the force control parameters through trial and error. However, tuning and setting these force control functions using physical testing is time-consuming and costly due to the manual trial and error required. Additionally, tuning parameters in an actual physical test system can be dangerous because the robot may not be compatible, potentially damaging the robot, the parts, or surrounding fixtures and structures due to unexpected forced contact between parts.
[0004] Visual servoing control systems are also known that use visual images of the operating environment to guide the robot's motion. Visual servoing may be used to guide the robot until contact between parts occurs, at which point force control takes over. However, in some types of assembly and other tasks, it is difficult for visual servoing systems to identify geometric features that can be used for pose detection and correction. Conventional methods often require manual training of the visual servoing system for feature recognition or special visual markers that enable more robust recognition of the robot's position and orientation.
[0005] In view of the above, there is a need for improved methods for visual pose estimation of robots, especially in tight tolerance applications. [Means for solving the problem]
[0006] The following disclosure describes a method and system for learning robotic skills using visual pose estimation. A robot arm performing a task, such as an assembly or workpiece positioning task, is equipped with a camera. The camera provides training images of the work scene from various positions and under various lighting conditions. For each image, the relative pose of the tool center point with respect to a target pose is recorded. The images are used in a supervised learning process to train a neural network to minimize the difference between the estimated pose and the relative pose. The trained neural network is used to calculate a relative target position used for visual servoing of the robot. The robot may also use a force controller for final workpiece positioning upon contact with a mating part.
[0007] Additional features of the present disclosure will become apparent from the following description and claims, taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 illustrates a robotic assembly operation being performed on tight tolerance parts and shows sources of part positioning uncertainty that pose challenges for the robotic assembly operation.
[0009] [Figure 2] 10A-10C illustrate parts being assembled by a robot that require alignment in a manner that causes the robot to perform hole searching in a plane perpendicular to the insertion axis.
[0010] [Figure 3] FIG. 1 is a block diagram of a system configured for robotic assembly operations using compliance control (i.e., force or admittance control), as known in the art.
[0011] [Figure 4]FIG. 1 is a diagram of a system for learning visual pose estimation for robotic tasks, including data collection for training a pose estimation neural network, according to an embodiment of the present disclosure.
[0012] [Figure 5] FIG. 5 is a block diagram of a system configured to train visual pose estimation in a neural network using image and relative pose data from the system of FIG. 4 according to an embodiment of the present disclosure.
[0013] [Figure 6] FIG. 6 is a block diagram that schematically illustrates the structure of the visual pose estimation neural network from FIG. 5, according to an embodiment of the present disclosure.
[0014] [Figure 7] FIG. 1 is a block diagram of a system configured for robotic workpiece positioning using a trained visual pose estimation neural network in a visual servoing controller, according to an embodiment of the present disclosure.
[0015] [Figure 8] FIG. 8 is a block diagram of a system configured for robotic workpiece positioning using the visual servo controller of FIG. 7 for pre-positioning and a compliance controller for final positioning, according to an embodiment of the present disclosure.
[0016] [Figure 9] FIG. 1 is a block diagram of a system configured for robotic workpiece positioning using joint compliance control and visual servoing with a visual pose estimation neural network according to an embodiment of the present disclosure.
[0017] [Figure 10]FIG. 1 is a flowchart diagram of a method for learning visual pose estimation for robotic tasks, including offline training of a neural network using images at different positions and lighting conditions, and online visual servoing using the trained neural network, according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0018] The following description of embodiments of the present disclosure directed to systems and methods for learning visual pose estimation in robotic tasks is merely exemplary in nature and is not intended to limit the disclosed techniques or their application or uses.
[0019] The use of industrial robots for a wide variety of manufacturing and assembly tasks is well known, and the present disclosure is directed to overcoming challenges encountered in many robotic tasks, such as part assembly, where visual positioning is performed and a great deal of precision in workpiece placement is required.
[0020] FIG. 1 is an illustration of a robotic assembly operation being performed on tight-tolerance parts, illustrating some sources of part positioning uncertainty that present challenges for robotic assembly operations. A robot 100 with a gripper 102 grasps a first part 110 to be assembled with a second part 120. In this example, the first part 110 is a peg part, and the second part 120 is a hole structure. The peg part 110 is to be inserted into a hole in the hole structure 120. The tolerances of the parts in a peg-in-hole assembly are typically quite tight, so the assembly can operate without excessive looseness after assembly. Some peg-in-hole assemblies have dual coaxial pegs on one part or dual parallel-axis pegs on one part, which must simultaneously insert into dual holes in the other part, further complicating the assembly operation. Many other types of mating part assemblies, such as electrical connectors and complex planar shapes, have similarly tight tolerances.
[0021] Assembly tasks of the type described above are often performed manually due to the difficulty of robots detecting and correcting complex misalignments that can occur in tight-tolerance assembly tasks. That is, due to slight deviations in the pose of a part, the robot cannot simply move the part to its nominal placement position; rather, it must "feel" that one part is aligned and fits with the other. There are many sources of error and uncertainty in part pose. First, the exact position and orientation (collectively "pose") of the peg part 110 grasped by the gripper 102 may vary slightly from the expected pose. Similarly, the exact pose of the hole part 120 in its fixture may also vary from the expected pose. In systems where a camera 130 is used to provide an image of the workspace scene for position identification, perceptual errors can also contribute to uncertainty in relative part positioning. Furthermore, calibration errors in the placement of the robot 100 and the fixture holding the part 120 within the workspace, as well as minor variations in robot joint position, can all contribute to part positioning uncertainty. These factors combine to make it impossible for the robot 100, controlled by the controller 140, to simply pick up the peg component 110 and insert it into the hole structure 120 in a single motion.
[0022] FIG. 2 is a diagram of a part being assembled by a robot, which requires alignment in a manner that causes the robot to perform a hole search in a plane perpendicular to the insertion axis. Gripper 202 grasps part 210, which needs to be inserted into hole 220, in the same manner as shown in FIG. 1. Distance 230, exaggerated for visual effect, represents the uncertainty in the lateral position of part 210 relative to hole 220. Proper alignment of part 210 with hole 220 may require the robot to search for the hole, in which case gripper 202 moves part 210 back and forth in a zigzag pattern 240 in a plane perpendicular to the axis of part 210. In a system with camera image input, distance 230 can be minimized by vision control, but the robot still cannot simply insert part 210 into hole 220 in a single motion.
[0023] 1 and 2 show examples of assembly operations where tight tolerances require high accuracy in workpiece placement. One known technique that has been developed for use in these types of robotic assembly operations is the use of force or compliance control devices, described below. Other types of robotic operations, such as placing a workpiece in process into a tool fixture or placing a finished part into a form-fit compartment of a container, similarly require varying degrees of workpiece placement accuracy. For some of these types of operations, compliance control is not appropriate, and accurate visual workpiece placement is preferred.
[0024] 3 is a block diagram of a system 300 configured for a robotic assembly task using compliance control (i.e., force or admittance control), as known in the art. In the physical world, a robot controller 310 communicates with a robot, such as the robot 100 of FIG. 1. The controller 310 sends joint movement commands to the robot 100 and receives state feedback from the robot 100, as is well known in the art and described below. As shown in FIG. 1, the robot 100 has a gripper 102 that grips a first part 110, and the controller 310 provides commands for the purpose of assembling the first part 110 into (within) a second part 120.
[0025] Block 320 represents the controller 310 and the robot 100 in block diagram form. The controller 310 is configured as a compliant controller, the function of which is described below. Block 330 provides a nominal target position for the first part 110. The nominal target position can be predefined and immutable for a particular robot work cell. Alternatively, the nominal target position can be provided by a vision system based on the observed position of a second part 120, for example, movable on a conveyor. For purposes of this discussion, it is assumed that the position of the second part 120 within the robot work cell is known, and the nominal target position from block 320 defines the position of the first part 110 for placing the first part 110 within the second part 120. The nominal target position of the first part 110 can then be transformed into gripper coordinates and then into robot joint positions using inverse kinematics in a known manner.
[0026] A summing junction 340 is included after block 330. Junction 340 does not have a second input in FIG. 3, but a second input can be added in some embodiments. Block 350 defines the motion limits of robot 100. The motion limits ensure that robot 100 does not take excessively large motion steps during the control process, which could result in a dangerous situation or force contact between robot 100 and / or parts 110 / 120 and each other or other objects in the workcell. If the difference between the target position and the current position is greater than the motion limit, the motion limit takes precedence and limits the size of the step.
[0027] Block 360 includes an admittance control function that interacts with the robot in block 370, which performs the assembly task in block 380. Blocks 370 and 380 represent the physical actions of the robot 100 in placing the first part 110 within the second part 120. The robot in block 370 provides state feedback on line 372 to the admittance control function in block 360. The state feedback provided on line 372 includes robot joint states (position / velocity) along with contact forces and torques. Alternatively, the position and velocity state data can be provided in Cartesian coordinates, which can be easily converted to joint coordinates or vice versa using the conversion calculations described above. Force and torque sensors (not shown) are required to measure contact forces between parts 110 and 120 during operation of robot 100; force and torque sensors can be placed between robot 100 and gripper 102, between gripper 102 and first part 110, or between second part 120 and its "ground" (fixture device). Contact forces and torques can also be estimated from robot joint torque sensors or from other signals such as motor currents.
[0028] The admittance control function in block 360 is known in the art and, as briefly described here, operates as follows: Impedance control (or admittance control) is an approach to dynamic control of force and position. It is often used in applications where a manipulator interacts with its environment and the relationship between force and position is of interest. Mechanical impedance is the ratio of force output to motion input. Controlling the impedance of a mechanism means controlling the resistance to external motion imposed by the environment. Mechanical admittance is the inverse of impedance and defines the motion resulting from a force input. The theory behind the impedance / admittance control method is to treat the environment as an admittance and the manipulator as an impedance.
[0029] JPEG2025126909000002.jpg46164
[0030] JPEG2025126909000003.jpg46164
[0031] Elements 330-360 are programmed as algorithms within the controller 310. The interaction between the controller 310 and the robot 100 occurs in the real world via cables (or wirelessly), and this interaction is represented in block 320 by a forward arrow (motion command) from block 360 to block 370 and a feedback line 372 (joint state and contact force).
[0032] As mentioned above, conventional compliance control techniques are not effective for all types of assembly tasks. For example, in tight-tolerance assembly tasks, such as those shown in FIG. 2, where the initial position is significantly offset from the target position, the robot in block 370 may spend a long time in a feedback loop with the admittance control function in block 360, ultimately failing to complete the assembly task, potentially resulting in part damage in the process. This situation can be somewhat alleviated by fine-tuning the impedance / admittance control parameters, but this is only effective in some situations and only for specific workpiece assembly tasks.
[0033] Another robot control technique developed for assembly and other high-precision motion applications is visual servoing. Rather than controlling a robot to move a workpiece to a nominal target position at a fixed, predefined location, visual servoing dynamically adjusts the required workpiece movement based on images of the workpiece and working environment. As the robot moves the workpiece, new images are provided and the adjusted target position is adjusted in a feedback control loop.
[0034] Similar to compliance control, visual servoing has proven effective in several robotic applications. However, visual servoing has drawbacks, including image occlusion, depth inaccuracy when using 2D cameras, and processing speed when using 3D cameras. Therefore, visual servoing techniques often fail to meet the challenges of robotic assembly and other precision motion applications.
[0035] The techniques of the present disclosure have been developed to address the shortcomings of existing visual servoing and compliance control techniques and are described in detail below.
[0036] 4 is a diagram of a system for learning visual pose estimation for robotic tasks, including data collection for training a pose estimation neural network, according to an embodiment of the present disclosure. Robot 400 is similar to robot 100 described above. Robot 400 includes an articulated robot arm with a gripper 402 at the end of the arm. Gripper 402 grasps a workpiece 410, which is placed within or assembled to workpiece 420. Gripper 402 and workpiece 410 are shown in multiple positions in FIG. 4 for reasons explained below.
[0037] The robot 400 communicates with a controller 440, which controls the movement of the robot 400 and receives data from the robot 400. The camera 450 is mounted on the outer robot arm so that its orientation relative to the gripper 402 is fixed. The field of view of the camera 450 preferably includes the workpiece 410 and the mating workpiece 420 in most of the images as the robot 400 moves. The camera 450 is preferably a two-dimensional (2D) camera with color (i.e., "RGB") imaging capability. A 2D camera is preferred because 2D camera images are much faster to process than three-dimensional (3D) camera data, and the techniques of the present disclosure eliminate the need for 3D camera data. The camera 450 provides its images to the controller 440 or to a separate computer (not shown) that communicates with the controller 440.
[0038] A number of lights (460, 462, 464) are positioned within the robot's workspace. Lights 460-464 function as controllable illumination sources in the workspace where robot tasks are being performed. Each of lights 460-464 may be any type of lighting fixture deemed suitable for illuminating the workspace, including ambient lighting fixtures that diffuse light over a wide area (e.g., overhead fluorescent lights), floodlights, etc. Lights 460-464 may simply be available lighting fixtures already installed in the building (e.g., factory or warehouse) in which the robot workspace is located. The purpose and operation of lights 460-464 are discussed further below.
[0039] According to the techniques of this disclosure, the system of FIG. 4 is used to collect data used to train a pose estimation neural network, which is then used for visual servoing robot control. The data used for training is in the form of multiple images from camera 450, with the robot pose recorded for each image. The recorded robot pose is the six-degree-of-freedom relative pose of the tool center point with respect to a target pose (when workpiece 410 is assembled with workpiece 420). Images are taken by camera 450, and the robot 400 is configured with the gripper 402 and workpiece 410 in various positions near the target pose, and lights 460-464 providing various lighting conditions. Each image, along with its corresponding relative pose, is recorded by controller 440 or a separate computer, and this data is used to train the neural network. Control of the robot position, lighting conditions, and image acquisition can be manual or automatic.
[0040] By training with multiple images depicting scenes with different relative poses and different lighting conditions, the pose estimation neural network becomes robust to variations in the part's grasp orientation and starting position, robust to camera calibration parameters, and also robust to variations in lighting conditions and shadows. Details of how the neural network is trained and the subsequent use of the trained neural network for visual servoing robot control are described below.
[0041] Figure 5 is a block diagram of a system 500 configured to train visual pose estimation in a neural network using image and relative pose data from the system of Figure 4, according to an embodiment of the present disclosure. The neural network training system of Figure 5 can run on the controller 440, a separate computer as described above with respect to Figure 4, or yet another computer that accesses input data. Input from the system of Figure 4 is provided in blocks 510 and 520.
[0042] Block 510 includes multiple images of the workpiece placement scene from camera 450. As described above, the images in block 510 are taken from various positions near the target pose and under various lighting conditions. The positions captured in the images can form a grid pattern surrounding the target pose or other geometric or prescribed pattern, in which case the movement of robot 400 and the triggering of camera 450 to capture each image can be defined in a control program. Alternatively, the movement of robot 400 and the acquisition of images can be manually controlled by an operator (e.g., using a teaching pendant). In either case, the positional change covers a range of offset distances in at least two lateral directions (and optionally vertically) and offset angles around all three axes. For example, the positional change range can be specified as + / - 60 mm in two lateral directions (e.g., X and Y), no change in the vertical (Z) direction, + / - 5 degrees for the pitch and roll axes, and + / - 10 degrees for the yaw axis (the gripper axis). These ranges can be selected as appropriate for a particular application. Lighting conditions can be changed while the robot maintains a given pose, and / or lighting conditions can be changed from one pose to the next. Changing lighting conditions can include any combination of turning lights 460-464 on, off, and / or dimming them individually. Images are taken at each unique combination of robot pose and lighting conditions. All images are provided in block 510.
[0043] Block 520 contains the relative pose recorded by the robot controller 440 corresponding to each image in block 510. The relative pose defines the location of the robot tool center point relative to the target workpiece position and is obtained for each image (each unique combination of robot pose and lighting conditions) as described above. The relative pose provided for each image is a vector defining all six degrees of freedom, e.g., x / y / z offset distances and yaw / pitch / roll offset angles. Each image in block 510 has a uniquely identified relative pose vector in block 520.
[0044] Block 530 is a visual pose estimation neural network. The structure of neural network 530 is described below with respect to FIG. 6. Block 540 is the estimated pose output from neural network 530 for each image. For each image, the relative pose from block 520 is compared to the estimated pose from block 540, and the difference is calculated as a numerical value in union 550. The difference value may be a weighted sum of the differences (relative pose minus estimated pose) in six dimensions (x, y, z, yaw, pitch, roll).
[0045] For each image, the difference value from union 550 is used in a cost function to train neural network 530 in a supervised learning process, indicated by dashed line 560 returning through neural network 530. This supervised learning is an automatic process, where a large difference (between the relative pose and the guessed pose) on line 560 tells neural network 530 that the pose estimation on that image was not very good, and conversely, a small difference on line 560 tells neural network 530 that the pose estimation on that image was good. With a sufficient number of training images, neural network 530 learns which parameter settings result in the most accurate estimate of the workpiece's relative pose with respect to the placement target pose.
[0046] Figure 6 is a block diagram that schematically illustrates the structure of the visual pose estimation neural network 530 of Figure 5, according to an embodiment of the present disclosure. In the preferred embodiment shown in Figure 6, the visual pose estimation neural network 530 includes a section 532 that includes layers from a pre-trained neural network and a section 534 that includes linear projection layers that are customized to match task data from a robotic task.
[0047] Section 532 contains several layers of neural networks that have been developed and pre-trained to be effective at extracting features from images. In these neural networks, both the structure (number of layers and nodes, node connectivity, etc.) and preliminary values of parameters (weights) are pre-trained for feature extraction effectiveness. Such pre-trained neural network packages are available from a variety of commercial public domain sources. By using some of the pre-trained neural networks in section 532, the training time for the visual pose estimation neural network 530 is dramatically reduced compared to starting with a "blank slate" neural network architecture.
[0048] The linear projection layer in section 534 performs linear matrix multiplication to project the high-dimensional discrete vectors (from section 532) into lower-dimensional continuous vectors. The exact matrix multiplication parameters are learned through a training process.
[0049] In Figure 6, a single image (510A) from block 510 is provided to a visual pose estimation neural network 530. Using the current parameters in feature extraction layer section 532 and linear projection layer 534, visual pose estimation neural network 530 outputs an estimated pose 540A corresponding to image 510A. This is as shown previously in Figure 5, and then supervised learning training is performed on neural network 530 using a cost function based on the difference between the relative pose (corresponding to image 510A) and the estimated pose 540A. This incremental training process is performed for each image in block 510. In a preferred embodiment, only the parameters of neural network 530 (in both feature extraction layer section 532 and linear projection layer 534) are modified in training, not the structure.
[0050] The structure of the visual pose estimation neural network 530 and the corresponding training method shown in Figures 4-6 and described above have proven advantageous for several reasons. First, the training process is easy and automatic using data collected as shown in Figure 4. No manual calibration of the neural network 530 is required, either in structure or parameters, and no physical features need to be identified or selected in the image. Furthermore, calibration of the camera 450 is not required, as any camera calibration or alignment inaccuracies are automatically corrected in the neural network training process.
[0051] Furthermore, neural network training and execution are fast. Actual trials have shown that training of neural network 530 converges to a highly accurate pose estimate in just a few minutes on readily available computing equipment. This training was performed using a large number of training images (in block 510), with hundreds of training images. Once trained, execution (inference mode) of neural network 530 is also very fast, on the order of a few milliseconds. This performance in inference mode is fast enough for use in manufacturing robotics operations, where neural network 530 receives camera images and infers a relative pose used in visual servo robot control for workpiece placement.
[0052] The visual pose estimation neural network 530 also provides accurate results (demonstrated at the sub-millimeter level) and is robust to changing lighting conditions, which is necessary for high-precision localization tasks in real-world environments where shadows and poor / varying lighting conditions are present.
[0053] The simple and straightforward data collection and fast and automatic training, along with the accuracy and robustness of the resulting pose estimation from the trained network, make the visual pose estimation neural network and the related training method described above very effective in visual servo control applications.
[0054] FIG. 7 is a block diagram of a system 700 configured for robotically positioning a workpiece using a visual pose estimation neural network trained in a visual servoing controller, according to an embodiment of the present disclosure. Block 710 is configured as a robotic visual servoing controller. Robot 400 (of FIG. 4) is shown in block 710. The robot and camera in the visual servoing control system of FIG. 7 need not be the same as the individual devices used for data collection in FIG. 4, as long as the robot type / model, camera type, and workpiece positioning application are the same between the data collection system (FIG. 4) and the production system (FIG. 7). Visual servoing controller block 710 can execute on controller 440 or a similar robot controller.
[0055] A camera 450 (mounted on the robot arm as described above) provides images of the workpiece positioning scene to block 710. Specifically, the camera 450 provides the images to a pre-trained visual pose estimation neural network 530. Upon receiving the images, the trained neural network 530, operating in inferencing mode, outputs an estimated relative target pose, as described above. Thus, on line 720, an estimated target pose of the workpiece relative to the robot's current position is provided for processing. The target pose on line 720 defines where the robot 400 must move to properly position the workpiece.
[0056] A target pose (relative to the current position) can be provided to an optional motion constraint module 730, which ensures that the robot 400 does not take excessively large motion steps during the control process. The target position, with motion constraints where appropriate, is provided to a position planner module 740. The position planner module 740 calculates the robot motions required to move the workpiece from the current position to the target position. For example, module 740 can calculate a spline function, which is configured to move the workpiece from its current position and orientation (six degrees of freedom), as provided by the neural network 530, to the target position and orientation. From the trajectory of the spline curve, the position planner module 740 can calculate the corresponding robot joint motions using inverse kinematics calculations in a known manner.
[0057] Position planner module 740 provides robot motion commands to robot 400, as indicated by the forward arrow. The robot returns state data (e.g., joint positions and velocities) to position planner module 740 on feedback line 742. The feedback control loop from position planner module 740 to robot 400 and back on line 742 can continue for several steps in real time as the robot moves along a calculated path (e.g., the spline curve trajectory described above). A feedback loop also exists on line 750, returning to pose estimation neural network 530. This is the essence of a visual servoing system: move the robot a certain distance toward a target, then acquire another image and update the control commands based on the new image. In the case of the system of FIG. 7, a new image is processed by pose estimation neural network 530, which estimates a new relative target pose, which is provided to planner module 740, which calculates new robot motion commands. This process continues until the robot reaches the goal position and the neural network 530 indicates that the goal position is equal to the current position.
[0058] If the workpiece positioning or assembly task is not particularly difficult, the visual servo control system of Figure 7 can complete the task, release the workpiece, and return to the starting position to grasp a new workpiece. However, in very tight tolerance applications, the visual servo control approach of Figure 7 may need to be integrated with or followed by a force control operation to complete the workpiece positioning or assembly. The combination of visual servo control and compliance control is shown in the following two figures and described below.
[0059] FIG. 8 is a block diagram of a system 800 configured for robotically positioning a workpiece using the visual servoing controller of FIG. 7 for pre-positioning and a compliance controller for final positioning, according to an embodiment of the present disclosure. The system of FIG. 8 can operate on controller 440 or a similar controller. The visual servoing controller of block 710 is shown at the top of FIG. 8 and operates exactly as described above with respect to FIG. 7. That is, block 710 receives images from camera 450, and visual pose estimation neural network 530 estimates a relative target pose that is used by the position planner to control the robot. Visual servoing is used to pre-position the workpiece in the first step (indicated by the circled 1) of the system of FIG. 8.
[0060] Following pre-positioning of the workpiece, final positioning (e.g., placement or assembly) of the workpiece occurs in the second overall step (indicated by circled 2). Final positioning of the workpiece is performed using a compliance controller, indicated by block 820. The compliance controller of block 820 receives the target position from neural network 530 on line 810. This transfer of the target position from visual servo control block 710 to compliance control block 820 is a one-time transfer and indicates what relative position movement is required to complete the placement of the workpiece. Thereafter, visual servo control block 710 is no longer active.
[0061] In compliance controller block 820, a relative target position is received in block 830. From there, compliance controller block 820 operates as previously described for block 320 of FIG. 3. That is, the target position is motion-constrained if necessary (block 840) and provided to admittance control block 850, which communicates with robot 400 to provide motion commands to robot 400 and receive state feedback (including forces and torques associated with robot motion) on line 852. The assembly operation is shown in block 860, with the interaction of robot 400 with the physical environment (workpiece contacting mating parts with resulting contact forces and torques) indicated by the double-headed arrow. Compliance controller block 820 operates in a feedback loop (admittance control block 850 and robot 400) until the workpiece is located at the final target position. In the system of FIG. 8, there is no feedback loop to block 830, which determines a new target position.
[0062] The two-stage control shown in Figure 8 is one embodiment of a system that combines visual servoing control (using a visual pose estimation neural network) and compliance control for precision workpiece placement applications. Another embodiment, in which visual servoing and compliance control are integrated in a nested loop, is shown in Figure 9.
[0063] 9 is a block diagram of a system 900 configured for robotic workpiece positioning using visual servoing with integrated compliance control and visual pose estimation neural network, according to an embodiment of the present disclosure. The system of FIG. 9 can operate on controller 440 or a similar controller. The integrated visual servoing and compliance controller is shown in block 910. Camera 450 provides images of the workpiece positioning scene to a trained visual pose estimation neural network 530, which operates as described above to estimate a relative target position.
[0064] The target position from the neural network 530 is motion constrained if necessary (block 920) and provided to an admittance control block 930 which communicates with the robot 400 to provide motion commands to the robot 400 and receive state feedback (including forces and torques associated with robot motion) on line 932. The assembly operation is shown in block 940, and the interaction of the robot 400 with the physical environment (the workpiece contacting the mating parts with resulting contact forces and torques) is indicated by the double-headed arrows as previously described. The admittance control block 930 and robot 400 operate in an inner feedback loop for several cycles, and then the robot position state is provided to the pose estimation neural network 530 on outer feedback loop line 950. Using new images from the camera 450 and the actual relative pose on line 950, the neural network 530 estimates a new relative target pose which is fed to the admittance control block 930.
[0065] The system of FIG. 9 provides integrated control that takes advantage of visual servoing control using a pose estimation neural network 530 (which is particularly effective for large movements as the workpiece is moved from an initial position toward a target position) and compliance control (which is particularly effective for fine positioning movements required when placing a workpiece with tight tolerances, such as when assembled with mating parts).
[0066] Whether used in a robot controller employing purely visual servo control or incorporated into a robot controller that also employs compliance control, the visual pose estimation neural network 530 provides the engine for workpiece placement. The visual pose estimation neural network 530 is easily trained using the techniques discussed above, is cost-effective using 2D cameras, is accurate and fast, and is robust to changing or suboptimal lighting conditions due to the variation contained in the training image set.
[0067] FIG. 10 is a flowchart diagram 1000 of a method for learning visual pose estimation for robotic tasks, including offline training of a neural network using images at different positions and lighting conditions, and online visual servoing using the trained neural network, according to an embodiment of the present disclosure.
[0068] In box 1002, a robot / camera system is provided for data collection. This is the system shown in Figure 4, comprising a robot 400 and a controller 440, set up to place a first workpiece 410 within a second workpiece 420. This is all located within a work cell with lights 460 / 462 / 464 that can be controlled to vary the lighting conditions.
[0069] In box 1004, training images are collected using the robot / camera system. As previously described, the robot 400 moves the gripper 402 along with the workpiece 410 to various positions near the target location, varying lighting conditions while acquiring many images. For each image captured, the controller 440 records the pose of the workpiece relative to the target pose. The images and poses are stored in a file or database 1006. The database 1006 may be located in the controller 440 or, optionally, in a separate computer (not shown in FIG. 4 ). The separate computer can, for example, store the training data (database 1006) and can also be used for the neural network training process.
[0070] In box 1008, a pre-configured pose estimation neural network is provided. This is neural network 530 shown in Figures 5 and 6. Neural network 530 is preferably pre-configured with the design shown in Figure 6 and described above. In box 1010, the pose estimation neural network is trained in a supervised learning process using image and relative pose data from database 1006. As discussed above with respect to Figure 5, training involves the neural network inferring the relative pose of the images, and the difference between the recorded and estimated relative poses is used in a cost function to train the parameters of the neural network, penalizing large differences and rewarding small differences.
[0071] This process described above is repeated until either a maximum count is reached or the neural network converges to a consistently accurate pose estimate and meets a threshold difference between the recorded and estimated relative poses. At decision diamond 1012, if the neural network performance has not yet converged satisfactorily, the process loops back to box 1010 for the next image. If the neural network performance has converged satisfactorily, the process proceeds from decision diamond 1012 to box 1014.
[0072] In box 1014, the trained pose estimation neural network is used in an inference mode for visual servoing control of the robot. This can include any of the control system architectures shown in FIGS. 7-9, where the trained visual pose estimation neural network 530 is used in a pure visual servoing control system (FIG. 7), or where the neural network 530 is used for visual servoing control for preliminary positioning, followed by compliance control for final positioning (FIG. 8). Alternatively, the neural network 530 is used in an integrated visual servoing / compliance control system (FIG. 9). In all of these system architectures, the trained neural network 530 provides an inferred or calculated value of the relative target pose from the most recent camera image. The method steps of FIG. 10 may be performed in the robot controller 440, as outlined above, and optionally in other robot controllers and separate computers. For example, neural network training can be performed on a computer other than the robot controller, and the trained neural network can then be provided to and used by a control module running on the robot controller.
[0073] The methods and systems disclosed herein enable simple, automated training of visual pose estimation neural networks that are then fast and accurate enough to be used in real-time visual servoing robot control, in combination with compliance control where appropriate. The trained neural networks are robust to changes in lighting conditions as a result of the training data image set, and the systems employed in both training and estimation modes use cost-effective 2D cameras. Thus, the disclosed methods and systems offer significant improvements over conventional image-based robot control systems.
[0074] The preceding discussion describes various computers and controllers. It should be understood that the software applications and modules of these computers and controllers are executed on one or more computing devices having processors and memory modules configured to train visual pose estimation in a robotic task. In particular, this includes the processor within the robot controller 440, as well as any separate computers used for data collection in FIG. 4, the computers used for the training process in FIG. 5, and the robot controller and any other computers performing the functions of visual servoing and compliance controller (including operation of the visual pose estimation neural network in inferential mode) shown in FIGS. 7-9.
[0075] The foregoing discussion discloses and describes merely exemplary embodiments of the present disclosure. Those skilled in the art will readily appreciate from such description, the accompanying drawings, and the claims that various changes, modifications, and variations can be made without departing from the spirit and scope of the present disclosure, as defined in the following claims.
Claims
1. 1. A system for learning visual pose estimation using a robot, comprising: a robot having a gripper configured to perform an operation on a workpiece; a camera mounted on the outer arm of the robot proximate the gripper to provide an image of a workpiece operation scene; One or more lights for illuminating a workpiece carried by the robot; at least one computing device in communication with the robot and the camera, the computing device configured with a neural network, the neural network being trained for visual pose estimation using a plurality of images including a plurality of workpiece positions and a plurality of lighting conditions and an actual relative pose for each image, the neural network being run after training in an inferencing mode for visual pose estimation used in visual servoing of the robot performing the task; A system comprising:
2. 2. The system of claim 1, wherein the workpiece operation scene in each image includes the workpiece in the gripper of the robot and at least a portion of a placement target area, and the actual relative pose in each image defines a relative position of the workpiece with respect to a target position determined from joint positions of the robot.
3. 2. The system of claim 1, wherein the at least one computing device is configured with a supervised learning algorithm that trains the neural network for visual pose estimation by computing, for each of the plurality of images, a difference between the image's inferred pose from the neural network and the image's actual relative pose, and applying a cost function that penalizes large differences and rewards small differences.
4. The system of claim 1 , wherein the multiple lighting conditions are achieved by turning on, off, and / or dimming one or more lights individually or collectively.
5. 2. The system of claim 1, wherein the neural network has a structure including a plurality of layers preconfigured for image feature extraction and a linear projection layer that receives outputs from the plurality of layers and outputs an inferred pose.
6. The system of claim 5 , wherein parameter values of the neural network are modified during training to improve the accuracy of the estimated pose, but the structure is not modified.
7. 2. The system of claim 1, wherein in the visual servoing control of the robot, the neural network receives camera images and calculates an estimated relative pose, and a position planning module calculates robot joint movements required to move the workpiece to a target position based on the estimated relative pose.
8. The system of claim 1 , wherein the visual servoing control is used in conjunction with compliance control of the robot performing the task.
9. 9. The system of claim 8, wherein the visual servo control is used to perform preliminary positioning of the workpiece and the compliance control is then used to perform final positioning of the workpiece, or the visual servo control operates in an outer feedback loop and the compliance control operates in an inner feedback loop when positioning the workpiece.
10. The system of claim 1 , wherein the camera is a two-dimensional (2D) camera.
11. The system of claim 1 , wherein the operation is to move the workpiece to a destination location or to place the workpiece on or inside a second workpiece.
12. 2. The system of claim 1, wherein the at least one computing device is a robot controller that controls the movement of the robot, receives joint state data from the robot, and performs training of the neural network.
13. 2. The system of claim 1, wherein the at least one computing device is a computer in communication with a robot controller, the computer receiving images from the camera and joint position data from the robot controller to train the neural network, and the robot controller performing the visual servoing using the trained neural network.
14. 1. A system for learning visual pose estimation using a robot, comprising: a robot having a gripper configured to perform an operation on a workpiece; a camera mounted on the outer arm of the robot proximate the gripper to provide an image of a workpiece operation scene; One or more lights for illuminating a workpiece carried by the robot; at least one computing device in communication with the robot and the camera; the at least one computing device is configured with a supervised learning algorithm that trains a neural network for visual pose estimation using a plurality of images including a plurality of workpiece positions and a plurality of lighting conditions and an actual relative pose for each image, the supervised learning algorithm calculating, for each of the plurality of images, a difference between the inferred pose of the image from the neural network and the actual relative pose of the image and applying a cost function that penalizes large differences and rewards small differences; The at least one computing device, after training, executes the neural network in an inferencing mode for visual pose estimation used in visual servo control of the robot performing the task, the neural network receives camera images and calculates an inferred relative pose, and a position planning module calculates the robot joint movements required to move the workpiece to a target position based on the inferred relative pose.
15. 1. A method for learning visual pose estimation in a robotic task, comprising: providing a robot having a gripper configured to perform an operation on a workpiece, one or more lights to illuminate a workpiece carried by said robot, and a camera mounted on an outer arm of said robot proximate to said gripper to provide an image of a workpiece operation scene; collecting, by a computing device in communication with the robot and the camera, training data including a plurality of images including a plurality of workpiece positions and a plurality of lighting conditions, and an actual relative pose for each image; training a neural network for visual pose estimation using the training data; running the neural network in an inferencing mode for visual pose estimation used in visual servoing of the robot performing the task; A method comprising:
16. 16. The method of claim 15, wherein the workpiece operation scene in each image includes the workpiece in the gripper of the robot and at least a portion of a placement target area, and the actual relative pose in each image defines a relative position of the workpiece to a target position determined from joint positions of the robot.
17. 16. The method of claim 15, wherein the computing device is configured with a supervised learning algorithm that trains the neural network for visual pose estimation by computing, for each of the plurality of images, the difference between the inferred pose of the image from the neural network and the actual relative pose of the image, and applying a cost function that penalizes large differences and rewards small differences.
18. The method of claim 15 , wherein the multiple lighting conditions are achieved by turning on, off, and / or dimming one or more lights individually or collectively.
19. 16. The method of claim 15, wherein the neural network has a structure including multiple layers preconfigured for image feature extraction and a linear projection layer that receives outputs from the multiple layers and outputs an estimated pose, and parameter values of the neural network are modified during training to improve the accuracy of the estimated pose.
20. 16. The method of claim 15, wherein in the visual servoing of the robot, the neural network receives camera images and calculates an estimated relative pose, and a position planning module calculates the robot joint movements required to move the workpiece to a target position based on the estimated relative pose.
21. The method of claim 15 , wherein the visual servoing control is used in conjunction with compliance control of the robot performing the task.
22. 22. The method of claim 21, wherein the visual servo control is used to perform preliminary positioning of the workpiece and the compliance control is then used to perform final positioning of the workpiece, or the visual servo control operates in an outer feedback loop and the compliance control operates in an inner feedback loop when positioning the workpiece.
23. The method of claim 15 , wherein the operation is moving the workpiece to a destination location or placing the workpiece on or within a second workpiece.