Device and method for controlling a robot to perform a task

The method uses stereoscopic imaging and force-torque feedback with contrasting learning to enhance robot control for plug-in tasks, addressing inefficiencies in existing methods by reducing data needs and improving speed and reliability.

DE102022202144B4Active Publication Date: 2025-07-10ROBERT BOSCH GMBH
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
DE102022202144
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-02
Publication Date
2025-07-10
Estimated Expiration
2042-03-02

AI Technical Summary

Technical Problem

Existing robot control methods are inefficient for plug-in tasks, particularly those involving complex shapes and variations, with visual techniques being slow and requiring additional sensors, while current training methods are data-intensive and prone to overfitting.

Method used

A method using two cameras positioned at a 45-degree angle on a robot's gripper to capture stereoscopic images, combined with force-torque feedback, enables efficient robot control through a machine learning model that utilizes contrasting learning and one-shot training to derive delta motions without additional sensors.

Benefits of technology

The method achieves high data efficiency and accurate robot control for plug-in tasks, reducing training data requirements and improving speed and reliability compared to traditional visual techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for controlling a robot (100) to perform a task, comprising: Training a machine learning model (503, 504) to derive delta movements from pairs of image data items by collecting a set of primary training items, each primary training item comprising an image data item comprising at least an image from the perspective of an end effector (104) at a respective position and a motion vector for moving the end effector (104) from the position to a desired position, generating secondary training items by selecting, for each secondary training item, two of the primary training items, wherein the image data items of the primary training items are included in the training item and a difference between the motion vectors of the primary training items is included as a ground truth motion vector, and training the machine learning model (503, 504) using the secondary training items in a supervised manner; Acquiring a target image data item comprising at least one target image from a perspective of the end effector (104) of the robot (100) at a target position of the robot (100) in which the robot (100) has performed the task; Acquiring an original image data item comprising at least one original image from the perspective of the end effector (104) of the robot (100) at an original position of the robot (100), Feeding the source image data item and the target image data item to the trained machine learning model (503, 504) to derive a delta movement between the current source position and the target position; and Controlling the robot (100) to move according to the delta movement to perform the task.
Need to check novelty before this filing date? Find Prior Art

Description

The present disclosure relates to apparatuses and methods for controlling a robot to perform a task.Assembly, such as electrical wiring assembly, is one of the most common manual operations in the industry. Examples are the mounting of switchboards and the mounting of switching installations in housings. Complicated assembly processes can generally be described as a sequence of two main activities: gripping and plugging in. Similar tasks are encountered, for example, in cable production, which typically involves the insertion of cables for validation and verification.While robot control schemes suitable for industrial gripping tasks are typically available, performance of plug-in or "pin-in-hole" tasks by robots is typically applicable only to small subsets of problems, primarily those involving simple fixed site shapes and where variations are not accounted for. In addition, the visual techniques present are slow, typically about three times slower than human operators.Therefore, efficient methods for training a controller for a robot to perform tasks, such as a plug-in task, are desirable.The publication DE 10 2021 109 332 A1 describes a method for controlling a robot for inserting an object into an insertion opening. The method includes controlling the robot to hold the object, generating an estimate of a target position for inserting the object into the insertion opening, controlling the robot to move to the estimated target position, capturing a camera image using a camera mounted on the robot after the robot is controlled to move to the estimated target position, feeding the camera image to a neural network trained to derive motion vectors indicating motions from the positions at which the camera images are captured to insert objects into insertion openings, from camera images, and controlling the robot to move according to the motion vector derived from the camera image by the neural network.The publication DE 102020 119 704 A1 describes a control device for a robot device in which a first feature portion of a first workpiece and a second feature portion of a second workpiece are defined in advance. A feature amount detection unit detects, in an image captured by a camera, a first feature amount relating to the position of the first feature portion and a second feature amount relating to the position of the second feature portion. A calculation unit calculates a difference between the first feature amount and the second feature amount as the relative position amount. A command generation unit generates a movement command for operating the robot on the basis of the relative position amount in the image captured by the camera and a relative position amount in a reference image set in advance.The publication DE 10 2019 106 458 A1 describes a method for determining a position of a workpiece, comprising: capturing image data of a workpiece via a camera, searching for a reference structure of the workpiece using the captured image data, determining a current position of at least one point of the structure, comparing the current position with a nominal position thereof, generating commands for placing the tool in a region of the workpiece to be machined and determining the current position by determining a current image size of the structure and by determining a distance of the structure from the camera.DE 10 2019 002 065 A1 describes a machine learning device having a state observation unit for observing, as state variables, an image of a workpiece picked up by a vision sensor and a movement amount of an arm end portion from an arbitrary position, the movement amount being calculated to bring the image closer to a target image, a determination data retrieval unit for retrieving the target image as determination data, and a learning unit for learning the movement amount to move the arm end portion or the workpiece from the arbitrary position to a target position. The target image is an image of the workpiece captured by the vision sensor when the arm end portion or the workpiece is disposed at the target position.Publication DE 11 2017 007 025 T5 describes a position control device including: an imaging unit that captures an image in which two objects are; a control parameter generation unit that inputs information regarding the captured image of two objects into an input layer of a neural network and outputs a position control amount for controlling the positional relationship of the two objects as an output layer of the neural network; a control unit that uses the output position control amount to control a current for controlling the positional relationship of the two objects; and a drive unit that uses the current for controlling the positional relationship of the two objects to move a position out of the positional relationship of the two objects.The publication EP 3 515 671 B1 describes the training and use of both a geometry network and a network for predicting a gripping result. The trained geometry network may be used to generate geometry outputs based on two-dimensional or two and a half-dimensional images that are geometry aware and that represent (e.g., high-dimensional) three-dimensional features captured from the images. In some implementations, the geometry output(s) include at least one encoding generated based on a trained encoding neural network trained to generate encodings representing three-dimensional features (e.g., shape). The trained network for predicting the grasp result may be used to generate a prediction of the grasp result for a possible grasp posture based on the geometry output(s) and additional data as input(s) to the network.The invention is based on the object of automatically moving a robot into the target position on the basis of image data of a target position without data from further sensors being required.The object is achieved by the subject matters of the independent claims.According to various embodiments, a method for controlling a robot for performing a task according to independent claim 1 is provided.The provided method enables efficient robot control by acquiring images for desired targets and, for a current position, supplying corresponding target image data and image data acquired for the current position for the machine learning model. Thus, the robot can be controlled only (or at least mainly) on the basis of image data and no further sensors are required. It further enables efficient multi-step control.Various examples are given below.Example 1 is the provided method for controlling a robot to perform a task.By selecting combinations of image data items and their associated motion vectors, a high number of training items of the machine learning model can be easily generated, since all possible pairs can be selected. Thus, high data efficiency of the training can be achieved.Example 2 is the method of example 1, wherein the machine learning model comprises an encoder network, the method comprising supplying the source image data element and the target image element to the encoder network, respectively, and supplying embeddings of the source image data element and the target image element to a neural network of the machine learning model configured to derive the delta motion from the embeddings.The use of encodings can reduce the complexity of the neural network, since it only needs to operate with two embeddings and not with two image data elements. The embeds for the source image data element and the target data element are generated by the same encoder network, which can be trained independently of the neural network in order to avoid overfitting. The neural network for deriving the delta motion may then be trained using a small amount of training data.Example 3 is the method according to one of Examples 1 to 2, comprising capturing the original image data element and the target image data element such that these respectively comprise two images, wherein the images for the same position of the end effector of the robot are recorded by two different cameras attached to the end effector.With two cameras (e.g. arranged on opposite sides of a gripper plane of the end effector), the problem of misrehensibility in only one image can be avoided and depth information can be extracted and at the same time occlusion can be avoided on the entire insertion trajectory (if the view of one camera is occluded, the other camera has free view). For example, each camera is positioned at a 45 degree angle with respect to its respective finger opening, allowing good vision of the scene and object between the fingers.Example 4 is the method of example 1, comprising controlling the robot for a task of inserting an object into an insert, wherein capturing the target image element comprises capturing the at least one target image by grasping, by the robot, a reference object that fits into the insert, bringing the robot into a position such that the reference object is inserted into the insert, and capturing the at least one target image at the position.The target data can therefore be easily generated by capturing images of the view that the end effector has (e.g. object and deployment) when the task is achieved. The target data may be gathered once before performing the task and used later multiple times as a target in different control scenarios.Example 5 is a robot controller configured to perform a method according to any one of Examples 1 to 4.Example 6 is a robot comprising a robot controller according to example 6 and an end effector having at least one camera configured to acquire the at least one origin image.Example 7 is a computer program comprising instructions that, when executed by a processor, cause the processor to perform a method according to any one of Examples 1 to 4.Example 8 is a computer readable medium storing instructions that, when executed by a processor, cause the processor to perform a method according to any one of Examples 1 to 4.It should be noted that the embodiments and examples described in connection with the robot also apply analogously to the method for controlling a robot and vice versa.In the drawings, like reference numerals refer to the same parts generally throughout the several views. The drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the invention. In the following description, various aspects will be described with reference to the following drawings, in which: FIG. 1 shows a robot. FIG. 2 shows a robot end effector in detail. FIG. 3 illustrates training of an encoder network, according to an embodiment. FIG. 4 shows the determination of a delta movement from an image data element and a force input. FIG. 5 shows the determination of a delta movement from two image data elements. FIG. 6 illustrates an example of a multi-step plug-in task. FIG. 7 is a flowchart illustrating a method of controlling a robot to perform a task.The following detailed description refers to the accompanying drawings, which illustratively show specific details and aspects of this disclosure in which the invention may be practiced. Other aspects may be utilized and structural, logical, and electrical changes may be made without departing from the scope of the invention. The various aspects of this disclosure are not necessarily mutually exclusive, as some aspects of this disclosure may be combined with one or more other aspects of this disclosure to form novel aspects.Various examples are described in more detail below.FIG. 1 shows a robot 100.The robot 100 includes a robot arm 101, for example an industrial robot arm for handling or mounting a workpiece (or one or more other objects). The robot arm 101 includes manipulators 102, 103, 104 and a base (or support) 105 by which the manipulators 102, 103, 104 are supported. The term "manipulator" refers to the movable components of the robot arm 101 whose actuation allows a physical interaction with the environment, e.g., to perform a task. For control, the robot 100 includes a (robot) controller 106 configured to implement the interaction with the environment according to a control program. The last component 104 (furthest from the support 105) of the manipulators 102, 103, 104 is also referred to as the end effector 104, and may include one or more tools, such as a welding torch, a gripping instrument, a paint shop, or the like.The other manipulators 102, 103 (located closer to the support 105) may form a positioning device such that, together with the end effector 104, the robot arm 101 is provided with the end effector 104 at its end. The robot arm 101 is a mechanical arm that can provide similar functions to a human arm (possibly with a tool at its end).The robotic arm 101 may include joint members 107, 108, 109 that connect the manipulators 102, 103, 104 to each other and to the support 105. A joint element 107, 108, 109 may comprise one or more joints, which may each provide a rotatable movement (i.e. rotational movement) and / or translatory movement (i.e. displacement) for associated manipulators relative to one another. The movement of the manipulators 102, 103, 104 may be initiated by actuators controlled by the controller 106.The term "actuator" may be understood as a component configured to effect a mechanism or process in response to its drive. The actuator may implement instructions (so-called activation) generated by the controller 106 into mechanical movements. The actuator, e.g., an electromechanical transducer, may be configured to convert electrical energy to mechanical energy in response to its drive.The term "controller" may be understood as any type of logic implementing entity that may include, for example, circuitry and / or a processor capable of executing software, firmware, or a combination thereof stored in a storage medium and capable of issuing instructions, e.g., an actuator in the present example. The controller may be configured, for example, by program code (e.g., software) to control the operation of a system, a robot in the present example.In the present example, the controller 106 includes one or more processors 110 and a memory 111 in which code and data according to which the processor 110 controls the robot arm 101 are stored. According to various embodiments, the controller 106 controls the robot arm 101 based on a machine learning model 112 stored in the memory 111.According to various embodiments, the machine learning model 112 is configured and trained to allow the robot 100 to perform a plug-in task (e.g., pin-in-hole), for example, plugging a plug 113 into a corresponding jack 114. For this purpose, the controller 106 records images of the plug 113 and the socket 114 by means of the cameras 117, 119. The plug 113 is, for example, a USB (Universal Serial Bus) plug or can also be a power plug. It should be noted that if the plug has multiple pins, such as a power plug, each pin may be considered an item to be inserted (the insert being a corresponding hole). Alternatively, the entire plug may be seen as the object to be plugged in (the insert being a socket). It should be noted that (depending on what is considered an object) the object 113 is not necessarily fully inserted into the insert. As in the case of the USB connector, the USB connector is deemed to be inserted when the metal contact portion 116 is inserted into the socket 114.The robot controller for performing a pen-in-hole task typically includes two main phases: searching and plugging. During the search, jack 114 is identified and located to obtain the essential information required for plugging in plug 113.The search for deployment may be based on vision or blind strategies including spiral paths, for example. Visual techniques depend heavily on the location of cameras 117, 119 and panel 118 (in which bushing 114 is placed in panel surface 115) as well as obstacles, and are typically about three times slower than human operators. Due to the limitations of visual methods, the controller 106 may take into account force-torque feedback and haptic feedback, either solely or in combination with visual information.Constructing a robot that reliably inserts various objects (e.g., plugs, motor gears) is a great challenge in the development of manufacturing, inspection, and home service robots. Minimizing action time, maximizing reliability, and minimizing contact between the gripped object and the target component are difficult due to the associated uncertainties in sensing, control, sensitivity to applied forces and occlusions.According to various embodiments, a data efficient, secure and monitored approach for detecting a robot strategy is provided. This enables learning of a control strategy, in particular for a multi-step plug-in task, with few data points, by using contrasting methodologies and one-shot learning techniques.According to various embodiments, the training and / or the robot control comprises one or more of the following:1) Use of two cameras to avoid the problem of misrehensibility in only one image and to extract depth information. This makes it possible, in particular, for the contact of the socket surface 115 no longer to be required when the object 113 is inserted.2) Integration of Contrasting Learning to Reduce the Amount of Labeled Data.3) A relationship network that enables one-shot learning and multi-step plug-in.4) Multi-Step Insertion Using This Relationship Network.FIG. 2 shows a robot end effector 201 in detail.The end effector 201 corresponds to, for example, the end effector 105 of the robot arm 101, e.g., a six-degree-of-freedom (Degres of Freedom, DoF) robot arm. According to various embodiments, the robot has two sensory inputs that the controller 106 can use to control the robot arm 101. The first is for stereoscopic perception provided by two (see item 1 above) wrist cameras 202 tilted at an angle of 45°, for example, and focused on a point between the end effector fingers (EEF) 203.The images 205 and 206 are examples of images captured by the first camera and the second camera, respectively, for a position of the end effector 201.In general, each camera 202 is oriented such that the captured images show, for example, a portion (here, a pin 207) of an object 204 gripped by the end effector 201 that is to be inserted and an area around it, such that, for example, the insert 208 is visible. It should be noted that the insert here refers to the hole for the pin, but it may also refer to the holes for the pins as well as the opening for the cylindrical part of the plug. The insertion of an object into an insert therefore does not necessarily mean that the object is completely inserted into the insert, but only one or more parts.Starting from a height H, a width W, and three channels for each camera image, an image data item for a current (or original) position of the arm provided by the robot is indicated as Img ∈ R H×W×6( with six channels because it includes the images from both cameras 202). The second sensory input is a force input, i.e., measurements of a force sensor 120, that measures a moment and force experienced by the end effector 105, 201 when pressing the object 113 onto a plane (e.g., the surface 115 of the plate 118). The force measurement can be performed by the robot or by an external force and torque sensor. The input force comprises, for example, an indication of force and an instant indication M=(m x, m y, m z) ∈R 3.For the following explanations, the observation of the robot for a current position is given as Obs=(Img; F; M). To accurately detect the contact forces and generate liquid movements, high-frequency communication (real-time data exchange) may be used between the sensor devices (cameras 202 and force sensor 120) and the controller 106.For example, force and torque (moment) measurements are sampled at 500 Hz and commands are sent to the actuators at 125 Hz.The end effector fingers 203 form a gripper, the position of which is indicated as L.In particular, the position of the gripper is and its position. The action of the robot in Cartesian space is defined by (Δ x, Δ y, Δ z, Δθ x, Δθ y, Δθ z) where Δx, Δy, and Δz are the desired corrections required for the EEF in Cartesian space with respect to the current location. This robot action indicates the movement of the robot from a current (or original) position (in particular a current position) to a target position (in particular a target position).By the two-camera scheme, i.e., there is an image data item including two images for each robot position under consideration, the distance between two points shown in the images can be obtained, i.e., the problem of misrehensibility of visual impressions occurring when attempting to obtain the distance between two points in world coordinates without depth information using a single image can be avoided.According to various embodiments, backward learning is used (e.g. by the controller 106), in which the images of the two cameras 202 are used to collect training data (in particular images) not only after touching the surface 115 but also along the movement trajectory. That is, to collect a training data item, the controller 106 places the robot arm in its final (target) position Lfinal, i.e., when the plug 116 is inserted into the jack 114 (or similarly for any intermediate target for which the machine learning model 112 is to be trained). (Note that using cameras 202, a target image data item may be captured in this position, which is used according to various embodiments as described below.). For each training data element, two points are then sampled from the probability distribution (e.g., a normal distribution): one is Thigh positioned at a random location above the socket and the second is Tlow positioned randomly around the height of the socket.A correction for this training data element is defined by where transdom is a high or low point (i.e., Thigh or Tlow). A training data set, indicated as D, is formed based on a set of training tasks, where there is a corresponding one for each task τi. For each task, a randomized set of points is generated that yields and yields the starting and ending points for the tasks and the starting random points (where j=high, low). A corresponding correction is defined for each DT i,j. Algorithm 1 shows a detailed example of data collection and general backward learning for task τ.According to Algorithm 1, force sensor data is acquired. This is not necessary according to various embodiments, in particular those which operate with target image data elements, as described further below with reference to FIG. 5.According to various embodiments, the machine learning model 112 includes a plurality of components, one of which is an encoder network that the controller 106 uses to determine encoding for each image data element Img.FIG. 3 illustrates the training of an encoder network 301, according to an embodiment.The encoder network 301 is, for example, a convolutional neural network, e.g., with a ResNet18 architecture.The encoder network 301 (implementing the function φ) is trained using a contrasting loss and one or both of a delta strategy loss and a relationship data loss, for example according to the following loss:These loss components will be described below.Contrasting learning, i.e. training based on contrasting losses, is a framework for learning representations following similarity or dissimilarity conditions in a dataset mapped to positive and negative labels, respectively. One possible contrasting learning approach is instance discrimination, where an example and image are a positive pair if they are data augmentations of the same instance and are a negative pair otherwise. A central challenge in contrasting learning is the selection of the negative examples, since it can influence the quality of the learned underlying representations.According to various embodiments, the encoder network 301 is trained (e.g. by the controller 106 or by an external device that is later stored in the controller 106) using a contrasting technique, such that it learns relevant features for the respective task without specific labels. An example is InfoMCE loss (NCE: Noise-Contrast Examination). At this time, by stacking two images of the two cameras 202 into one image data item 302, depth registration of the connector 113 and the socket 114 is obtained. This depth information is used to augment the image data item in various ways.Pairs of image data elements are scanned from the original (i.e. non-augmented) image data element and one or more augmentations obtained in this way, one element of the pair being fed to the encoder network 301 and the other to another version 303 of the encoder network. The other version 303 realizes the function φ', which has, for example, the same parameters as the encoder network 301 and is updated using a Polyak averaging according to φ'=φ'+ (1-m)φ with m=0.999 (wherein φ, φ' were used to represent the weighting of the two encoder network versions 301, 302).The two encoder network versions 301, 302 each output a representation (i.e., embedding) of the magnitude L for the input data element 302 (i.e., a 2 x L output of the pair). If this is carried out for a stack of N image data elements (i.e. one pair of augmentations or origin and augmentation is formed for each image data element), i.e. for training input image data of size N×6×H×W, this yields N pairs of representations that are output by the two encoder network versions 301, 302 (i.e. representation output data of size 2×N×L). Using these N pairs of representations, the contrasting loss of the encoder network 301 is calculated by forming positive pairs and negative pairs from the representations included in the pairs. Here, two representations are a positive pair when generated from the same input data item 302 and a negative pair when generated from different input data items 302. This means that a positive pair contains two augmentations of the same original image data element or contains one original image data element and one augmentation thereof. All other pairs are negative pairs.For determining the contrasting loss 304, a similarity function sim( .) is used, which measures the equality (or the distance between two embeddings). It may use the Euclidean distance (in the latent space, i.e. in the space of the embeddings), but more complicated functions may also be used, e.g. using a kernel. The contrasting loss is then given, for example, by the sum over i, j of where ziare the embeddings. τ is here a temperature normalization factor (not to be confused with the task τ used above).FIG. 4 shows the determination of a delta movement 405 from an image data element 401 and a force input 402.The delta motion Δθ=(Δθ, Δθ, 0, Δθ x, Δθ y, Δθ z) is the robot action Δθ=(Vx, Δy, Δθ, Δθ x, Δθ y, Δθ z) without the motion in the z direction, because the controller 106 according to various embodiments controls the motion in the z direction Δθ independently of other information, for example, using the height of the table on which the base 114 is placed from the prior knowledge or from the depth camera.In this case, the encoder network 403 (corresponding to the encoder network 301) generates embedding for the image data item. The embedding is forwarded together with the force input 402 to a neural network 404 which provides the delta motion 405. The neural network 404 is, for example, a convolutional neural network, which is referred to as a delta network and is intended to implement a delta (control) strategy.The delta loss Idelta for training the encoder network 403 (as well as the delta network 404) is determined by having ground truth delta motion labels for training input data items (including an image data item 401 and a force input 402, i.e., what was particularly indicated above in Algorithm 1 with Obs=(Img; F; M)). The training data set D generated by the algorithm 1 includes the training input data elements Obs and the ground truth labels for delta loss.FIG. 5 shows the determination of a delta movement 505 from two image data elements 501, 502.In this case, the encoder network 503 (corresponding to the encoder network 301) generates an embedding for each image data element 501, 502. The embeds are passed to a neural network 504, which provides the delta motion 505. The neural network 504 is, for example, a convolutional neural network called a relationship network, which is intended to implement a relationship (control) strategy.The relationship loss Irelationfor training the encoder network 503 (as well as the relationship network 504) is determined by having ground truth delta motion labels for pairs of image data items 501, 502. The ground truth delta motion label for a pair of image data items 501, 502 may be generated, for example, by determining the difference between the ground truth delta motion labels (i.e., the actions) included for the image data items in the dataset D generated by the algorithm 1.That is, for the training with the relationship loss, the data set D is used to calculate the delta motion between two images of the same plug-in task, Imgiand Imgj, where j≠k by calculating the ground truth by the difference Δτk,j=Δτi,j-Δτi,k. If Imgiand Imgjare augmented for this training, augmentations that are consistent are used.The relationship loss Irelation facilitates one-shot learning, enables multi-step plug-in, and improves utilization of the collected data.When trained, the encoder network 403 and the delta network 404 which, as described with reference to FIG. 4, are used to derive a delta movement from an image data element 401 and a force input 402, implement the so-called delta (control) strategy. Similarly, the trained encoder network 503 and the relationship network 504, which are used as described with reference to FIG. 5 to derive delta motion from an image data element 501 (for a current position) and an image data element 502 (for a target position), implement the so-called relationship (control) strategy.The controller may use the delta strategy or the relationship strategy as a residual strategy π Residual in combination with a main strategy π Main. For the inference, the controller 106 thus uses the encoder network 403, 503 for the delta strategy or the relationship strategy. This can be decided depending on the application. For example, for one-shot or multi-step plug-in tasks, the relationship strategy (and relationship architecture of FIG. 5 ) is used because it can be more generalized to these tasks. For other tasks, the delta strategy (and delta architecture of FIG. 4 ) is used.Following the main strategy, the controller 106 approximates the location of the hole, i.e., locates, for example, the holes, sockets, threads, etc., in the scene, for example, from images, and uses, for example, a PD controller to follow a trajectory calculated from the approximation.It then activates the residual strategy, for example at a specific R of the plug 113 from the surface 115 and carries out the actual plugging-in according to the residual strategy. One action of the residual strategy is delta motion Δθ = (Δ x, Δ y, 0, Δθ x, Δθ y, Δθ z).Algorithm 2 is a detailed example of this procedure. According to various embodiments, various augmentations for training data elements may be used to improve robustness as well as generalizing over color and shape. Both the order and the properties of each augmentation have a great influence on generalization. For visual augmentation (i.e., augmentation of training image data), for training based on delta loss and relationship loss, this may include resizing, random clipping, color jitter, shifting, rotation, deleting, and random convolution. For the contrasting loss, example augmentations are such as random size matches of the clip, strong displacement, strong rotation, and erasure. Similar augmentations may be used for training data items within the same stack. In force augmentation (i.e., augmentation of training force input data), the direction of the vectors (F, M), rather than their magnitude, is typically the more important factor. Thus, according to various embodiments, the force input 402 to the delta network is the direction of the force and moment vectors. These may be augmented (e.g., by jitter) for training.As mentioned above, the relationship strategy can be used in particular for a multi-step task, e.g. multi-step plugging. It should be noted that in multi-step plug-in tasks, such as closing a door, it is typically more difficult to collect training data and verify that each step can be completed.According to various embodiments, in a multi-step task, images are pre-stored for each target, including one or more intermediate targets and a final target. Then, for each target to be currently reached (depending on the current step, i.e. according to a sequence of intermediate targets and, as the last element, the final target), an image data element for the current position is captured (e.g. by capturing images from both cameras 202) and fed into the encoder network 503 together with the image (or images) for the target, and a delta motion is derived by the relationship network 504 as described with reference to FIG. 5.FIG. 6 illustrates an example of a multi-step plug-in task.In the example of FIG. 6, a task for closing a door, the task consists of three steps starting from a starting position 601: insertion of the key 602, rotation of the lock 603 and then turning back 604. For each step, an image of the condition is captured with the two (e.g., 45-degree) cameras 202 and stored in advance. Execution of the task follows Algorithm 2 with the relational strategy and similarity function to switch between steps. For example, an intermediate target is considered to be reached when the similarity between the captured images and the pre-stored target images for the current step is below a threshold.In this manner, even if periodic backward data collection (e.g., according to Algorithm 1) is used and the lock and plug in states are not achieved even during training (only seeking states above or touching the hole surface), the controller 106 may successfully perform the task.In summary, according to various embodiments, a method as illustrated in FIG. 7 is provided.FIG. 7 shows a flow diagram 700 illustrating a method of controlling a robot to perform a task.At 701, a target image data element, which comprises at least one target image from a perspective of an end effector of the robot at a target position of the robot in which the robot has performed the task, is acquired.At 702, an original image data element is acquired that includes at least one original image from the perspective of the end effector of the robot at an original position of the robot.At 703, the source image data element and the target image data element are provided to a machine learning model configured to derive a delta motion between the current source position and the target position.At 704, the robot is controlled to move according to the delta motion to perform the task.The method of FIG. 7 may be performed by one or more computers including one or more computing devices. The term "data processing unit" may be understood as any type of entity that enables the processing of data or signals. For example, the data or signals may be processed according to at least one (i.e., one or more than one) specific function performed by the data processing unit. A computing device may include or be formed from an analog circuit, a digital circuit, a mixed signal circuit, a logic circuit, a microprocessor, a microcontroller, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a programmable gate array (FPGA), an integrated circuit, or any combination thereof. Any other type of implementation of the respective functions can also be understood as a data processing unit or logic circuitry. It should be appreciated that one or more of the method steps described in detail herein may be performed (e.g., implemented) by a computing device via one or more specific functions performed by the computing device.Various embodiments may receive and use image data from various visual sensors (cameras) such as video, radar, LiDAR, ultrasound, thermal image, etc. The embodiments may be used to train a machine learning system and to autonomously control a robot, e.g., a robot manipulator, to perform different plug-in tasks in different scenarios. Note that the neural network can be trained for a new plug-in task after training for a plug-in task, which reduces training time to bottom-on (transfer learning capabilities) as compared to training. The embodiments are applicable in particular for controlling and monitoring the execution of manipulation tasks, e.g. in assembly lines.According to one embodiment, the method is computer implemented.

Claims

A method of controlling a robot (100) to perform a task, comprising: training a machine learning model (503, 504) to derive delta movements from pairs of image data items by collecting a set of primary training items, each primary training item comprising an image data item comprising at least one image from the perspective of an end effector (104) at a respective position and a motion vector for moving the end effector (104) from the position to a desired position; generating secondary training items by selecting, for each secondary training item, two of the primary training items, wherein the image data items of the primary training items are included in the training item and a difference between the motion vectors of the primary training items is included as a ground truth motion vector; and training the machine learning model (503, 503, 504) using the secondary training elements in a monitored manner; acquiring a target image data element comprising at least one target image from a perspective of the end effector (104) of the robot (100) at a target position of the robot (100) at which the robot (100) has performed the task; acquiring an original image data element comprising at least one original image from the perspective of the end effector (104) of the robot (100) at an original position of the robot (100), supplying the original image data element and the target image data element to the trained machine learning model (503, 504) to derive a delta motion between the current original position and the target position; and controlling the robot (100) to move according to the delta motion to perform the task.The method of claim 1, wherein the machine learning model (503, 504) comprises an encoder network (503), the method comprising supplying the source image data item and the target image item to the encoder network (503), respectively, and supplying embeddings of the source image data item and the target image item to a neural network (504) of the trained machine learning model (503, 504).Method according to one of claims 1 to 2, comprising capturing the original image data element and the target image data element such that they each comprise two images, wherein the images for the same position of the end effector (104) of the robot (100) are recorded by two different cameras (117, 119) attached to the end effector (104).The method of claim 1, comprising controlling the robot (100) for a task of inserting an object (113) into an insert (114), wherein capturing the target image element comprises capturing the at least one target image by grasping a reference object that fits into the insert (114) by the robot (100), bringing the robot (100) into a position such that the reference object is inserted into the insert (114), and capturing the at least one target image at the position.A robot controller (106) configured to perform a method according to any one of claims 1 to 4.A robot (100) comprising a robot controller according to Example 5 and comprising an end effector (104) having at least one camera (117, 119) configured to acquire the at least one origin image.A computer program comprising instructions which, when executed by a processor, cause the processor to perform a method according to any one of claims 1 to 4.A computer readable medium storing instructions that, when executed by a processor, cause the processor to perform a method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Machine learning device, robot control device and robot vision system that uses a machine learning device, and machine learning method

    DE102019002065A1

  • Method for controlling an industrial robot

    DE102019106458A1

  • Robot control system

    DE102019122790A1

  • CONTROL DEVICE OF A ROBOT DEVICE THAT CONTROLS THE POSITION OF A ROBOT

    DE102020119704A1

  • Device and method for controlling a robot to insert an object into an insertion point

    DE102021109332A1