Apparatus and methods for training neural networks for controlling robots
Patent Information
- Application Number
- CN202210277446.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-22
- Filing Date
- 2022-03-21
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2042-03-21
Smart Images

Figure CN115107020B_ABST
Abstract
Description
Technical Field
[0001] Various embodiments generally relate to apparatus and methods for training neural networks for controlling robots. Background Technology
[0002] To enable the flexible manufacture or processing of objects by robots, it is desirable for the robot to handle objects regardless of their orientation within the robot's workspace. Therefore, the control device should be able to correctly generate control parameters for the robot based on the object's corresponding posture (i.e., position and orientation), allowing the robot to, for example, grasp the object at the correct position to, for example, attach the object to another object, or otherwise process (weld, paint, etc.) the object at the correct position. This means the control device should be able to generate appropriate control parameters from sensor information reflecting the object's posture, such as from one or more camera images recorded by cameras fixed to or near the robot. Summary of the Invention
[0003] According to various embodiments, a method for training a neural network for controlling a robot is provided, comprising: determining a camera pose in a robot unit and an uncertainty range around the determined camera pose; determining an object region in the robot unit, the object region containing the location of an object to be processed by the robot; generating training camera images, wherein for each training camera image, a training camera image camera pose is randomly set within the uncertainty range around the determined camera pose; a training camera image object pose is randomly set in the object region and the training camera image is generated, such that the training camera image displays an object having the training camera image object pose from the angle of a camera having the training camera image camera pose; generating training data from the training camera images, wherein one or more training robot control parameters are assigned to each training camera image for processing an object under the training camera image object pose of the training camera image; and using the training data to train the neural network, wherein the neural network is trained to output a description of one or more robot control parameters from a camera image displaying the object for processing the object.
[0004] The method described above utilizes properties of a robot arrangement used in practice (particularly camera position) to generate a training dataset for a robot controller (e.g., for grasping an object). What is achieved here is that the simulated camera images closely approximate camera images that occur or will occur in practice (for the corresponding object pose). Using these camera images as a training dataset for a neural network enables the neural network to produce accurate predictions of control parameters, such as the grasping position in a scenario where the robot should remove an object from an object region (e.g., a box).
[0005] Various embodiments are described below.
[0006] Example 1 is a method for training a neural network for controlling a robot as described above.
[0007] Example 2 is based on the method of Example 1, wherein the uncertainty range of the determined camera pose has the uncertainty range of the determined camera position and / or the uncertainty range of the determined camera orientation.
[0008] Therefore, both positioning and orientation uncertainties can be considered during camera calibration. Since both types of uncertainty typically occur in practice, training data can be generated to train neural networks that are robust to both types of uncertainty.
[0009] Example 3 is based on the method of Example 2, wherein the random selection of the camera pose of the training camera image is based on a normal distribution corresponding to the uncertainty range of the determined camera position and / or on a normal distribution corresponding to the uncertainty range of the determined camera orientation.
[0010] The normal distribution provides a physically real variation around the determined camera pose, which is typically derived in the case of calibration errors.
[0011] Example 4 is a method according to one of Examples 1 to 3, including determining an uncertainty range around the determined object region pose, and wherein the setting of the training camera image object pose includes randomly selecting a camera image object region pose within the object region uncertainty range and randomly setting the training camera image object pose in the object region to the selected camera image object region pose.
[0012] Therefore, uncertainties in the calibration target area are also taken into account.
[0013] Example 5 is based on the method of Example 4, wherein the random selection of the pose of the object region in the training camera image is performed according to a normal distribution corresponding to the uncertainty range of the determined object region pose.
[0014] Similar to the case of cameras, the normal distribution provides a physically real variation around the pose of the determined object region.
[0015] Example 6 is a method according to one of Examples 1 to 5, wherein the setting of the training camera image object pose is performed by setting the training camera image object pose according to the uniform distribution of the object pose in the object region.
[0016] This makes it possible to train the neural network uniformly for all object poses that occur during operation.
[0017] Example 7 is a method according to one of Examples 1 to 6, wherein the setting of the pose of the training camera image object is performed by setting the pose of the training camera image object according to the transformation to the object region.
[0018] Within the object region, the poses (positions and orientations) of the objects in the training camera images can be randomly selected, for example, selected in a uniform distribution. This allows the neural network to be trained uniformly for all object orientations that may appear during operation.
[0019] Example 8 is a method for controlling a robot, comprising: training a neural network according to one of Examples 1 to 7, receiving a camera image displaying an object to be processed by the robot, feeding the camera image to the neural network, and controlling the robot according to robot control parameters described by the output of the neural network.
[0020] Example 9 is an apparatus configured to perform a method according to one of Examples 1 to 8.
[0021] Example 10 is a computer program including program instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to Examples 1 to 8.
[0022] Example 11 is a computer-readable storage medium having stored program instructions thereon that, when executed by one or more processors, cause the one or more processors to perform a method according to one of Examples 1 to 8. Attached Figure Description
[0023] Embodiments of the invention are illustrated in the accompanying drawings and explained in more detail below. In the drawings, the same reference numerals generally refer to the same parts in multiple views. The drawings are not necessarily drawn to scale, but generally focus on illustrating the principles of the invention.
[0024] Figure 1 The robot is shown.
[0025] Figure 2 The components of a robotic unit related to the simulated generation of training camera images are shown.
[0026] Figure 3 A flowchart is shown for a method used to generate training camera images.
[0027] Figure 4 A flowchart is shown for a method of training a neural network for controlling a robot. Detailed Implementation
[0028] Various implementations, particularly those described below, can be achieved using one or more circuits. In one implementation, "circuit" can be understood as any type of logic implementation entity, which can be hardware, software, firmware, or a combination thereof. Thus, in one implementation, "circuit" can be hardwired logic circuitry or programmable logic circuitry, such as a programmable processor, such as a microprocessor. "Circuit" can also be software implemented or executed by a processor, such as any type of computer program. Any other means of implementing the corresponding functions as described in more detail below can be understood as "circuit" in accordance with the alternative implementations.
[0029] According to various embodiments, an industrial robot system consisting of an industrial robot with a grasping end effector, a camera, and a box with a grasping object is used. See below for reference. Figure 1 and Figure 2 Describe an example. The type and implementation of the robot kinematics are not specified here. (For example, it could be serial or parallel kinematics with different types and numbers of joints).
[0030] Figure 1 Robot 100 is shown.
[0031] Robot 100 includes a robotic arm 101, such as an industrial robot arm, for handling or mounting workpieces (or one or more other objects). Robotic arm 101 includes arm elements 102, 103, and 104 and a base (or support) 105 for supporting the arm elements 102, 103, and 104. The term "arm element" refers to a movable part of the robotic arm 101, the actuation of which enables physical interaction with the environment to, for example, perform a task. For control, robot 100 includes a (robot) control device 106 designed to enable interaction with the environment according to a control program. The final assembly 104 of arm elements 102, 103, and 104 (the assembly furthest from the base 105 in the kinetic chain), also called an end effector 104, may include one or more tools, such as a welding torch, a gripping instrument, a painting device, etc.
[0032] Other arm elements 102, 103 (closer to support 105) can form a positioning device, thereby providing a robotic arm 101 and an end effector 104 at its end, together with the end effector 104. The robotic arm 101 is a mechanical arm that can provide functions similar to a human arm (possibly with a tool at its end).
[0033] The robotic arm 101 may include joint elements 107, 108, and 109 that connect arm elements 102, 103, and 104 to each other and are connected to a support 105. Joint elements 107, 108, and 109 may include one or more joints, each of which can provide rotational movement (i.e., rotational motion) and / or translational movement (i.e., translational motion) to the associated arm element relative to each other. Movement of arm elements 102, 103, and 104 may be initiated by means of an actuator controlled by a control device 106.
[0034] The term "actuator" can be understood as a component constructed to realize a mechanism or process in response to its actuation. An actuator can realize mechanical motion from instructions (so-called activation) created by control device 106. Actuators such as electromechanical converters can be designed to convert electrical energy into mechanical energy in response to their actuation.
[0035] The term "control device" can be understood as any type of logical implementation entity, which may include, for example, circuitry and / or a processor capable of executing software, firmware, or a combination thereof stored in a storage medium and issuing instructions, for example, to the actuator in this example. The control device may be configured, for example, by program code (e.g., software) to control the operation of the system, which in this example is a robot.
[0036] In this example, the control device 106 includes one or more processors 110 and a memory 111, the memory 111 storing code and data, based on which the processors 110 control the robotic arm 101. According to various embodiments, the control device 106 controls the robotic arm 101 based on a machine learning model 112 stored in the memory 111.
[0037] According to various implementations, the machine learning model 112 is designed and trained to describe (or predict) control parameters for the robotic arm 101 in response to one or more camera images fed to the machine learning model 112.
[0038] The robot 100 and / or the robot unit arranged therein may be equipped with, for example, one or more cameras 114, which are configured to record images of the workspace of the robot 100 and images of an object 113 (or multiple objects) arranged in the workspace.
[0039] For example, the machine learning model is a neural convolutional network, to which one or more images (e.g., depth images) are fed. Before practical use, the convolutional network is trained with training data, thereby allowing the control device to specify appropriate control parameters for different poses (positions and orientations) of the object 113.
[0040] For example, the training data has a large number of training data elements, where each training data element contains a camera image (or multiple camera images) and associated (i.e., ground-based) control parameters. These control parameters, for example, include the position that the end effector 104 should take in order to process object 113. The neural network can then be trained using supervised learning.
[0041] According to various implementations, different training data images are generated for different object poses and associated control parameters. For example, if a specific random object pose is assumed for the training data images, the correct control parameters can be directly calculated from them (e.g., because the position where the end effector 104 will be moved or which position on the object will be grasped is known from the object pose). Therefore, the main task of generating training data elements can be viewed as generating suitable camera images for a specific (random) object pose (for the corresponding camera in the case of multiple cameras), so that if the object has said object pose in practice, the (corresponding) camera 114 will record that image.
[0042] According to various embodiments, a method is provided for calculating the transformation between camera pose and object pose to generate a camera image (e.g., for a training dataset with virtual depth images, used as input to a convolutional neural network to control the grasping of an object by a robotic arm 101).
[0043] According to various implementations, calibration parameters (typically applicable to robot cells in practice) are used herein, and uncertainties (e.g., uncertainties in camera pose) are considered in a physically reasonable manner. This provides a method for generating training datasets based on simulations, the training datasets being hardware-aware, i.e., taking into account the properties of existing hardware (such as cameras).
[0044] Objects can have any geometry. For example, a camera image is a depth image. For example, if the camera is a pinhole camera, the camera image is the corresponding image.
[0045] According to various implementations, a robustness metric is also calculated during the training data generation process. This robustness metric represents the probability of successful access (pickup) of a specific end effector of the robot (i.e., gripper type). The position and orientation of a successfully accessed gripper in the camera images (or multiple camera images) are encoded by the corresponding translation, rotation, and size of the image content (i.e., the specific object to be recorded) in the camera images. During training, a neural network (e.g., a deep convolutional network) is trained, for example, to predict robustness metric values for different gripping positions in the camera images. These robustness metrics can be viewed as descriptions of the robot's control parameters, for example, as the gripping position where the neural network outputs the highest robustness metric value.
[0046] Following training, i.e. during operation (in real-world applications), the trained neural network is used for reasoning, i.e., predicting robustness values from real-world recorded camera images for various access possibilities (e.g., end effector poses). The control device can then, for example, select the access possibility with the maximum robustness value and control the robot accordingly.
[0047] When the camera generates (simulated) camera images for different object poses, the calibration results of the camera in the actual robot arrangement (i.e., the robot cell) are used. Camera calibration is typically performed when building the robot system in the robot cell or before the processes that should be performed by the robot to improve the accuracy of the understanding of the camera pose, and thus also improve the accuracy of the robot actions (robot control) determined from the images generated by the camera.
[0048] According to various implementations, non-inherent camera calibration values are used to generate training camera images, namely, three-dimensional (e.g., Cartesian) translation and rotation values of the camera relative to the robot unit's world coordinate system (i.e., the robot unit's global reference coordinate system). These parameter values can be obtained, for example, through typical optimization-based calibration methods. Furthermore, according to various implementations, uncertainty parameters that can be determined during the calibration process are used in the camera calibration. These uncertainty parameters represent the accuracy of the translation and rotation values determined during calibration. Additionally, inherent camera parameters (camera matrix and noise values) of the real camera used in the robot unit can be used to ensure that the generated virtual camera images are as close as possible to images from the real camera during simulation. Furthermore, information regarding the pose and size of the object region (e.g., a box) where object 113 may be located is determined and said information is used as input for generating the training camera images.
[0049] Figure 2 The robot unit 200 and its components related to the simulation-generated training camera images are shown.
[0050] Camera 201 has a coordinate system K K Object 202 has a coordinate system K O The object region (e.g., a box or container) 203 has a coordinate system K. B The robot has a coordinate system K in its base. R The end effector (gripper) moved by the robot has a coordinate system K. G The robot unit 200 has a global coordinate system K. WThis global coordinate system represents a fixed reference system within robot unit 200. The coordinate systems of components 201, 202, and 203 are fixed for their respective components, thus defining the pose of the corresponding components within the robot unit. To generate training camera images, the camera is modeled as a virtual pinhole camera with a perspective projection of the surface of the measured object's surface model.
[0051] Camera 201 has a field of view 204.
[0052] Figure 3 A flowchart 300 is shown for a method of generating training camera images.
[0053] This method can be executed by the control device 106 or by an external data processing device that generates the training data. The control device 106 or the external data processing device can then use the training data to train the neural network 112.
[0054] The input data used to generate camera images is the external camera calibration parameter value p. K = (x K , y K , z K A K B K C K These describe the camera coordinate system K. K In the global coordinate system K W The pose (position and orientation) and the associated uncertainties of translation and rotation values, such as u K =(t K , r K For simplicity, this example assumes a translation t for all coordinates or angles. K Uncertainty and rotation r K The uncertainties are all the same. These uncertainties are typically a result of the camera calibration process. Angle A K B K C K For example, Euler angles can be used, but other representations of spatial orientation can also be used.
[0055] The additional input data consists of inherent camera parameters in the form of a camera matrix K, which can typically be read from the camera after inherent camera calibration, and the noise value σ in the camera's viewing direction. K The noise value is typically specified by the camera manufacturer.
[0056] The other input data is the expected object region location p. B = (x B , y B , z B AB B B C B The object region location describes the position of the center of the cuboid object region in the global coordinate system in this example, and the size of the object region b. B =(h B , w B , d B These parameter values are typically given by the configuration of the robot cell.
[0057] In 301, based on the camera calibration parameter value p K Calculate the initial camera transformation That is, mapping the standard camera pose to the calibration parameter value p K Translation and rotation of a given camera pose.
[0058] In 302, the additional camera rotation Rot is calculated. K It calculates the uncertainty of rotation during camera calibration. Rotation Rot K This is calculated, for example, based on a convention of rotation axis and rotation angle, which is based on the fact that every rotation in three-dimensional space can be represented by a suitable rotation axis and a rotation angle about that axis. To calculate the rotation Rot... K The rotation axis is sampled from a uniform distribution of unit vectors in three-dimensional space. This is based on a mean and standard deviation r that are both zero. K The rotation angle is sampled from a normal distribution (i.e., the uncertainty of camera orientation). Then the transformation is calculated. .
[0059] In 303, the additional camera translation Trans is calculated. K The camera translation takes into account the uncertainty of camera calibration translation. To this end, the translation direction is sampled from a uniform distribution of unit vectors in 3D space. This is achieved using vectors with a mean and standard deviation of zero, t. K The translation length is sampled from a normal distribution (i.e., the uncertainty of the camera position). Then the final camera transformation is calculated. .
[0060] In 304, based on the object region pose p B Computational object region transformation That is, mapping the pose of a standard object region to a value p. B Translation and rotation of a given object region pose. Similar to the case of camera pose, uncertainties in the object region pose can also be additionally considered.
[0061] In 305, the additional object rotation Rot is calculated. OTo this end, rotations are sampled from a uniform distribution of rotations to account for random rotations of object 202 within object region 203. This is relevant, for example, in the case where objects are arranged in a stack, where each object can rotate randomly in three-dimensional space. Fixed orientations can also be used if it can be expected that objects will always lie on a flat surface (rather than in a stack). Furthermore, object translation Trans is calculated by sampling from a uniform distribution over the object region. O : U([-hB / 2, +hB / 2] x [-wB / 2, +wB / 2] x [dB / 2, +dB / 2]), to account for the possibility that object 202 may have a random position within object region 203. In the case of an object heap, the object rotations and object translations of the individual objects in the heap with respect to the object region can also be determined alternatively through physical simulation. Therefore, the transformation is derived from object rotations and object translations. .
[0062] In 306, the pose of object region 203 is used to transform object 202. Here, by means of... Calculate the object transformation with respect to the world coordinate system.
[0063] In 307, the camera transformation with respect to the object coordinate system is calculated as follows: The transformation is then used to locate the virtual camera in the simulated space. The object (i.e., its surface model) is located at the origin of this simulated space. Camera images are generated through virtual measurements by rendering the simulated scene produced in this way (specifically using the located virtual camera) (e.g., to generate a depth image). For this, a camera matrix K is used, and a noise parameter σ is applied to the measurement data. K The normally distributed noise.
[0064] This method of generating camera images was performed multiple times in order to produce a large number of camera images for the training dataset 308.
[0065] Since the camera poses and object poses used physically represent parameter values from real robot units with associated uncertainties, this results in the training dataset corresponding to the real physical configuration of the robot units.
[0066] In summary, a method is provided according to various implementation methods, as referenced below. Figure 4 As described.
[0067] Figure 4 A flowchart 400 is shown, illustrating a method for training a neural network for controlling a robot.
[0068] In 401, the camera pose in the robot cell and the range of uncertainties around the determined camera pose are determined.
[0069] In step 402, an object region within the robot unit is determined, the object region containing the location of objects to be processed by the robot.
[0070] At 403, training camera images are generated.
[0071] Here, for each training camera image, in 404 the camera pose of the training camera image is randomly set within an uncertainty range around the determined camera pose; in 405 the object pose of the training camera image is randomly set in the object region; and in 406 a training camera image is generated such that the training camera image displays an object having the object pose of the training camera image from a camera angle having the camera pose of the training camera image.
[0072] In step 407, training data is generated from the training camera images, wherein one or more training robot control parameters are assigned to each training camera image for processing objects in the training camera image object pose.
[0073] In step 408, the training data is used to train a neural network, wherein the neural network is trained to output descriptions of one or more robot control parameters for processing the object from camera images displaying the object.
[0074] This method generates synthetic sensor data. This sensor data can correspond to various optical sensors such as stereo cameras, time-of-flight cameras, and laser scanners.
[0075] These implementation methods can be used to train machine learning systems and autonomously control robots to perform different manipulation tasks in various scenarios. Specifically, they can be applied to control and monitor the execution of manipulation tasks, such as on an assembly line. These implementation methods can, for example, be seamlessly integrated into traditional GUIs for control processes.
[0076] "Robot" can be understood as any physical system (with mechanical parts whose movement is controlled), such as computer-controlled machines, vehicles, household appliances, power tools, manufacturing machines, personal assistants, or access control systems.
[0077] Neural networks can be convolutional networks used for data regression or data classification.
[0078] Various implementations can receive and utilize sensor signals from various optical sensors to obtain sensor data, such as information about the state of a demonstration or system (robot and object), configuration, and scene. This sensor data can be processed. This may include classifying the sensor data or performing semantic segmentation on the sensor data to, for example, identify the presence of an object (in the environment from which the sensor data was obtained). Implementations can be used to train machine learning systems and autonomously control robots to perform different manipulation tasks in different scenarios. In particular, implementations can be applied to control and monitor the execution of manipulation tasks, such as on an assembly line. These implementations can, for example, be seamlessly integrated into conventional GUIs for control processes.
[0079] According to one implementation, the method is computer-implemented.
[0080] Although the invention has been shown and described, particularly with reference to specific embodiments, those skilled in the art will understand that various changes in design and detail may be made without departing from the spirit and scope of the invention as defined by the following claims. Therefore, the scope of the invention is determined by the appended claims and is intended to include all variations falling within the literal meaning or equivalent scope of the claims.
Claims
1. A method for training a neural network for controlling a robot, comprising: Determine the camera pose in the robot cell and the range of uncertainties around the determined camera pose; Determine the object region within the robot unit, the object region containing the location of the object to be processed by the robot; Generate training camera images, For each training camera image The camera pose of the training camera images is randomly set within a range of uncertainty around the determined camera pose; The pose of training camera image objects is randomly set in the object region; The training camera image is generated such that the training camera image displays the object having the object pose of the training camera image from the perspective of a camera having the camera pose of the training camera image; Training data is generated from the training camera images, wherein one or more training robot control parameters are assigned to each training camera image for processing objects in the training camera image object pose of the training camera image; as well as The training data is used to train the neural network, wherein the neural network is trained to output descriptions of one or more robot control parameters from camera images displaying the object for processing the object.
2. The method according to claim 1, wherein, The uncertainty range of the determined camera pose includes the uncertainty range of the determined camera position and / or the uncertainty range of the determined camera orientation.
3. The method according to claim 2, wherein, The random selection of the camera pose for the training camera images is based on a normal distribution corresponding to the uncertainty range of the determined camera position and / or a normal distribution corresponding to the uncertainty range of the determined camera orientation.
4. The method according to any one of claims 1 to 3, comprising determining an uncertainty range around the determined object region pose, wherein setting the training camera image object pose comprises randomly selecting a camera image object region pose within the object region uncertainty range and randomly setting the training camera image object pose in the object region to the selected camera image object region pose.
5. The method according to claim 4, wherein, The random selection of the pose of the object region in the training camera image is based on a normal distribution corresponding to the uncertainty range of the determined object region pose.
6. The method according to any one of claims 1 to 3, wherein, The pose of the training camera image object is set by setting the position of the training camera image object according to the uniform distribution of the object position in the object region.
7. The method according to any one of claims 1 to 3, wherein, The pose of the training camera image object is set by setting the pose of the training camera image object according to the transformation to the object region.
8. A method for controlling a robot, comprising: Train the neural network according to any one of claims 1 to 7; Receive and display camera images of objects to be processed by the robot; The camera images are fed into the neural network; as well as The robot is controlled based on robot control parameters described by the output of the neural network.
9. An apparatus configured to perform the method according to any one of claims 1 to 8.
10. A computer program product comprising program instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having stored thereon program instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Mechanical arm control method and device based on deep learning
CN109531584A
Robot intelligent grabbing control method and system based on 3D vision
CN111275063A