Device and method for training control policy for manipulating object

By using semi-supervised learning to generate pseudo-labels and combining it with reinforcement learning, the method improves the efficiency and adaptability of robot grasping techniques, overcoming the limitations of sparse reward feedback and traditional supervised learning.

JP2025096228APending Publication Date: 2025-06-26ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024218307
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-12-13
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing grasping techniques in robot operation rely on supervised learning and offline training, which limits their ability to adapt to unseen objects or new environmental conditions, and are hindered by sparse reward feedback.

Method used

A semi-supervised learning method is employed to generate pseudo-labels for unlabeled data, combining with reinforcement learning to train a control policy, thereby improving learning efficiency and adaptability.

Benefits of technology

This approach addresses the issue of sparse reward feedback by utilizing unlabeled data in online grasping learning, enhancing data efficiency and enabling better adaptation to new objects and environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025096228000001_ABST
    Figure 2025096228000001_ABST
Patent Text Reader

Abstract

To provide a method for training a control policy for manipulating an object.SOLUTION: A method includes: receiving an input data element including image data representing a shape of an object to be manipulated and a position of the object in a scene for each of one or more objects in each of one or more scenes; generating, for each input data element, one or more training data elements by generating augmentation of the image data and pseudo labels for the augmented image data according to a semi-supervised learning scheme and training a control policy using the generated training data elements.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an apparatus and a method for training a control policy for operating an object.

Background Art

[0002] The core task in robot operation is grasping, which is a basic skill that opens the door to more complex actions such as pick-and-place or bin picking. In bin picking, the goal is to remove objects from a container and place them in specific locations, which has a wide range of applications. However, bin picking is difficult due to problems such as noise perception, object obstacles, and collisions in planning. Therefore, a robust approach for effectively handling this task is needed. Recent grasping techniques often rely on deep learning methods, which allow each machine learning model to predict grasping actions without relying on a pre-defined model, thereby making the machine learning model applicable to a wide range of objects in an unstructured environment. However, typical approaches rely on supervised learning and offline training, which potentially limits the ability of the machine learning model to adapt to unseen objects or new environmental conditions. Therefore, an approach that enables efficient online grasping (or generally operation) learning is desired.

Summary of the Invention

Means for Solving the Problems

[0003] According to various embodiments, a method for training a control policy (represented by a machine learning model, that is, training the control policy includes training the machine learning model) for operating an object, The method is, ·Receiving an input data element including image data representing the shape of an object to be manipulated and the position of the object within a scene for each of one or more objects in each of one or more scenes; ·Generating one or more training data elements by, for each input data element, generating an augmentation of the image data and a pseudo-label for the augmented image data according to a semi-supervised learning scheme; ·Training a control policy using the generated training data elements (and, optionally, other training data elements, in particular other training data elements including received input data elements without augmentation, e.g., observed rewards (and thus observed labels rather than pseudo-labels)); A method is provided that includes the above.

[0004] By the above method, it is possible to address the problem of sparse reward feedback by using unlabeled data in online grasping learning using semi-supervised learning (SSL), thereby improving learning efficiency. Various SSL methods can be incorporated into reinforcement learning, e.g., into convolutional soft actor-critic (ConvSAC), to obtain a scheme referred to herein as SSL-ConvSAC. In particular, an SSL method based on curriculum learning can be used.

[0005] This makes it possible to address the problem that the balance between the amount of labeled data and the amount of unlabeled data, which typically occurs in grasping applications (and similar applications) and can cause divergence in online training, is extremely skewed.

[0006] The operation can in particular mean grasping and picking up (e.g., gripping or suction in the case of a suction pad). The method is also applicable to other tasks such as turning a key, pressing a button, pulling a lever, etc.

[0007] Examples are described below.

[0008] Example 1 is a method for training a control policy as described above.

[0009] Example 2 is the method described in Example 1, including training a control policy using reinforcement learning.

[0010] By using a semi-supervised learning scheme in the context of reinforcement learning to train a control policy (in other words, an agent) for operating an object, it becomes possible to address the problem of sparse rewards in such settings.

[0011] Example 3 is the method described in Example 2, where training the control policy includes training an actor and a critic using the generated training data elements.

[0012] Both the actor and the critic can be trained using their respective losses that use the generated training data elements. By generating training data elements, the number of training data elements increases (compared to an approach that uses only training data elements directly corresponding to the input data elements), so training using losses based on the generated training data elements results in losses that cover a wider range of states.

[0013] Example 4 is the method described in any one of Examples 1 to 3, where training the control policy includes training a neural network representing the control policy.

[0014] By using the generated training data elements, the neural network can be efficiently trained using backpropagation.

[0015] Example 5 is the method according to any one of Examples 1 to 4, in which for each of the generated training data elements, pseudo-labels for each of a plurality of operation postures are generated, and each operation posture includes an operation position corresponding to each pixel in each of the extended image data.

[0016] Thereby, for example, dense pseudo-labels are generated, and such dense pseudo-labels significantly increase data efficiency compared to sparse rewards (which are only received for successful operation postures).

[0017] Example 6 is the method according to any one of Examples 1 to 5, in which training a control policy includes determining a loss (e.g., for each actor and critic, each machine learning model (e.g., actor and critic, or other machine learning model that at least partially represents the control policy) is adapted to reduce its respective loss) that includes a loss term for the generated training data elements, each loss term being soft-weighted in the loss function by applying a softmax function to the confidence of the pseudo-label of the operation posture of each training data element, and / or the loss term being filtered out of the loss function if the confidence of the pseudo-label of the operation posture of each training data element is below a predetermined threshold.

[0018] Thereby, it becomes possible to improve generalization and reduce confirmation bias.

[0019] Example 7 is the method according to any one of Examples 1 to 6, in which the threshold is a per-pixel threshold.

[0020] Thereby, it becomes possible to further improve generalization and further reduce confirmation bias.

[0021] Example 8 is a method for controlling a robot device. This method includes training the control policy described in any one of Examples 1 to 7, receiving image data representing a scene for which the robot device is to be controlled, supplying the acquired image data to the control policy, and generating a control signal for the robot device according to the output generated by the control policy in response to the acquired image data.

[0022] Example 9 is a data processing device (particularly, a control device of a robot device) configured to implement the method described in any one of Examples 1 to 8.

[0023] Example 10 is a computer program including instructions for causing a computer to implement the method described in any one of Examples 1 to 8 when executed by the computer.

[0024] Example 11 is a computer-readable medium including instructions for causing a computer to implement the method described in any one of Examples 1 to 8 when executed by the computer.

[0025] In the drawings, like reference numerals generally refer to the same parts throughout the plurality of different drawings. The drawings are not necessarily to scale; instead, emphasis is generally placed on illustrating the principles of the present invention. In the following description, various aspects will be described with reference to the following drawings.

Brief Description of the Drawings

[0026]

Figure 1

Figure 2

Figure 3

Best Mode for Carrying Out the Invention

[0027] The following detailed description refers to the accompanying drawings, which show specific details and aspects of the present disclosure for implementing the present invention. Other aspects can also be used, and structural, logical, or electrical changes can be made without departing from the scope of the present invention. Since some aspects of the present disclosure can be combined with one or more other aspects of the present disclosure to form new aspects, the various aspects of the present disclosure are not necessarily mutually exclusive.

[0028] Hereinafter, various embodiments will be described in more detail.

[0029] FIG. 1 shows a robot 100.

[0030] The robot 100 includes a robot arm 101, for example, an industrial robot arm for processing a workpiece (or one or more other objects 113) or for assembling. The robot arm 101 includes manipulators 102, 103, 104 and a base (or support) 105 that supports these manipulators 102, 103, 104. The term "manipulator" refers to the movable members of the robot arm 101, and by operating these movable members, physical interaction with the environment becomes possible, for example, to perform a certain task. For control, the robot 100 includes a (robot) control device 106, which is configured to implement interaction with the environment according to a control program. The last member 104 (the one farthest from the support 105) among the manipulators 102, 103, 104 is also referred to as an end effector 104 and includes a gripping tool (which may be a suction gripper).

[0031] Other manipulators 102, 103 closer to the support 105 can constitute a positioning device, whereby a robot arm 101 having the end effector 104 at its end is provided together with the end effector 104. The robot arm 101 is a mechanical arm that can provide functions similar to those of a human arm.

[0032] The robot arm 101 may include joint elements 107, 108, 109, and these joint elements 107, 108, 109 interconnect the manipulators 102, 103, 104 with each other and interconnect the manipulators 102, 103, 104 and the support 105. The joint elements 107, 108, 109 may have one or more joint mechanisms, and each of these joint mechanisms can bring about a rotatable movement (i.e., rotational movement) and / or a translational movement (i.e., displacement) relative to each other for the associated manipulator. The movement of the manipulators 102, 103, 104 can be initiated using actuators controlled by the control device 106.

[0033] The term "actuator" can be understood to be a component configured to act on a mechanism or process in response to being driven. The actuator can carry out instructions (so-called activation) issued by the control device 106 to cause a mechanical movement. The actuator, for example an electromechanical transducer, can be configured to convert electrical energy into mechanical energy in response to being driven.

[0034] The term "control device" can be understood as any type of entity that implements logic, and this entity can include, for example, software, firmware stored in a storage medium, or a circuit and / or a processor capable of executing a combination of these, which in this embodiment, for example, can issue commands to an actuator. For example, a control device can be configured by program code (e.g., software) to control the operation of a system, which in this embodiment is the operation of a robot device.

[0035] In this embodiment, the control device 106 includes one or more processors 110 and a memory 111 for storing code and data, and the processor 110 controls the robot arm 101 based on this code and data. According to various embodiments, the control device 106 controls the robot arm 101 based on a machine learning model 112 (e.g., including one or more neural networks) stored in the memory 111.

[0036] Reinforcement learning (RL) can be used to train a machine learning model, for example, using actor-critic RL. For example, the machine learning model can use a fully convolutional network (FCN) to learn dense pixel-by-pixel grasping quality prediction, that is, to train the critic. Pixel-by-pixel parameterization can also be used for grasping primitives, that is, for the actor. However, during online learning, the agent (i.e., the control device 106) only receives sparse feedback of grasping success or failure at only one pixel position (of the input image indicating the object to be grasped) selected by the agent according to the control policy used by the agent (i.e., according to the actor). Therefore, for example, the corresponding neural network (implementing the actor and the critic) is only updated through backpropagation via the respective losses at this single pixel point.

[0037] Accordingly, according to various embodiments, an approach is provided that can utilize backpropagation through all pixel points of an input image. In particular, according to various embodiments, the advantages of semi-supervised learning (SSL) and the advantages of RL-based online grasping learning are combined. Pixel points with reward feedback are used as labeled data, while on the other hand, the remaining pixels without reward feedback are considered as unlabeled data but are utilized using semi-supervised learning (by generating pseudo-labels for these pixels) to improve training and overall performance.

[0038] For example, a semi-supervised learning-based fully convolutional soft actor-critic (SSL-ConvSAC) that combines both the true reward and the pseudo-labeled reward for grasping policy learning is used.

[0039] Various embodiments including the SSL-ConvSAC scheme will be described in detail below with respect to online grasping learning in bin picking applications.

[0040] Specifically, when a scene's RGB-D (i.e., including depth in addition to color) image I ∈ R H×W×4 (e.g., of the workspace of robot 100) is given, a grasping policy π (represented by the neural network of machine learning model 112) should be learned. The grasping policy π (generally a control policy) is a mapping from the image space to the output map space R H×W×4 that ideally maximizes the total grasping success rate over a long period. For the image I, the grasping policy output (action map) is a multi-channel map of a one-dimensional grasping quality Q per pixel and a three-dimensional grasping configuration A per pixel representing gripper rotation, for example, by Euler angles. Then, the control device 106 selects an action (i.e., a grasping position) having (h * , w * ) = argmax h’,w’ Q[h’, w’], and the grasping configuration for this grasping position from the action map A[h* , w * is extracted as such, and the robot arm 101 can be correspondingly controlled to grip the object 113. After each gripping attempt (index t), the reward r t is 1 if the gripping attempt is successful in picking up the object, and 0 otherwise. The objective is to optimize the policy π so that the total reward Σ t r t is maximized. For each input image, the reward feedback r t is associated only with the selected gripping, i.e., the gripping at the pixel (h * , w * ), while on the other hand, other pixel positions

Number

Number

[0041] The online grasping learning problem can be formulated as a Markov decision process (MDP) given by the tuple (S, A, P, R), where S is the state space, A is the action space, P is the transition probability function, and R is the reward function. At each control step, the state is observed, an action is executed according to the control policy, and a reward is received from the environment. The subsequent state follows according to the transition probability function.

[0042] FIG. 2 shows a fully convolutional Soft Actor-Critic based on SSL (SSL-ConvSAC) according to one embodiment.

[0043] The machine learning model 200 (e.g., corresponding to the machine learning model 112) includes a fully convolutional neural network (FCN) (e.g., having an architecture used by ConvSAC and HACMan (Hybrid Actor-Critic Maps for Manipulation)) as an actor network 201 for inferring the dense grasping configuration map A φ (s), and as a critic network 202 for approximating the dense grasping quality map Q θ (s, A φ (s)). The machine learning model 200 uses a pixel encoder network 203 that generates an action pixel encoding 204 and a state pixel encoding 205 to infer an embedding for each pixel position of the input state s. The actor 201 convolves the action pixel encoding 204 and infers the grasping configuration for each pixel. The output of the actor 201 is, for example, a Gaussian distribution for each pixel, where the mean is the predicted (i.e., inferred) grasping orientation at each pixel, and its variance is the uncertainty used for exploration in learning. The action (i.e., the grasping orientation, and the position of the action is calculated by transforming each pixel position into the world coordinate system) is concatenated 206 with the corresponding state pixel embedding of these actions and evaluated by the critic module 202, and as a result, a dense grasping Q-value map is obtained.

[0044] The state s is represented by 7 - dimensional input data composed of a color image, a normal surface map, and a height map. That is, the t - th state is the triple s t =(I c ,I n ,I d ) t and,

Number

Number

Number

[0045] The critic network and the actor network can be updated by formulating a critic loss using the labeled pixels. The critic loss is formulated as a classification task using a reward label r∈{0,1} indicating grasping failure and grasping success respectively. The critic uses, for example, the BCE loss, and the episode range ends after each grasping attempt. Specifically, the critic loss and the actor loss for the labeled pixels are

Number

[0046] This update should be noted to backpropagate the loss only through a single pixel (h t , w t ) in both the dense grasping quality map Q and the action map A.

[0047] Specifically, for each input image, the amount of labeled data is N l = 1 only, while on the other hand, the amount of unlabeled data is N u = (H × W) - 1, which are the remaining state pixels. This setting is due to the fact that by rearranging the scene to its previous state to collect grasping samples at other pixel positions, as a result, other different states may occur in the real - world setting. The approach provided according to various embodiments enables handling the realistic setting where online learning operates without interruption in industrial picking cells.

[0048] According to various embodiments, as described above, this problem that the reward feedback in online grasping learning based on RL is sparse (and thus the set of labeled data is small and the set of unlabeled data is large) is addressed by semi - supervised learning.

[0049] According to various embodiments, as will be described in more detail below, SSL techniques such as FixMatch and curriculum - learning - based SSL such as FlexMatch and FreeMatch are applied to the online grasping learning problem, i.e., incorporated as SSL into, for example, SSL - ConvSAC. According to other embodiments, SSL based on context - curriculum learning is used.

[0050] SSL - ConvSAC uses consistency regularization for SSL to train the actor A φ and the critic Q θRewrite the loss. The critic 201 and the actor 202 are updated using a common objective based on labeled and unlabeled data. The update using labeled data is defined in Equation (1). The update using unlabeled data is given by Equation (2) when a data sample (s,a,r) is given, where the action a encodes the pixel (h,w) labeled by the reward r, while the unlabeled pixels are

Number

Number

Number

Number

Number

Number

Number

[0051] According to various embodiments, the SSL objective includes calculating a per-pixel loss 213 (e.g., including BCE loss without reduction, e.g., applying an augmentation 214 to the pseudo-labels (i.e., pseudo-label map) 211 to associate these pseudo-labels 211 with the strongly augmented input (image) data 210 and calculating the BCE loss 212), whereby, in order to simultaneously process the losses for all N u unlabeled data points, it should be noted that it is possible to utilize parallel computing in a fully convolutional network to determine the per-pixel loss 213.

[0052] As a result, the common objectives of the actor and the critic are, respectively

Number

Number

[0053] The SSL objective for unlabeled data is calculated per pixel, and thus the final loss is the sum of the losses over all pixels. When there is an argmax operation or a max operation for the grasping quality map Q ∈ R H×W at any time, for the binary class, Q ∈ R H×W×2Specifically, it should be noted that Q[·,·,1]=Q, which is the Q-value for the class of success, and Q[·,·,0]=1.0 - Q, which is the Q-value for the class of failure, are implicitly assumed. Furthermore, the argmax operation or the max operation is applied over the last axis (i.e., regarding whether grasping success or grasping failure exists).

[0054] Furthermore, a pseudo-label mask 215 (for the pseudo-labeled pixels) used in the actor loss can be provided, that is, it should be noted that the actor loss is backpropagated only at the pixels where the pseudo-labels exist.

[0055] In the following, several options regarding the SSL scheme used in SSL-ConvSAC will be described.

[0056] 1) SSL-ConvSAC Based on FixMatch The first option is to utilize FixMatch for SSL-ConvSAC. For this purpose, a certain threshold τ is defined, and based on this certain threshold τ, the pseudo-labels with high confidence are retained. In particular, the weighting function is

Equation

[0057] 2) Curriculum-Based SSL-ConvSAC For example, one of FlexMatch and FreeMatch, which are SSL frameworks based on two curriculums, can be used. Instead of using a fixed constant threshold τ, FlexMatch and FreeMatch introduce curriculum learning to adjust τ and control the method by which pseudo-labels from individual classes are retained. The main idea can be seen as filtering out noisy pseudo-labels with a high threshold and leaving only high-quality pseudo-labels. Specifically, the adaptive threshold that can be used to recalculate the weighting function for each class c is

Number

[0058] a) SSL-ConvSAC based on FlexMatch According to FlexMatch, in each training step t, the model learning effect σ t (c), c ∈ {0, 1} is defined, where the class with few samples having a predicted confidence reaching the threshold is considered to have a greater learning difficulty or a worse learning state. Assuming the size of the replay buffer is |B|, the total number of unlabeled pixels is N u ×|B|. The learning effect is

Number

[0059] Here (as is the case in most situations), the operations within the identity function are performed pixel by pixel. The sum is taken over all unlabeled pixels and over the samples in the replay buffer. As a result, by normalizing σ t (c) within the range [0, 1], the adaptive threshold τ t (c) is

Number

[0060] b) SSL-ConvSAC based on FreeMatch Instead of adjusting the confidence threshold according to only the information of the current step as in FlexMatch, FreeMatch involves self-adapting this value as the model learning progresses. In particular, the self-adaptive global threshold is in Equation (7) to globally track the overall learning state across all classes among the unlabeled data

Number

Number

Number

Number

Number

[0061] It may or may not use the fairness regularization of FreeMatch.

[0062] 3) SSL-ConvSAC Based on Context Curriculum The above variations of SSL-ConvSAC mainly utilize existing SSL methods for the framework of the soft actor-critic and per-pixel grasping prediction. Compared with standard SSL, the main issue in this setting can be seen as the extreme imbalance between the labeled data and the unlabeled data as described above. This can rapidly lead to the problem of confirmation bias. If the mini-batch includes a 1:100 ratio between the labeled data and the unlabeled data, most SSL methods may be affected by this problem. According to various embodiments, three countermeasures (or at least one or two of these countermeasures) are taken to improve generalization and reduce the confirmation bias as described below. In particular, the reliability of the machine learning model is reduced: 1) Lower confidence threshold: This is useful for filtering out pseudo-labels with low confidence for the curriculum-based method. In particular, the lower limit of the adaptive threshold is τ t = max{τ t , τ lb} is introduced, where τ lb is, for example, a predetermined lower confidence threshold for filtering out overly low-confidence labels when the threshold τ t is made overly small. 2) Soft weighting function: The hard weighting λ t in Equation (4) treats both low-confidence and high-confidence pseudo-labels equally as long as their confidence exceeds the threshold. According to various embodiments, soft weighting via the softmax function

Equation

Number

Number

Number

Number

[0063] The calculations of equations (6) and (9) are also performed pixel-wise. As a result, the weighting function in equation (4) completely includes pixel-wise terms.

[0064] Weak augmentation and strong augmentation are, for example, the following, i.e., i) Color conversions including automatic contrast, brightness, contrast, equalization, posterization, sharpness, and / or solarization, ii) Geometric transformations: rotation and shift operations that change information regarding normal vectors within a 7-channel image, iii) Application of noise operators: uniform distribution noise and binary noise including one or more of.

[0065] For example, the color conversion is applied to the RGB channels, the uniform distribution noise is applied to the depth channel, the binary noise is applied to the normal vector channel, and the geometric transformation is applied to the entire 7 channels. In the case of weak augmentation, the random rotation set is, for example, [-10, 10] degrees, the random shift is [-10, 10] pixels, and there is color jitter in the RGB channels. In the case of strong augmentation, the random rotation is, for example, within the range of [-180, 180] degrees, and the random shift shifts by [-30, 30] pixels across the entire 7 channels. Uniform distribution noise in the range of 5 mm is used for the depth channel, and for example, 10% zeroing out in the normal vector is used.

[0066] In summary, according to various embodiments, the method is provided as shown in FIG. 3.

[0067] FIG. 3 shows a flowchart 300 illustrating a method for training a control policy (represented by a machine learning model, i.e., training the control policy includes training the machine learning model) for operating an object.

[0068] At 301, for each of one or more objects in each of one or more scenes, an input data element including image data representing the shape of the object to be operated and the position of the object within the scene is received (from one or more sensors and / or a sensor fusion device).

[0069] In 302, for each input data element, one or more training data elements are generated by generating an augmentation of the image data and a pseudo-label for the augmented image data according to a semi-supervised learning scheme.

[0070] In 303, a control policy is trained using the generated training data elements (and, optionally, other training data elements, in particular, other training data elements including received input data elements without augmentation, for example, observed rewards (thus, observed labels rather than pseudo-labels)).

[0071] Various embodiments can receive and use image data (i.e., digital images) from various visual sensors (cameras) such as video, radar, LiDAR, ultrasonic, thermal imaging, motion, sonar, etc. as a basis for obtaining input data (representing respective states). The image data may be a color image (e.g., an RGB image), or a black-and-white image, or a grayscale image, but it should be noted that the term "image data" herein also includes other "dense" data such as height maps, depth images, or normal surface maps (i.e., data having one or more values (one per channel) for each of the pixels in the array).

[0072] For example, the modalities RGB and depth are captured by a stereo sensor in a view looking down on the object bin from above. Each control action is defined, for example, as the final grasping pose and grasping position of a suction gripper, for example, with respect to the origin of the robot coordinates, for example, the base link of the robot arm, and corresponds to a three-dimensional orientation represented by Euler angles (α t , β t , γ t ) and Cartesian coordinates (x t , y t , z t ). The z is directly obtained from the height map tcan be extracted, and in the case of an axially symmetric suction gripper, γ t is unnecessary, and in this way, the gripping action can be defined as a t =(x t , y t , α t , β t ). The reward r t is 1 if the execution of the gripping is successful and is treated as 0 otherwise.

[0073] Using the approach of FIG. 3, a new search strategy for online learning in, for example, bin picking can be provided. For example, a policy network is used that maps from an RGB-D image to a per-pixel gripping map that predicts both the gripping quality (from 0 being least grippable to 1 being most grippable) and the gripping configuration (the orientation of the gripper) at all pixels. Using the approach of FIG. 3, a search strategy can be provided, for example, for better learning a new scene setting for a new object portfolio, camera setting, and bin setting, or for adapting to a new scene setting for a new object portfolio, camera setting, and bin setting.

[0074] A control policy can be used to control a robotic device (i.e., generate a control signal for the robotic device). The robotic device can be understood to relate to any technical system such as a computer-controlled machine, such as a robot, vehicle, household appliance, power tool, manufacturing machine, personal assistant, or access control system (having mechanical parts whose operations are controlled). According to various embodiments, a policy for controlling a technical system can be learned and then the technical system can be operated accordingly.

[0075] The method of FIG. 3 can be implemented by one or more data processing apparatuses (e.g., a computer or a microcontroller) having one or more data processing units. The term "data processing unit" can be understood to mean any kind of entity that enables the processing of data or signals. For example, data or signals can be processed according to at least one (i.e., one or more than one) specific function implemented by the data processing unit. The data processing unit can include, or can be composed of, analog circuits, digital circuits, logic circuits, microprocessors, microcontrollers, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), or any combination thereof. Any other means for implementing each function described in more detail herein can also be understood to include a data processing unit or a logic circuit device. One or more of the method steps described in more detail herein can be implemented (e.g., implemented) by the data processing unit via one or more specific functions implemented by the data processing unit.

[0076] Thus, according to one embodiment, the method is computer-implemented.

Claims

1. A method for training a control policy for manipulating an object (113), comprising: The method comprises: receiving (301) for each of one or more objects (113) in each of one or more scenes, an input data element (207) including image data representing a shape of the object (113) to be manipulated and a position of the object within the scene; generating (302) one or more training data elements by generating, for each input data element (207), an augmentation (208, 209) of the image data and a pseudo label (211) for the augmented image data (208, 209) according to a semi-supervised learning scheme; training (303) the control policy using the generated training data elements; Including, training the control policy includes determining a loss including a loss term for the generated training data elements; Each loss term is soft-weighted in a loss function by applying a softmax function to the confidence of the pseudo label (211) of the pose of each of the training data elements, and / or a loss term is filtered out (212) from the loss function if the confidence of the pseudo label (211) of the pose of each of the training data elements is below a predefined threshold.

2. The method includes training the control policy using reinforcement learning. The method of claim 1.

3. The method includes training the control policy using actor-critic reinforcement learning; training the control policy includes training actors (201) and critics (202) using the generated training data elements. The method of claim 2.

4. training the control policy includes training a neural network (200) representing the control policy.

4. The method according to claim 1 .

5. For each generated training data element, a pseudo label for each of a plurality of operating poses is generated; Each of the manipulation postures includes a manipulation position corresponding to a respective pixel in each of the expanded image data (208, 209).

5. The method according to any one of claims 1 to 4.

6. The threshold is a pixel-by-pixel threshold.

6. The method according to any one of claims 1 to 5.

7. A method for controlling a robotic device (101), comprising: The method comprises: Training a control policy according to any one of claims 1 to 6; receiving, for a scene in which the robotic device (101) is to be controlled, further image data representative of said scene; providing the acquired further image data to the control policy and generating control signals for the robotic device (101) according to outputs generated by the control policy in response to the acquired further image data; A method comprising:

8. A data processing apparatus (106) configured to perform the method according to any one of claims 1 to 7.

9. A computer program comprising instructions which, when executed by a computer, cause the computer to carry out the method according to any one of claims 1 to 7.

10. A computer readable medium comprising instructions which, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 7.