Operating a robot using data processing and training this data processing.
By augmenting image sub-regions to generate additional training information, the method addresses the labor-intensive nature of existing training methods, resulting in improved precision and reliability of machine learning-based data processing for robot pose determination.
Patent Information
- Application Number
- DE102024104640
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-02-20
- Publication Date
- 2025-05-15
- Estimated Expiration
- 2044-02-20
AI Technical Summary
The existing methods for training machine learning-based data processing for determining robot poses from image data are labor-intensive and limit the quality of the trained data processing, requiring a large amount of training effort or compromising on data quality.
The method involves generating additional training information by augmenting image sub-regions in the output images, allowing for the creation of more training data without increasing the recording effort, and improving the precision and reliability of the data processing.
This approach simplifies and accelerates the generation of training data, enhancing the precision, reliability, hit rate, flexibility, robustness, and speed of the machine learning-based data processing for determining robot poses.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The present invention relates to a method and a system for training a data processing system based at least partially on machine learning for determining robot poses on the basis of image data of a robot environment or for operating a robot using this (trained) data processing system, as well as a computer program or computer program product for carrying out a method described here.
[0002] A robot can advantageously be operated by using data processing based at least partially on machine learning to determine target robot poses on the basis of image data of an environment of the robot and controlling the robot's drives on the basis of these determined target robot poses.
[0003] To train the data processing system appropriately, a large amount of training data is regularly required, each of which contains training input data with associated robot poses and image data. If all this training data is collected using robots, especially hand-guided robots, and cameras, this requires a very large training effort and, conversely, limits the quality of the data processing system trained in this way.
[0004] DE 10 2023 206 009 B3 relates to a method for operating a robot, comprising the steps of: providing image data of an environment of the robot associated with robot poses, predicting or updating a robot target pose or determining new robot target poses by means of data processing based at least partially on machine learning, determining or updating a robot target movement or determining a new robot target movement on the basis of the target pose, and controlling drives of the robot to execute the robot target movement.
[0005] DE 10 2023 124 117 A1 describes systems and techniques relating to training one or more machine learning models for use in controlling a robot, wherein in at least one embodiment one or more machine learning models are trained at least based on simulations of the robot and renderings of such simulations, which may be performed using one or more ray tracing algorithms, operations or techniques.
[0006] According to various embodiments of DE 10 2022 201 719 A1, a method for training a machine learning model for generating descriptor images for images of one or more objects is described, comprising capturing a plurality of camera images, each camera image showing one or more objects, generating one or more augmented versions of the camera image for each camera image by applying, for each augmented version of the camera image, a respective augmentation to the camera image, the augmentation comprising a change in the position of pixel values of the camera image, and generating training image pairs, each comprising the camera image and an augmentation of the camera image or two augmented versions of the camera image, and training the machine learning model by means of contrast loss using the training image pairs, wherein descriptor values that the machine learning model generates for corresponding pixels,as positive pairs and descriptor values that the machine learning model generates for pixels that do not correspond to each other are used as negative pairs.
[0007] At least one embodiment of DE 11 2020 005 020 T5 relates to processing resources used to perform and enable artificial intelligence, for example, at least one embodiment relates to processors or computing systems used to train neural networks according to various novel methods described herein.
[0008] DE 10 2017 105 174 B4 relates to a method for generating training data for an artificial neural network, in which image data of a monitoring area of a safety application are recorded, wherein in the safety application a source of danger in the monitoring area is monitored by at least one safe sensor which, upon detection of a dangerous situation, triggers a safety-related safeguarding of the source of danger, wherein the image data are stored as training data together with a safety-related assessment, wherein the image data are assessed as safety-critical or not safety-critical, depending on whether the safe sensor triggers a safety-related safeguarding or not at the time the image data is recorded.
[0009] EP 4 238 721 A1 relates to a control device for a robot with a three-dimensional sensor, comprising a definer that defines a scanning area, which is an area measurable by the three-dimensional sensor, and an object-free area, which is an area in which the robot is allowed to move to measure the scanning area, and an operation controller that moves the three-dimensional sensor to measure the scanning area by controlling an operation of the robot to cause the robot to move within the object-free area.
[0010] An object of an embodiment of the present invention is therefore to improve the training of a data processing system based at least partially on machine learning for determining robot poses on the basis of image data of a robot environment or the operation of a robot.
[0011] This object is achieved by a method having the features of claims 1 and 7, respectively. Claims 9 and 10 protect a system, computer program, or computer program product for implementing a method described herein. The subclaims relate to advantageous developments.
[0012] According to one embodiment of the present invention, a method for training a data processing system based at least partially on machine learning (symbolized below with f or F for explanation) for this purpose or in such a way that this data processing system determines robot poses (symbolized with p for explanation) on the basis of image data (symbolized with b for explanation) of a robot environment, in particular input data (symbolized with e for explanation), which each comprise image data of an environment of a robot (e(b)), in particular can be this (e = b) or can comprise additional data (e = {b, y}), maps it to robot poses (p = f(e(b)), on the basis of which target poses p* for controlling the robot are (can be) determined or which are already such target poses, or which is set up or used for this purpose (p* = p*(p) = p*(f(e(b)), in one embodiment p* = F(e(b)), the step: - Providing one or preferably several initial training information items (for explanation with a i symbolized, preferably i = 1, 2,...,n), where: the or one initial training information (each): - a training robot pose (for explanation with p' i symbolized); and - Training input data (for explanation with e' i symbolized), the image data (for explanation with b' i symbolized) one of these training robot poses p' i assigned output image (for explanation with B' i symbolized) of a robot environment (a i = {e' i (b' i ), p' i}); and in or for the initial image or preferably in or for several initial training information a locked image sub-area (for explanation with G' jsymbolized) or released image section (for explanation with R' j symbolized) is identified, in particular determined and / or specified.
[0013] In one embodiment, in a further development, the or one or more of the initial training information is determined by manually guiding the robot, wherein the robot is guided to a robot pose or out of a robot pose by manually exerting forces on its structure and a corresponding, in particular compliant, control, which is (thereby) defined as a training robot pose, and in one or more poses when guiding into or out of the pose, images of an environment of the robot are taken and used as initial images (assigned to this training robot pose).
[0014] According to one embodiment of the present invention, the method comprises the step of adding one or more additional training information items (illustrated with z) to one or more of the initial training information items. i,j symbolized, preferably i = 1, 2,...,n and / or j = 1, 2,...,m), whereby additional training information is - the training robot pose p' i the initial training information a i at which this additional training information is generated; and - Training input data, which includes image data of an additional image of a robot environment assigned to this training robot pose (for explanation with e' i,j (b' i,j (B' i,j )) symbolized), whereby this additional image is created by augmenting at least a part of the released image sub-area or the unlocked image sub-area - in the initial image B' ithe initial training information a i , to which this additional training information z i,j is generated, or - in the additional image B' i,k≠j a further additional training information (already) generated to this initial training information z i,k is generated; has.
[0015] According to one embodiment of the present invention, the method comprises the step: - Training a data processing system based at least partially on machine learning for determining robot poses on the basis of image data of a robot environment;wherein the data processing is based on one or more of the additional training information, e.g. i,j , preferably with i = 1, 2,...n and particularly preferably with j = 1, 2,...m.
[0016] As a result, in one embodiment, (more or many) training data can be made available easily and / or quickly for training a data processing system based at least partially on machine learning for determining robot poses on the basis of image data of a robot environment, or the training or the trained data processing or the determination of robot poses on the basis of image data of a robot environment can be improved by means of this training data or additional training information, in particular a precision, reliability, hit rate, flexibility, robustness and / or speed can be increased.
[0017] In one embodiment, the robot has at least one robot arm. In one embodiment, the robot, preferably the or one or more of the robot arms, has at least three, preferably at least six, in one embodiment at least seven, joints or (movement) axes, preferably rotary joints or axes, which are adjusted by, or are configured to adjust by, preferably motorized, in one embodiment electric motorized, drives of the robot.
[0018] The data processing that is at least partially based on machine learning includes, in one embodiment, a regression method and / or at least one, preferably deep, artificial neural network (“(deep) artificial neural network”) for determining the robot pose(s) based on image data, in particular based on input data that can be these image data or, in addition to these image data, still further data, in particular kinematic data of the robot and / or predefined environmental data and / or sensorially detected environmental data of the robot environment.
[0019] In one embodiment, a robot pose includes a, preferably one-, two- or three-dimensional, position and / or a, preferably one-, two- or three-dimensional, orientation of an end effector of the robot.
[0020] In one embodiment, the data processing is also trained based on one or preferably more of the output training information.
[0021] As a result, in one embodiment, (even) more training data can be made available easily and / or quickly for training a data processing system based at least partially on machine learning for determining robot poses on the basis of image data of a robot environment, or the training or the trained data processing or the determination of robot poses on the basis of image data of a robot environment can be further improved by means of this training data or additional training information, in particular precision, reliability, hit rate, flexibility, robustness and / or speed can be further increased.
[0022] In one embodiment, the one or more of the blocked image sub-regions and / or the one or more of the released image sub-regions of images of a robot environment mentioned here, in particular of the one or more source images, are (each) determined based on a transformation of a spatial sub-region of this robot environment, preferably specified by a user or automatically. Blocked or released image sub-regions of additional images that are used to generate further additional images can, in particular, correspond (or be defined) to the underlying source images. In one embodiment, the method comprises the corresponding transformation.
[0023] If, in one embodiment, image data of an initial image is or are provided using at least one camera, preferably on the robot or in the environment, this transformation is or will be determined in a further development on the basis of a calibration of this camera(s) and transforms the predetermined spatial sub-area into the corresponding initial or additional image. In a particularly advantageous embodiment, boundary points of the predetermined spatial sub-area are transformed into corresponding image points in accordance with the calibration of the camera(s), and these are connected to form a boundary of the blocked or released image sub-area. In this case, a limited tolerance or enlargement of the blocked image sub-area orA reduction of the released sub-area compared to the exact transformation of the specified spatial sub-area can be achieved, for example, a more complex boundary contour of the transformation can be converted into a simpler boundary contour of the locked or released image sub-area. This advantageously reduces inaccuracies in the specification of the spatial sub-area.
[0024] In one version, the room section is or will be - based on user input or automatically; and / or - using a virtual model of the robot environment, in further training using a visual or graphic representation of the robot environment; and / or - based on a specification of one or more reference points, in a further development of a central or geometric center of gravity and / or one or more edge points, in particular corner points, of the spatial sub-area; and / or - based on a specification of the dimensions and / or shape of the room sub-area; and / or - based on at least one predetermined or recognized element of the, in particular, object in the, robot environment, in a further development of at least one predetermined or recognized obstacle and / or at least one predetermined or recognized object to be handled by the robot; and / or - based on a robot pose, in a further development of the training robot pose to which the image of the robot environment is assigned, in an embodiment such that the spatial sub-area has or contains the (position of the) training robot pose, for example is centered thereto;
[0025] In this way, blocked or released image sub-regions can be particularly advantageously identified, particularly in combination with two or more of these features. In one embodiment, an identified blocked or released image sub-region is identified, defined, predefined, or specified by a predefined edge, particularly predefined edge points.
[0026] In one embodiment, augmentation comprises shifting, rotating, and / or deleting image regions and / or changing the color and / or resolution of image regions and / or removing imaged objects, preferably identified by image recognition, and / or adding virtual objects. In one embodiment, this allows additional images particularly suitable for training to be generated particularly easily, quickly, and / or robustly, without the augmentation or the present invention being limited thereto.
[0027] As already indicated, input data for the data processing, on the basis of which the data processing determines robot poses, in particular training input data on the basis of which the data processing is trained, can comprise, in addition to the image data of a robot's environment, further data, in particular kinematic data of this robot and / or predetermined environmental data of this environment and / or sensor-detected environmental data of this environment, wherein the data processing for determining robot poses is trained on the basis of such training input data or one or more target robot pose(s) is determined on the basis of such input data with the aid of the trained data processing, and drives of the robot are controlled on the basis of these determined target robot pose(s). Control within the meaning of the present invention also includes, in particular, regulation.
[0028] This allows the operation of the robot to be further improved, in particular by determining suitable target poses more reliably.
[0029] In one embodiment, the kinematic data of a robot comprise a current, preferably one-, two- or three-dimensional, position and / or orientation of an end effector of the robot and / or current joint positions of the robot and / or time derivatives of this position, orientation or joint positions.
[0030] Predefined environmental data may include, in particular, geometries of the environment, in particular geometries and / or poses of workpieces to be approached or transported, obstacles to be bypassed or avoided, or the like.
[0031] Sensor-captured environmental data can include, in particular, poses of workpieces to be approached or transported and / or obstacles to be circumvented or avoided and / or joint and / or drive loads of the robot and / or audio data or the like. Accordingly, the sensors for (sensor-based) capturing the environmental data can include, in particular, force and / or torque sensors, distance sensors, radar sensors, microphones, and the like, whereby a combination of two or more sensors or a combination of environmental data captured (sensor-based) by different sensors can be particularly advantageous.
[0032] By additionally considering such kinematic and / or environmental data, the robot's operation can be further improved, in particular, more reliable determination of suitable target poses. For example, data processing can determine different, particularly suitable target poses for the same image of a robot environment with different robot joint positions and / or different known or sensor-detected obstacles and / or joint and / or drive loads of the robot.
[0033] According to one embodiment of the present invention, a method for operating a robot comprises the steps of: - Providing image data of an environment of the robot, wherein the image data is or will preferably be provided by means of at least one camera, in particular an environment-side or robot-side camera, - Determining one or more target robot pose(s), in a further development for gripping and / or setting down and / or processing an object with the robot or the like, on the basis of these image data with the aid of a data processing system that is at least partially based on machine learning and that is trained according to a method described here, in a further development before and / or during operation of the robot; and - Controlling the robot's drives on the basis of these determined target robot pose(s), in particular for approaching or assuming the target robot pose(s).
[0034] A particularly advantageous operation of a robot, wherein image data of a robot's environment are provided and robot target poses are predicted by means of data processing based at least partially on machine learning, a desired robot movement is determined on the basis of these target poses and thus on the basis of the image data with the aid of the data processing, and robot drives are controlled to execute the desired robot movement and thus on the basis of the determined desired robot pose(s), as well as a particularly advantageous training of the data processing, in which output images are collected during movements performed by a robot or demonstrator, is described in German patent application 10 2023 206 009.4, to which reference is made accordingly and the disclosure of which is incorporated into the present disclosure. In particular, one or more of the methods described in this German patent application 10 2023 206 009.4 can also be used or realized particularly advantageously in the present invention.
[0035] According to one embodiment of the present invention, a system for training a data processing system based at least partially on machine learning for determining robot poses on the basis of image data of a robot environment and / or for operating a robot comprises the data processing system or a data processing system based at least partially on machine learning for determining robot poses on the basis of image data of a robot environment.
[0036] According to one embodiment of the present invention, the system, in particular hardware and / or software, in particular program technology, is configured to carry out a method described here.
[0037] According to one embodiment of the present invention, the system comprises: - means for providing one or more initial training information, wherein an initial training information each - a training robot pose; and - training input data comprising image data of an output image of a robot environment associated with this training robot pose; and in at least one output image, a locked or unlocked image sub-area is identified; - Means for generating one or more additional training information items for one or more initial training information items, wherein one additional training information item - the training robot pose of the initial training information for which this additional training information is generated; and - training input data comprising image data of an additional image of a robot environment assigned to this training robot pose, wherein this additional image is created by augmenting at least a part of the released image sub-area or the unlocked image sub-area - in the initial image of the initial training information for which this additional training information is generated, or - the additional image of a further additional training information generated in addition to this initial training information; and - Means for training the data processing for determining robot poses on the basis of image data of a robot environment on the basis of one or more of the additional training information, in a further development also on one or more initial training information.
[0038] According to an embodiment of the present invention, the system additionally or alternatively comprises: - means for providing image data of an environment of the robot; - means for determining at least one target robot pose on the basis of these image data using data processing; and - Means for controlling drives of the robot on the basis of this determined at least one target robot pose.
[0039] In one embodiment, the system or its means comprises: - Means for determining at least one locked or released image sub-area of an image of a robot environment on the basis of a transformation of a predetermined spatial sub-area of this robot environment, in a further development means for specifying the spatial sub-area - based on user input or automatically; and / or - using a virtual model of the robot environment; and / or - based on a specification of at least one reference point and / or a dimension and / or shape of the spatial sub-area; and / or - based on at least one specified or recognized element of the robot environment; and / or - Means for shifting, rotating and / or deleting image areas and / or changing the color and / or resolution of image areas and / or removing depicted objects and / or adding virtual objects for augmentation to generate at least one additional image; and / or - Means for providing the image data of at least one output image and / or an environment of the robot using at least one camera, in particular this camera(s) and / or an image processing unit and / or a corresponding memory.
[0040] A means in the sense of the present invention can be formed in terms of hardware and / or software, in particular at least one, preferably a digital processing unit (CPU), graphics card (GPU) or the like, which is data- or signal-connected, in particular to a memory and / or bus system, and / or have one or more programs or program modules. The processing unit can be configured to process commands implemented as a program stored in a memory system, to capture input signals from a data bus and / or to output output signals to a data bus. A memory system can have one or more, in particular different, storage media, in particular optical, magnetic, solid-state and / or other non-volatile media. The program can be configured such that it embodies the methods described herein oris capable of executing, so that the processing unit can carry out the steps of such methods and thus in particular train the data processing or operate the robot. In one embodiment, a computer program product can have, in particular be, a storage medium, in particular a computer-readable and / or non-volatile one, for storing a program or instructions or with a program or instructions stored thereon. In one embodiment, execution of this program or these instructions by a system or a controller, in particular a computer or an arrangement of several computers, causes the system or the controller, in particular the computer(s), to carry out a method described here or one or more of its steps, or the program or the instructions are configured to do so.
[0041] In one embodiment, one or more, in particular all, steps of the method are fully or partially computer-implemented or one or more, in particular all, steps of the method are fully or partially automated, in particular by the system or its means.
[0042] In one embodiment, the system comprises the robot.
[0043] Further advantages and features emerge from the subclaims and the exemplary embodiments. The following shows, partially schematically: Fig. 1: a system according to an embodiment of the present invention; and Fig. 2: a method according to an embodiment of the present invention; and Fig. 3: Determining a blocked image sub-area of an image of a robot environment based on a transformation of a predetermined spatial sub-area of this robot environment in a method according to an embodiment of the present invention.
[0044] Fig. 1 shows a system according to an embodiment of the present invention, comprising a robot with a controller 1 and a robot arm 10 with an end effector 11, as well as at least one camera 20 guided by the robot and / or at least one environment-side camera 21. In modifications not shown, the camera 20 or 21 may be omitted and / or additional cameras guided by the robot and / or environment-side cameras may be provided and / or the robot arm may have a different configuration, for example, more or fewer than the six joints or (motion) axes shown.
[0045] An example is Fig. 1 in solid lines a robot start pose and dashed lines a robot target pose, in a dash-dotted line a movement of the robot based on a “LIN” movement command, which is a straight line of a Fig. 1 by coordinate systems indicated TCPs of the robot in Cartesian space, dash-double-dotted lines illustrate a movement of the robot based on a "CIRC" movement command, which causes a circular path of the TCP in Cartesian space, and double-dash-dotted lines illustrate a movement of the robot based on a "PTP" movement command.
[0046] Fig. 2 shows a method according to an embodiment of the present invention.
[0047] In order to train a data processing system based at least partially on machine learning for predicting robot target poses on the basis of image data, environmental data, for example known geometries of workpieces to be handled by the robot 10 and / or obstacles to be avoided or the like, are specified in a step S10.
[0048] Then, in step S20, a movement of the robot 10 or a demonstrator, for example, only the loose end effector 11, is performed from a starting pose to reach a target pose. Additionally or alternatively, movements from a target pose to reach a starting pose can also be performed.
[0049] The movements are compared with different movement trajectories (cf. Fig. 1) and / or under different environmental conditions.
[0050] For these movements, in step S20, image data of an environment 30 and the respective target pose for machine learning are collected for several poses of the robot or demonstrator assumed during the respective movement, each of which is associated with this pose.
[0051] The respective target pose and the respective image data each represent a training robot pose p' i and image data b' i an output image B' assigned to this training robot pose i an initial training information a i = {e' i (b' i ), p' i} represents.
[0052] In a preferred development, kinematic data which are assigned to these poses or image data in time, preferably via time stamps or the like, and which indicate the respective robot pose, and / or environmental data recorded by sensors, for example sensor data from force-torque sensors, joint torque sensors, radar sensors, audio microphones, tracking systems 22 for detecting obstacles 31 or the like, are collected, which together with the image data b' i and the environmental data specified in step S10, training input data e' i can form.
[0053] Based on a user input or specification, a spatial sub-area 100 is specified that is relevant for the data processing to be trained. Similarly, a spatial sub-area can also be specified that is not relevant for data processing.
[0054] For this purpose, the user specifies, for example, a center and an orientation and size or a corner point of a corresponding cuboid (cf. the spatial sub-area 100 in Fig. 1), within which an object to be grasped or processed, or a storage location for a robot-guided object, is located. Advantageously, the spatial sub-area can contain the (position of) the respective target pose(s) and, in a further development, can be centered for this purpose.
[0055] Based on a corresponding transformation of the given spatial sub-area, the output images B' iIn each case, an image sub-area blocked for augmentation is determined or identified, preferably by transforming a spatial sub-area specified as relevant for the data processing to be trained into the corresponding source image or by eliminating an image sub-area resulting from transforming a spatial sub-area specified as not relevant for the data processing to be trained into the corresponding source image. Similarly, again with the aid of a corresponding transformation of the specified spatial sub-area, in the source images B' iIn each case, an image sub-area released for augmentation is determined or identified, for example, a spatial sub-area specified as not relevant for data processing is mapped to an image sub-area released for augmentation using a corresponding transformation, or an image sub-area that results from the transformation of a spatial sub-area specified as relevant for the data processing to be trained into the corresponding original image is eliminated.
[0056] The identification of the corresponding image sub-area, in particular a corresponding boundary in an image of a training information, is stored for the respective initial training information.
[0057] If enough data has been collected (S25: “Y”), in a step S30 for one or more of the initial training information a i one or more additional training information zi,j determined.
[0058] For this purpose, one or more additional images B' i,j generated by at least a part of the released image sub-area or the unlocked image sub-area of the corresponding output image B' i is augmented, for example image areas of the released or unlocked image sub-area of the corresponding output image B' i shifted and / or rotated and / or image areas of the released or unlocked image sub-area of the corresponding output image B' i are deleted and / or colors and / or resolutions of image areas of the released or unlocked image sub-area are changed and / or objects depicted in the released or unlocked image sub-area, preferably recognized by means of image recognition, are removed and / or virtual objects are added or the like.
[0059] Additionally or alternatively, one or more additional images B' i,j generated by at least a part of the released image sub-area or the unlocked image sub-area of an already assigned to the corresponding output image B' i generated additional image B' i,k≠j is augmented.
[0060] The image data b' i,j which leads to an initial training information a i additional images B' generated by the augmentation of released or unlocked image areas as described above i,j form, if necessary together with the kinematic data and / or sensor-recorded environmental data and / or predefined environmental data, training input data e' i,j , which together with the corresponding training robot pose p"i each contain additional training information z i,j = {e' i,j (b' i,j ), p' i} form.
[0061] Indicatively speaking, in step S20 for at least one target pose p' i and several poses taken by the robot when approaching them each produce an image B' i recorded and corresponding image data of these images, if necessary together with kinematic data and / or sensory recorded environmental data and / or predefined environmental data, as input data together with the respective target pose as output training information a i = {e' i (b' i ), p' i} stored, whereby in each case on the basis of a transformation of a spatial sub-area 100 of the robot environment 30 specified as relevant for the data processing to be trained into the corresponding image B' i a part of the image in this image B' that is blocked for augmentation iis identified or based on a transformation of a spatial sub-area of the robot environment 30, which is not specified as relevant for the data processing to be trained, into the corresponding image B' i a part of the image released for augmentation in this image B' i is identified or a released / blocked image sub-area is determined or identified by eliminating an image area that results from the transformation of a spatial sub-area that is specified as relevant / not relevant for the data processing to be trained.
[0062] In step S30, additional training information is i,j = {e' i,j (b' i,j ),p' i} is generated by augmenting released image sub-regions (which are considered not relevant for data processing or its training) or by not augmenting blocked image sub-regions (which are considered relevant for data processing or its training).
[0063] Then, in step S30, the data processing is carried out based on this initial training information. i = {e' i (b' i ),p' i} and additional training information z i,j = {e' i,j (b' i,j ). If necessary, the generation of additional training information and / or the training can also take place at least partially in parallel with the collection of further data based on other movements performed.
[0064] The additional training information can be used to improve data processing and training, in particular, a large amount of additional image data can be made available or used, thereby improving data processing and training, without requiring any additional effort to capture the relevant images and without changing the image area relevant for data processing and training, thus preventing unpredictable influences on the assignment of corresponding robot poses.
[0065] To operate the robot 10, environmental data, for example known geometries of workpieces to be handled by the robot 10 or the like, are specified in a step S100; start kinematics data indicating the robot start pose are provided in a step S110; start image data of an environment 30 of the robot 10 associated with this robot start pose as well as sensor-detected environmental data, for example sensor data from force-torque sensors, joint torque sensors, radar sensors, audio microphones, tracking systems 22 for detecting obstacles 31 or the like (not shown), are provided in a step S120; in a step S130, a first robot target pose is predicted by means of the trained data processing based on the provided environmental data, start image data and start kinematics data; in a step S140, a robot target movement is determined based on this first target pose;and in a step S150, drives of the robot 10 are controlled to execute the robot target movement, of which three drives are provided with the reference numeral 12 as an example.
[0066] During this control in step S150, analogous to step S110, multiple current kinematic data are provided, analogous to step S140, the robot target movement is updated on the basis of the first target pose and this current kinematic data, and then the drives 12 of the robot 10 are controlled to execute this updated robot target movement instead of the previously executed robot target movement.
[0067] While the robot target movement continues to be executed, in a step S210, as previously in step S110, updated kinematic data indicating the current robot pose are provided, in a step S220, as previously in step S120, new image data of the environment of the robot 10 associated with this current robot pose as well as sensor-detected environmental data are provided, in a step S230, as previously in step S130, an updated robot target pose is predicted by means of the trained data processing on the basis of the provided environmental data, new image data and updated kinematic data, in a step S240 the robot target movement is updated on the basis of this updated target pose, and in a step S250, as previously in step S150, the drives 12 of the robot 10 are now controlled to execute this updated robot target movement.
[0068] Also during this control in step S250, analogous to step 150 or S210, current kinematic data are provided several times, analogous to step S240, the robot target movement is updated on the basis of the target pose updated in step S230 and this current kinematic data, and then the drives 12 of the robot 10 are controlled to execute this updated robot target movement.
[0069] As long as no termination condition is met (S255: “N”), for example the robot 10 has reached a workpiece or the like, the system or method returns to step S210, otherwise (S255: “Y”) the method is terminated (step S260).
[0070] In a modification, which can preferably be implemented in addition to or particularly preferably as an alternative to the above-explained aspect of predicting an updated robot target pose, updating the robot target movement on the basis of the updated target pose and controlling drives of the robot to execute the updated robot target movement, not only or not the final target or end pose is predicted, but additionally or alternatively (in each case) one or more new robot target pose(s) are predicted on the basis of the new image data respectively assigned to a current robot pose.
[0071] As already described above, in steps S10, S20, environmental data are specified and movement of the robot 10 or a demonstrator is carried out, and image data of the environment 30 and the respective target pose are collected for machine learning of the prediction. Here, the target pose is not the final target or end pose, but rather the target pose approached in a subsequent time step, or a sequence with several consecutive such target poses is collected. Here, too, kinematic data indicating the respective robot pose, which is assigned to these poses or image data over time, preferably via timestamps or the like, and / or environmental data acquired by sensors, for example, sensor data from force-torque sensors, joint torque sensors, radar sensors, audio microphones, tracking systems 22 for detecting obstacles 31, or the like, can be collected.
[0072] In step S30, the data processing unit, or possibly a further data processing unit, is trained based on these collected, associated image data and target poses, and possibly kinematic and / or environmental data. However, this training is not aimed at predicting the final target or end pose, but rather at predicting the target pose(s) to be approached in the next time step(s). In this case, additional training information is again generated, as described above, and used for training purposes together with the collected image data.
[0073] To operate the robot 10, environmental data, for example known geometries of workpieces to be handled by the robot 10 or the like, are specified in step S100; start kinematics data indicating the robot start pose are provided in a step S110; start image data of an environment 30 of the robot 10 associated with this robot start pose as well as sensor-detected environmental data, for example sensor data from force-torque sensors, joint torque sensors, radar sensors, audio microphones, tracking systems 22 for detecting obstacles 31 or the like (not shown) are provided in a step S120; in a step S130, a first robot target pose is predicted by means of the trained data processing based on the provided environmental data, start image data and start kinematics data; in a step S140, a robot target movement is determined based on this first target pose;and in a step S150, drives of the robot 10 are controlled to execute the robot target movement, of which three drives are provided with the reference numeral 12 as an example.
[0074] After or preferably already during this control in step S150, in a step S210, as previously in step S110, updated kinematic data indicating the current robot pose are provided; in a step S220, as previously in step S120, new image data of the environment of the robot 10 associated with this current robot pose as well as sensor-detected environmental data are provided; in a step S230, as previously in step S130, a new robot target pose to be approached in the next time step or several consecutive such robot target poses (to be approached in successive time steps) are predicted by means of the trained data processing on the basis of the provided environmental data, new image data and updated kinematic data; in a step S240, as previously in step S140, a new robot target movement is determined based on these new target pose(s); and in a step S250, as previously in step S150,the drives 12 of the robot 10 are controlled to execute this new robot target movement.,
[0075] Also during this control in step S250, current kinematics data are provided multiple times, analogous to step 150 or S210, a new robot target movement is determined analogously to step S240 on the basis of the new target pose determined in step S230 and this current kinematics data, and then the drives 12 of the robot 10 are controlled to execute this new robot target movement.
[0076] As long as no termination condition is met (S255: "N"), for example, the new target pose deviates only (sufficiently) slightly from the last determined target pose or the final or end pose is reached, or the like, the system or method returns to step 210; otherwise (S255: "Y"), the method is terminated (step S260). Advantageously, upon reaching the final or end pose, the artificial intelligence or data processing predicts the current pose as the new target pose, so that the robot stops automatically.
[0077] The training or operation explained above includes in particular features as described in German patent application 10 2023 206 009.4, for or with which the present invention can be used or combined particularly advantageously.
[0078] However, it is not limited to this. In a simple, Fig.In the example illustrated in Figure 3, for example for a grasping pose of a robot, several images B' i an environment of the robot, a user specifies a relevant spatial sub-area 100 of the environment for the gripping pose, such as a cuboid, a sphere or the like, for example a spatial sub-area in which an object 110 to be grasped is located. This spatial sub-area is converted into a correspondingly locked image sub-area G' based on the calibration of the recording camera 200 i Now released or unlocked image areas G' i = B' i \G' iThe images are augmented and, together with the respective grasping pose, form additional training data, on which a data processing system is trained. This data processing system can then determine target grasping poses based on images when the robot is operating, which the robot is then controlled to approach and grasp an object.
[0079] On the one hand, this allows for a large amount of additional training information to be obtained or made available for training the data processing system easily, quickly, and without great effort. On the other hand, by limiting augmentation to image regions that are authorized or not blocked for this purpose, augmentation of the image regions relevant for data processing near the grasping pose prevents the performance of the trained data processing system from being impaired.
[0080] In the present disclosure, "has an X" generally does not imply an exhaustive list, but is a shortened form of "has at least one X" and also includes "has two or more Xs" and "has Y in addition to X." Although exemplary embodiments have been explained in the preceding description, it should be noted that a variety of modifications are possible. Furthermore, it should be noted that the exemplary embodiments are merely examples that are not intended to limit the scope, applications, or construction in any way. Rather, the preceding description provides a guide to the person skilled in the art for implementing at least one exemplary embodiment, wherein various changes can be made, particularly with regard to the function and arrangement of the described components. List of reference symbols 1 control 10 robots 11 End effector 12 drive 20-22 Camera 30 surroundings 31 Obstacle 100 room sub-area 110 Item 200 camera G' i locked image area B' i (Source) image
Claims
[1] A method for training a data processing system based at least partially on machine learning for determining robot poses on the basis of image data of a robot environment, the method comprising the steps of: - Providing (S20) one or more initial training information, wherein an initial training information each - a training robot pose; and - training input data comprising image data of an output image of a robot environment associated with this training robot pose; and in at least one output image (B' j ) a blocked (G' j ) or released image sub-area is identified; - generating (S30) one or more additional training information items for one or more initial training information items, wherein one additional training information item - the training robot pose of the initial training information for which this additional training information is generated; and - training input data comprising image data of an additional image of a robot environment assigned to this training robot pose, wherein this additional image is created by augmenting at least a part of the released image sub-area or the unlocked image sub-area - in the initial image of the initial training information for which this additional training information is generated, or - in the additional image of a further additional training information generated in addition to this initial training information; and - Training (S30) a data processing system based at least partially on machine learning for determining robot poses based on image data of a robot environment; wherein the data processing system is trained based on one or more of the additional training information. [2] Method according to claim 1, characterized by that the data processing is also trained on the basis of one or more initial training information. [3] Method according to one of the preceding claims, characterized by that at least one locked or released image sub-area of an image of a robot environment is determined on the basis of a transformation of a predetermined spatial sub-area (100) of this robot environment. [4] Method according to claim 3, characterized bythat the spatial sub-area is or will be specified on the basis of a user input or automatically and / or with the aid of a virtual model of the robot environment and / or on the basis of a specification of at least one reference point and / or a dimension and / or shape of the spatial sub-area and / or on the basis of at least one specified or recognized element of the robot environment and / or on the basis of a robot pose. [5] Method according to one of the preceding claims, characterized by that augmentation comprises moving, rotating and / or deleting image areas and / or changing the color and / or resolution of image areas and / or removing depicted objects and / or adding virtual objects. [6] Method according to one of the preceding claims, characterized bythat training input data, in addition to the image data of an environment of a robot, also includes kinematic data of this robot and / or predetermined environmental data of this environment and / or sensor-recorded environmental data of this environment and the data processing for determining robot poses is trained on the basis of such input data. [7] A method for operating a robot (10), the method comprising the steps of: - Providing (S120, S220) image data of an environment of the robot; - determining (S130, S230) at least one target robot pose on the basis of these image data using data processing based at least partially on machine learning, which is trained according to a method according to one of the preceding claims; and - Controlling (S150, S250) drives (12) of the robot on the basis of this determined at least one target robot pose. [8] Method according to one of the preceding claims, characterized by that the image data of at least one output image and / or an environment of the robot are or will be provided by means of at least one camera (21, 22, 200). [9] System for training a data processing system based at least partially on machine learning and configured to determine robot poses on the basis of image data of a robot environment and / or for operating a robot (10), wherein the system comprises the data processing system, wherein the system is configured to carry out a method according to one of the preceding claims. [10] A computer program or computer program product, the computer program or computer program product containing instructions which, when executed by one or more computers or a system according to claim 9, cause the computer or computers or the system to carry out a method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method for generating training data for monitoring a hazard source
DE102017105174B4
Device and method for training a machine learning model to generate descriptor images of objects
DE102022201719A1
TRAINING OF MACHINE LEARNING MODELS USING SIMULATIONS FOR ROBOT SYSTEMS AND APPLICATIONS
DE102023124117A1
Method and system for training a data processing system, at least partially based on machine learning, to predict robot target poses and / or to operate a robot.
DE102023206009B3
Positioning using one or more neural networks
DE112020005020T5
Cited By
Humanoid machine control method and system based on model prediction
CN120715915A