Operating a robot using data processing and training said data processing
The method enhances robot pose determination by using locked and released image sub-areas for training data processing, improving precision and reliability through machine learning, addressing the inefficiencies of existing training methods.
Patent Information
- Application Number
- PCT/EP2024/085686
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-20
- Filing Date
- 2024-12-11
- Publication Date
- 2025-08-28
AI Technical Summary
Existing methods for training data processing systems to determine robot poses using machine learning require significant effort and can limit the quality of the data processing, especially when using hand-guided robots and cameras.
A method for training a data processing system that involves providing initial training information with locked and released image sub-areas, generating additional training information by augmenting these sub-areas, and using machine learning to improve the determination of robot poses based on image data, incorporating kinematic and environmental data for enhanced precision and reliability.
Enables efficient and accurate training of robot pose determination with increased precision, reliability, flexibility, and speed, reducing the need for extensive data collection and maintaining the quality of the training process.
Smart Images

Figure EP2024085686_28082025_PF_FP_ABST
Abstract
Description
[0001]2023P00022 WO 1 / 28 Kuka Deutschland GmbH Description Operating a robot using data processing and training this data processing The present invention relates to a method and a system for training a data processing system based at least partially on machine learning for determining robot poses based on image data of a robot environment or for operating a robot using this (trained) data processing system, as well as a computer program or computer program product for carrying out a method described here. A robot can advantageously be operated by determining target robot poses using data processing system based at least partially on machine learning based on image data of a robot environment, and controlling the robot's drives based on these determined target robot poses. In order to train the data processing system accordingly,A large amount of training data is regularly required, each of which contains training input data with associated robot poses and image data. If all this training data is collected using robots, especially hand-guided robots, and cameras, this requires a very large training effort or, conversely, limits the quality of the data processing trained in this way. One object of an embodiment of the present invention is therefore to improve the training of a data processing system based at least partially on machine learning for determining robot poses based on image data of a robot environment, or the operation of a robot. This object is achieved by a method having the features of claim 1 and 7, respectively. Claims 9,10 protect a system, computer program, or computer program product for carrying out a method described here. The subclaims relate to advantageous developments. According to one embodiment of the present invention, a method for training a data processing system based at least partially on machine learning (symbolized below with f or F for explanation) has the following characteristics: this data processing system determines robot poses (symbolized with p for explanation) based on image data (symbolized with b for explanation) of a robot environment, in particular input data (symbolized with e for explanation), which each comprise image data of an environment of a robot (e(b)), in particular can be this (e = b) or comprise additional data (e = {b, y}), maps it to robot poses (p = f(e(b)),on the basis of which target poses p* for controlling the robot are (can be) determined or which are already such target poses, or which is set up or used for this purpose (p* = p*(p) = p*(f(e(b)), in one embodiment p* = F(e(b)), the step of: - providing one piece of initial training information or preferably several pieces of initial training information (symbolized by ai for explanation, preferably i = 1, 2,...,n), where: the or one piece of initial training information (each): - a training robot pose (symbolized by p'i for explanation); and - training input data (symbolized by e'i for explanation), which comprises image data (symbolized by b'i for explanation) of an output image (symbolized by B'i for explanation) of a robot environment assigned to this training robot pose p'i (ai = {e'i(b'i),p'i}); and in which or for the initial image or preferably in or for several pieces of initial training information, a locked image sub-area (symbolized by G'j for explanation) or released image sub-area (symbolized by R'j for explanation) is identified, in particular determined and / or specified. 2023P00022 WO 3 / 28 Kuka Deutschland GmbH In one embodiment, in a further development, the or one or more of the pieces of initial training information are determined by manually guiding the robot, wherein the robot is guided to or from a robot pose by manually exerting forces on its structure and a corresponding, in particular compliant, control, which is (thereby) specified as the training robot pose,and in one or more poses, images of the robot's environment are captured during guidance into or out of the pose and used as initial images (assigned to this training robot pose). According to one embodiment of the present invention, the method comprises the step of generating one or more additional training information items (symbolized for explanation by zi,j, preferably i = 1, 2,...,n and / or j = 1, 2,...,m) for one or more of the initial training information items, wherein each additional training information item comprises - the training robot pose p'i of the initial training information ai for which this additional training information is generated; and - training input data comprising image data of an additional image of a robot environment assigned to this training robot pose (symbolized for explanation by e'i,j(b'i,j(B'i,j))).wherein this additional image is generated by augmenting at least a portion of the released image sub-area or the unlocked image sub-area - in the initial image B'i of the initial training information ai, for which this additional training information zi,j is generated, or - in the additional image B'i,k^j of a further additional training information zi,k (already) generated for this initial training information; According to one embodiment of the present invention, the method comprises the step: 2023P00022 WO 4 / 28 Kuka Deutschland GmbH - Training a data processing system based at least partially on machine learning for determining robot poses based on image data of a robot environment; wherein the data processing is based on one or more of the additional training information zi,j, preferably with i = 1, 2,...n and particularly preferably with j = 1, 2,...m,is trained. In one embodiment, this allows (more or many) training data to be made available easily and / or quickly for training a data processing system based at least partially on machine learning for determining robot poses based on image data of a robot environment, or this training data or additional training information can be used to improve the training or the trained data processing or the determination of robot poses based on image data of a robot environment, in particular to increase precision, reliability, hit rate, flexibility, robustness and / or speed. In one embodiment, the robot has at least one robot arm. In one embodiment, the robot, preferably the or one or more of the robot arms, has at least three, preferably at least six, in one embodiment at least seven, joints or (movement) axes, preferably rotary joints or axes, which are connected by,preferably motor-driven, in one embodiment electro-motor-driven, drives of the robot are adjusted or are configured to do so. The data processing, which is based at least partially on machine learning, comprises in one embodiment a regression method and / or at least one, preferably deep, artificial neural network (“(deep) artificial neural network”) for determining the robot pose(s) on the basis of image data, in particular on the basis of input data, which may be this image data or, in addition to this image data, may contain further data, in particular kinematic data of the robot and / or predetermined environmental data and / or sensor-detected environmental data of the robot environment. 2023P00022 WO 5 / 28 Kuka Deutschland GmbH In one embodiment, a robot pose comprises a, preferably one-, two-, or three-dimensional, position and / or a, preferably one-, two-, or three-dimensional,Orientation of an end effector of the robot. In one embodiment, the data processing is also trained on the basis of one or preferably several of the initial training information items. As a result, in one embodiment, (even) more training data can be made available easily and / or quickly for training a data processing system based at least partially on machine learning for determining robot poses based on image data of a robot environment, or the training or the trained data processing or the determination of robot poses based on image data of a robot environment can be further improved by this training data or additional training information, in particular a precision, reliability, hit rate, flexibility,Robustness and / or speed can be further increased. In one embodiment, the one or more of the blocked image sub-areas and / or the one or more of the released image sub-areas of images of a robot environment mentioned here, in particular of the one or more source images, are (each) determined based on a transformation of a spatial sub-area of this robot environment, preferably specified by a user or automatically. Blocked or released image sub-areas of additional images used to generate further additional images can, in particular, be defined correspondingly to the underlying source images. In one embodiment, the method comprises the corresponding transformation. If, in one embodiment, image data of an output image is or are provided using at least one camera, preferably on the robot or environment side,This transformation is or will be determined in a further development based on a calibration of this camera(s) and transforms the specified spatial sub-area into the corresponding output or additional image. In a particularly advantageous embodiment, boundary points of the specified 2023P00022 WO 6 / 28 Kuka Deutschland GmbH spatial sub-area are transformed into corresponding image points according to the calibration of the camera(s), and these are connected to form a boundary of the blocked or released image sub-area. In this case, a limited tolerance or enlargement of the blocked image sub-area or reduction of the released sub-area can advantageously be realized compared to the exact transformation of the specified spatial sub-area.For example, a more complex boundary contour of the transformation can be converted into a simpler boundary contour of the locked or unlocked image sub-area. This advantageously reduces inaccuracies in the specification of the spatial sub-area. In one embodiment, the spatial sub-area is or will be defined - based on user input or automatically; and / or - using a virtual model of the robot environment, in a further development using a visual or graphic representation of the robot environment; and / or - based on a specification of one or more reference points, in a further development a central or geometric center of gravity and / or one or more edge points, in particular corner points, of the spatial sub-area; and / or - based on a specification of a dimension and / or shape of the spatial sub-area; and / or - based on at least one specified or recognized element of the, in particular object in the, robot environment,In a further development, at least one predetermined or recognized obstacle and / or at least one predetermined or recognized object to be handled by the robot; and / or - based on a robot pose, in a further development of the training robot pose to which the image of the robot environment is assigned, in an embodiment such that the spatial sub-area has or contains the (position of the) training robot pose, for example, is centered thereon; predetermined. 2023P00022 WO 7 / 28 Kuka Deutschland GmbH In this way, blocked or released image sub-areas can be particularly advantageously identified, in particular in combination with two or more of these features. An identified blocked or released image sub-area is or is, in one embodiment, defined by a predetermined edge, in particular predetermined edge points.identified or defined or specified or fixed. In one embodiment, augmentation comprises shifting, rotating and / or deleting image areas and / or changing the color and / or resolution of image areas and / or removing imaged objects, preferably identified by image recognition, and / or adding virtual objects. In one embodiment, additional images particularly suitable for training can thereby be generated particularly easily, quickly and / or robustly, without the augmentation or the present invention being limited thereto. As already indicated, input data for the data processing, on the basis of which the data processing determines robot poses, in particular training input data, on the basis of which the data processing is trained, can include, in addition to the image data of a robot's environment, further data,in particular kinematic data of this robot and / or predetermined environmental data of this environment and / or sensor-detected environmental data of this environment, wherein the data processing for determining robot poses is trained on the basis of such training input data or one or more target robot pose(s) are determined on the basis of such input data with the aid of the trained data processing and drives of the robot are controlled on the basis of these determined target robot pose(s). Controlling in the sense of the present invention also includes, in particular, regulating. This allows the operation of the robot to be further improved, in particular more reliable and suitable target poses to be determined. The kinematic data of a robot, in one embodiment, comprise a current, preferably one-, two-, or three-dimensional,Position and / or orientation of a 2023P00022 WO 8 / 28 Kuka Deutschland GmbH end effector of the robot and / or current joint positions of the robot and / or time derivatives of this position, orientation, or joint positions. Specified environmental data can include, in particular, geometries of the environment, in particular geometries and / or poses of workpieces to be approached or transported, obstacles to be avoided or avoided, or the like. Environmental data acquired by sensors can include, in particular, poses of workpieces to be approached or transported and / or obstacles to be avoided or avoided and / or joint and / or drive loads of the robot and / or audio data or the like. Accordingly, the sensors for (sensory) acquisition of the environmental data can include, in particular, force and / or torque sensors, distance sensors, radar sensors, microphones, and the like.wherein a combination of two or more sensors or a combination of environmental data acquired (sensorily) by different sensors can be particularly advantageous. By additionally taking such kinematic and / or environmental data into account, the operation of the robot can be further improved, in particular, more suitable target poses can be determined more reliably. For example, the data processing can determine different, particularly suitable target poses for the same image of a robot environment with different robot joint positions and / or different known or sensor-detected obstacles and / or joint and / or drive loads of the robot. According to one embodiment of the present invention, a method for operating a robot comprises the steps of: - Providing image data of an environment of the robot, wherein the image data is preferably generated using at least one,in particular an environment-side or robot-side camera are or will be provided; - determining one or more target robot pose(s), in a further development for gripping and / or placing and / or processing an object 2023P00022 WO 9 / 28 Kuka Deutschland GmbH with the robot or the like, based on these image data with the aid of data processing that is at least partially based on machine learning, which is trained according to a method described here, and in a further development is (further) trained before and / or during operation of the robot; and - controlling the robot's drives based on these determined target robot pose(s), in particular for approaching or assuming the target robot pose(s). A particularly advantageous operation of a robot, wherein image data of an environment of the robot is provided and robot target poses are predicted by means of data processing that is at least partially based on machine learning,a robot target movement is determined based on these target poses and thus on the basis of the image data with the help of data processing and drives of the robot are controlled to execute the robot target movement and thus on the basis of the determined target robot pose(s), as well as a particularly advantageous training of the data processing, in which output images are collected during movements performed by a robot or demonstrator, is described in the German patent application 102023206009.4,to which reference is made accordingly and the disclosure of which is incorporated into the present disclosure. In particular, one or more of the features described in this German patent application 102023206009.4 can also be used or implemented particularly advantageously in the present invention. According to one embodiment of the present invention, a system for training a data processing system based at least partially on machine learning for determining robot poses on the basis of image data of a robot environment and / or for operating a robot comprises the data processing system based at least partially on machine learning for determining robot poses on the basis of image data of a robot environment. 2023P00022 WO 10 / 28 Kuka Deutschland GmbH According to one embodiment of the present invention, the system, in particular in terms of hardware and / or software, in particular programming,to carry out a method described here. According to one embodiment of the present invention, the system comprises: - means for providing one or more pieces of initial training information, wherein each piece of initial training information comprises - a training robot pose; and - training input data comprising image data of an initial image of a robot environment associated with this training robot pose; and in at least one initial image, a locked or unlocked image sub-area is identified; - means for generating, for each piece of initial training information, one or more pieces of additional training information, wherein each piece of additional training information comprises - the training robot pose of the initial training information for which this additional training information is generated; and - training input data,the image data of an additional image of a robot environment assigned to this training robot pose, wherein this additional image is generated by augmenting at least a part of the released image sub-area or the unlocked image sub-area - in the initial image of the initial training information for which this additional training information is generated, or - the additional image of another additional training information generated for this initial training information; and 2023P00022 WO 11 / 28 Kuka Deutschland GmbH - means for training the data processing for determining robot poses based on image data of a robot environment based on one or more of the additional training information,In a further development, also one or more pieces of initial training information. According to one embodiment of the present invention, the system additionally or alternatively comprises: - means for providing image data of the robot's environment; - means for determining at least one target robot pose based on this image data using data processing; and - means for controlling the robot's drives based on this determined at least one target robot pose. In one embodiment, the system or its means comprises: - means for determining at least one locked or enabled image sub-area of an image of a robot's environment based on a transformation of a predetermined spatial sub-area of this robot's environment,In a further development, means for specifying the spatial sub-area - based on a user input or automatically; and / or - using a virtual model of the robot environment; and / or - based on a specification of at least one reference point and / or a dimension and / or shape of the spatial sub-area; and / or - based on at least one specified or recognized element of the robot environment; and / or - means for shifting, rotating, and / or deleting image areas and / or changing the color and / or resolution of image areas and / or removing imaged objects and / or adding virtual objects for augmentation to generate at least one additional image; and / or - means for providing the image data of at least one original image and / or an environment of the robot using at least one camera,2023P00022 WO 12 / 28 Kuka Deutschland GmbH, in particular these cameras and / or an image processing unit and / or a corresponding memory. A means within the meaning of the present invention can be designed in hardware and / or software, in particular at least one, preferably data- or signal-connected, particularly digital, processing unit, in particular a microprocessor unit (CPU), graphics card (GPU) or the like, and / or one or more programs or program modules. The processing unit can be designed to execute commands implemented as a program stored in a memory system, to acquire input signals from a data bus, and / or to output signals to a data bus. A memory system can comprise one or more, in particular different, storage media, in particular optical, magnetic,Solid-state and / or other non-volatile media. The program can be such that it embodies or is capable of executing the methods described here, so that the processing unit can execute the steps of such methods and thus, in particular, train the data processing or operate the robot. In one embodiment, a computer program product can have, in particular be, a storage medium, in particular a computer-readable and / or non-volatile one, for storing a program or instructions or with a program or instructions stored thereon. In one embodiment, execution of this program or these instructions by a system or a controller, in particular a computer or an arrangement of several computers, causes the system or the controller, in particular the computer(s), toto carry out a method described here or one or more of its steps, or the program or the instructions are set up for this purpose. In one embodiment, one or more, in particular all, steps of the method are fully or partially computer-implemented, or one or more, in particular all, steps of the method are fully or partially automated, in particular by the system or its means. In one embodiment, the system comprises the robot. Further advantages and features emerge from the subclaims and the exemplary embodiments. In this regard,Partially schematic: Fig. 1: a system according to an embodiment of the present invention; and Fig. 2: a method according to an embodiment of the present invention; and Fig. 3: determining a blocked image sub-area of an image of a robot environment based on a transformation of a predetermined spatial sub-area of this robot environment in a method according to an embodiment of the present invention. Fig. 1 shows a system according to an embodiment of the present invention, which comprises a robot with a controller 1 and a robot arm 10 with an end effector 11, as well as at least one camera 20 guided by the robot and / or at least one environment-side camera 21. In modifications not shown, the camera 20 or 21 can be omitted and / or additional cameras guided by the robot and / or environment-side cameras can be provided and / or the robot arm can have a different configuration.For example, more or fewer than the six joints or (motion) axes shown. For example, in Fig. 1, solid lines represent a robot start pose and dashed lines represent a robot target pose. A dash-dotted line represents a robot movement based on a "LIN" movement command, which causes a straight line of a robot TCP indicated in Fig. 1 by coordinate systems in Cartesian space. A dash-double-dotted line represents a robot movement based on a "CIRC" movement command, which causes a circular path of the TCP in Cartesian space.and a double-dash-dotted line illustrates a movement of the robot based on a "PTP" movement command. 2023P00022 WO 14 / 28 Kuka Deutschland GmbH Fig. 2 shows a method according to an embodiment of the present invention. To train a data processing system based at least partially on machine learning for predicting robot target poses based on image data, environmental data, for example known geometries of workpieces to be handled by the robot 10 and / or obstacles to be avoided or the like, are specified in a step S10. Then, in step S20, a movement of the robot 10 or a demonstrator, for example only the loose end effector 11, is determined.from a starting pose to reach a target pose. Additionally or alternatively, movements from a target pose to reach a starting pose can also be performed. The movements are performed with different movement trajectories (cf. Fig. 1) and / or under different environmental conditions. For these movements, in step S20, image data of an environment 30 associated with each pose of the robot or demonstrator during the respective movement and the respective target pose are collected for machine learning. The respective target pose and the respective image data each represent a training robot pose p'i and image data b'i of an output image B'i of an output training information ai = {e'i(b'i), p'i} associated with this training robot pose. In a preferred development, these poses or image data are assigned temporally, preferably via timestamps or the like.Each associated kinematic data indicating the respective robot pose and / or sensor-detected environmental data, for example, sensor data from force-torque sensors, joint torque sensors, radar sensors, audio microphones, tracking systems 22 for detecting obstacles 31, or the like, are collected, which, together with the image data b'i and the environmental data specified in step S10, can form training input data e'i. 2023P00022 WO 15 / 28 Kuka Deutschland GmbH Based on a user input or specification, a spatial sub-area 100 is specified that is relevant for the data processing to be trained. Similarly, a spatial sub-area that is not relevant for the data processing can also be specified. For this purpose, the user specifies, for example, a center, an orientation, and a size or a corner point of a corresponding cuboid (cf. the spatial sub-area 100 in Fig. 1).within which an object to be grasped or processed or a storage location for a robot-guided object is located. Advantageously, the spatial sub-area can have the (position of the) respective target pose(s), and in a further development, can be centered for this purpose. Based on a corresponding transformation of the given spatial sub-area, an image sub-area blocked for augmentation is determined or identified in the output images B'i, preferably by transforming a spatial sub-area specified as relevant for the data processing to be trained into the corresponding output image or by eliminating an image sub-area resulting from the transformation of a spatial sub-area specified as irrelevant for the data processing to be trained into the corresponding output image. Likewise, again with the aid of a corresponding transformation of the given spatial sub-area,In the output images B'i, an image sub-area released for augmentation is determined or identified, for example, a spatial sub-area specified as not relevant for data processing is mapped to an image sub-area released for augmentation using a corresponding transformation, or an image sub-area resulting from the transformation of a spatial sub-area specified as relevant for the data processing to be trained into the corresponding output image is eliminated. The identification of the corresponding image sub-area, in particular a corresponding boundary in an image of a piece of training information, is stored for the respective piece of output training information. 2023P00022 WO 16 / 28 Kuka Deutschland GmbH If sufficient data has been collected (S25: "Y"), in a step S30, one or more additional training information items zi are created for one or more of the pieces of output training information ai.j is determined. For this purpose, one or more additional images B'i,j are generated by augmenting at least a part of the released image sub-area or the unlocked image sub-area of the corresponding source image B'i, for example, image areas of the released or unlocked image sub-area of the corresponding source image, B‘i vshifted and / or rotated and / or image areas of the released or unlocked image sub-area of the corresponding source image B'i are deleted and / or colors and / or resolutions of image areas of the released or unlocked image sub-area are changed and / or objects depicted in the released or unlocked image sub-area, preferably recognized by image recognition, are removed and / or virtual objects are added or the like. Additionally or alternatively, one or more additional images B'i,j are generated by augmenting at least a portion of the released image sub-area or the unlocked image sub-area of an additional image B'i,k^j already generated for the corresponding source image B'i. The image data b'i,j of the additional images B'i,j generated for an initial training information ai by the augmentation of released or unlocked image sub-areas as described above, optionally together with the kinematic data and / or sensor-recorded environmental data and / or predetermined environmental data, form training input data e'i,j , which together with the corresponding training robot pose p'i form additional training information zi,j = {e'i,j(b'i,j), p'i}. Illustratively speaking, in step S20, an image B'i is recorded for at least one target pose p'i and several poses taken by the robot when approaching them, and corresponding image data of these images, optionally together with kinematic data and / or sensor-detected environmental data and / or predetermined environmental data, are used as input data together with the respective target pose as output training information ai = {e'i(b'i),p'i} is stored, wherein, based on a transformation of a spatial sub-area 100 of the robot environment 30 specified as relevant for the data processing to be trained into the corresponding image B'i, an image sub-area blocked for augmentation is identified in this image B'i, or, based on a transformation of a spatial sub-area of the robot environment 30 specified as not relevant for the data processing to be trained into the corresponding image B'i, an image sub-area released for augmentation is identified in this image B'i, or, a released / blocked image sub-area is determined or identified by eliminating an image area that results from the transformation of a spatial sub-area specified as relevant / not relevant for the data processing to be trained. In step S30, additional training information zi,j = {e'i,j(b'i,j), p'i} is generated,by augmenting released image sub-regions (which are considered irrelevant for data processing or its training) or by not augmenting blocked image sub-regions (which are considered relevant for data processing or its training). Then, in step S30, the data processing is trained based on this initial training information ai = {e'i(b'i), p'i} and additional training information z1,j = {e'i,j(b'i,j). If necessary, the generation of additional training information and / or the training can also take place at least partially in parallel with the collection of further data based on further movements performed. The additional training information can improve the data processing or its training. In particular, a large amount of additional image data can be made available or used, thereby improving the data processing or its training.wherein no additional effort is required to capture corresponding images, and the image area relevant for data processing or its training is not changed, thus preventing the assignment of corresponding robot poses from being unpredictably influenced. To operate the robot 10, environmental data, for example known geometries of workpieces or the like to be handled by the robot 10, are specified in a step S100; starting kinematics data indicating the robot starting pose are provided in a step S110; starting image data of an environment 30 of the robot 10 associated with this robot starting pose, as well as sensor-detected environmental data, for example sensor data from force-torque sensors, joint torque sensors, radar sensors, audio microphones, tracking systems 22 for detecting obstacles 31, or the like (not shown), are provided in a step S120.In a step S130, a first robot target pose is predicted using the trained data processing based on the provided environmental data, start image data, and start kinematics data; in a step S140, a robot target movement is determined based on this first target pose; and in a step S150, drives of the robot 10 are controlled to execute the robot target movement, of which three drives are provided with the reference numeral 12 as an example. During this control in step S150, current kinematics data are provided multiple times, analogously to step S110; analogously to step S140, the robot target movement is updated based on the first target pose and this current kinematics data, and then the drives 12 of the robot 10 are controlled to execute this updated robot target movement instead of the previously executed robot target movement. While the robot target movement continues to be executed, in a step S210, as previously in step S110,updated kinematic data indicating the current robot pose are provided; in a step S220, as previously in step S120, new image data of the environment of the robot 10 associated with this current robot pose, as well as sensor-detected environmental data, are provided; in a step S230, as previously in step S130, an updated robot target pose is predicted using the trained data processing based on the provided environmental data, new image data, and updated kinematic data; in a step S240, the robot target movement is updated based on this updated target pose; 2023P00022 WO 19 / 28 Kuka Deutschland GmbH; and in a step S250, as previously in step S150, the drives 12 of the robot 10 are now controlled to execute this updated robot target movement. Also during this control in step S250, current kinematic data is provided multiple times, analogous to step S150 or S210.Analogous to step S240, the robot target movement is updated based on the target pose updated in step S230 and this current kinematic data, and then the drives 12 of the robot 10 are controlled to execute this updated robot target movement. As long as no termination condition is met (S255: "N"), for example, the robot 10 has reached a workpiece or the like, the system or method returns to step S210; otherwise (S255: "Y"), the method is terminated (step S260). In a modification, which can preferably be implemented in addition to or particularly preferably alternatively to the above-explained aspect of predicting an updated robot target pose, updating the robot target movement based on the updated target pose, and controlling the robot's drives to execute the updated robot target movement, not only or not only the final target or end pose is predicted,but additionally or alternatively (in each case) one or more new robot target pose(s) based on the new image data assigned to a current robot pose. As already described above, in steps S10, S20, environmental data are specified and movement of the robot 10 or a demonstrator is carried out, and image data of the environment 30 and the respective target pose are collected for machine learning of the prediction, wherein the target pose is not the final target or end pose but the target pose approached in a subsequent time step, or a sequence with several consecutive such target poses is collected. Here, too, kinematic data assigned to these poses or image data, which indicate the respective robot pose, and / or sensor-detected environmental data, for example sensor data from force-torque sensors, joint torque sensors, radar sensors,Audio microphones, 2023P00022 WO 20 / 28 Kuka Deutschland GmbH tracking systems 22 for detecting obstacles 31 or the like. In step S30, the or, if applicable, further data processing is trained based on these collected, mutually associated image data and target poses, and, if applicable, kinematic and / or environmental data, however, not to predict the final target or end pose, but rather the target pose(s) to be approached in the next time step(s). In this case, as described above, additional training information is again generated and used together with the collected image data for training. To operate the robot 10, environmental data, for example, known geometries of workpieces or the like to be handled by the robot 10, are specified in step S100, and start kinematic data, which specify the robot's start pose, are provided in a step S110.In a step S120, start image data of an environment 30 of the robot 10 assigned to this robot start pose as well as environmental data detected by sensors, for example sensor data from force-torque sensors, joint torque sensors, radar sensors, audio microphones, tracking systems 22 for detecting obstacles 31 or the like (not shown), are provided, In a step S130, a first robot target pose is predicted by means of the trained data processing on the basis of the provided environmental data, start image data and start kinematics data, In a step S140, a robot target movement is determined on the basis of this first target pose, and In a step S150, drives of the robot 10 are controlled to execute the robot target movement, of which three drives are provided with the reference number 12 as an example. After or preferably already during this control in step S150, in a step S210, as previously in step S110, updated kinematic data,which indicate the current robot pose, in a step S220, as previously in step S120, new image data of the environment of the robot 10 associated with this current robot pose as well as sensor-detected environmental data are provided, in a step S230, as previously in step S130, a new robot target pose to be approached in the next time step or several consecutive such robot target poses (to be approached in successive time steps) are predicted by means of the 2023P00022 WO 21 / 28 Kuka Deutschland GmbH trained data processing on the basis of the provided environmental data, new image data and updated kinematics data, in a step S240, as previously in step S140, a new robot target movement is determined on the basis of these new target pose(s), and in a step S250, as previously in step S150,The drives 12 of the robot 10 are controlled to execute this new robot target movement. Also during this control in step S250, current kinematics data are provided multiple times, analogous to step S250 or S210, a new robot target movement is determined analogously to step S240 based on the new target pose determined in step S230 and this current kinematics data, and then the drives 12 of the robot 10 are controlled to execute this new robot target movement. As long as no termination condition is met (S255: "N"), for example, the new target pose deviates only (sufficiently) slightly from the last determined target pose or the final or end pose has been reached or the like, the system or method returns to step S210; otherwise (S255: "Y"), the method is terminated (step S260). Advantageously, when the final or end pose is reached, the artificial intelligence or data processing predicts the current pose as the new target pose,so that the robot stops by itself. The training and operation explained above particularly includes features as described in German patent application 102023206009.4, for or with which the present invention can be used or combined particularly advantageously. However, it is not limited thereto. In a simple example illustrated with reference to Fig. 3, for example for a grasping pose of a robot, several images B'i of an environment of the robot are recorded. A user specifies a relevant spatial sub-area 100 of the environment for the grasping pose, such as a cuboid, a sphere, or the like, for example a spatial sub-area,in which an object 110 to be grasped is located. This spatial sub-area is transformed into a correspondingly locked image sub-area G'i based on the calibration of the recording camera 200. Now, enabled or unlocked image areas G'i = B'i\G'i of the images are augmented and, together with 2023P00022 WO 22 / 28 Kuka Deutschland GmbH of the respective grasping pose, form additional or supplementary training data, on the basis of which a data processing system is trained. This can then, when operating the robot, determine target gripping poses based on images, on the basis of which the robot is then controlled to approach these and grasp an object. In this way, simple,A great deal of additional training information can be obtained or made available quickly and without great effort for training the data processing. On the other hand, by limiting the augmentation to image subregions that are enabled or unlocked for this purpose, it is advantageously prevented that an augmentation of the image subregions relevant for data processing near the grasping pose impairs the performance of the trained data processing. In the present disclosure, "has an X" generally does not imply an exhaustive list, but is a shortened form of "has at least one X" and also includes "has two or more Xs" and "has Y in addition to X." Although exemplary embodiments have been explained in the preceding description, it should be noted that a multitude of modifications are possible. Furthermore, it should be noted that the exemplary embodiments are merely examples.which are not intended to limit the scope of protection, the applications, and the structure in any way. Rather, the preceding description provides the person skilled in the art with a guide for implementing at least one exemplary embodiment, whereby various changes, in particular with regard to the function and arrangement of the described components, can be made without departing from the scope of protection as it results from the claims and their equivalent feature combinations. 2023P00022 WO 23 / 28 Kuka Deutschland GmbH List of reference symbols 1 Controller 10 Robot 11 End effector 12 Drive 20-22 Camera 30 Environment 31 Obstacle 100 Spatial sub-area 110 Object 200 Camera G'i Blocked image area B'i (Input) image,
Claims
2023P00022 WO 24 / 28 Kuka Deutschland GmbH Patent claims 1. Method for training a data processing system based at least partially on machine learning for determining robot poses on the basis of image data of a robot environment, the method comprising the steps of: - providing one or more pieces of initial training information, wherein each piece of initial training information comprises - a training robot pose; and - training input data comprising image data of an initial image of a robot environment assigned to this training robot pose; and a locked or released image sub-area is identified in at least one initial image; - generating one or more additional training information items for each piece of initial training information, wherein each piece of additional training information comprises - the training robot pose of the initial training information for which this additional training information is generated;and - training input data comprising image data of an additional image of a robot environment assigned to this training robot pose, wherein this additional image is generated by augmenting at least a part of the released image sub-area or the unlocked image sub-area - in the output image of the initial training information for which this additional training information is generated, or - in the additional image of a further additional training information generated for this initial training information; and; 2023P00022 WO 25 / 28 Kuka Deutschland GmbH - Training a data processing system based at least partially on machine learning for determining robot poses based on image data of a robot environment; wherein the data processing system is trained based on one or more of the additional training information items.
2. The method according to claim 1, characterized in that the data processing system is also trained based on one or more initial training information items.
3. The method according to one of the preceding claims, characterized in that at least one blocked or released image sub-region of an image of a robot environment is determined based on a transformation of a predetermined spatial sub-region of this robot environment. 4.Method according to claim 3, characterized in that the spatial sub-area is or will be specified based on a user input or automatically and / or with the aid of a virtual model of the robot environment and / or based on a specification of at least one reference point and / or a dimension and / or shape of the spatial sub-area and / or based on at least one specified or recognized element of the robot environment and / or based on a robot pose.
5. Method according to one of the preceding claims, characterized in that augmentation comprises shifting, rotating and / or deleting image regions and / or changing the color and / or resolution of image regions and / or removing depicted objects and / or adding virtual objects. 6.Method according to one of the preceding claims, characterized in that training input data, in addition to the image data of an environment of a robot, also comprises kinematic data of this robot and / or predetermined environmental data of this environment and / or sensor-detected environmental data of this environment and the data processing for determining robot poses is trained on the basis of such input data. 2023P00022 WO 26 / 28 Kuka Deutschland GmbH 7. A method for operating a robot, the method comprising the steps of: - providing image data of an environment of the robot; - determining at least one target robot pose on the basis of this image data with the aid of data processing based at least partially on machine learning, which is trained according to a method according to one of the preceding claims;and - controlling drives of the robot on the basis of this determined at least one target robot pose.
8. Method according to one of the preceding claims, characterized in that the image data of at least one initial image and / or an environment of the robot are or will be provided with the aid of at least one camera.
9. System for training a data processing system based at least partially on machine learning and set up to determine robot poses on the basis of image data of a robot environment and / or for operating a robot, wherein the system comprises the data processing system, wherein the system is set up to carry out a method according to one of the preceding claims and / or comprises: - means for providing one or more items of initial training information, wherein an item of initial training information each comprises - a training robot pose;and - training input data comprising image data of an output image of a robot environment assigned to this training robot pose; and in at least one output image, a locked or unlocked image sub-area is identified; - means for generating one or more additional training information items for each of one or more items of output training information, wherein each item of additional training information comprises the training robot pose of the output training information for which this additional training information is generated; and; 2023P00022 WO 27 / 28 Kuka Deutschland GmbH - training input data which comprises image data of an additional image of a robot environment assigned to this training robot pose, wherein this additional image is generated by augmenting at least a part of the released image sub-area or the unlocked image sub-area - in the output image of the initial training information for which this additional training information is generated, or - in the additional image of further additional training information generated for this initial training information; and - means for training the data processing to determine robot poses on the basis of image data of a robot environment on the basis of one or more of the additional training information; and / or comprises: - means for providing image data of an environment of the robot; - means for determining at least one target robot pose on the basis of this image data with the aid of the data processing;and - means for controlling the robot's drives based on this determined at least one desired robot pose.
10. A computer program or computer program product, wherein the computer program or computer program product contains instructions, in particular stored on a computer-readable and / or non-volatile storage medium, which, when executed by one or more computers or a system according to claim 9, cause the computer(s) or the system to perform a method according to one of claims 1 to 8.
Citation Information
Patent Citations
Device and method for training a neural network to control a robot for a placement task
DE102021109333A1
Device and method for controlling a robot to perform a task
DE102022202144A1
DE102023206009A1