Method for determining poses of objects for a robot

The method uses sensor data and an externally specified selection algorithm to determine optimal object poses for robots, addressing the challenge of selecting easy-to-grasp objects and enhancing grasping efficiency by adapting to specific conditions and needs.

WO2025242856A1PCT designated stage Publication Date: 2025-11-27ROBERT BOSCH GMBH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/064252
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-23
Filing Date
2025-05-23
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Determining optimal poses for objects in a work environment to facilitate efficient grasping and movement by robots is challenging, especially when dealing with large numbers of objects, as selecting easy-to-grasp objects can be difficult due to varying environmental conditions and object locations.

Method used

A method utilizing environmental information from sensors like cameras and lidars to determine object poses, combined with an externally specified selection algorithm that applies multiple selection filters to adapt robot operations to specific needs, allowing for flexible configuration and dynamic filtering of object poses.

Benefits of technology

Enables efficient and adaptable robot operation by selecting optimal object poses, reducing the risk of collisions and improving grasping efficiency through customizable filtering strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025064252_27112025_PF_FP_ABST
    Figure EP2025064252_27112025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for determining one or more poses of one or more objects in a working environment, said method being intended for use in the process of determining a sequence of movements for a robot in order to grip and / or move the one or more objects by means of an end effector of the robot and having the steps of: providing (314) environment information (312) acquired from the working environment by means of at least one sensor; determining (320), on the basis of the environment information (312), a plurality of poses (322) for one or more objects in the working environment in order to obtain an initial object-pose data set (324); determining (330), from the initial object pose data set (324), one or more selected poses (332) for the one or at least one of the plurality of objects using an externally specified selection algorithm (302) in order to obtain a selected-object-pose data set (334); and providing (340) the selected-object-pose data set (334) in order to determine the movement sequence for the robot.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method for determining the poses of objects for a robot

[0002] Description

[0003] The present invention relates to a method for determining one or more poses of one or more objects in a working environment, for use in determining a movement sequence for a robot in order to grasp and / or move one or more of the objects by means of an end effector of the robot, a computing unit and a computer program for carrying it out, and a robot.

[0004] Background of the invention

[0005] Robots, often also referred to as kinematics, can be used in production facilities, logistics, and other systems to, for example, transport or assemble parts. Typical types of such robots include Cartesian robots, SCARA robots, and articulated robots.

[0006] Disclosure of the invention

[0007] According to the invention, a method for determining one or more poses, a computing unit and a computer program for its execution, as well as a robot with the features of the independent claims, are proposed. Advantageous embodiments are the subject of the dependent claims and the following description.

[0008] The invention relates generally to robots that can be used, for example, in production plants, logistics, or other facilities. Types of such robots, often also referred to as kinematics, include, for example, Cartesian robots, SCARA robots, and articulated robots. Such robots can be used to grasp and / or move objects (i.e., items or parts). A typical application is to remove an object from a box containing, for example, a large number of (identical or different) objects and, for example, place it somewhere or possibly assemble it. For this purpose, such a robot has, for example, a gripper or, more generally, an end effector. Such robots can also be referred to as manipulators or manipulation devices.

[0009] In order for the robot to grasp and / or move an object in a work environment, e.g. in a box, using the end effector, the robot, in particular its end effector, must perform a movement sequence to move the end effector to the position of the object and then, e.g. after the object has been grasped, to move the end effector away from the position of the object again, e.g. to a desired storage position.

[0010] However, especially when there are a large number of objects, for example in a box or elsewhere, it can be difficult to select an object that is, for example, as easy to grasp as possible.

[0011] This process begins by providing environmental information captured from the work environment using at least one sensor. This includes, in particular, a camera and / or a lidar or other depth sensor. The environmental information then specifically comprises 2D and / or 3D environmental information. While the 2D environmental information can be images, such as color, grayscale, or black-and-white images, the 3D environmental information can be point clouds or depth maps, and possibly also 3D images. The sensors can be mounted on the robot or be part of the robot; alternatively, they can be provided and used separately to capture environmental information about the work environment, such as a box containing objects.

[0012] Based on the environmental information, a pose is then determined for one or more of several objects in the work environment. In typical use cases, multiple object poses are available in a so-called initial object pose dataset; this can be a list of recognizable object poses. Based on the initial object pose dataset, one or more selected poses are then determined for one or more of the objects to obtain a selection object pose dataset, which is then provided. In particular, a movement sequence can then be determined based on this, and the robot can be instructed to execute the movement sequence. Specifically, for example, control information for moving the robot can be determined based on the movement sequence and then provided, or used to move the robot.

[0013] The selection (or filtering) of object poses is performed using an externally defined selection algorithm. This allows for specific adaptation of the robot's operation to particular cases and applications. It has been shown that the requirements for selecting or filtering object poses can vary depending on the application; this can depend, for example, on the type of objects typically present or the types of boxes used.

[0014] For example, objects located at the edge of a box can be filtered out, as these are difficult to grasp. However, the type of filtering required may vary depending on the type of box used. Objects might also be located not in a box, but on a carrier or pallet without a rim. In this case, the edge issue is less of a concern, but objects stacked in layers should ideally be picked from the top (or very top) of the stack. It's also conceivable that objects in certain locations or areas should not be picked, for example, because a barcode or other code is located there that needs to be read while the object is being picked. However, this is only practical if such types of objects or such reading are actually supported by the system.

[0015] By enabling the use of an externally specified selection algorithm, the operation of the robot can be adapted to individual needs, situations, applications and user wishes.

[0016] In one embodiment, the selection algorithm comprises several selection filters, each determining the object poses. These selection filters are distinct from one another and can—but do not necessarily have to—build upon each other. These selection filters can also be referred to as selection or filter strategies. For example, one selection filter might specify that object poses at the edge (e.g., the edge of a box) are excluded (i.e., not selected), while another might specify that object poses from the topmost, directly accessible objects are selected. These are just a few examples of such selection filters; ultimately, any number of such filters can be implemented, depending, for example, on user requirements.

[0017] In one embodiment, the selection algorithm is provided externally, and is applied at runtime to the initial object pose data set to determine the selection object pose record. This can be done, for example, via a plug-in or a configuration file. When the process reaches the appropriate point, i.e., when the initial object pose data set is available, the selection algorithm is called and executed accordingly.

[0018] In one embodiment, this is done statically, meaning the multiple selection filters of the selection algorithm are executed in a predefined order. It is also conceivable that a selection option is included in this way, whereby, depending on the current application (then at runtime), only certain selection filters from the predefined order are actually applied.

[0019] In one embodiment, however, it is also provided that the use of the multiple selection filters is dynamically determined depending on the result of applying one of the multiple selection filters and / or an external event. For example, when applying a first selection filter in a specific case, only one or even no object pose may remain, so that no other selection filter can be applied. It is also conceivable, for example, that a different sequence of selection filters is chosen depending on the situation.

[0020] In one embodiment, the selection algorithm for determining the selection object pose data set is implemented by compilation before execution. In other words, the selection algorithm specified by, for example, the user, can be implemented once before its first use by compilation.

[0021] The proposed approach, using an interface for an externally specified selection algorithm, allows for flexible system configuration and decoupling of individual selection filters. It enables expansion through the aforementioned plug-in system, and even parallel processing of the selection filters (provided this is possible given the type of selection filters). Furthermore, it allows for, for example, the rapid testing of new selection filters without affecting existing functionality.

[0022] A computing unit according to the invention (i.e., generally a system for data processing), e.g., a control unit or a control unit of a robot, or a central server or other computing system, is, in particular in terms of programming, equipped to carry out a method according to the invention.

[0023] The invention also relates to a robot that is configured to receive control information as described above. In addition, or alternatively, the robot has a computing unit according to the invention.

[0024] Furthermore, the robot includes, in particular, a control unit and a drive unit for moving the robot. In addition, the robot may have at least one sensor for capturing environmental information, e.g., a camera and / or a lidar sensor.

[0025] Implementing a method according to the invention in the form of a computer program or computer program product with program code for carrying out all method steps is also advantageous, as this incurs particularly low costs, especially if an executing control unit is already available for other tasks. Finally, a machine-readable storage medium is provided with a computer program stored on it as described above. Suitable storage media or data carriers for providing the computer program are, in particular, magnetic, optical, and electrical storage media, such as hard drives, flash memory, EEPROMs, DVDs, etc. Downloading a program via computer networks (Internet, intranet, etc.) is also possible. Such a download can be wired or wireless (e.g., via a WLAN network, a 3G, 4G, 5G, or 6G connection, etc.).

[0026] Further advantages and embodiments of the invention will become apparent from the description and the accompanying drawing.

[0027] The invention is schematically illustrated in the drawing using an exemplary embodiment and is described below with reference to the drawing.

[0028] Brief description of the drawings

[0029] Figure 1 schematically shows a robot to illustrate the invention.

[0030] Figure 2 schematically shows a box containing objects to illustrate the invention.

[0031] Figure 3 schematically shows a process flow in one implementation form.

[0032] Figures 4, 5 and 6 schematically show the sequence of processes in further implementation forms.

[0033] embodiment(s) of the invention

[0034] Figure 1 schematically illustrates a robot 100 to explain the invention. By way of example, the robot 100 has stand and arm components 102, 104, 106, so-called axes, which are each movably and movably connected by means of joints 112, 114.

[0035] Furthermore, the robot 100 has an end effector 108, e.g., a gripper. The end effector 108 is movably and reversibly connected to the arm component 106 by means of a joint 116.

[0036] Furthermore, the robot 100 has a drive system 120, shown only schematically here, as well as a computing unit 122 designed, for example, as a control unit. The control unit can be, for example, part of the computing unit 122 or be provided separately. This allows the drive system 120 to be controlled, for example, using control information, in order to move the robot according to a desired sequence of movements. This can include, for example, moving the axes relative to each other by means of the joints, but also rotating the axes themselves, provided that corresponding drives are available.

[0037] It should be noted that the robot 100 is only used here as an example for illustrative purposes. A robot for grasping and / or moving objects in containers can also be constructed differently, for example, using an end effector that can only move linearly along several different rails.

[0038] Furthermore, a box 132 is shown in a work environment 130, containing, for example, an object 140. The robot 100 can now be controlled, for example, in such a way that it grasps and / or moves the object 140 using the end effector 108, in particular also taking it out of the box and, for example, placing it somewhere else.

[0039] Furthermore, an example sensor 124 is shown, which could be, for example, a camera or a lidar sensor. Both types of sensors can also be provided. Likewise, several identical sensors can be provided. The sensor 124 can, for example, be arranged in a suitable manner in the working environment, e.g., on a ceiling and thus separately from the robot 100. However, the sensor 124 could also, for example, be part of the robot and be arranged, for example, on the arm component 106 or the end effector 108.

[0040] The sensor 124 can now detect the working environment 130 and, in particular, the box 132 and its interior, including the object 140. Based on the environmental information and sensor data obtained in this way, a movement sequence can be created for the robot to grasp and / or move the object 140 using the end effector 108.

[0041] Figure 2 shows a box 232, comparable to box 132 in Figure 1, containing various objects 240, 241, 242, 243, and 244. The goal is to grasp object 240 and remove it from the box. For this purpose, an approach path, or first movement sequence 251, is shown, along which the end effector moves to object 240 in order to grasp it. Additionally, a retraction path, or second movement sequence 252, is shown, along which the end effector, with object 240, can be moved out of the box.

[0042] Together, the first partial movement sequence 251 and the second partial movement sequence 252 form a complete movement sequence. It can be seen that the poses of the objects in box 232 are relevant for determining the movement sequence, and not all object poses are equally suitable. The first partial movement sequence 251 runs very close to a side wall of box 232 (not shown here), so there is a risk of collision.

[0043] Therefore, for example, object poses close to the edge of box 232 could be filtered out. However, object poses that are on top, such as object 244, could also be selected.

[0044] This will be explained in more detail below.

[0045] Figure 3 schematically illustrates the sequence of a process in one embodiment. The process, or the steps described herein, can be carried out, for example, on the computing unit 122 according to Figure 1 or otherwise.

[0046] In step 300, a user provides a selection algorithm 302 of their choosing; the selection algorithm 302 includes, for example, several selection filters 304a, 304b, 304c, based on which object poses can be selected. This can be done in advance, i.e., before the actual use of the robot. It should be noted that only three selection filters are shown here as examples; there could be more or fewer. The specific type of selection filters is not important here, but for a more detailed explanation of examples, please refer to the preceding explanations.

[0047] In step 310 – for example, during robot operation – environmental information 312 is captured from the work environment using at least one sensor. Specifically, this could be, for example, a scene recording of the box or the work environment in which objects are to be recognized. The sensor captures, for example, one or more 3D and / or 2D images, which are processed in subsequent steps. This captured environmental information is then made available in step 314 for further use or processing.

[0048] In step 320, based on the environment information, a pose 322 is determined for one or each of several objects in the work environment in order to obtain an initial object pose record 324.

[0049] Specifically, image processing can involve processing images or other environmental information. Such image processing can range from simple surface segmentation to complex algorithms using neural networks that classify and segment complex objects. The result might be, for example, a vector containing object poses. These poses represent all localized objects.

[0050] In step 330, one or more selected poses 332 are then determined from the initial object-pose data set 324, using the externally specified selection algorithm 302, for one or at least one of the several objects, in order to obtain a selection-object-pose data set 334.

[0051] In this case, the selection algorithm 302 is called and executed at runtime, similar to a plug-in. As mentioned above, it would also be conceivable to compile the software required for operating the robot before use in order to implement the provided and predefined selection algorithm.

[0052] In step 340, the selection object pose data set is then made available for determining the robot's motion sequence. In step 342, for example, the motion sequence is determined for the robot as shown in Figure 2, in order to grasp and / or move the object in the work environment using one or more of the robot's end effectors. Figure 4 schematically illustrates a further implementation of a procedure. Here, an externally specified selection algorithm 402 is intended to include, by way of example, the selection filters 404a, 404b, and 404c. The selection algorithm 402 is to be implemented statically and in advance by compiling. The order is static as specified, even in later application: selection filter 404a, then 404b, and then 404c.

[0053] Figure 5 schematically illustrates the flow of a procedure in another implementation. Here, an externally defined selection algorithm 502 is intended to include, by way of example, the selection filters 504a, 504b, 504c, 504d, and 504e. The selection algorithm 502 is to be statically defined but called and applied at runtime. This can be done using an (external) configuration file 510, which, for example, allows only the selection filters required for the current case to be called; in the example shown, these are the selection filters 504a, 504c, and 504e. The order of the selection filters themselves, however, does not change.

[0054] In this case, the user specifies, for example, in the external configuration file (e.g., via a sequence in the configuration file), which algorithms or selection filters are used for filtering. As explained below, this can also be done dynamically without user input, using selection algorithms (which are stored in the transition objects mentioned below). Similar to the game "Mikado," the next stick—or object—to be picked up can usually be selected from several positions. However, there are simple positions—the stick or object lies freely without being touched—and more difficult positions—the stick or object lies on top, but also on top of several other sticks or objects, which may be moved when one is picked up.

[0055] Figure 6 schematically illustrates the flow of a procedure in another implementation. Here, an externally specified selection algorithm 602 is intended to include, by way of example, the selection filters 604a, 604b, 604c, 604d, and 604e. The selection algorithm 602 is to be called and applied dynamically at runtime. This can be done using a selection mechanism 610, e.g., a so-called transition object. Any predefined order of the selection filters is irrelevant here. Rather, this allows for a dynamic combination at runtime, dependent, for example, on the results of a strategy or on control by external events 620. In the example shown, selection filter 604c is chosen based on the external event 620. Next, one of the other selection filters can be used, for example.

[0056] This transition object 610 specifies, for example, which strategy will be used next. Any calculations can be performed within this transition object, or external events can be reacted to in order to select the next strategy to be executed. Parallel filtering processes can also be triggered via these transitions.

[0057] The transition object contains, for example, an algorithm that, depending on input signals from one or more preceding processing steps, determines the selection for the next processing step(s). In principle, the object answers the question of what should be calculated next and with what tool the calculation should be performed.

Claims

Claims 1. Method for determining one or more poses of one or more objects in a work environment (130), for use in determining a motion sequence for a robot (100) to grasp and / or move one or more of the objects (140) by means of an end effector (108) of the robot, comprising: Providing (314) environmental information (312) that has been acquired from the working environment by means of at least one sensor (124); Determine (320), based on the environment information (312), several poses (322) for one or more objects in the work environment to obtain an initial object pose record (324); Determine (330), from the initial object-pose record (324), using an externally specified selection algorithm (302), one or more selected poses (332) for one or at least one of the multiple objects, in order to obtain a selection-object-pose record (334); and Providing (340) the selection object pose data set (334) for determining the motion sequence for the robot.

2. Method according to claim 1, wherein the selection algorithm (302) comprises several selection filters (304a, 304b, 304c) on the basis of which object poses are selected.

3. Method according to claim 1 or 2, further comprising: providing (300) the externally specified selection algorithm (302); wherein the selection algorithm for determining the selection object pose data set is applied at runtime to the initial object pose data set.

4. Method according to claims 2 and 3, wherein the use of the multiple selection filters is dynamically determined depending on a result of an application of one of the multiple selection filters and / or an external event (620).

5. Method according to claim 1 or 2, wherein the selection algorithm for determining the selection object pose data set is implemented by compilation prior to execution.

6. Method according to one of the preceding claims, wherein the selection algorithm is or has been specified by a user.

7. Method according to any of the foregoing claims, further comprising: Determine (342) the movement sequence (251 , 253) based on the selection object poses dataset (334).

8. The method of claim 7, further comprising: Determine, based on the movement sequence, control information for moving the robot, and Providing control information and / or moving the robot based on the control information.

9. Computing unit (122) comprising means for carrying out the method according to any of the preceding claims.

10. Robot (100) configured to receive control information determined by a method according to claim 8, and / or with a computing unit (122) according to claim 9, and with a drive system (120) and a control or regulation unit for controlling the drive system, with one or more, in particular several selectable, end effectors for grasping and / or moving an object, and preferably with at least one sensor, in particular a camera, for capturing environmental information of a working environment.

11. Computer program comprising instructions which, when the program is executed by a computer, cause it to execute the procedure according to claim 1 to 8.

12. Computer-readable storage medium on which the computer program according to claim 11 is stored.

Citation Information

Patent Citations

  • Methods for object localization and pose estimation for an object of interest

    DE102015113434A1

  • Computerized system and method using different image views to find grasp locations and trajectories for robotic pick up

    EP3695941A1

  • Autonomous unknown object pick and place

    WO2020205837A1

  • Systems and methods for robotic control under contact

    WO2021025955A1