Method for determining object poses
The method combines classical image processing and AI-based object recognition algorithms to determine object poses, improving robotic grasping efficiency by selecting the most promising poses for robotic movement sequences, addressing inefficiencies in existing technologies.
Patent Information
- Application Number
- PCT/EP2025/070979
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-23
- Filing Date
- 2025-07-22
- Publication Date
- 2026-01-29
AI Technical Summary
Existing methods for determining object poses for robotic grasping and movement are inefficient, especially when prior knowledge or product data about the objects is lacking, leading to inaccurate and time-consuming object recognition and movement sequences.
A method utilizing multiple object recognition algorithms, including classical image processing and AI-based models, to determine object poses, followed by a ranking system that evaluates and selects the most promising pose for robotic grasping, using environmental information like 3D point clouds to optimize movement sequences.
Enhances the precision and efficiency of robotic object grasping by leveraging the strengths of different algorithms, ensuring accurate and efficient retrieval of objects even without prior knowledge, thereby optimizing movement sequences.
Smart Images

Figure EP2025070979_29012026_PF_FP_ABST
Abstract
Description
[0001] Methods for determining object poses
[0002] Description
[0003] The present invention relates to a method for determining object poses, for use in determining a movement sequence for a robot to grasp and / or move an object in a working environment by means of an end effector of the robot, a computing unit and a computer program for carrying it out, and a robot.
[0004] Background of the invention
[0005] Robots, often also referred to as kinematics, can be used in production facilities, logistics, and other systems to, for example, transport or assemble parts. Typical types of such robots include Cartesian robots, SCARA robots, and articulated robots.
[0006] Disclosure of the invention
[0007] According to the invention, a method for determining object poses, a computing unit and a computer program for its execution, as well as a robot with the features of the independent claims, are proposed. Advantageous embodiments are the subject of the dependent claims and the following description.
[0008] The invention relates generally to robots that can be used, for example, in production plants, logistics, or other facilities. Types of such robots, often also referred to as kinematics, include, for example, Cartesian robots, SCARA robots, and articulated robots. Such robots can be used to grasp and / or move objects (i.e., items or parts). A typical application is to remove an object from a box containing, for example, a large number of (identical or different) objects and, for example, place it somewhere or possibly assemble it. For this purpose, such a robot has, for example, a gripper or, more generally, an end effector. Such robots can also be referred to as manipulators or manipulation devices. In order for the robot to grasp an object in a work environment, such as a box, or in or on other supports (e.g., a tray), the robot is equipped with a gripper or, more generally, an end effector.To be able to grasp and / or move an object (e.g., a container or pallet) using the end effector, the robot, in particular its end effector, must perform a movement sequence to move the end effector to the position of the object and then, e.g., after the object has been grasped, to move the end effector away from the position of the object, e.g., to a desired storage position.
[0009] For a successful gripping operation, the (preferably) exact pose for the gripper – hereinafter also referred to as the gripping pose – must first be determined. Environmental information, especially 3D environmental information such as a point cloud of the work environment (scene), particularly of a support structure such as a box, can be acquired and provided for this purpose. Sensors such as lidar sensors (laser scanners), cameras, or other depth sensors can be used for this. The environmental information then also includes, in particular, one or more objects located in the work environment, especially in or on a support structure.
[0010] Based on the 3D environment information, particularly the point cloud, object poses can first be determined for each—or at least some—of the multiple objects. This can involve, for example, segmenting the point cloud to find suitable object poses. An object pose specifies, in particular, the position and orientation of an object in space. An object pose thus characterizes an object. Based on the object pose, one or possibly several gripping poses can then be determined for the robot to pick up the object in question.
[0011] In many use cases, there is a large number of objects, for example in a box, from which objects are to be taken one by one. For example, it might be intended that objects are taken from a box until the box is empty.
[0012] The so-called "object segmentation" for determining object poses is a process step that typically occurs after the acquisition of environmental information. This environmental information is usually a 3D scan, which results in a 3D point cloud. The goal of object segmentation is then to identify objects and their position and orientation (location) in space (pose, object pose) within the point cloud.
[0013] Object segmentation can distinguish between two groups: those with prior knowledge and product data, and those without prior knowledge and product data.
[0014] In cases where prior knowledge and product data are available, object segmentation is performed using, for example, CAD matching. A CAD model of the object must be available for this purpose and can then be found and assigned within the point cloud using a matching algorithm. This allows the poses of objects, such as those located in a box, to be determined.
[0015] In the case of no prior knowledge or product data, however, no (extensive) information or even CAD data about the objects is available; therefore, information must be obtained about the arrangement of the points within the point cloud to enable the determination of the object poses.
[0016] Various object recognition algorithms can be used to determine object poses for objects in the work environment, e.g., in a box. As it turns out, using several different object recognition algorithms, e.g., executed in parallel, allows for a more precise and targeted determination of object poses, and in particular, an evaluation of which object pose allows for the fastest and / or simplest possible retrieval.
[0017] Object recognition algorithms can be based on classical image processing or on machine learning models, i.e., artificial intelligence (AI). For example, an area detection algorithm can be used to detect areas within the surrounding information. Another example is a cylinder detection algorithm, which can detect cylindrical shapes within the surrounding information. Both of these object recognition algorithms are based on classical image processing. Generally, such algorithms attempt to identify primitive shapes (areas, cylinders, etc.) in the point cloud, i.e., to find arrangements of points that represent primitive shapes. The object poses, and therefore the gripping points, of these primitives are known.An object recognition algorithm based on AI is, for example, a recognition algorithm based on a machine learning model (e.g., a neural network) that can detect predefined shapes and / or objects within the environment. The predefined shapes and / or objects depend, in particular, on the training of the machine learning model. For this purpose, the machine learning models should generally be trained beforehand with point clouds and image data of the objects that correspond to the expected object categories, such as cuboid boxes, cylindrical objects, blister packs, etc.
[0018] In principle, at least two different object recognition algorithms can be used, but it is advisable to use at least one object recognition algorithm based on classical image processing and at least one based on AI.
[0019] First, environmental information, captured by a sensor from the work environment, is provided, as already mentioned. This environmental information includes, in particular, a 3D point cloud.
[0020] Based on environmental information, object data records are then determined for objects detected in the work environment using one of several different object recognition algorithms. Each object data record for an object detected by an object recognition algorithm comprises one object pose of that object.
[0021] It should be noted that while ideally every object recognition algorithm would recognize every object, this is not necessarily the case. For example, each of three object recognition algorithms could recognize each of ten objects; then there would be 30 object records, with three object records corresponding to each actual object. In practice, however, not every object recognition algorithm will recognize every object; this will be discussed in more detail later. In general, this means there can be more recognized objects than actual objects.
[0022] In addition to the object pose, an object record can also include further information. In one embodiment, an object record further includes a value for one or more types of meta-object information. These types can be, for example, a measure of the accuracy of the object pose recognition using the respective object recognition algorithm, as well as a geometric property of the object.
[0023] For object recognition algorithms based on classical image processing, the quality measure can be, for example, a so-called "root-mean-square error" of the object fit, i.e., the quality of the match between, for example, an area and the points found in the point cloud. For AI-based object recognition algorithms, this can generally be the quality of the object match.
[0024] In object recognition algorithms based on classical image processing, the geometric property can be, for example, the underlying geometric shape, such as the size of an area or the length and / or radius of a cylinder or cylindrical shape. In AI-based object recognition algorithms, this can be, for example, the extent or size of a bounding box of the detected object.
[0025] Based on the object data records, a ranking value is then determined for at least some of the recognized objects, based on their respective object poses. This can be the case for all recognized objects, but it may also be possible to filter out objects that, although recognized, are not promising, for example, based on their meta-object information.
[0026] Information about the ranking values is then provided for determining a suitable object pose and, in particular, for determining the robot's movement sequence. Based on an object pose dataset—that is, a multitude of possible object poses (e.g., from all objects recognized as described above)—and the ranking value information, a suitable object pose can then be determined. Such a ranking value could be, for example, a number or a point value, which is higher (or lower) the more promising the object pose.
[0027] Using multiple object recognition algorithms allows you to leverage the advantages of each algorithm, thereby compensating for the disadvantages of any single one. A simple example is when only one object is recognized by all the algorithms. In this case, it can be assumed that the object's pose is the most promising for successful capture. Further variations will be explained below.
[0028] Based on the object pose to be used, and if necessary after determining a required gripping pose, the movement sequence can then be determined and the robot controlled accordingly.
[0029] In one embodiment, determining, based on the object data sets, for at least some of the detected objects, a ranking value for each object pose comprises: Determining, for each object detection algorithm, based on the object data sets of the objects detected by the respective object detection algorithm, one or more lists, each containing a sorting of the object poses. An intermediate ranking value is then determined for each object pose in the respective list, so that the ranking values of the object poses can be determined based on the intermediate ranking values.
[0030] One example of such a list is a quality measure list sorted according to the values of the quality measure, meaning that the object positions are sorted in descending order with respect to the value of the quality measure. The higher the position in the list, the higher (or, for example, lower) the intermediate ranking value of the respective object position.
[0031] One example of such a list could be a geometry list sorted according to the values of the geometric property, meaning that the object poses are sorted in descending order with respect to the value of the geometric property (e.g., area size, cylinder length). The higher the position in the list, the higher (or, for example, lower) the intermediate ranking value of the respective object pose.
[0032] One example of such a list could be a height list, sorted according to the height of objects in the workspace. This means that the object poses are sorted in descending order based on their position in the z-direction (information derived from the object's pose). The higher the position in the list, the higher (or lower, for example) the intermediate ranking value of the respective object pose. The height, or z-direction, is defined in relation to, or along, a direction of gravity. The rationale behind this is that the higher an object is located, the easier it is to access its pose.
[0033] In this way, various other properties can be used to assess the object poses.
[0034] In one embodiment, determining, based on the object data sets, for at least some of the detected objects, a ranking value for each object pose comprises: determining, for each detected object, a recognition count by different object recognition algorithms by which the respective object was detected. As already mentioned, not every object is necessarily detected by every object recognition algorithm. However, if an object is detected by two or more object recognition algorithms (the recognition count is then two or more), this indicates that this object is easy to detect and therefore also easy to retrieve, as it will presumably be relatively unobstructed. The ranking values of the object poses are then determined based on the recognition counts.
[0035] A computing unit according to the invention (i.e., generally a system for data processing), e.g., a control unit or a control unit of a robot, or a central server or other computing system, is, in particular in terms of programming, equipped to carry out a method according to the invention.
[0036] The invention also relates to a robot configured to receive control information as described above. Furthermore, or alternatively, the robot comprises a computing unit according to the invention.
[0037] Furthermore, the robot includes, in particular, a control unit and a drive unit for moving the robot. In addition, the robot may have at least one sensor for capturing environmental information, e.g., a camera and / or a lidar sensor.
[0038] Implementing a method according to the invention in the form of a computer program or computer program product with program code for carrying out all method steps is also advantageous, as this incurs particularly low costs, especially if an executing control unit is already available for other tasks. Finally, a machine-readable storage medium is provided with a computer program stored on it as described above. Suitable storage media or data carriers for providing the computer program are, in particular, magnetic, optical, and electrical storage media, such as hard drives, flash memory, EEPROMs, DVDs, etc. Downloading a program via computer networks (Internet, intranet, etc.) is also possible. Such a download can be wired or wireless (e.g., via a WLAN network, a 3G, 4G, 5G, or 6G connection, etc.).
[0039] Further advantages and embodiments of the invention will become apparent from the description and the accompanying drawing.
[0040] The invention is schematically illustrated in the drawing using an exemplary embodiment and is described below with reference to the drawing.
[0041] Brief description of the drawings
[0042] Figure 1 schematically shows a robot to illustrate the invention.
[0043] Figure 2 schematically shows a container to illustrate the invention.
[0044] Figure 3 schematically shows a process flow in one embodiment.
[0045] embodiment(s) of the invention
[0046] Figure 1 schematically illustrates a robot 100 to explain the invention. By way of example, the robot 100 has stand and arm components 102, 104, 106, so-called axes, which are each movably and movably connected by means of joints 112, 114.
[0047] Furthermore, the robot 100 has an end effector 108, e.g., a gripper. The end effector 108 is movably and reversibly connected to the arm component 106 by means of a joint 116.
[0048] Furthermore, the robot 100 has a drive system 120, shown only schematically here, as well as a computing unit 122 designed, for example, as a control or regulation unit. This allows the drive system 120 to be controlled, for example, using control information, in order to move the robot according to a desired sequence of movements. This can include, for example, moving the axes relative to each other by means of the joints, but also rotating the axes themselves, provided that appropriate drives are available.
[0049] It should be noted that the robot 100 is only used here as an example for illustrative purposes. A robot for grasping and / or moving objects in containers can also be designed differently, for example, using an end effector that can only move linearly along several different rails.
[0050] Furthermore, a box 132 is shown in a work environment 130, containing, for example, an object 140. The robot 100 can now be controlled, for example, in such a way that it grasps and / or moves the object 140 using the end effector 108, in particular also taking it out of the box and, for example, placing it somewhere else.
[0051] Furthermore, an example sensor 124 is shown, which is, for example, a 3D sensor, in particular a lidar sensor. Using this sensor 124, 3D environmental information, such as a point cloud, can be captured. The sensor 124 can, for example, be arranged in a suitable manner in the work environment, e.g., on a ceiling and thus separately from the robot 100. However, the sensor 124 could also, for example, be part of the robot and be arranged, for example, on the arm component 106 or the end effector 108.
[0052] The sensor 124 can now detect the working environment 130 and, in particular, the box 132 and its interior, including the object 140. Based on the environmental information or sensor data obtained in this way, object poses can first be determined (within the framework of so-called object segmentation), and based on this, a movement sequence can be created for the robot to grasp and / or move the object 140 using the end effector 108, as will be explained in more detail below.
[0053] Figure 2 shows a box 232, comparable to box 132 in Figure 1, containing various objects 240, 241, 242, 243, and 244. A typical task for a robot is to remove as many objects as possible from the box. It is advantageous to first remove the easiest or safest object to grasp and so on, while adhering to certain guidelines, such as not covering forbidden areas with the end effector. It can also be seen that the objects can differ from one another, for example, being cuboid or cylindrical, and that the objects can also be randomly arranged.
[0054] An example movement sequence is shown for grasping object 240 and removing it from the box. This includes an approach path, or first part of the movement sequence 251, by which the end effector moves towards object 240 in order to grasp it. Additionally, a retraction path, or second part of the movement sequence 252, is shown, along which the end effector can be moved with object 240 out of the box.
[0055] Together, the first partial movement sequence 251 and the second partial movement sequence 252 form a complete movement sequence. It can be seen that the movement sequence is such that the end effector must be in a specific pose – a grasping pose – in order to properly grasp the object 240.
[0056] It can also be seen that, in principle, any object could be grabbed next; however, it is more practical to grab first the object where the grab is most promising.
[0057] Object poses are shown for objects 240, 241, 242, 243, and 244, using a small coordinate system superimposed on the respective object. For object 244, one such object pose is labeled 264.
[0058] To find the object where the sampling is most promising—or at least more promising than with other objects—the object poses of as many objects as possible in box 232 should be determined. Environmental information can be collected for this purpose, as indicated here by an example using a 3D point cloud 270.
[0059] The detailed procedure will be explained in more detail below.
[0060] Figure 3 schematically illustrates the sequence of a process in one embodiment. Reference is also made to Figures 1 and 2.
[0061] In step 300, environmental information 302 is first provided, which has been acquired from the working environment by means of a sensor - e.g. the sensor 124 according to Figure 1. The environmental information 302 can in particular be a 3D point cloud, as shown e.g. in Figure 2 with 270.
[0062] In step 310, based on the environment information 302, object data sets 314 of objects detected in the working environment are determined using one of several different object recognition algorithms. The following three different object recognition algorithms will be used as examples.
[0063] A surface detection algorithm 312a, by means of which surfaces can be detected in the environment information, a cylinder detection algorithm 312b, by means of which cylindrical shapes can be detected in the environment information, and a machine learning model-based detection algorithm 312c, by means of which predefined shapes and / or objects can be detected in the environment information, wherein the predefined shapes and / or objects depend in particular on a training of the machine learning model.
[0064] The surface detection algorithm 312a can also be referred to as a plane fit, e.g. for the detection of surfaces within the point cloud; this is particularly suitable for cuboid-shaped items such as boxes (see objects 240, 243 in Figure 2).
[0065] The cylinder detection algorithm 312b can also be referred to as a cylinder fit, e.g. for the detection of cylinders within the point cloud; this is particularly suitable for cylindrical articles (see objects 241, 242, 244 in Figure 2).
[0066] The recognition algorithm 312c, which is based on a machine learning model, can also be referred to as an AI network, e.g. for the detection of individually shaped objects, but also planar and cylindrical objects (depending on the training data).
[0067] As mentioned, other object recognition algorithms and also a higher number than three, e.g. four, five or six different object recognition algorithms, can be used.
[0068] It can be assumed that no information about the objects in the box is available. Using each object recognition algorithm, object data sets for objects in the work environment are then determined or identified based on the 3D point cloud, specifically for those objects that are also recognized as objects by the respective object recognition algorithm.
[0069] Each object data record then includes an object pose of the object, as well as, for example, a value for a measure of accuracy in recognizing the respective object pose using the respective object recognition algorithm, and a value for a geometric property of the object. The object pose can, for example, include the position (x, y, z) and the orientation (rotation about x, rotation about y, rotation about z) (see also the coordinate axes x, y, z in Figure 2).
[0070] For the area detection algorithm 312a, or area fit, the goodness-of-fit measure can be the root-mean-square error of the object fit (i.e., the quality of the match), and the geometric property can be the size of the detected area. For the cylinder detection algorithm 312b, or cylinder fit, the goodness-of-fit measure can be the root-mean-square error of the object fit (i.e., the quality of the match), and the geometric property can be the radius and / or length of the detected cylinder. For the detection algorithm 312c, or the Kl mesh, the goodness-of-fit measure can be the quality of the object match, and the geometric property can be the extent of the bounding box or bounding box region of the detected object.
[0071] The identified object poses are then evaluated according to their properties and sorted, clustered, or classified accordingly. The goal is to find the most promising object pose for a grab in order to minimize errors with the robot or its end effector.
[0072] The following describes various sorting and clustering methods, as well as possible combinations thereof.
[0073] Each object recognition algorithm initially provides a set of object data records. It should be noted that it is also possible for an object recognition algorithm to fail to detect any objects. The object poses within these object data records can then be sorted in various ways. In this way, multiple lists can be created for each object recognition algorithm. One such list can be a quality measure list, in which the object poses are sorted according to the values of the quality measure, specifically in descending order; that is, the better the value, the higher up the object pose appears in the list. Quality measure lists for the three different object recognition algorithms mentioned are designated 316a.1, 316b.1, and 316c.1.
[0074] A list can be a geometry list in which the object poses are sorted according to the values of the geometric property, specifically in descending order; that is, the higher the value, e.g., the larger the area or the length of the cylinder, the higher up in the list the object pose is. Geometry lists are designated 316a.2, 316b.2, and 316c.2 for the three different object recognition algorithms mentioned.
[0075] A list can be a height list, in which the object poses are sorted according to the height of the objects in the workspace, specifically in descending order, meaning the higher an object is located, the higher its pose appears in the list. Such different heights, i.e., different positions along the z-direction, are shown in Figure 2. For example, the height of the object's center or center of gravity can be used, or even its highest point. Height lists for the three different object recognition algorithms mentioned are designated 316a.3, 316b.3, and 316c.3.
[0076] Assuming there are n objects in the box and each object recognition algorithm ideally detects all objects, this results in 3n object poses for n objects with different properties (e.g., z-height, quality of match, area size from surface fit, etc.). Based on the sorting, the properties of an object pose can be evaluated and quantified, creating a ranking. Each object pose receives an intermediate ranking value, e.g., a score, depending on its position in the property-based sort. For example, each object pose in each list can be assigned a score based on its position in the list.
[0077] Within the object poses, each object pose can usually be uniquely assigned to an object, for example, by comparing the Cartesian pose information (x, y, z, rotation about x, rotation about y, rotation about z) and clustering it accordingly. This allows us to determine how many object recognition algorithms have recognized an object. Thus, in step 320, a recognition count of 322 is determined. In the example with three different object recognition algorithms, the recognition count can be one, two, or three.
[0078] In step 330, the ranking values 332 for the object poses are then determined. This is done based on the intermediate ranking values 318 and the recognition counts 322.
[0079] By combining all measures, object poses for n objects result, where each object or object pose has the properties
[0080] - Quality / probability of the match,
[0081] - geometric property (size of the area, bounding box, length of the cylinder),
[0082] - Z-height,
[0083] The number of object recognition algorithms that have found the object is displayed. The ranking of the object poses will depend on the evaluation of the properties. The object pose with the best ranking value is then selected, for example, for the next retrieval.
[0084] In step 340, information about the ranking values is provided, so that an object pose to be used and, in particular, the movement sequence for the robot can be determined.
Claims
Claims 1. Method for determining object poses, for use in determining a motion sequence for a robot (100) to grasp and / or move an object (140) in a work environment by means of an end effector (108) of the robot, comprising: Providing (300) environmental information (302) acquired from the work environment by means of a sensor (124), wherein the environmental information includes in particular a 3D point cloud (270); Determine (310), based on the environment information, using each of several different object recognition algorithms (312a, 312b, 312c), object data sets of objects detected in the work environment, wherein an object data set (314) for an object detected by means of an object recognition algorithm comprises an object pose of the object; Determine (330), based on the object data records (314), for at least some of the detected objects, a ranking value (332) of the respective object poses; and Providing (340) information about the ranking values for determining an object pose to be used and, in particular, for determining the movement sequence for the robot.
2. The method of claim 1, wherein an object data record for an object that has been recognized by means of an object recognition algorithm further comprises: a value for one or more types of meta-object information.
3. Method according to claim 2, wherein the one or more types of meta-object information are one or more of the following types: a measure of the accuracy of the recognition of the respective object pose by means of the respective object recognition algorithm, a geometric property of the object.
4. Method according to one of the preceding claims, wherein the determination, based on the object data sets, for at least a part of the The detected objects, each with a ranking value for the respective object poses, include: Determine, for each object recognition algorithm, based on the object data sets of the objects recognized by the respective object recognition algorithm, one or more lists, each containing a sorting of the object poses; Determine one intermediate ranking value (318) of each object pose of each list; wherein the ranking values (322) of the object poses are determined based on the intermediate ranking values.
5. The method according to claims 3 and 4, wherein, for each object recognition algorithm, the list is one of the following lists, or wherein, for each object recognition algorithm, the multiple lists are at least two of the following lists: a quality measure list sorted according to the values of the quality measure (316a.1, 316b.1, 316c.1), a geometry list sorted according to the values of the geometric property (316a.2, 316b.2, 316c.2), a height list sorted according to a height of the objects in the working environment (316a.3, 316b.3, 316c).
6. Method according to one of the preceding claims, wherein the determination, based on the object data sets, for at least some of the identified objects, comprises: Determine (320) for each detected object, a recognition count (322) of different object recognition algorithms by which the respective object has been detected; wherein the ranking values (332) of the object poses are determined based on the recognition counts.
7. A method according to any of the preceding claims, wherein the several different object recognition algorithms comprise at least two of the following object recognition algorithms: an area recognition algorithm (312a) by means of which areas can be detected in the environment information, a cylinder detection algorithm (312b) by means of which cylinder shapes can be detected in the environment information, a machine learning model-based detection algorithm (323c) by means of which predefined shapes and / or objects can be detected in the environment information, wherein the predefined shapes and / or objects depend in particular on a training of the machine learning model.
8. A method according to any of the foregoing claims, further comprising: Determine, based on an object pose dataset and information about the ranking values, the object pose to use, and Determining the sequence of movements, taking into account the object pose to be used.
9. The method of claim 8, further comprising: Determine, based on the movement sequence, control information for moving the robot, and Providing control information and / or moving the robot based on the control information.
10. Computing unit comprising means for carrying out the method according to any of the preceding claims.
11. Robot configured to receive control information determined by a method according to claim 9, and / or with a computing unit according to claim 10, and with a drive system and a control or regulation unit for controlling the drive system, with one or more, in particular several selectable, end effectors for grasping and / or moving an object, and preferably with at least one sensor, in particular a camera, for capturing environmental information of a working environment.
12. Computer program comprising instructions which, when the program is executed by a computer, cause it to execute the method according to claims 1 to 9.
13. Computer-readable storage medium on which the computer program according to claim 12 is stored.
Citation Information
Patent Citations
Measurement parameter optimization method and device as well as computer control program
DE102021103726A1
Robot system, control method, image processing apparatus, image processing method, method of manufacturing products, program, and recording medium
EP4070922A2
Autonomous unknown object pick and place
WO2020205837A1
Transformation for covariate shift of grasp neural networks
WO2022250658A1