Object position detection by automatic feature extraction and / or feature correspondence
By taking images from different perspectives and determining descriptor images using pre-trained artificial intelligence, and automating feature extraction and pose estimation, the inefficiency problems of feature extraction and corresponding in the prior art are solved, and the rapid and accurate identification and capture of similar objects are achieved.
Patent Information
- Application Number
- CN202380085633.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-13
- Filing Date
- 2023-11-15
- Publication Date
- 2025-07-22
AI Technical Summary
In the prior art, the feature extraction and feature correspondence processes lack automation, resulting in inefficient object recognition, and it is difficult to quickly and accurately determine the positioning position and grab position in multiple object scenarios.
By using the shooting device to capture images from different perspectives, pre-trained artificial intelligence determines the descriptor image, automatically extracts and corresponding features, and combines deep data and neural networks to achieve automatic selection of features and pose estimation.
It realizes fast and robust feature extraction and pose determination of similar objects, reduces the need for manual labeling and calculation distortion, and improves the efficiency of object recognition and grabbing.
Smart Images

Figure CN120359544A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for automatic feature extraction and / or feature correspondence, a system for automated feature extraction, and a computer program or computer program product. Background Art
[0002] Feature-based object recognition generally consists of two steps: The first step is feature extraction, which typically includes manually selecting features such as circles, edges, corners, etc., and selecting or implementing algorithms and / or methods for feature recognition, especially for example the Hough transform for circle recognition. For edge detection, for example, appropriate parameters must be found, especially when using the Canny filter. The second step is usually feature correspondence and subsequent pose recognition. Accordingly, the pose of the object is derived based on the distortion and / or displacement in different shots or scene views of the scene containing the object. Summary of the Invention
[0003] The object of the present invention is to improve feature-based object recognition, in particular to automate feature extraction, and more particularly to automate feature correspondence.
[0004] The object of the present invention is achieved by a method having the features described in claim 1. The dependent claims give preferred extensions.
[0005] In one embodiment of the present invention, a method for automatic feature extraction and / or feature correspondence is provided.
[0006] In one embodiment, the method includes: taking an image from the perspective of a scene including at least one object by means of an imaging device. In one embodiment, the method includes: determining a descriptor image based on the taken image. In one embodiment, the determination of the descriptor image is based on an image, especially an image taken by the imaging device. In one embodiment, the determination of the descriptor image is based on a pre-trained artificial intelligence, or based on the following method, that is, using or by means of a trained artificial intelligence to especially determine or be able to determine the descriptor image. In one embodiment, the scene includes at least one object, especially at least one object for automatic feature extraction and / or feature correspondence, and more particularly at least one object that should be grasped or is to be grasped based on automated feature extraction and / or feature correspondence. In one embodiment, the perspective can be predetermined or specified in advance, especially in a fixedly installed imaging device, or the perspective can be at least substantially random or randomly occupied.
[0007] In one embodiment, the method particularly includes, before the above-mentioned method for automated feature extraction and / or feature correspondence, a step of training an artificial intelligence, which includes a step, particularly a first step, of capturing a first image from a first perspective of a scene using an imaging device, particularly a first imaging device. In one embodiment, the scene includes at least one object, particularly at least one object for automated feature extraction and / or feature correspondence. In one embodiment, the method includes a step, particularly a second step: capturing a second image from a second perspective of the scene using the first or the same imaging device and / or a second imaging device. In one embodiment, the method, particularly in a third step, includes: determining the correspondence between the pixels of the first image and the pixels of the second image, particularly based on the first perspective and the second perspective. In one embodiment, the perspective can be predetermined or determined in advance, particularly in the case where the imaging device is fixedly installed, or the perspective can be at least substantially random or randomly occupied. In one embodiment, particularly, the first perspective is made different from the second perspective. In one embodiment, the method, particularly in a fourth step, includes: determining the correspondence between the pixels of the first image and the pixels of the second image, particularly based on depth data, wherein in one embodiment, the first and / or second imaging device is designed to detect or capture depth data, particularly to determine depth data from the first image and / or the second image. In one embodiment, the method, particularly in a fifth step, includes: determining a descriptor image based on the first and / or second image, particularly based on the inference of the first and / or second image. In one embodiment, the descriptor image is determined based on an image, particularly the first image and / or the second image. In one embodiment, the image, particularly the first and / or second image, is a two-dimensional image. In one embodiment, the method, particularly in a sixth step, includes: determining at least one feature of at least one object based on the determined descriptor image.
[0008] The term "scene" as used herein should particularly be understood as a snapshot (Momentaufnahme) of the (relevant) environment, which includes a scene with at least one object, particularly in it. In one embodiment, the scene has: at least one object, which can particularly be in a container; and an environmental part of the container, particularly a robotic part, more particularly the part of the system including the robot. In one embodiment, the irrelevant parts of the scene can be (automatically) filtered out.
[0009] The term "perspective" as used herein should particularly be understood as the visual direction of the imaging device, which is particularly oriented such that at least part of the scene described herein is located in the field of view of the imaging device, particularly the entire scene, particularly when automated feature extraction and / or feature correspondence is applied to determine the pose and / or the grasping position in a (robotic) grasping application.
[0010] The term "descriptor image" as used herein should be understood in particular as a (learned) dense visual descriptor map that exploits the spatial R W×H×3 to map an image, in particular an RGB image, more particularly at full resolution, into the dense descriptor space R W×H×D , where in particular there is a D-dimensional descriptor vector (D) for each pixel; in one embodiment, W and H refer to the position of the pixel in the image, which includes the width ("width") and height ("height") of the corresponding pixel in the image, in particular in relation to the width and height of the image. In one embodiment, this is done pixel by pixel such that each pixel has in particular its own descriptor vector. In one embodiment, this (automated) process, in particular the determination of (at least one) descriptor vector, may be referred to as inference, or is inference.
[0011] The term "inference" as used herein should be understood in particular as "a conclusion automatically generated by a formal system", more particularly as an inference automatically generated by an inference engine (English: "inference engine", i.e., the system and its devices described hereinafter), which herein in particular includes the determination of correspondences and the derivation of descriptor vectors on a pixel-by-pixel basis.
[0012] Thereby, in one embodiment, automatic feature selection can be advantageously achieved, in particular based on the descriptor image, in particular (more) time-saving compared to manual feature selection. Furthermore, in one embodiment, automatic extraction and / or correspondence of features of multiple objects with a similar appearance, in particular of similar objects, can thus be advantageously achieved. Thereby, in one embodiment, it is possible to advantageously avoid defining new features for other, in particular similar, objects.
[0013] Advantageously, in one embodiment, the same or similar features of similar objects can be determined (more) simply based on the determined descriptor images of the respective objects, in particular by automatically or automatically finding the correspondences between the respective images, in particular by means of or based on the descriptor image. In one embodiment, if various features are pre-determined or selected based on a reference image, in particular after training a neural network (designed for this purpose), and in particular three-dimensional pose estimation or determination can be performed based on these features, and these are in particular different features arranged distributively with respect to the object, in particular with respect to the object of the reference image, then these features can be advantageously determined (automatically) in the image, in particular when dealing with objects that are different from the reference image, in particular similar objects.
[0014] In one embodiment, an image, particularly a first image and / or a second image, may be derived from or extracted from a video captured by a capturing device, particularly a first and / or a second capturing device, especially when the capturing device has changed its perspective with respect to the scene or the perspective has changed. In one embodiment, the image may be derived from or extracted from a video captured by the capturing device. In one embodiment, the first image may be derived from or extracted from a video captured by the first capturing device, and the second image may be derived from or extracted from a video captured by the second capturing device.
[0015] In one embodiment, the method further includes: determining the pose of at least one object, particularly based on at least three determined features of the object.
[0016] Advantageously, in one embodiment, the pose of at least one object can be determined in (less) time, particularly compared to corresponding to manual features.
[0017] In one embodiment, the method includes: determining a grasping position on at least one object based on the determined pose of the object. In one embodiment, the method includes: determining the grasping position of a grasping robot, particularly based on the determined pose of the object.
[0018] Accordingly, the grasping robot can grasp the object in the scene (more) quickly in an advantageous manner, or the grasping position of the grasping robot can be determined.
[0019] In one embodiment, a descriptor image is determined by means of at least one artificial neural network. In one embodiment, other steps of the method, particularly the step of determining the correspondence and / or the step of determining at least one feature, may be performed or executed by (an) artificial neural network.
[0020] Thereby, it can be advantageously allowed to identify or be able to identify the features of a reference object, particularly the features of a reference image, among similar objects, particularly (more) quickly, especially when there are multiple objects in the scene. In one embodiment, it can also be allowed thereby to automatically extract or identify and / or correspond to the selection of features of at least one object, particularly multiple objects, in the scene.
[0021] In one embodiment, determining the pose of at least one object is based on at least three predetermined features of a reference image, wherein in one embodiment, the features of the reference image are predetermined, and in one embodiment are predetermined according to the determined descriptor image of the reference object. In one embodiment, the reference image corresponds to the descriptor image of the reference object, in particular corresponds to having the descriptor image of the reference object. In one embodiment, at least three different features of the object (in the scene) correspond to at least three predetermined features of the reference image.
[0022] Thereby, in one embodiment, it is advantageously allowed to: determine or be able to determine the pose of at least one object (more) quickly and / or more robustly. In addition, in one embodiment, it is possible to omit determining the pose of at least one object by calculating, in particular manually calculating, the distortion and / or displacement based on a manually marked reference image. In one embodiment, it is possible to advantageously omit the statistical detection of the features of different (but of the same kind) objects, in particular in one or more reference images. In particular, in one embodiment, determining at least one feature is not limited to the same object and / or objects having the same appearance, but can be applied to objects of the same kind and, accordingly, (as described herein) determine the pose of at least one object, in particular the poses of multiple (of the same kind) objects. The term "of the same kind" used herein should be particularly applied or understood as the same species, genus, family, and / or order, etc., in particular similar to biology.
[0023] In one embodiment, if the scene has multiple objects, in particular objects of the same kind, the method includes the step of determining a probability, where the probability describes the similarity of the combination of the features determined in the scene with the features of the predetermined reference image and / or the reference object, in particular for determining the association of the features with a single object, or for excluding combinations of features that are distributed in a combined manner over multiple different objects. In one embodiment, the relative arrangement between the features is determined by the particularly predetermined features of the reference image and / or the particularly predetermined features of multiple reference images.
[0024] Thus, in one embodiment, objects in a scene with multiple objects, especially multiple homogeneous objects, can be identified or determined in a (more) robust and / or (more) rapid manner, and in one embodiment, the pose of an object can be identified or determined (more) robustly and / or (more) rapidly. In addition, in one embodiment, feature combinations that belong to not only one object or whose features are distributed over multiple different objects can be advantageously at least substantially excluded. Thus, in one embodiment, feature mis-correspondence to the objects in the scene can be at least substantially prevented according to probability, and in particular, the determined features can be made to automatically correspond to the objects (more) clearly. Without loss of generality, in one embodiment, the object, especially the object to be grasped, can be a fish, which is stored in a box together with other kinds of fish. According to the determined descriptor image, and especially based on the features pre-determined in the reference image, the same features of different fish can be identified or determined, especially at least substantially. In one embodiment, the feature can be identified or determined on the object independently of the number of objects. With the aid of probability, especially the probability described herein, it can be determined or ascertained whether the determined feature belongs to one object or different objects, for example, whether it belongs to one fish or multiple fish. In one embodiment, if the feature belongs to one fish, especially if the probability that the feature belongs to the object (here the fish) is greater than a pre-determined probability, the pose of the fish can be determined, more particularly by means of the determined and corresponding features. In one embodiment, according to the determined pose, a grasping pose or a grasping position matching the fish can be determined, and the fish can be grasped through this grasping pose or grasping position.
[0025] In one embodiment, the method includes: determining a grasping pose based on probability. Thus, a grasping pose matching the determined features can be advantageously found, especially compared with the method of determining the position based on a pre-determined, especially manually selected, grasping position on the object and according to the determined image.
[0026] Thus, in one embodiment, it can be advantageously allowed that: a (more) favorable grasping position, especially a grasping pose, can be determined or ascertained, which can especially match or be matched with the (multiple) homogeneous objects in the scene.
[0027] In one embodiment, the (first and / or second) imaging device is arranged on at least one robot, in particular on the flange of the robot. In one embodiment, in order to capture images, in particular a first image and a second image, the robot is moved and / or oriented, in particular between capturing the first image and capturing the second image, and more particularly in order to adjust the viewing angle, in particular a first viewing angle and / or a second viewing angle. In one embodiment, the method comprises the steps of moving and / or orienting the robot, in particular the flange of the robot, and more particularly the imaging device arranged on the flange of the robot.
[0028] Thereby, in one embodiment, it is advantageously possible to allow: to automatically achieve or capture or be able to capture the basis for pixel correspondence as described herein. In one embodiment, in particular even if the imaging device is not mounted on the robot but on a (different) predetermined point, this can still be carried out automatically, in particular by observing the scene, so that a first image with a first viewing angle and a second image with a second viewing angle, in particular different from the first viewing angle, can be captured.
[0029] Particularly preferably, the imaging device comprises an imaging device for capturing digital and / or two-dimensional, in particular three-dimensional images, and may in particular comprise at least one 2D camera, 3D camera and / or at least two cameras spaced apart from each other in space and / or at least one scanner, preferably for three-dimensional scanning. In one embodiment, the images described herein (respectively) have point clouds and / or color information, preferably three-dimensional point clouds, and more particularly point clouds with color information, and in particular the points belonging to the point cloud may in particular be the point cloud or be constituted by the point cloud. Correspondingly, in particular, the three-dimensional point cloud and / or color information captured by means of the imaging device is referred to as the (image captured by means of the imaging device), which in one embodiment is typically a three-dimensional image and / or an image with color information.
[0030] In one embodiment of the present invention, a system for automatic feature extraction and / or for operating a multi-axis machine, in particular a grasping robot, is provided. In one embodiment, the system is designed to perform the method described herein. In one embodiment, the system includes at least one robotic arm, in particular a gripper guided by the robotic arm. In one embodiment, the system includes an imaging device, in particular an imaging device as described herein. In one embodiment, the imaging device is arranged on the flange of at least one robotic arm. In one embodiment, the system and / or its device includes a device for determining a descriptor image, in particular based on (automatically determined) reasoning about an image captured by the imaging device, in particular based on a first image captured by a first imaging device and / or a second image captured by a second imaging device. In one embodiment, the system includes a device for determining the correspondence between pixels of a first image (captured by the imaging device) and pixels of a second image (captured by the imaging device), in particular when the system is designed or configured for (artificial) neural network training).
[0031] In an extended embodiment, the system and / or its device includes a device for determining the pose of at least one object. In an extended embodiment, the system and / or its device includes a device for determining a grasping position on at least one object, in particular based on the determined pose of the at least one object. In an extended embodiment, the system and / or its device includes a device for determining a probability.
[0032] In one embodiment of the present invention, a method for grasping an object using a grasping robot includes the step of determining a grasping position as described herein. Additionally, in one embodiment, the method includes: grasping the object using the grasping robot based on the (determined) grasping position.
[0033] Thereby, in one embodiment, it is advantageously possible to allow the grasping robot to grasp (more) quickly on the object or to perform the grasping (more) quickly, in particular compared to methods of the prior art.
[0034] The system and / or device in the sense of the present invention can be implemented in hardware technology and / or software technology, in particular having: at least one processing unit, preferably data-connected or signal-connected to a memory system and / or a bus system, in particular a digital processing unit, in particular a microprocessor unit (CPU), a graphics card (GPU), etc., and / or one or more programs or program modules. The processing unit can be designed for this purpose to: process the instructions of a program implemented as stored in the storage system; collect input signals from the data bus; and / or send output signals to the data bus. The storage system can have one or more, in particular different, storage media, in particular optical, magnetic, solid-state and / or other non-volatile media. The program can be provided such that it can embody or implement the method described herein, such that the processing unit can execute the steps of the method and thereby, in particular, can operate or monitor a multi-axis machine, in particular a gripping robot.
[0035] In one embodiment, the computer program product can have, in particular can be computer-readable and / or non-volatile, a storage medium for storing a program or instructions, or a storage medium having a program stored thereon or having instructions stored thereon. In one embodiment, the execution of the program or the instructions causes the system or controller, in particular a computer or an array of multiple computers, to execute the program or the instructions, such that the system or controller, in particular the one or more computers, executes the method described herein or one or more of its steps, or the program or instructions are designed for this purpose.
[0036] In one embodiment, one or more, in particular all, steps of the method are performed fully or partially automatically, in particular by a controller or its means. In one embodiment, the system includes a robot. Description of the Drawings
[0037] More advantages and features are given by the dependent claims and the embodiments. For this purpose, some are schematically shown:
[0038] Figure 1 : The system according to one embodiment of the present invention; and
[0039] Figure 2 : Multiple like objects in a scenario according to one embodiment;
[0040] Figure 3 : An object having the determined features according to one embodiment;
[0041] Figure 4 : Multiple like objects in a scenario, having the determined features, feature combinations and gripping positions;
[0042] Figure 5:Block diagram of a method according to one embodiment. Detailed implementation
[0043] Figure 1 A system 1 with a robot 2 is shown, and a gripper 3 is arranged on the flange of the robot. Also shown in Figure 1 is a photographing device 4, which has different perspectives on a scene 10 containing a plurality of objects 5 and is data-connected to the robot 2 through a processing unit. In one embodiment, the photographing device 4' is arranged on the robot 2, particularly on the robot arm, and more particularly on the flange of the robot arm (shown in dashed lines in Figure 1 ). The object 5 is shown in a container 6. The object 5 is detected by the photographing device 4, where the photographing device 4 takes an image from a perspective of the scene 10, particularly a first image from a first perspective of the scene 10 and a second image from a second perspective of the scene 10, and the second perspective is different from the first perspective. Based on the different perspectives, the pixels of one image are corresponded to the pixels of the second image, so that the correspondence between the first image and the second image can be determined, particularly the correspondence between their pixels, particularly the correspondence between each other, particularly during the training of an (artificial) neural network or for the training of an (artificial) neural network. In order to be able to automatically take the first image and the second image, when the photographing device 4 is installed or arranged on the flange of the robot, particularly on the flange of its robot arm, the photographing device is moved by the robot, particularly by the robot controller, or in another pose of the photographing device achieved by the movement of the robot and / or the robot arm, particularly taking the second image from another (second) perspective with respect to the scene 10. In one embodiment, a descriptor image can be determined according to the taken images, and in one embodiment, the descriptor image can be used for and / or is used to determine features.
[0044] Figure 2 The scene 10 with a plurality of homogeneous objects 5 is schematically shown, and these objects can be particularly placed in a container (not shown). These (homogeneous) objects 5 are represented by patterns representing descriptor images, where particularly each pixel corresponds to a feature vector. The regions with the same pattern exemplarily represent regions with the same features here. The demarcation between these regions is clear here; in an embodiment, this can include a particularly continuous transition between the regions, particularly, the descriptors in the regions may be (slightly) different from each other, but in Figure 2In the following figures, they are exemplarily assigned to a region. Thus, although the same type of objects 5 may differ especially in terms of size, length, width, height, etc. and / or similar features and / or the manifestation forms of individual features, they at least basically have the same features. Then, based on the determined descriptor image, features can be determined especially for each object, and these features at least basically have the same features in the descriptor image (refer to Figure 3 for the feature illustration).
[0045] Figure 3 Figure 5 is schematically shown, in which features M1 to M3 are determined, especially according to the features of a reference image with pre-determined features. Features M1 to M3 have especially characteristic attributes. In addition, the connections between features M1 to M3 are also schematically shown, especially their relative arrangements to each other. For example, when feature M1 and feature M3 are connected by a (virtual) line (shown by a dashed line), feature M2 is arranged deviating from this line (shown by a dashed line perpendicular to the connection line of M1 and M3). In addition, features M1 to M3 can also be directly connected by a (virtual) line. This is especially studied in Figure 4 . Figure 4 Figure 10 shows three objects 5 of the same type, which respectively exemplarily display or have the determined features M1 to M3. In addition, in Figure 4 , each feature M1 is connected to feature M2 and feature M3. W1 to W3 represent the respective connection probabilities of features M1 to M3, and these probabilities describe the similarity between the arrangements of features M1 to M3 in the reference image and the pre-determined features M1 to M1. It can be seen from this that the similarity of the arrangements of features M1 to M3 in the Figure 4 two objects 5 shown on the left is lower than that in the object shown on the right in the figure. From this, it can be especially inferred that: the feature combinations shown on the left (probably) do not belong to the feature combination M1 to M3 of a single object 5 because their similarity to the feature arrangement M1 to M3 of the reference object, especially the reference image, is smaller.
[0046] In addition, in Figure 4 , the grasping position G of a gripper especially for a grasping robot is also schematically shown, which is exemplarily indicated by three grippers. According to features M1 to M3, the pose of object 5 can be determined, and in one implementation, the (more) optimal grasping position G on the object can be determined according to this pose.
[0047] In Figure 5FIG. 0 schematically shows a block diagram of method 20, which shows the steps of method 20. Among them, S10 exemplarily describes taking a first image from a first perspective; S20 schematically describes taking a second image from a second perspective different from the first perspective; S30 schematically describes determining the correspondence between the pixels of the first image and the pixels of the second image; S40 schematically describes determining a descriptor image based on the captured first image, particularly based on reasoning; and S50 schematically describes determining at least one feature M1, M2, M3 based on the determined descriptor image. S20 and S30 are shown in dashed lines to schematically show that determining the descriptor image in S40 can be performed based on one, particularly just one, captured image, and S20 and S30 can be considered particularly in or for training an (artificial) neural network.
[0048] Although exemplary embodiments have been set forth in the foregoing description, it should be noted that there may still be many variations. It should also be noted that the exemplary embodiment is merely an example and should not form any limitation on the scope of protection, application, and construction. On the contrary, through the foregoing description, those skilled in the art can be taught to implement the conversion of at least one exemplary embodiment, wherein various changes, particularly regarding the functions and arrangements of the components, can be achieved without departing from the scope of protection of the present invention, for example, can be obtained according to the claims and their equivalent feature combinations.
[0049] List of reference numerals
[0050] 1 System
[0051] 2 Robot
[0052] 3 Gripper
[0053] 4 Imaging device
[0054] 5 Object
[0055] 6 Container
[0056] 7 Processing device
[0057] 10 Scene
[0058] 20 Method
[0059] M1, M2, M3 Determined features
[0060] W1, W2, W3 Determined probabilities
[0061] S10 Take the first image
[0062] S20 Take the second image
[0063] S30 Determine the correspondence
[0064] S40 Determine the descriptor image
[0065] S50 Determine at least one feature.
Claims
1. A method (20) for automatic feature extraction and / or feature correspondence, comprising the following steps: - Taking (S10) an image from the perspective of a scene (10) including at least one object (5) by means of an imaging device (4); - Determining (S40) a descriptor image based on the captured image; and - Determining (S50) at least one feature (M1, M2, M3) of the at least one object (5) based on the determined descriptor image.
2. The method (20) according to claim 1, wherein The method (20) includes: determining the pose of the at least one object (5) based on the at least one feature (M1, M2, M3).
3. The method (20) according to the preceding claim, characterized in that, Determining the pose (50) of the at least one object is based on the features (M1, M2, M3) of at least three predetermined reference images, wherein at least three different determined features (M1, M2, M3) of the object (5) are corresponded to the features (M1, M2, M3) of the at least three predetermined reference images.
4. The method (20) according to any one of the preceding claims, characterized in that, Determining the descriptor image by means of artificial intelligence, in particular by means of at least one artificial neural network.
5. The method (20) according to the preceding claim, characterized in that, If the scene has multiple, in particular homogeneous, objects (5), the method (20) further includes the step of determining probabilities (W1, W2, W3), wherein the probabilities (W1, W2, W3) describe the similarity between the combination of the determined features (M1, M2, M3) in the scene and the features (M1, M2, M3) of the predetermined reference images and / or reference objects.
6. The method (20) according to any one of the preceding claims, characterized in that, The imaging device, in particular the first and / or second imaging device (4, 4'), is arranged on at least one robot, in particular on the flange of the robot.
7. The method (20) according to any one of the preceding claims, characterized in that, The method (20) includes: determining a grasping position (G) on the at least one object (5), in particular determining a grasping position (G) for a grasping robot, in particular based on the determined pose of the object (5).
8. A method for grasping an object by a grasping robot, comprising the following steps: - Determining the grasping position (G) according to the foregoing claims; and - Grasping the object (5) by means of the grasping robot based on the grasping position (G).
9. A system for automatic feature extraction and / or for operating a multi-axis machine, in particular a grasping robot, the system being designed to perform the method according to any one of the foregoing claims, and / or comprising: - A robot arm and a gripper guided by the robot arm; - At least one imaging device, wherein the imaging device is in particular arranged on the flange of the robot arm; - A device for determining a descriptor image.
10. A computer program or computer program product, wherein, The computer program or computer program product contains instructions, in particular stored on a computer-readable and / or non-volatile storage medium, which when executed by one or more computers or by the system according to claim 9 cause the one or more computers or the system to perform the method according to any one of claims 1 to 7 and / or 8.