Object location detection using automatic feature extraction and / or feature assignment process
Patent Information
- Application Number
- EP2023809139
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-13
- Filing Date
- 2023-11-15
- Publication Date
- 2025-10-22
AI Technical Summary
Feature-based object recognition in robotics is inefficient due to the manual selection of features and the need for parameter tuning, which hinders automated feature extraction and assignment, especially in scenarios with multiple similar objects.
A method utilizing a trained artificial intelligence to generate a descriptor image from captured images, allowing for automated feature extraction and assignment by determining pixel descriptors and inferring feature vectors, enabling quick and robust pose estimation and grip position determination for gripping robots.
This approach automates feature selection and assignment, reducing time and error in recognizing and gripping similar objects, allowing for efficient and accurate robotic grasping without manual intervention.
Smart Images

Figure 1.1
Abstract
Description
[0001] Description
[0002] Object position detection with automated feature extraction and / or feature assignment
[0003] The present invention relates to a method for automated feature extraction and / or feature assignment, a system for automated feature extraction, and a computer program or computer program product.
[0004] Feature-based object recognition typically consists of two steps: A first step with feature extraction, which usually involves a manual selection of features such as circles, edges, corners, etc., and a selection or implementation of algorithms and / or methods for feature recognition, such as the Hough transform for circle detection, etc. For edge detection, for example, suitable parameters must be found, such as a Canny filter. A second step, which usually involves feature assignment followed by pose recognition, follows. Accordingly, the object's pose is derived from the distortion and / or displacement of the object in different images or perspectives of a scene containing the object.
[0005] The object of the present invention is in particular to improve feature-based object recognition, in particular to automate feature extraction, and further in particular to automate feature assignment.
[0006] This object is achieved by a method having the features of claim 1. The subclaims relate to advantageous developments.
[0007] In one embodiment of the present invention, a method for automated feature extraction and / or feature assignment is provided.
[0008] In one embodiment, the method comprises capturing an image with a perspective of a scene with at least one object using a capturing device. In one embodiment, the method comprises determining a descriptor image based on the captured image. In one embodiment, the determination of the descriptor image is based on the image captured, in particular by the capturing device. In one embodiment, the determination of the descriptor image is based on a previously trained artificial intelligence or on a method that uses a trained artificial intelligence to determine, in particular, a descriptor image or with the aid of which the descriptor image is or can be determined.In one embodiment, the scene comprises at least one object, in particular at least one object for automated feature extraction and / or feature assignment, further in particular at least one object that is to be grasped or is grasped based on the automated feature extraction and / or feature assignment. In one embodiment, a perspective can be predetermined or predefined, particularly in the case of permanently installed recording device(s), or a perspective can be, at least substantially, randomized or adopted in a randomized manner.
[0009] In one embodiment, the method, in particular upstream of the method for automated feature extraction and / or feature assignment, in particular as described above, comprises training an artificial intelligence with a step, in particular a first step, of capturing a first image with a first perspective of a scene using a, in particular first, recording device. In one embodiment, the scene comprises at least one object, in particular at least one object for the automated feature extraction and / or feature assignment. In one embodiment, the method has a step, in particular a second step, of capturing a second image with a second perspective of the scene using the first or the same recording device and / or a second recording device.In one embodiment, the method comprises, in particular in a third step, determining an assignment of pixels of the first image to pixels of the second image, in particular based on the first perspective and the second perspective. In one embodiment, a perspective can be predetermined or set in advance, in particular in the case of permanently installed recording devices, or a perspective can be, at least substantially, randomized or adopted in a randomized manner, in particular such that, in one embodiment, the first perspective differs from the second perspective. In one embodiment, the method comprises, in particular in a fourth step, determining an assignment of pixels of the first image to pixels of the second image, in particular based on depth data, wherein the first and / or the second recording device is configured, in one embodiment, to acquire or record depth data.to record, in particular depth data is determined from the first image and / or the second image. In one embodiment, the method comprises, in particular in a fifth step, determining a descriptor image based on, in particular an inference, of the first and / or the second image. In one embodiment, the determination of the descriptor image is based on an image, in particular the first image and / or the second image. In one embodiment, the image, in particular the first and / or the second image, is a two-dimensional image. In one embodiment, the method comprises, in particular in a sixth step, determining at least one feature of the at least one object based on the determined descriptor image.
[0010] The term "scene," as used herein, should be understood in particular as a snapshot of a (relevant) environment, which comprises a scenery with at least one object, in particular in a container. In one embodiment, the scene comprises the at least one object, which may in particular be in a container, and parts of the container's environment, in particular parts of a robot, further in particular parts of a system that includes the robot. In one embodiment, irrelevant parts of the scene can be (automatically) filtered out.
[0011] The term “perspective” as used herein should be understood in particular as the viewing direction of the recording device, which is in particular oriented such that the scene, as described herein, is at least partially in the field of view of the recording device, in particular the entire scene, in particular in applications of automated feature extraction and / or feature assignment for determining a pose and / or a grip position in (robotic) gripping applications.
[0012] The term “descriptor image” as used herein is intended in particular to be understood as a (learned) dense visual descriptor map representing an image, in particular an RGB image, further in particular with full resolution, with space R WxHx3 , to a dense descriptor space, R WxHxD, wherein in particular a D-dimensional descriptor vector (D) is present for each pixel; W and H refer in one embodiment to a position of the pixel in the image with a width (“width”) and a height (“height”) of the respective pixel in the image, in particular related to the width and height of the image. In one embodiment this is done pixel by pixel, so that each pixel in particular has its own descriptor vector. This (automated) procedure, in particular determining the (at least one) descriptor vector, can be referred to as inference in one embodiment or is an inference.
[0013] The term "inference" as used herein shall be understood in particular as "inference automatically generated from a formal system", further in particular as conclusion(s) drawn automatically by an inference engine, a system described below and / or its means, which includes in particular the pixel-by-pixel assignment and the derivation of the descriptor vector.
[0014] This advantageously allows, in one embodiment, an automated feature selection to be implemented, in particular based on the descriptor image, which is less time-consuming, particularly compared to manual feature selection. Furthermore, in one embodiment, features of multiple objects with a similar appearance, in particular of similar objects, can be extracted and / or assigned automatically. In one embodiment, this advantageously eliminates the need to define new features of another, in particular similar, object.
[0015] In one embodiment, identical or similar features of similar objects can advantageously be determined more easily on the basis of determined descriptor image(s) for the respective objects, in particular correspondences between the respective images can be found, in particular automatically or automatically, in particular with the aid of or based on the descriptor image(s). If, in one embodiment, different features are predetermined or selected on the basis of a reference image, in particular after (completed) training of a neural network set up (for this purpose), in particular in such a way that a, in particular three-dimensional, pose estimation or determination based on the features is possible, and these, in particular different features are arranged distributed over the object, in particular over the object of the reference image, these orThese are advantageously determined in the image (automatically), especially if the object is different from the reference image, but especially if it is of the same type.
[0016] In one embodiment, the image, in particular the first image and / or the second image, can originate from or be extracted from a video recorded by the recording device, in particular the first and / or second recording device, in particular if the recording device has changed its perspective on the scene or the perspective has been changed. In one embodiment, the image can originate from or be extracted from a video recorded by a recording device. In one embodiment, the first image can originate from a video recorded by a first recording device and the second image can originate from or be extracted from a video recorded by a second recording device.
[0017] In one embodiment, the method further comprises determining a pose of the at least one object, in particular based on at least three determined features of the object.
[0018] In one embodiment, this advantageously allows a pose of the at least one object to be determined with less time expenditure than, in particular with manual feature assignment.
[0019] In one embodiment, the method comprises determining a grip position on the at least one object based on a determined pose of the object. In one embodiment, the method comprises determining a grip position for a gripping robot, in particular based on a determined pose of the object.
[0020] This advantageously allows an object in the scene to be grasped more quickly by a gripping robot or a gripping position for the gripping robot to be determined.
[0021] In one embodiment, the descriptor image is determined using at least one artificial neural network. In one embodiment, further steps of the method, in particular determining an assignment and / or determining at least one feature, can be performed using or are performed by the (artificial) neural network.
[0022] This advantageously makes it possible for features of a reference object, in particular features of a reference image, to be recognized or recognizable more quickly in similar objects, especially when multiple objects are present in the scene. In one embodiment, this also makes it possible for a selection of features on the at least one object, in particular on multiple objects, in a scene to be automatically extracted or recognized and / or assigned or to be assigned.
[0023] In one embodiment, the determination of the pose of the at least one object is based on at least three predetermined features of a reference image, wherein the features of the reference image, in one embodiment, are predetermined based on a determined descriptor image of a reference object. In one embodiment, the reference image corresponds to a descriptor image of a reference object, in particular a descriptor image that has a reference object. In one embodiment, at least three different features of the object (in the scene) are assigned to the at least three predetermined features of the reference image.
[0024] In one embodiment, this advantageously makes it possible to determine the pose of the at least one object more quickly and / or more robustly. Furthermore, in one embodiment, the pose of the at least one object can be determined by calculating a distortion and / or a displacement from a manually marked reference image, in particular manually. Advantageously, in one embodiment, statistical recording of features of different (but similar) objects, in particular in one reference image or multiple reference images, can be dispensed with. In particular, in one embodiment, the determination of at least one feature is not limited to identical objects and / or objects with the same appearance, but can be applied to similar objects and accordingly (as described herein) a pose of the at least one object, in particular of the plurality of (similar) objects, can be determined.“Similar”, as used herein, shall be used or understood in particular as being of the same species, genus, family and / or order or the like, particularly by analogy with biology.
[0025] In one embodiment, if the scene has a plurality of, in particular similar, objects, the method comprises a step of determining a probability, wherein the probability describes a similarity of a combination of determined features in the scene to the predetermined features of the reference image and / or the reference object, in particular in order to determine the affiliation of the features to a single object or to exclude combinations of features which, in combination, are distributed across a plurality of different objects. In one embodiment, the relative arrangement of the features to one another is determined from the, in particular predetermined, features of the reference image and / or a plurality of, in particular predetermined, features of a plurality of reference images.
[0026] In one embodiment, this advantageously allows an object in a scene with multiple objects, in particular with multiple similar objects, to be detected or determined more robustly and / or quickly, and in one embodiment, the pose of the object can be detected or determined more robustly and / or quickly. Furthermore, in one embodiment, feature combinations that do not belong to just one object or in which features are distributed across multiple different objects can advantageously be excluded, at least substantially. In one embodiment, this makes it possible to at least substantially prevent an incorrect assignment of features to an object in the scene based on the probability, in particular, determined features can be assigned to an object more clearly and automatically.Without limiting the generality, in one embodiment an object, in particular an object to be grasped, can be a fish that is stored in a box with other fish of a different fish species. Using a determined descriptor image and in particular based on predetermined features in a reference image, identical features of the different fish can be recognized or determined, in particular at least substantially. In one embodiment the features are recognized or determined on the objects independently of the plurality of objects. By means of a probability, in particular as described herein, it can be determined or is determined whether the determined features belong to one object or to different objects, in the example, to one or more of the fish.If, in one embodiment, the features in the example are assigned to a fish, in particular if the probability that the features belong to the object, in this case a fish, is greater than a predetermined probability, a pose of the fish can be determined, in particular using the determined and assigned features. Based on the determined pose, in one embodiment, a grip pose or grip position tailored to the fish can be determined, with which the fish can be or is grasped.
[0027] In one embodiment, the method comprises determining a gripping pose based on the probability. This advantageously allows a gripping pose to be found that is tailored to the determined features, particularly in comparison with methods that are based on a predetermined, in particular manually selected, gripping position on the object and determine this position based on acquired images.
[0028] In one embodiment, this advantageously makes it possible to determine or is determined a more advantageous grip position, in particular a grip pose, which can be or is adjusted in particular to (several) similar objects in a scene.
[0029] In one embodiment, the (first and / or second) recording device is arranged on the at least one robot, in particular on a flange of the robot. In one embodiment, the robot is moved and / or aligned to record the image, in particular the first image and the second image, in particular between recording the first image and recording the second image, further in particular to adjust the, in particular first and / or second, perspective. In one embodiment, the method comprises a step of moving and / or aligning the robot, in particular the flange of the robot, further in particular the recording device arranged on the flange of the robot.
[0030] This advantageously makes it possible, in one embodiment, for a basis for the assignment of pixels, as described herein, to be carried out or recorded automatically. In one embodiment, this can be carried out automatically, in particular, even if the recording device is not attached to the robot, but rather, in one embodiment, at (different) predetermined points, in particular with a view of the scene, so that a first image can be or is recorded from a first perspective and a second image from a second perspective, in particular one different from the first perspective.
[0031] The recording device particularly preferably comprises a recording device for recording digital and / or two-dimensional, in particular three-dimensional, images. In particular, it can have at least one 2D camera, 3D camera and / or at least two spatially spaced cameras and / or at least one scanner, preferably for three-dimensional scanning. In one embodiment, an image mentioned here (in each case) has a point cloud and / or color information, preferably a three-dimensional point cloud, further in particular a point cloud with color information, in particular assigned to the points of the point cloud, and can in particular be such or consist of such.Accordingly, in particular a three-dimensional point cloud and / or color information recorded with the aid of a recording device is referred to as an image (recorded with the aid of the recording device), which in one embodiment is generally a three-dimensional image and / or an image with color information.
[0032] In one embodiment of the present invention, a system for automated feature extraction and / or for operating a multi-axis machine, in particular a gripper robot, is provided. In one embodiment, the system is configured to carry out a method described herein. In one embodiment, the system comprises at least one robot arm, in particular a gripper guided by a robot arm. In one embodiment, the system comprises a receiving device, in particular a receiving device as described herein. In one embodiment, the receiving device is arranged on a flange of the at least one robot arm.In one embodiment, the system and / or its means comprise means for determining a descriptor image, in particular based on an (automatically determined) inference of the image recorded by the recording device, in particular based on the first image recorded by the first recording device and / or the second image recorded by the second recording device. In one embodiment, the system comprises means for determining an assignment of pixels of a first image (recorded by the recording device) to pixels of a second image (recorded by the recording device), in particular if the system is configured or designed to train an (artificial) neural network, as described herein.
[0033] In one development, the system and / or its means comprise means for determining a pose of the at least one object. In one development, the system and / or its means comprise means for determining a grip position on the at least one object, in particular based on a determined pose of the at least one object. In one development, the system and / or its means comprise means for determining a probability.
[0034] In one embodiment of the present invention, a method for gripping an object with a gripping robot comprises a step of determining a gripping position, as described herein. Furthermore, in one embodiment, the method comprises gripping the object with the gripping robot based on the (determined) gripping position.
[0035] Advantageously, in one embodiment, this makes it possible for the gripping robot to find a grip on the object more quickly or to carry out this grip more quickly, in particular in comparison to prior art methods.
[0036] A system and / or means within the meaning of the present invention can be designed in hardware and / or software, in particular at least one, in particular digital, processing unit, in particular a microprocessor unit (CPU), graphics card (GPU) or the like, preferably connected to a memory and / or bus system for data or signals, and / or one or more programs or program modules. The processing unit can be designed to execute instructions implemented as a program stored in a memory system, to detect input signals from a data bus, and / or to output signals to a data bus. A memory system can have one or more, in particular different, storage media, in particular optical, magnetic, solid-state, and / or other non-volatile media. The program can be designed in such a way that it embodies the methods described here oris capable of carrying out, so that the processing unit can carry out the steps of such methods and thus in particular can operate or monitor the multi-axis machine, in particular the gripping robot.
[0037] In one embodiment, a computer program product can comprise, in particular be, a storage medium, in particular a computer-readable and / or non-volatile one, for storing a program or instructions or with a program or instructions stored thereon. In one embodiment, execution of this program or these instructions by a system or a controller, in particular a computer or an arrangement of multiple computers, causes the system or the controller, in particular the computer(s), to carry out a method described here or one or more of its steps, or the program or the instructions are configured to do so.
[0038] In one embodiment, one or more, in particular all, steps of the method are carried out fully or partially automatically, in particular by the controller or its means. In one embodiment, the system comprises the robot.
[0039] Further advantages and features emerge from the subclaims and the exemplary embodiments. The following shows, partly schematically:
[0040] Fig. 1 : a system according to an embodiment of the present invention; and
[0041] Fig. 2: several similar objects in a scene after an execution;
[0042] Fig. 3: an object with determined features after an execution;
[0043] Fig. 4: Several similar objects in a scene with determined features, combinations of features, and a grip position; Fig. 5: A block diagram of a method according to one embodiment.
[0044] Fig. 1 shows a system 1 with a robot 2, on whose flange a gripper 3 is arranged. Furthermore, Figure 1 shows a recording device 4, which has different perspectives on the scene 10 with several objects 5, and which is data-connected to the robot 2 via a processing unit. In one embodiment, the recording device 4' is arranged on the robot 2, in particular on the robot arm, further in particular on the flange of the robot arm (shown in dashed lines in Figure 1). The objects 5 are shown in a container 6. The objects 5 are captured by the recording device 4 in that the recording device 4 records an image from one perspective on the scene 10, in particular a first image from a first perspective on the scene 10 and a second image with a second perspective, which differs from the first perspective, on the scene 10.Based on the different perspectives, pixels of one image are assigned to pixels of the second image, so that in particular an assignment can be determined between the first image and the second image, in particular their pixels, in particular to one another, in particular during or for training an (artificial) neural network. In order to automatically capture a first image and a second image, if the recording device 4 is attached or arranged on the flange of the robot, in particular on the flange of its robot arm, it can be moved by the robot, in particular by a robot controller, or capture a second image in a different pose of the recording device achieved by the movement of the robot and / or the robot arm, in particular with a different (second) perspective on the scene 10.In one embodiment, a descriptor image can be determined from the recorded image, which descriptor image can be used or can be used and / or is used in one embodiment to determine features.
[0045] Figure 2 schematically shows a scene 10 with several similar objects 5, which may in particular be in a container (not shown). The (similar) objects 5 are represented with patterns that are exemplary of the descriptor image, in which, in particular, each pixel is assigned a feature vector. Areas with the same pattern are intended to indicate, by way of example, areas with the same features. A separation between the areas is sharply delineated here; in embodiments, this can comprise a, in particular continuous, transition between the areas. In particular, descriptors in the areas can differ (slightly) from one another, but in Figure 2 and the following figures are assigned to an area by way of example. Thus, similar objects 5, which differ in particular in size, length, width, height and / or the like and / or the characteristics of individual features, can have at least substantially the same features.Based on the determined descriptor image, a feature can then be determined, in particular on each object 5, which has, at least essentially, the same features in the descriptor image (reference is made to a representation of features in Figure 3).
[0046] Figure 3 schematically shows an object 5 in which features M1 to M3 were determined, in particular based on features of a reference image with predetermined features. The features M1 to M3 have, in particular, characteristic properties. Furthermore, a relationship, in particular a relative arrangement to one another, between the features M1 to M3 is shown schematically. For example, if feature M1 and feature M3 are connected by a (virtual) line (shown in dashed lines), feature M2 is arranged offset from this line (shown with a dashed line that is perpendicular to the connecting line between M1 and M3). Furthermore, the features M1 to M3 can also be directly connected via (virtual) lines. This can be taken up in particular in Fig. 4, which shows three similar objects 5, each of which shows or has determined features M1 to M3 by way of example. Furthermore, in Figure 4, a feature M1 is each connected to a feature M2 and a feature M3.W1 to W3 denote probabilities of the respective connections of the features M1 to M3, which describe a similarity to the arrangement of the features M1 to M3 in a reference image with predetermined features M1 to M1. From this, it can be seen that the similarity of the arrangement of the features M1 to M3 is lower for the two objects 5 shown on the left side of Figure 4 than for the object shown on the right in the figure. From this, it can be deduced in particular that the feature combinations shown on the left (probably) do not belong to the feature combination M1 to M3 of an individual object 5, because a similarity to the feature arrangement M1 to M3 of a reference object, in particular a reference image, is lower. Furthermore, Figure 4 schematically shows a gripping position G, in particular for a gripper of a gripping robot, which is shown as an example with three gripping fingers.Based on the features M1 to M3, a pose of the object 5 can be determined, which can be used to determine an optimal grip position G on the object in one embodiment.
[0047] Figure 5 schematically shows a block diagram of a method 20 that shows the steps of the method 20, wherein S10 exemplifies the recording of a first image from a first perspective, S20 the recording of a second image from a second perspective that differs from the first perspective, S30 determines an assignment of pixels of the first image to pixels of the second image, S40 determines a descriptor image based on the recorded first image, in particular based on an inference, and S50 determines at least one feature M1, M2, M3 based on the determined descriptor image. S20 and S30 are dashed to schematically show that determining S40 a descriptor image is or can be based on one, in particular exactly one, recorded image and in particular S20 and S30 are used in or for training an (artificial) neural network.can be.
[0048] Although exemplary embodiments have been explained in the preceding description, it should be noted that a multitude of modifications are possible. Furthermore, it should be noted that the exemplary embodiments are merely examples and are not intended to limit the scope of protection, applications, or structure in any way. Rather, the preceding description provides the skilled person with a guide for implementing at least one exemplary embodiment, whereby various changes, particularly with regard to the function and arrangement of the described components, can be made without departing from the scope of protection as it results from the claims and equivalent combinations of features.
[0049] 1 system
[0050] 2 robots 3 grippers
[0051] 4 Mounting device
[0052] 5 Object
[0053] 6 containers
[0054] 7 Processing device 10 Scene
[0055] 20 procedures
[0056] M1, M2, M3 determined characteristics
[0057] W1 , W2, W3 determined probabilities
[0058] S10 Taking a first picture
[0059] S20 Taking a second picture
[0060] S30 Determining an assignment
[0061] S40 Determining a descriptor image
[0062] S50 Determination of at least one characteristic
Claims
Patent claims Method (20) for automated feature extraction and / or feature assignment, comprising the steps: - capturing (S10) an image with a perspective of a scene (10) with at least one object (5) using a recording device (4); - determining (S40) a descriptor image based on the captured image; and - Determining (S50) at least one feature (M1, M2, M3) of the at least one object (5) based on the determined descriptor image. Method (20) according to claim 1, characterized in that the method (20) comprises determining a pose of the at least one object (5) based on the at least one feature (M1, M2, M3). Method (20) according to the preceding claim, characterized in that determining the pose of the at least one object (50) is based on at least three predetermined features (M1, M2, M3) of a reference image, wherein at least three different determined features (M1, M2, M3) of the object (5) are assigned to the at least three predetermined features (M1, M2, M3) of the reference image. Method (20) according to one of the preceding claims, characterized in that the descriptor image is determined by means of artificial intelligence, in particular by means of at least one artificial neural network.Method (20) according to the preceding claim, characterized in that the method (20), if the scene has a plurality of, in particular similar, objects (5), further comprises a step of determining a probability (W1, W2, W3), wherein the probability (W1, W2, W3) describes a similarity of a combination of determined features (M1, M2, M3) in the scene to the predetermined features (M1, M2, M3) of the reference image and / or a reference object.
6. Method according to one of the preceding claims, characterized in that the, in particular first and / or second, receiving device (4, 4') is arranged on the at least one robot, in particular on a flange of the robot.
7. Method (20) according to one of the preceding claims, characterized in that the method (20) comprises determining a grip position (G) on the at least one object (5), in particular a grip position (G) for a gripping robot, in particular based on the determined pose of the object (5).
8. Method for gripping an object with a gripping robot comprising the steps: - Determining a grip position (G) according to the preceding claim; and - Gripping the object (5) with the gripping robot based on the grip position (G).
9. System for automated feature extraction and / or for operating a multi-axis machine, in particular a gripping robot, which is designed to carry out a method according to one of the preceding claims and / or comprises: - a robot arm and a gripper guided by the robot arm; - at least one receiving device, wherein the receiving device is arranged in particular on a flange of the robot arm; - Means for determining a descriptor image.
10. A computer program or computer program product, wherein the computer program or computer program product contains instructions, in particular stored on a computer-readable and / or non-volatile storage medium, which, when executed by one or more computers or a system according to claim 9, cause the computer(s) or the system to carry out a method according to one of claims 1 to 7 and / or 8.