Method and device for automatically locating objects suitable for removal from an object cluster

EP4655753A1Pending Publication Date: 2025-12-03ISRA VISION GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2023825411
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-24
Filing Date
2023-12-08
Publication Date
2025-12-03

AI Technical Summary

Technical Problem

Existing methods for automatic localization of objects in object accumulations are complex and unreliable, particularly when objects are stacked or arranged randomly, as they require multiple camera recordings and extensive data processing, making them inefficient for real-time removal or manipulation.

Method used

A method using a trained artificial neural network (NN algorithm) that processes a height map of the object accumulation to identify and determine the 3D position and orientation of objects, allowing for the control of gripping and transport elements or robot movements to remove or manipulate objects, with the NN algorithm being trained on specific object and container models.

Benefits of technology

The method enables robust, reliable, and fast identification and removal of objects from object accumulations, requiring only a height map for input, and can adapt to various object types, container shapes, and movements, improving the efficiency and accuracy of object localization and manipulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2023084983_02082024_PF_FP
    Figure EP2023084983_02082024_PF_FP
Patent Text Reader

Abstract

The invention relates to a safe and reliable method for automatically locating objects in an observed area of an object cluster in an open container, consisting of a plurality of objects of at least one object type, comprising the following steps: receiving a data model of the at least one object type, e.g. in the form of a coloured coordinate model; receiving a trained NN algorithm, wherein the NN algorithm is trained based on the data model of the at least one object type; receiving an elevation map of the observed area of the object cluster; identifying a plurality of objects lying at the top of the observed area of the object cluster and determining the three-dimensional position and orientation of these objects based on the NN algorithm trained in relation to the at least one object type using the data model of the at least one object type and exclusively using the elevation map of the observed area of the object cluster with regard to the observed area; and providing the three-dimensional position and orientation of the plurality of objects lying at the top of the observed area of the object cluster at an interface of the data processing unit. The invention also relates to a corresponding device and a system for controlling a gripper and / or transport element.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method and device for the automatic localization of objects suitable for removal from an object accumulation

[0002] The invention relates to a method and a device for the automatic localization of objects from an object accumulation consisting of a plurality of objects of at least one object type arranged one above the other and / or next to one another, for example for removing objects from an object accumulation present in an open container, using an artificial neural network, a system for the corresponding control of a gripping and / or transport element for removing or manipulating such an object from the object accumulation or for controlling a robot element for assuming a predetermined position relative to an object from the object accumulation, and the corresponding training of the artificial neural network.

[0003] The automatic identification and localization of objects from a cluster of objects arranged in a container represents an important application of robotics. For example, it is desired that components that have been randomly poured into an open-top component container or stacked in a predetermined manner can be automatically removed individually using a gripper and transport arm and transported to a predetermined location (e.g., a conveyor belt or a machine) where they can then be reused. The removal of objects from the container is also referred to as bin picking.

[0004] The identification and localization of objects present in an accumulation of objects is also of interest for line tracking of objects or visual servoing. In visual servoing, the movement of a robot element relative to a moving object is controlled based on information obtained from an image sensor (visual feedback). In contrast, in line tracking, a moving object is observed using optical means and a movable gripping and / or transport element is controlled accordingly to manipulate the object. The automatic detection of 3-dimensional objects using artificial intelligence and specifically using artificial neural networks (hereinafter referred to as NN algorithms) has already been described many times. Documents DE 10 2022 107 311 A1 and DE 10 2022 107 228 A1 disclose methods in which objects are detected in a container and removed from the container.The comparatively complicated methods based on the use of 2D RGB color images of the object cluster, which are subjected to an image segmentation process using a neural network.

[0005] Based on the above prior art, the object is to provide a simple, safe, and reliable method for automatically locating objects from an object cluster consisting of a plurality of objects of at least one object type arranged one above the other and / or next to one another, and to create a device that implements such a method. Furthermore, the object is to create a safe and reliable system for controlling the movement of a gripping and / or transport element that removes or manipulates an object from such an object cluster, or for controlling the movement of a robot element so that it assumes a predetermined position relative to an object from the object cluster.

[0006] The above object is achieved by methods for the automatic localization of objects from an object accumulation consisting of a plurality of objects of at least one object type arranged one above the other and / or next to one another, with the features of claim 1, a corresponding device with the features of claim 9, a method and a system for controlling a movable gripping and / or transport element for removing or manipulating an object from the object accumulation with the features of claims 5, 7, 12, 14 or a method and a system for controlling a movable robot element so that it assumes a predetermined position relative to an object of the object accumulation, with the features of claim 8,15 and by a method for the machine training of an NN algorithm for the identification and localization of objects suitable for removal from an accumulation of objects present in an open container by means of a data processing unit having the features of claim 16.

[0007] In particular, the above object is achieved by a method for the automatic localization of objects of an observed area of ​​an object accumulation consisting of a plurality of objects of at least one object type arranged one above the other and / or next to each other, comprising the following steps:

[0008] • Receiving a data model of the at least one object type, for example in the form of a colored coordinate model,

[0009] • Receiving an NN algorithm trained for the specific at least one object type, wherein the NN algorithm is trained based on the data model of the at least one object type,

[0010] • Receiving a height map of the observed area of ​​object accumulation,

[0011] • Identifying a plurality of objects arranged in the observed area of ​​the object accumulation and determining the 3-dimensional position and orientation of this plurality of objects in a predetermined coordinate system by means of a data processing unit based on the NN algorithm trained with respect to the at least one object type using the data model of the at least one object type and using exclusively the height map of the observed area of ​​the object accumulation with respect to the observed area,

[0012] • Providing the 3-dimensional position and orientation of the plurality of objects arranged in the observed area of ​​the object accumulation at an interface of the data processing unit.

[0013] The data model of at least one object type and the NN algorithm trained for that at least one object type are generally received once at the beginning of the process's use in a machine or on a conveyor belt. If necessary, the NN algorithm is updated from time to time during further use of the process (e.g., after further training), or the data model and the NN algorithm are supplemented / modified when the object types change. The elevation map of the observed area of ​​the object accumulation, on the other hand, is received repeatedly—as shown below, for example, at predetermined times after it has been determined and transmitted, for example, using a 3D scanner.Accordingly, on the basis of each height map, the plurality of objects arranged in the area of ​​the object accumulation is identified and their 3-dimensional position and orientation are determined and provided at the interface of the data processing unit.

[0014] In one embodiment, the above method is used for the automatic localization of objects suitable for removal from an observed area of ​​an object accumulation present in an open container consisting of a plurality of objects of at least one object type lying one above the other, with the following steps:

[0015] • Receiving a data model of the at least one object type, for example in the form of a colored coordinate model,

[0016] • Receiving an NN algorithm trained for the specific at least one object type, wherein the NN algorithm is trained based on the data model of the at least one object type,

[0017] • Receiving a height map of the observed area of ​​object accumulation,

[0018] • Identifying a plurality of objects lying on top of the observed area of ​​the object accumulation and determining the 3-dimensional position and orientation of this plurality of objects lying on top of the object accumulation in a predetermined coordinate system by means of a data processing unit based on the NN algorithm trained with respect to the at least one object type using the data model of the at least one object type and using exclusively the height map of the observed area of ​​the object accumulation with respect to the observed area, providing the 3-dimensional position and orientation of the plurality of objects lying on top of the observed area of ​​the object accumulation at an interface of the data processing unit.

[0019] The above method is usually realized as a computer-implemented method by means of a data processing unit (e.g. a computer) which has a processor and a memory unit readable by the processor.

[0020] The method is intended to identify 3-dimensional objects (e.g., components) arranged in an observed area (a section) of an object cluster, for example, in an object cluster accessible from one side (e.g., from above) in a container, and to determine their position and orientation (pose) in a predetermined coordinate system (e.g., the coordinate system of a movable gripping and / or transport element, a movable robot element, or a 3D scanner) (localization). A control device of the movable gripping and / or transport element or of the movable robot element can receive this data and determine control signals, which it then uses to remove or manipulate the identified object from the container and, if necessary, transport it to a location, or to move the movable robot element relatively.

[0021] The method includes the steps of transmitting a data model of at least one object type (i.e., possibly multiple object types) and, if applicable, a data model of the container, for example data from a CAD model. In addition, a specific NN algorithm trained for the at least one object type and, if applicable, the respective container is received. If there are multiple object types, an object-type-specific NN algorithm can be used for each object type (i.e., for example, 2 or 3 NN algorithms for 2 or 3 object types) or a single NN algorithm trained jointly for the respective object types can be used. This data is stored in the memory unit and made available for the respective processing of the method. The data model of an object type or, if applicable, of the container is the software-based implementation of the dimensions of a (virtual) object or container.When using several different objects (e.g. components), the object accumulation includes in particular one or more objects of each object type, although cases may also occur in which one or more object types are not present.

[0022] Furthermore, a 3D scanner (which, for example, has at least two cameras with appropriate data processing, which view the observed area from different viewing directions) can be used to determine a height map of an observed area of ​​a current object cluster. The height map contains a plurality of pixels in two dimensions (x / y), with the position of the respective point on the surface of the object cluster and the distance from the 3D scanner encoded as a corresponding gray value for each pixel. For example, pixels without an object or with the greatest distance from the 3D scanner can be shown in white, and object pixels with the smallest distance from the 3D scanner can be shown in black in the height map. Such a height map is shown, for example, in Fig. 20 and provided with the reference numeral 91.

[0023] Additionally, in one embodiment of the method, a data model of the container in which the objects are typically arranged, for example in the form of a colored coordinate model, can also be included and received accordingly. Accordingly, the NN algorithm can additionally be trained based on the data model of the at least one container. In this embodiment, this NN algorithm, which is trained on the at least one object type and the at least one container, can be used as a basis for identifying the plurality of objects lying on top of the observed area of ​​the object accumulation and for determining the 3-dimensional position and orientation of this plurality of objects lying on top of the object accumulation in the predetermined coordinate system by means of the data processing unit.

[0024] According to the method according to the invention, the data model of the at least one object type and, if applicable, the container, as well as the currently determined height map of the observed area are used in the NN algorithm to identify a plurality of objects arranged in the observed area of ​​the object cluster, e.g., objects lying on top of the object cluster (whereby the identification may include the recognition of the respective object type if different object types may be present in the object cluster) and to determine the position and orientation of the identified objects. Further camera recordings or other data of the observed area are not required.With regard to the observed area, only the height map of the respective observed area is required to identify the multitude of objects in the observed area and their 3D position and orientation as a "measurement parameter" or measured / current input. With a continuous removal of objects from a container / manipulation of an object from an object cluster / movement of a robot element relative to an object from an object cluster, only the current height map of the observed area needs to be generated as input for the NN algorithm, possibly repeatedly at specified times.Objects (for example, the removable objects located on the top) can be continuously identified using the height map and the method according to the invention, and their 3-dimensional position (in the specified coordinate system) and their 3-dimensional orientation can be determined easily and quickly (the terms 3-dimensional position and orientation include the 3-dimensional position and orientation in 3 dimensions). This enables the gripping and / or transport element to move in such a way that the gripping and / or transport element picks up an object and transports it to the desired location or manipulates the object. Accordingly, a robot element is enabled to move in such a way that it assumes at least one specified position relative to an object from the object accumulation.Once the 3-dimensional position and orientation of the identified objects suitable for removal have been determined, they are provided to a corresponding interface of the data processing unit. This data can then be further processed by a control device of the gripping and / or transport element or the robot element into corresponding control signals for the gripping and / or transport element or the robot element. If multiple object types may be present in the observed area, the information available at the interface of the data processing unit naturally also includes an indication of the object type of the respective identified object for which the 3-dimensional position and orientation were determined.

[0025] As will be shown in more detail below, the method has proven to be very robust, reliable, fast, and safe in determining the position and / or orientation of objects in a given 3-dimensional coordinate system. Furthermore, it is sufficient to provide only the elevation map of the observed area. The method can be used for any object type, container shape, and container dimensions, as well as movements of the objects in the object cluster. Furthermore, the method can also be used when the environment of the object cluster changes. The NN algorithm is trained separately for each object type or for each plurality of object types, so that for the above method, an NN algorithm trained for the respective plurality of object types is used. The method is therefore adaptable to different object types and the number of object types.The procedure has proven to be variable with regard to possible container shapes and sizes and movements of the objects.

[0026] The NN algorithm used represents an assignment based on a neural network. The NN algorithm uses the neural network to assign a plurality of objects arranged one above the other and / or next to one another in the observed area, e.g., a plurality of objects lying on top of the observed area of ​​the object cluster, to the received height map, along with their 3-dimensional position and orientation in the specified coordinate system. In one embodiment, the neural network represents a convolutional neural network (CNN for short), which, for example, generates a map from the height map that illustrates the object coordinates of the objects in the observed area in a coordinate system of the respective object. For example, a neural network based on the U-Net architecture is used for this purpose; this network has proven to be very suitable for such image-to-image translation problems.In this embodiment, the U-network of the CNN can consist of three sections: the contraction section, the.

[0027] The bottleneck section and the expansion section. The contraction section contains many contraction blocks. Each block takes an input and applies, for example, two 3x3 convolutions + nonlinearity, followed by a 2x2 max pooling. The number of kernels or feature maps can double after each block, allowing the architecture to effectively learn complex structures. The bottom layer mediates between the contraction layer and the expansion layer. For example, it uses two 3x3 CNN layers followed by a 2x2 upsampling layer. Similar to the contraction layer, the expansion layer can also consist of multiple expansion blocks. Each block passes the input to two 3x3 CNN layers followed by a 2x2 upsampling layer. Here, too, after each block, the number of feature maps used by an upsampling layer is halved to maintain symmetry.However, the input of the underlying layer is also augmented with the feature maps of the corresponding contraction layer. This allows the neural network to obtain information directly from the contraction layer at the same resolution, eliminating the need to transmit high-frequency data through the low-resolution bottleneck at the bottom. The number of expansion blocks is the same as the number of contraction blocks. Finally, the resulting high-resolution feature map is downprojected to the required output dimension using another convolutional layer. In one embodiment, the neural network contains predefined blocks for all contraction and expansion layers, referred to as "encoder" and "decoder" layers. These can be concatenated to achieve the desired network depth.The parameters of the network architecture of the NN (number and type of layers, size of feature maps, interconnection) can be easily changed using a configuration structure to enable rapid experiments.

[0028] In one embodiment, in order to identify and determine the position and orientation of the plurality of objects arranged in the observed area, for example the plurality of objects lying on top of the observed area of ​​the object cluster, an object coordinate map of the observed area of ​​the object cluster and a segmentation mask of the observed area of ​​the object cluster are determined (i.e. predicted) from the height map of the object cluster using the NN algorithm, wherein the object coordinate map illustrates the position of the respective pixel in object coordinates of the respective object and the segmentation mask for each pixel contains information about the affiliation of the pixel to a segment of a plurality of segments or object boundaries.As already described above, the NN algorithm involves training on a height map and, in this embodiment, uses its pattern recognition capability to generate a new image (the object coordinate map) in which the object instances can, for example, be colored to represent correspondences to the data model of the respective object type (e.g., 3D CAD model of the respective object type). Using a transformation, the coloring of the predicted object coordinate map can be assigned to object coordinates of the coordinate system of the respective object. Such a transformation is based on the assumption used during training of the respective NN algorithm that each point on the surface of the object's data model corresponds to a unique color value, which is derived from the color value present at the same location within an RGB cube (see color coding of the data model described below).With the segmentation mask, the NN algorithm predicts the boundaries of the objects present in the observed area from the height map and stores them as a segmentation mask in an image with corresponding gray values. For example, the boundaries of every visible object, for example, every object visible from above, are represented in dark gray values, and areas without objects are represented in light gray values ​​or white.

[0029] The color coding of the object's data model (see Fig. 8a) for the object coordinate map can, as mentioned above, be defined such that each surface point of the data model of the respective object type is assigned a unique color. For example, a simple RGB cube-based coloring scheme can be used. First, a transformation is computed to map the oriented bounding box of the data model (Fig. 8a) to the range [0-255, 0-255, 0-255] so that it fits exactly into the RGB cube shown in Fig. 8b. For example, the major axis of the (centered) object type data model can be used to derive the rotation part. Scaling and translation can then be easily computed. With this transformation, each surface point of the object data model of an object type can be transformed into the RGB cube, and the corresponding color is used to color the surface point.Each surface point of the data model of the respective object type is therefore assigned a unique color, whereby the color uniquely represents the position of the surface point in 3D (see Fig. 9).

[0030] In one embodiment, the segmentation mask is applied to the object coordinate map, by means of this application individual objects are identified as being located in the observed area, for example lying on top of the observed area of ​​the object accumulation, and pixels of the object coordinate map and the height map belonging to the corresponding object are determined, wherein from this the 3-dimensional coordinates in a predetermined coordinate system and the 3-dimensional coordinates in the coordinate system of the respective object are determined for each pixel belonging to an identified object, and from the coordinate pairs thus determined of all pixels of the respective identified object the 3-dimensional position and orientation of these identified objects in the predetermined coordinate system is determined.

[0031] For the above procedure, in one embodiment, a binary mask is first generated from the segmentation mask, which highlights the boundaries of segments in the mask using predetermined threshold values. In particular, for each segment formed by a circumferential boundary of the segmentation mask, i.e. for each connected structure, which usually corresponds to an object, a number of pixels in the segment is calculated. Segments with a number of pixels below a predetermined pixel threshold value are not taken into account in the subsequent, further analysis. Here, it is assumed that they are heavily obscured objects. Segments with a number of pixels equal to or above the pixel threshold value are referred to as recognized segments and correspond to an object.

[0032] In this embodiment, the segmentation of the segmentation mask thus processed can then be transferred to the height map and the object coordinate map. In this embodiment, each pixel of the height map of each segment can also be assigned 3-dimensional coordinates in the predefined coordinate system of the 3D scanner. From the object coordinate map determined by the NN algorithm, each pixel of a segment can also be assigned a 3-dimensional coordinate in the coordinate system of the object corresponding to the respective segment, based on the transformation matrix of the color values ​​of the data model used in training the NN algorithm, so that as a result, for each pixel of each detected segment, a coordinate pair consisting of 3-dimensional coordinates in a predefined coordinate system (e.g.of the 3D scanner) and 3-dimensional coordinates in the coordinate system of the object corresponding to the respective segment. The coordinate pairs thus determined for each detected segment can be collected in a list, which may contain outliers due to an imperfect prediction of the segmentation mask and / or object coordinate map. To correct this situation, in one embodiment, a GPU-RANSAC algorithm can be used to find a transformation with respect to the position and orientation (pose) of the respective object type that matches the most coordinate pairs. It has been found that the predictions of the NN algorithm are generally quite good, and only a small number of outliers are to be expected. The use of robust estimation methods such as RANSAC in the embodiment also helps to improve the results of the NN algorithm in difficult situations.The RANSAC algorithm is a well-known resampling algorithm for estimating a model within a series of measured values ​​with outliers and gross errors. Alternatively, so-called M-estimators can also be used. The rough pose provided by the RANSAC algorithm can then be re-estimated using a weighted least squares approach, taking all outlier pixels into account, using an Orthogonal Procrustes algorithm based on the singular value decomposition of the 3x3 cross-covariance matrix of the centered points. The poses of the detected segments thus estimated, each corresponding to an object, are also called candidate poses. The Orthogonal Procrustes algorithm can register two point clouds with known correspondences.

[0033] In one embodiment, the segmentation mask can additionally contain information for each pixel regarding the prediction confidence level with respect to the pixel's position in object coordinates. The fact that the NN algorithm also predicts the prediction confidence information for the pixels of the object coordinates is taken into account and trained in this embodiment during training of the NN algorithm. This is described in more detail below. This prediction confidence can be used as a weighting in the above-described re-estimation of the pose using the Orthogonal Procrustes algorithm to determine candidate poses in a given coordinate system (e.g., in the coordinate system of the 3D scanner).

[0034] In one embodiment, the candidate poses of the detected segments can be further refined in a post-processing step, and refined poses can be determined using an iterative closest point (ICP) algorithm on the GPU of the data processing unit, e.g., in the coordinate system of the 3D scanner. During the ICP, coordinate correspondences determined using the above method in each segment are discarded, and new correspondences are dynamically calculated for each iteration of the algorithm to further improve the result. This exploits the fact that a height map for the observed area is available, which enables rapid calculation of the corresponding coordinates by projection along the view rays. The (projective) distance error of these dynamic correspondences can be minimized using a Levenberg-Marquard algorithm. Outliers are estimated using Tukey weighting within the least squares solver.After ICP refinement, the correspondences are evaluated again to calculate the registration rms score and the fraction of occluded pixels for each object / segment. The registration rms score is a measure of registration quality, i.e., the remaining error between the data model of the respective object type and the specific point cloud.

[0035] These refined poses (position and orientation in a given coordinate system) of the identified objects (segments) can then be provided at an interface of the data processing unit for transmission to a control element of a movable gripping and / or transport element or a movable robot element and used by this for controlling the movement of the gripping and / or transport element or the robot element.

[0036] In one embodiment, the height map of the observed area of ​​the object accumulation is determined using a 3D scanner and transmitted to the data processing unit. 3D scanners are systems with two or more cameras that capture the observed area from at least two viewing angles and determine the height map of the observed area of ​​the object accumulation from the image data thus obtained. This determination of the height map can be repeated at predetermined times at predetermined intervals (e.g., every minute, every second, or every tenth of a second) and / or depending on the progress in removing or manipulating the objects from the accumulation (e.g., after removing five objects), the movement of the objects, and / or depending on the occurrence of further events (e.g., filling the container with further objects, a signal from a monitoring device, e.g., a light barrier), and passed on to the data processing unit accordingly.This allows the determination of the position and orientation of removable objects and, accordingly, the work of a gripping and / or transport element or robot element to be adapted continuously or as needed to the changing conditions of the object accumulation.

[0037] In one embodiment, the method is set up to localize objects moving in the observed area, in which the elevation map of the observed area of ​​the object cluster is received at a predetermined time or several predetermined times, the plurality of objects arranged in the observed area is identified, and their 3-dimensional position and orientation is determined and provided at the interface of the data processing unit. This method is used for line tracking or visual servoing. The objects of the object cluster move, for example, along a conveyor belt. In this case, the object cluster represents, for example, a plurality of objects arranged one above the other and / or next to one another, for example a row of objects of one or more object types arranged one behind the other on a conveyor belt.At a given time / specified time points (this also includes a given temporal sequence with equal or unequal intervals between the time points, for example, every tenth of a second or every second), the visible objects in the cluster are identified, and the 3-dimensional position and orientation of the objects are determined as described above. The interval between the time points can be predetermined or changed based on the movement of the objects (for example, the time interval between the time points can be increased if the speed of movement of the objects increases, or conversely, the time interval between the time points can be decreased if the speed of movement of the objects decreases).Accordingly, following the specified time(s), the 3-dimensional position and orientation of the objects as well as the information about the identified objects are also provided at the interface of the data processing unit.

[0038] In one embodiment, the predetermined time / times can be determined by means of a monitoring device, for example a light barrier and / or by means of a scale integrated into the substructure of a conveyor belt and / or by means of camera monitoring. The monitoring device detects the time of a predetermined position of an object along its movement. Based on the time detected by the monitoring device, in this embodiment one or more times are specified for the determination, e.g. by means of a 3D scanner, transmission and corresponding reception of the height map of the observed area. At the respective time of determining the height map, at least one object from the object cluster is then generally present in the observed area (field of view).Accordingly, the elevation map of the observed area of ​​the object cluster is determined, transmitted, and received at a specified time or several specified times. This synchronizes the determination of the 3-dimensional position and orientation of the objects located in the observed area with the movement of the objects. In the case of moving objects, this improves the accuracy of the data on the 3-dimensional position and orientation of the object provided at the interface of the data processing unit.

[0039] Analogous to the method described above, the above task is achieved by a device for the automatic localization of objects from an observed area of ​​an object accumulation consisting of a plurality of objects of at least one object type arranged next to one another and / or one above the other, with a data processing unit which is set up in such a way that it:

[0040] • receives a data model of at least one object type,

[0041] • receives an NN algorithm trained for the specific at least one object type, wherein the NN algorithm is trained based on the data model of the at least one object type,

[0042] • receives a height map of the observed area of ​​the object accumulation,

[0043] • identifies a plurality of objects arranged on top of the observed area of ​​the object cluster and determines the 3-dimensional position and orientation of this plurality of objects in a predetermined coordinate system based on the NN algorithm trained with respect to the at least one object type using the data model and exclusively using the height map of the observed area of ​​the object cluster with respect to the observed area, and

[0044] • provides the 3-dimensional position and orientation of the plurality of objects arranged on top of the object cluster in the observed area of ​​the object cluster at an interface of the data processing unit. In one embodiment, a device for automatically locating objects suitable for removal from an observed area of ​​an object cluster present in an open container, consisting of a plurality of superimposed objects of at least one object type, is implemented with a data processing unit which is configured to:

[0045] • receives a data model of at least one object type,

[0046] • receives an NN algorithm trained for the specific at least one object type, wherein the NN algorithm is trained based on the data model of the at least one object type,

[0047] • receives a height map of the observed area of ​​the object accumulation,

[0048] • identifies a plurality of objects lying on top of the observed area of ​​the object accumulation and determines the 3-dimensional position and orientation of this plurality of objects lying on top of the object accumulation in a predetermined coordinate system by means of a data processing unit based on the NN algorithm trained with respect to the at least one object type, using the data model of the at least one object type and using exclusively the height map of the observed area of ​​the object accumulation with respect to the observed area, and

[0049] • provides the 3-dimensional position and orientation of the plurality of objects lying on top of the observed area of ​​the object accumulation to an interface of the data processing unit.

[0050] Analogous to the above method, the devices can be configured such that they additionally receive a data model of a container, use an NN algorithm that is additionally trained on the data model of the container and identifies a plurality of objects arranged in the observed area of ​​the object accumulation, for example objects lying on top of the observed area of ​​the object accumulation, and determines the 3-dimensional position and orientation of this plurality of objects in a predetermined coordinate system by means of a data processing unit additionally based on the NN algorithm trained with respect to the at least one container, using exclusively the height map of the observed area of ​​the object accumulation and the data model of the at least one object type.

[0051] The advantages of the above-mentioned devices, each of which has a data processing unit, can be seen from the above description of the corresponding method. The devices also have a storage unit in which at least the data relating to the data model of the at least one object type and, if applicable, the container, as well as the NN algorithm, can be stored.

[0052] In one embodiment, the data processing unit is configured such that, in order to identify and determine the position and orientation of the plurality of objects arranged in the observed area of ​​the object cluster, for example the objects lying on top of the observed area of ​​the object cluster, it determines an object coordinate map of the object cluster and a segmentation mask of the object cluster solely from the height map of the object cluster by means of the NN algorithm trained on the at least one object type and the container, wherein the object coordinate map illustrates the position of pixels in object coordinates and the segmentation mask for each pixel illustrates information about the affiliation of the pixel to a segment of a plurality of segments.

[0053] In one embodiment, the data processing unit is configured to apply the segmentation mask to the object coordinate map, to use the application to identify individual objects as being located in the observed area of ​​the object cluster, for example lying on top of the observed area of ​​the object cluster, and to determine pixels of the object coordinate map and the height map belonging to the corresponding object, from which it determines the 3-dimensional coordinates in a higher-level coordinate system and the 3-dimensional coordinates in the coordinate system of the respective object for each pixel belonging to an identified object, and to determine the 3-dimensional position and orientation of these identified objects in the predetermined coordinate system from the coordinate pairs thus determined of all pixels of the respective identified object.

[0054] In one embodiment, the segmentation mask additionally contains information for each pixel about the degree of prediction certainty with respect to the position of the pixel in object coordinates.

[0055] In one embodiment, the device additionally comprises a 3D scanner, which determines the elevation map of the observed area of ​​the accumulation and transmits it to the data processing unit. For this purpose, the 3D scanner has at least two cameras that view the observed area of ​​the object accumulation from different viewing directions. Each camera is designed, for example, as a digital camera, e.g., a matrix camera, whose position and orientation in space are known. The position and orientation of each camera can be determined using calibration. The 3D scanner can determine the elevation map from the images of the surface of the object accumulation taken by the at least two cameras.

[0056] In one embodiment, the device is configured to locate objects moving in the observed area. The device is configured to repeatedly receive the elevation map of the observed area of ​​the object cluster at a predetermined time or at several predetermined times, identify the plurality of objects arranged in the observed area, and determine their 3-dimensional position and orientation and provide them at the interface of the data processing unit. This device is particularly suitable for line tracking or visual servoing.

[0057] The embodiments explained above in connection with the localization method are also used analogously in a device according to the invention.

[0058] The above object is further achieved by a system for controlling a movable gripping and / or transport element for removing an object from an observed area of ​​an object accumulation present in an open container, consisting of a plurality of objects of at least one object type lying one above the other (bin picking), wherein the system comprises a device as described above and a control device, wherein the data processing unit of the device transmits the 3-dimensional position and orientation of the objects lying on top of the observed area of ​​the object accumulation to the control device,wherein the control device determines, from the transmitted 3-dimensional position and orientation of the objects, at least for a subset of the objects lying on the upper side, corresponding control signals for the movement of the gripping and / or transport element for removing the corresponding object from the container and moves the gripping and / or transport element in accordance with the determined control signals.

[0059] Furthermore, the above object is achieved by a line tracking system for controlling a movable gripping and / or transporting element for manipulating an object from an object accumulation consisting of a plurality of objects of at least one object type arranged one above the other and / or next to one another, wherein the system has a device as described above for locating objects moving in the observed area and a control device, wherein a data processing unit of the device is set up to transmit the 3-dimensional position and orientation of the objects arranged in the observed area of ​​the object accumulation to the control device of the gripping and / or transporting element, wherein the control device is set up to determine from the transmitted 3-dimensional position and orientation of the objects at least for a subset of the objects, e.g.one object, two objects or three objects, whereby the object(s) usually represent the most trackable or easiest to manipulate, to determine corresponding control signals for the movement of the gripping and / or transport element for manipulating the corresponding object and to move the gripping and / or transport element in accordance with the determined control signals. The manipulation can involve a predetermined movement of the gripping and / or transport element, which is carried out according to a predetermined start state. After completion of the movement, the respective object has reached an end state and the system identifies, after the next height map of the observed area has been received, a new object in the observed area for removal, manipulation and / or relative movement of the robot element or gripping and / or transport element.

[0060] For example, the gripping and / or transport element manipulates the localized object in the object accumulation in such a way that it grips the object (starting state) and then performs the specified movement, namely, lifting the object from a first conveyor belt, transporting it to a second conveyor belt, and placing it on the second conveyor belt at a specified distance from the preceding object. For this purpose, the gripping and / or transport element comprises, for example, a gripping hand, a suction cup, a two-jaw gripper, a clamping gripper, a clamp gripper, a magnetic gripper, or the like.Further manipulations may include, for example, the assembly of a component in and / or on the object, the inspection of the object or a specified section of the object, the transport of the object into a disposal container or on a disposal conveyor belt, the painting of the object or a specified section of the object or the joining of a component to the object by means of a joining element.

[0061] The above object is also achieved by a visual servoing system for controlling such a movable robot element, so that it assumes a predetermined position relative to an object from an object accumulation consisting of a plurality of objects of at least one object type arranged one above the other and / or next to one another, wherein the system has a device as described above for locating objects moving in the observed area and a control device, wherein a data processing unit of the device is configured to transmit the 3-dimensional position and orientation of the objects arranged in the observed area of ​​the object accumulation to the control device of the robot element, wherein the control device is configured toFrom the transmitted 3-dimensional position and orientation of the objects, to determine, at least for a subset of the objects, corresponding control signals for the movement of the robot element to assume the specified position relative to the corresponding object, and to move the robot element according to the determined control signals. The specified position may include several positions to be assumed consecutively in a specified sequence (i.e., a movement of the robot element). If, in the latter case, all positions have been passed through, a final state with respect to the respective object is also reached in this system.

[0062] A given position relative to the object can be

[0063] Furthermore, the above task is solved by a method for the machine training of a NN algorithm for the identification and localization of objects from an observed area of ​​an object accumulation, for example of objects suitable for removal from an observed area of ​​an object accumulation present in an open container, by means of a data processing unit with the following steps: a) Using a physical model to arrange a predetermined number of objects of at least one object type, each of which corresponds to the data model of this object type, in a container such that a ground-truth object accumulation is created from the predetermined number of objects, b) Rendering an ideal height map of the observed area of ​​the object accumulation based on the ground-truth object accumulation,c) Rendering an ideal object coordinate map of the observed area of ​​the object cluster based on the ground-truth object cluster, d) Rendering an ideal segmentation mask of the observed area of ​​the object cluster based on the ground-truth object cluster, e) Modifying the ideal height map using a first modification function to create a training height map, f) Determining a predicted object coordinate map and a predicted segmentation mask from the training height map using the current version of the NN algorithm for the observed area of ​​the object cluster, g) Determining damage using a loss function from a comparison of the predicted object coordinate map with the ideal object coordinate map and from a comparison of the predicted segmentation mask with the ideal segmentation mask,h) Modifying the NN algorithm to minimize the value of the loss function, i) Repeating steps e) to h) with different change functions that differ from the first change function, and / or steps a) to h) with a different number of objects of the same at least one object type and / or different orientation of the container and / or different observation areas of the container until a termination criterion is reached, j) Storing the last version of the trained NN algorithm in a data memory of the data processing unit.

[0064] In one embodiment of the training method or the localization method, the NN algorithm is based on a U-Net architecture.

[0065] The above method is also usually realized as a computer-implemented method by means of a data processing unit (e.g. a computer) which has a processor and a computer-readable storage unit.

[0066] In one embodiment, to generate training height maps, the ideal height map can be modified using a first modification function to create a training height map. For example, the modification function can include clipping, mirroring, scaling, brightness modification, color magnification, saturation modification, contrast modification, translation, and / or rotation. Additionally or alternatively, noise and / or artifacts corresponding to those in real data can be introduced into the ideal height map. For example, points of the ideal height map with a normal vector greater than a randomly defined threshold can be removed from the height map. This example is similar to the behavior of 3D scanners that incorrectly capture surface points with a large angle because the light is reflected away from the camera. This is particularly relevant for highly reflective materials.Another example of simulating real-world data as training heightmaps leverages the finding that measurement noise in a 3D scanner decreases with the number of cameras that can see a point. This behavior can be modeled in the training heightmaps for the NN algorithm by evaluating the cameras' visibility image, removing all points with a visibility number of less than two, and creating a noise intensity map that decreases for higher visibility numbers. Another example of adapting training heightmaps to real-world effects simulates the fact that measurement noise often correlates between neighboring pixels of the depth map.To create appropriate training height maps, multiscale noise (similar to Perlin noise) with a predefined spectrum is added to the depth values, and the amplitude is modulated by the previously created noise intensity map. Another example can be used to simulate the fact that strong specular reflections occasionally cause small areas of the height map to display erroneous values. Therefore, small pixel islands (approximately 1-100 pixels) with random heights per island are added to the ideal height map using the change function to make the NN algorithm robust against such artifacts. Additionally or alternatively, limited resolution and suboptimal camera sharpness lead to slightly blurred height images.To train this using training heightmaps, a bilateral Gaussian filter with random sigma is used when creating a training heightmap to smooth the depth in homogeneous regions but prevent blurring at large depth differences. In another example of simulating real heightmaps for training the NN algorithm, randomly selected rotated rectangles are deleted from the ideal heightmap (i.e., they appear completely black or completely white and are also called blackouts). This alteration of the data does not correspond to any real-world effect when acquiring the heightmap, e.g., using a 3D scanner. However, it serves two important purposes: First, a CAD model created from 3D scans may contain imperfections that are not present in the real objects and that could help the detector resolve symmetries or difficult poses.Such regions are occasionally overdrawn by the blackout, so the NN algorithm cannot rely exclusively on these features. Second, the blackout is applied only to the training height map, not to the ideal object coordinate map, so the NN algorithm is forced to supplement these regions only from the context. This helps determine robust and generally applicable parameters of the NN algorithm.

[0067] In one embodiment of the training method, any symmetry present in an object type is taken into account when determining the damage using the loss function. The damage is determined as the loss function damage for which the object coordinate map loss function component is minimal for different positions of the respective object with respect to the symmetry. For this purpose, it must be determined in advance when creating the data model of at least one object type whether and, if so, which axis(es) of symmetry the data model has. Each possible position of the respective object represents a different appearance in the object coordinate map.

[0068] In one embodiment of the training method, when determining the damage using the loss function, the pixels of the object boundaries of the predicted segmentation mask are weighted more heavily than the remaining pixels by multiplying them with a mask weight function. Such a function is explained in more detail below. The boundary pixels of an object thus have a greater influence on the overall loss function, which favors a correct prediction of the edge areas of an object in the segmentation mask and avoids segmentation errors. The result is normalized with respect to the number of pixels in the segmentation mask. The mask weight function can, for example, be calculated from a modification of the ideal segmentation mask of the observed area, in which each object, and in particular the respective pixels of the object, is assigned a unique scalar value (an ID) (object ID map).

[0069] In one embodiment of the training method, the predicted segmentation mask additionally contains prediction confidence information for each pixel, which results from the object coordinate map loss function portion of the loss function. As described in detail below, such an embodiment has proven advantageous in which a prediction confidence of the pixels of the object coordinate map is incorporated into the training process. The prediction confidence information can, for example, be stored in a part of the information assigned to each pixel of the segmentation mask. For example, the information relating to the segmentation of the mask can be stored in a value range [0.0;0.5[ (i.e., excluding 0.5), and information relating to the prediction confidence can be stored in a value range [0.5;1,0].The prediction confidence for the respective pixel can, for example, be derived directly from a difference between the ideal object coordinate map and the object coordinate map predicted for the respective training height map.

[0070] The data processing unit for identifying the removable objects, their position and orientation, and for training the NN algorithm comprises a processor, which represents a functional module that interprets and executes algorithm instructions / commands, as well as a command control unit, an arithmetic unit, and a logic unit. The processor may comprise at least one of a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA - a digital integrated circuit into which a logic circuit can be programmed), a discrete logic circuit, and any combination of these components. The data processing unit may also comprise a memory unit, an input module (e.g., keyboard or touchpad), a power supply module (e.g., battery), and a display module (e.g., display).The data processing unit can be implemented as a real hardware resource, for example, a smartphone, desktop computer, server, notebook, cluster / warehouse-scale computer, embedded system, or the like, or as a virtualized computer resource. Furthermore, the data processing unit can have a transmitter / receiver (transceiver) for exchanging data with the 3D scanner. The data processing unit also has an interface for exchanging data with the control element of the gripping and / or transport element. As already explained above, the methods explained above can each be implemented, for example,be realized as a computer program or computer-implemented method comprising instructions which, when executed, cause a processor of the data processing unit to carry out the steps of the above method, wherein the computer program includes a combination of the steps and data definitions described above which enable the computer hardware to carry out computing or control functions, and / or which represents a syntactical unit which conforms to the rules of a specific programming language and which consists of declarations and statements or instructions which are required for the functions, tasks or problem solutions explained above.

[0071] Furthermore, a computer program product is disclosed comprising instructions that, when executed by the processor of the data processing unit, cause the device to perform the steps of one or all of the methods defined above. Accordingly, a computer-readable medium storing such a computer program product is disclosed. The computer program product may be a software routine.

[0072] Further advantages, features, and possible applications of the present invention will become apparent from the following description of exemplary embodiments and the drawings. All described and / or illustrated features, individually or in any combination, constitute the subject matter of the invention, regardless of their summary in the claims or their references.

[0073] They show schematically

[0074] Fig. 1 shows an embodiment of a device according to the invention for the automatic localization of objects suitable for removal from an observed area of ​​an object accumulation present in an open container or an embodiment of a system for controlling a gripping and / or transport element for removing an object from the observed area of ​​an object accumulation present in the open container consisting of a plurality of objects of at least one object type lying one above the other,

[0075] Fig. 2 shows an embodiment of a method according to the invention for the automatic localization of objects suitable for removal from an observed area of ​​an object accumulation present in an open container consisting of a plurality of objects of at least one object type lying one above the other as a flow diagram,

[0076] Fig. 3 shows a section of the process according to Fig. 2 as a flow chart,

[0077] Fig. 4, 5 an example segmentation mask and a further processing of the segmentation mask used in the method according to Fig. 2,

[0078] Fig. 6 shows an embodiment of the generation of training data for a method for the machine training of a NN algorithm for the identification and localization of objects suitable for removal from an observed area of ​​an object accumulation present in an open container as a flow chart,

[0079] Fig. 7 shows an embodiment of a method for machine training of a NN algorithm for identifying and locating objects suitable for removal from an observed area of ​​an object accumulation present in an open container, in which the training data of the method according to Fig. 6 can be used, as a flow chart,

[0080] Fig. 8,9 an embodiment for the creation of a data model for an object type for a method according to Fig. 2 or Fig. 7 using image data or perspective views from the side,

[0081] Fig. 10 - 12 an embodiment for the creation of a data model of the container for a method according to Fig. 2 or Fig. 7 based on image data or perspective views from the side, Fig. 13 - 16 examples of objects with and without axes of symmetry in perspective views from the side,

[0082] Fig. 17 shows an example of the operation of a physical model for the method according to Fig. 7,

[0083] Fig. 18 shows an example of a ground-truth object cluster in a container and an arbitrary environment for the training method according to Fig. 7 in a perspective view from the side,

[0084] Fig. 19 shows an example of a camera visibility image rendered from the ground-truth object cluster according to Fig. 18 for an observed area,

[0085] Fig. 20 shows an example of an ideal height map rendered from the ground-truth object cluster according to Fig. 18 for an observed area,

[0086] Fig. 21 shows an example of an ideal object coordinate map rendered from the ground-truth object cluster according to Fig. 18 for an observed area,

[0087] Fig. 22 shows an example of an ideal object ID map rendered from the ground-truth object cluster according to Fig. 18 for an observed area,

[0088] Fig. 23 shows an example of an ideal map rendered from the ground-truth object cluster according to Fig. 18 for an observed area, showing the uncovered portion of each object,

[0089] Fig. 24 shows another example of an ideal height map rendered from a ground-truth object cluster for an observed area,

[0090] Fig. 25 -28 Examples of training height maps for the method according to Fig. 7, which were generated from the ideal height map according to Fig. 24,

[0091] Fig. 29 shows an example of an NN algorithm for the method according to Fig. 1, 2 and 7, Fig. 30 - 31 show examples of the determination of the damage using a loss function and the corresponding adaptation of the NN algorithm of the method according to Fig. 7 as flow diagrams,

[0092] Fig. 32 - 33 Examples of the boundary of a binary mask or a mask weighting function for determining the loss function according to Fig. 30 or 31 and

[0093] Fig. 34 shows a further embodiment of a device for the automatic localization of objects from an observed area of ​​an object accumulation or an embodiment of a system for controlling a gripping and / or transport element for manipulating an object or a robot element for positioning the robot element relative to the object from the observed area of ​​an object accumulation consisting of the plurality of objects of at least one object type lying one above the other or next to one another.

[0094] In the figures explained below, color representations actually used are represented in shades of gray (i.e., shades of gray between white and black). Those skilled in the art will appreciate that, particularly where expressly stated, different colors correspond to the different shades of gray.

[0095] Fig. 1 shows an object 10 and a device 20 for automatically locating such objects 10 from an observed area of ​​an object accumulation 40 present in an open container 30. Once objects 10 suitable for removal have been identified in the object accumulation 40 and their position and orientation (pose) have been determined by the device 20, the data on the position and orientation of the objects 10 can be transmitted via a corresponding interface 22 of the device to a gripping and / or transport element 50, e.g., a robot end effector, with a control device 60. The control device 60 further processes the transmitted position and orientation data of the objects 10 identified for removal from the container 30 and determines movement trajectories for the gripping and / or transport element 50.If the gripping and / or transport element 50 moves along these trajectories, the gripping and / or transport element can remove identified objects 10 individually from the container 30 and transport them to a predetermined location, for example a conveyor belt or a machine for further processing of the object 10. The object accumulation 40 is located in the container 30 and consists of a large number of objects 10 lying one above the other in a random or ordered arrangement. In the embodiment shown in Fig. 1, only a single object type is contained as object 10 in the object accumulation 40, namely the component shown in the top left of Fig. 1. However, different object types can also be used in the object accumulation 40, wherein no object, one object, or several objects of each object type can be present in the object accumulation 40.

[0096] For the automatic localization of objects 10 suitable for removal in the object accumulation, a data model (CAD model) of the object 10 is transmitted to the device 20 (see arrow 11). The CAD model is also used to train an NN algorithm 70 by means of a corresponding training device 170, which is also transmitted to the device 20 after training is completed (see arrow 12). The system also includes a 3D scanner 80 with four cameras, which view the surface of the object accumulation 40 in the container 30 from different viewing directions. The 3D scanner generates a height map 91 of an observed section of the current object accumulation 40, as it is currently located in the container 30. Examples of such a height map 91 can be found in Figs. 2 and 3.This height map 91 is also transmitted to the device 20 (see arrow 13), which has a data processing unit 23 and a storage unit 24. The data processing unit 23 can, as shown in Fig. 1, be implemented separately from the training device 170, or the training device 170 can be integrated into the data processing unit 23. Using the trained NN algorithm 70, which is stored in the storage unit 24, objects 10 on the surface of the object accumulation 40 are identified and their position and orientation are determined. This is described in more detail below. This data is provided at an interface 22 and can be transmitted to the control device 60 for controlling the gripping and / or transport element 50 for removing the identified objects, for example, for controlling the translational or yaw, pitch and / or roll movement of the gripping and / or transport element (see arrow 14).The gripping and / or transport element (e.g., the end effector of a robot arm) is controlled so that an identified object 10 is grasped and transported from the container 30 to a conveyor belt. The object 10 is placed there in a predetermined position and conveyed for further processing by means of the conveyor belt.

[0097] An exemplary embodiment of a method for automatic localization is described below with reference to Fig. 2. In a step 101, the 3D scanner 80 first generates a height map 91 of an observed area of ​​the object cluster 40 from which objects 10 are to be removed. The height map 91 is transmitted to the device 20. Based on the transmitted height map, the NN algorithm 70, trained on the specific object 10, predicts an object coordinate map 92 and a segmentation mask 94 in step 102. These are further processed by the data processing unit 23 (e.g., CPU / GPU) in step 103 based on the known data models of the object 10 and the container 30. The further processing results in a plurality of identified objects and their position and orientation in a predetermined coordinate system (step 104).In a subsequent step 105, this data can be further refined using a fast ICP (iterative closest point) algorithm, for example, using the GPU of the data processing unit 23, and for this purpose, consistency evaluations and occlusion estimates can be performed. As a result of this step, the data on the extractable objects, their position and orientation, and possibly further data (evaluations, occlusion estimates) are provided at the interface 22. Such an evaluation and estimation is important for the subsequent verification of the data and trajectory planning in the control device 60 in steps 106 and 107, for example, to decide in which order the identified objects are extracted from the object accumulation 40.

[0098] In further processing step 103, a binary mask in which the object boundaries are highlighted is generated from the segmentation mask 94 in the manner described above. An example of a predicted segmentation mask 94 for an object cluster with object types that partially differ from the object shown in Fig. 1 is shown in Fig. 4. Fig. 32 shows the progression of values ​​of the binary mask along a line in this mask, wherein the line intersects an object boundary that appears as a lowered step. In the mask, the object boundary can be shown in black accordingly. From this, contiguous / connected segments are determined. Segments that have a number of pixels below a predetermined threshold are discarded, i.e., they are not identified as objects.A resulting mask 95 showing connected segments, each segment being assigned a specific gray or color value, is shown in Fig. 5, the mask of Fig. 5 being determined from the predicted segmentation mask 94. Each identified connected segment is assumed to correspond to a single object in the object cluster.

[0099] Likewise, in further processing step 103, as already described above, a 3D coordinate 112 in the coordinate system of the 3D scanner 80 is determined for each pixel of the height map 91 using the calibration of the 3D scanner 80. Furthermore, for each pixel from the predicted object coordinate map 92, the corresponding 3D coordinate 116 in the object coordinate system is determined using a transformation matrix 114, which assigns a coordinate in the object coordinate system to the color values, so that, as a result, each pixel of the height map 91 is assigned a 3D coordinate pair 118 (from the coordinates 112, 116). Using the identified connected segments (see mask 95), the coordinate pairs 18 can be assigned to a segment. Using the RANSAC algorithm already described above, the GPU of the data processing unit 23 and using the Orthogonal Procrustes algorithm, this data can be used for each segment orObject the position and orientation (also referred to above as candidate poses) of identified objects in a given coordinate system are calculated.

[0100] In the further step 105, the refinement of the determined candidate poses of the identified objects (segments) is performed by the GPU of the data processing unit 23 using the ICP algorithm, as already explained above. As a result, refined poses of the identified objects are determined in the specified coordinate system.

[0101] The following describes an exemplary embodiment for training the NN algorithm 70, which is illustrated using the flow charts in Figs. 6 and 7. Fig. 6 illustrates the automatic generation of training data for training, and Fig. 7 illustrates the training process. Fig. 31 also shows the determination of the loss function.

[0102] The creation of training data (see Fig. 6) includes, in step 121, the provision of a data model for the at least one object type (e.g., CAD model). The data model is preprocessed, for example, by rescaling, transforming (step 122), and color coding the data model (see above and Figs. 8, 9) of the object 10 using an RGB cube 13 (in step 123). Unlike the grayscale representation in Figs. 8, 9, the color coding 15 contains the corresponding RGB colors of the RGB cube 13, which correspond to the respective object coordinates on the surface of the object 10. The preprocessed data model of the at least one object type is transmitted to the physics simulation, which is represented by step 130 in Fig. 6.

[0103] After providing the data model for the container 30 in step 124 and preprocessing the container data model in step 125, the physics simulation 130 also receives data from a processed container data model in step 126. For example, the container data model is analyzed to find its wall and floor planes for the physics simulation. Furthermore, the size of the container is adjusted to the desired dimensions. This can also be performed jointly for multiple container types. For example, a plurality of standard container data models can be provided as a list, with one embodiment of a container 30 being shown in Fig. 10. The container data model can be randomly selected for each training example to make the NN algorithm robust to changes in the container shape. During the analysis of the container data model (see Fig.10), a pixel cloud 31 is first created from each container data model (see Fig. 11), which represents the surface of the container 30. The pixels are then assigned to the four walls or a floor of the container based on their position and their normal vectors. For each wall (and floor) surface, a position range along the X, Y, and Z directions is derived from the extent of the bounding box. In addition, a normal angle range is defined by specifying a maximum angle between the pixel normal and the dominant wall direction (X, Y, Z axes). The results of this assignment algorithm are visualized in Fig. 11 using different colors (shades of gray) of the walls or floor.A robust iterative container plane-adaptation algorithm is then used to approximate the respective wall vertices by a plane by minimizing the square error between all interior pixels and the plane (see container data model 32 shown in Fig. 12). This data model (possibly for a plurality of containers 30) is also transmitted to the physics simulation 130.

[0104] In the physics simulation in step 130, the preprocessed data models of the object types (here, only a single object type 10) and the container 30 are used to simulate a filling process of the container with virtual objects 10. Various simulation parameters (e.g., number of parts, friction, direction of gravity) can be randomized. The physics simulation is a software component for simulating an artificial environment for depositing objects 10 in a container 30 (for example, pybullet can be used as a simulation library). As can be seen in Fig. 17, the objects 10 are created above a container 30 and then released to fall. This procedure guarantees a random distribution of the objects 10 within the container 30. The arrangements of the objects within the container resulting from the physics simulation, which are created in step 131, are then rendered in step 135.

[0105] In one embodiment of the physics simulation, the positions and orientations during the creation of objects above the container can be restricted and adjusted to cover applications requiring ordered filling. For this purpose, the user can specify a list of pose generators (more precisely, their parameters) along with a probability for the selection of the respective generator during the creation of the training example. The random selection of generators with the specified probabilities leads to diverse training datasets and thus to a more robust NN algorithm.

[0106] Overall, the physics simulation can be designed in one embodiment such that it can be adapted to the object accumulation to be trained. For example, it is possible to randomly determine the number of objects inside the container. A maximum number of objects can be calculated based on the container volume and the volume of the objects. For each training example, the number of objects can be randomly chosen between 0 and the maximum value. In this way, different fill levels of the container are covered. It is also possible to control gravity and friction in order to control the packing of the sub-units in the container. For example, if the direction of gravity is skewed, all objects are pushed to a specific wall or corner of the container and tightly packed.These variation options result in a randomization of the training data in order to obtain a NN algorithm that is trained as comprehensively as possible.

[0107] In step 135, rendering is performed based on the arrangement of the objects in the container obtained by the physics simulation to generate the training height maps and other data. The images resulting from the physics simulation 130 form the ideal height maps (ground truth height maps) and other representations for training the NN algorithm and are then saved. The renderer uses a scene with the information from the physics simulation results, virtually places the 3D scanner at a random location with a random viewing direction (all within predetermined, reasonable limits so that the container is in the field of view), and renders a series of different virtual images from the perspective of a reference camera of the 3D scanner. For this purpose, the renderer can use the sensor calibration of the scanner type used to localize the objects (e.g., the MiniPICK scanner from ISRA VISION GmbH).The following images can be created for each scene by rendering (see scene in Fig. 18):.

[0108] (a) Camera visibility image (see Fig. 19): Each pixel encodes the number of cameras in which it is visible (i.e., not occluded). This information is used to determine the pixels that a real scanner could reconstruct, since 3D triangulation requires at least two cameras.

[0109] (b) ideal height map (see Fig. 20): The pixels encode the height of the observed surface with 16 bits corresponding to the height range covered after 3D scanner calibration.

[0110] (c) ideal object coordinate map (see Fig. 21): The pixels encode the object coordinates according to the color coding of the data model described above as an ideal object coordinate map, whereby a representation of the colors is not possible within the scope of the present application and is replaced by a gray value representation.

[0111] (d) Object ID map (see Fig. 22): Grayscale image with 32 bits per pixel, in which each object (sub-instance) is uniformly colored with a shade of gray corresponding to a unique object ID. From this, an ideal segmentation mask can be derived.

[0112] (e) Occlusion ratio map (see Fig. 23): Grayscale image in which each object is evenly colored in proportion to the amount it is occluded by other objects. The darker the part, the less occluded it is. This information is used to determine which objects can be selected for evaluation by the NN algorithm.

[0113] The mentioned images / maps can then be stored in step 140 so that the information can be retrieved during the training of the NN algorithm.

[0114] Before the training data generation is completed, the symmetry of the object's data model is determined for each object type in step 136. In particular, rotational symmetries of the objects can be found. This information is later required by the loss function of the training method. The obtained symmetry information is stored in step 137 and made available to the training data stored in step 140. Figs. 13 to 15 show examples of different objects, wherein the object 10a illustrated in Fig. 13 has no rotational symmetry, the object 10b shown in Fig. 14 has discrete axes of symmetry, and the cylinder shown in Fig. 15 (object 10c) has continuous symmetry. A determined axis of symmetry of another object (gearwheel 10d) is illustrated in Fig. 16 using small black squares arranged on the axis of symmetry.The object exhibits a discrete symmetry at certain rotation angles, which is illustrated by the bright squares distributed around its circumference. Symmetry determination can be facilitated by specifying a number of symmetry axes for each object type.

[0115] For the training, which will be explained below using the flowchart in Fig. 7, in one embodiment of the training method, training data simulating real data can now be generated from the ideal data obtained above (ideal height map, ideal object coordinate map), as already described above. For this purpose, the ideal data is loaded from the memory in step 141 and modified in step 142 using one or more of the modification functions listed and explained above. The resulting training images, in particular training height maps, are then used in the training of the NN algorithm (illustrated by box 170 in Fig. 7). Examples of training height maps that have been formed from an ideal height map shown in Fig. 24 using a modification function are shown in Figs. 25 to 28.

[0116] Information about the NN algorithm is also required for training. This information about an NN algorithm, which can be based, for example, on a U-Net architecture (see Fig. 29), is provided in step 151. In step 151, the NN algorithm is created and made available to the training process 170 in step 155. The NN algorithm can, for example, have the structure described in more detail above as a CNN. Training includes predicting the desired images / maps from the training height maps using the NN algorithm in step 171, determining the loss function for the predicted images / maps in step 172, and optimizing the parameters of the neural network underlying the NN algorithm in step 173 to minimize the damage calculated with the loss function.Training is then continued with the correspondingly modified NN algorithm (illustrated by arrow 174), with the NN algorithm adapted each time until a predefined termination criterion is reached. Training is performed with respect to the at least one object type that is to be placed in and removed from the container in the bin-picking task. The trained NN algorithm is then saved in step 75 and made available to the localization process for predicting the object coordinate map and the segmentation mask based on the current height map of the observed area of ​​the object cluster.

[0117] The determination of the loss function during training of the NN algorithm 70 is illustrated in Fig. 30 and in more detail in Fig. 31.

[0118] From a training elevation map 191, an object coordinate map 92a and a segmentation map 94a are predicted using the NN algorithm 70. These maps 92a, 94a are compared with the ideal object coordinate map 92gt and the ideal segmentation map 94gt, and the damage is determined using the loss function (step 172). In optimization step 173, the parameters of the NN algorithm that need to be changed to minimize the damage determined with the loss function are determined. The NN algorithm 70 is adjusted accordingly, and training continues with the next training elevation map 191.

[0119] Fig. 31 shows an example of calculating the damage from the loss function. In this embodiment, the total damage 172t is composed of a mask damage 172m, a prediction certainty damage 172cf, and an object coordinate damage 172oc. The calculation of the loss function in Fig. 31 refers to an embodiment in which, as explained above, the segmentation mask for each pixel contains not only information about its membership in a segment but also information about the degree of prediction certainty with respect to the pixel's position in object coordinates.Input variables of the mask loss function 172f1 are therefore formed from the ideal segmentation mask 94gt and the predicted segmentation mask 94a, specifically the information part 94a1 of the predicted segmentation mask 94a, which contains the pixel's affiliation to a segment and is contained, for example, in the value range [0.0; 0.5]. Furthermore, an input variable forms the weighting function 75a shown in Fig. 33. Mask damage is an important factor for distinguishing objects in a 3D scan. This is achieved by identifying connected image elements. The damage is calculated as the sum of the squared difference between the predicted segmentation mask 94a (from information part 94a1) and the ideal segmentation mask 94gt.The pixel-based errors are further multiplied by the weighting function 75a, which causes the object's boundary pixels to have a greater influence on the mask damage 172m and thus on the total damage 172t. This configuration is advantageous for ensuring correct prediction of an object's edge regions and avoiding segmentation errors that can quickly degrade detection performance. The result is then normalized with respect to the number of pixels in the image.

[0120] To determine the prediction confidence loss 172cf, the loss function 172f3 is derived from the prediction confidence portion 94a2 of the predicted segmentation mask 94a and the value originating from the loss function of the object coordinate map 172f2, which, as indicated in step 176, is calculated per pixel.

[0121] Finally, the proportion of object coordinate damage 172oc is calculated from a comparison of the ideal object coordinate map 92gt and the predicted object coordinate map 92a. The segment of each object is iterated, and the coloring error of the predicted object coordinate map is calculated across all pixels belonging to the respective object (the so-called L2 norm). In this case, the challenge exists that symmetric objects can occur. These objects can have multiple congruent poses for which the geometry is the same but the coloring is different. Therefore, the object coordinate loss function 172f2 checks the damage for each symmetry of the respective object 10 separately and uses only the smallest damage for the respective object as the object coordinate damage 172oc and to calculate the total damage 172t.

[0122] The total damage 172t results from the sum of the mask damage 172m, the prediction security damage 172cf and the object coordinate damage 172oc.

[0123] After processing each training elevation map 191, such a total damage 172t is determined. Based on the total damage 172t, the parameters of the NN algorithm 70 are adjusted to minimize the damage 172t.

[0124] With an NN algorithm 70 trained in this way, the device 20 can easily, reliably, and quickly identify objects 10 in an object cluster 40 that can be removed from the cluster. Furthermore, their position and orientation (pose) can be determined, which can be transmitted to a control device 60 of the gripping and / or transport element 50 via the interface 22. Based on the poses of the respective objects, the control device can determine trajectories of the movement of the gripping and / or transport element 50 along which the gripping and / or transport element can be moved for removal.

[0125] Fig. 34 shows an object 10 and a device 220 for automatically locating such objects 10 from an observed area of ​​an object accumulation 240 present on a conveyor belt 230. The objects 10 are schematically shown on the conveyor belt 230 as simple cylinders, but these are intended to have the shape shown in the top left of this image. The method operates, for example, when the first of the objects 10 lying on the conveyor belt 230 passes a light barrier (not shown). A corresponding signal from this light barrier determines the height map in the observed area, transmits it, and receives it by the device 220. The identification of the objects 10 in the object accumulation 240 and the determination of their respective 3-dimensional position and orientation is carried out analogously to the bin-picking procedure described above for each received height map.Once objects 10 have been identified in the object cluster 240 and their respective 3-dimensional position and orientation (pose) have been determined by the device 220, the data relating to the position and orientation of the objects 10 can be transmitted via a corresponding interface 222 of the device to a gripping and / or transport element 250, e.g., a robot end effector, having a control device 260. For example, each 1The 3-dimensional position and orientation of the objects in the observed area of ​​the conveyor belt 230 is determined every 5 / 2 second (every five-tenths of a second) in order to take into account the movement of the objects on the conveyor belt as well as the movement of the gripping and / or transport element. The control device 260 further processes the transmitted position and orientation data of the objects 10 identified in the observed area of ​​the conveyor belt 230 and determines movement trajectories for the gripping and / or transport element 250, e.g., in the object coordinate system.If the gripping and / or transport element 250 moves along these trajectories, the gripping and / or transport element can remove identified objects 10 individually from the conveyor belt 230 and, for example, transport them onto a parallel second conveyor belt (not shown) and place them there, for example, at a predetermined distance from the preceding object 10, for example for further processing of the object 10. The object accumulation 40 is located on the conveyor belt 230 and consists of a plurality of objects 10 standing or lying one above the other or next to one another in a random or ordered arrangement. In the embodiment shown in Fig. 34, only a single object type is contained as object 10 in the object accumulation 240, namely the component shown at the top left in Fig. 1.However, different object types can also be used in the object accumulation 240, whereby no object, one object, or multiple objects of each object type can be present in the object accumulation 240. Alternatively, instead of the gripping and / or transport element 250, a robot element (actuator) 250 can be used, which assumes one or more predetermined positions relative to an object in relation to objects in the object accumulation 240, wherein the plurality of positions form a movement sequence. For example, the robot element can be configured in such a way, and its movement can be controlled by the control device 260 in such a way that it mounts a component at a specific position of the object or joins a component to the object.In the procedure described above, the control device 260 determines the object 10 that is currently being transported by means of the gripping and / or transport element or with respect to which the robot element assumes predetermined positions relative to the object at the respective time. As a rule, the control device 260 will select the object 10 for transport or with respect to the relative movement of the robot element that has the most favorable position with respect to the activity of the gripping and / or transport element or the robot element.

Claims

Patent claims 1. A method for the automatic localization of objects (10) of an observed area of ​​an object accumulation (40) consisting of a plurality of objects of at least one object type arranged one above the other and / or next to one another, comprising the following steps: • Receiving a data model of the at least one object type, for example in the form of a colored coordinate model, • Receiving an NN algorithm (70) trained for the specific at least one object type, wherein the NN algorithm is trained based on the data model of the at least one object type, • Receiving a height map (91, 191) of the observed area of ​​the object accumulation, • Identifying a plurality of objects arranged in the observed area of ​​the object accumulation and determining the 3-dimensional position and orientation of this plurality of objects in a predetermined coordinate system by means of a data processing unit (23) based on the NN algorithm trained with respect to the at least one object type using the data model of the at least one object type and using exclusively the height map of the observed area of ​​the object accumulation with respect to the observed area, • Providing the 3-dimensional position and orientation of the plurality of objects arranged in the observed area of ​​the object accumulation at an interface (22) of the data processing unit.

2. Method according to claim 1, characterized in that for the identification as well as position and orientation determination of the plurality of objects arranged in the observed area of ​​the object cluster, an object coordinate map (92) of the object cluster and a segmentation mask (94) of the object cluster are determined from the height map of the object cluster by means of the NN algorithm, wherein the object coordinate map contains the position of pixels in object coordinates of the respective object and the segmentation mask for each Pixel contains information about the pixel's affiliation to a segment of a plurality of segments.

3. Method according to claim 2, characterized in that the segmentation mask is applied to the object coordinate map, by means of this application individual objects are identified as being located in the observed area of ​​the object cluster, and pixels of the object coordinate map and the height map belonging to the corresponding object are determined, whereby for each pixel belonging to an identified object, the 3-dimensional coordinates in a predetermined coordinate system and the 3-dimensional coordinates in the coordinate system of the respective object are determined, and from the coordinate pairs thus determined of all pixels of the respective identified object, the 3-dimensional position and orientation of these identified objects in the predetermined coordinate system is determined, whereby, for example, the segmentation mask for each pixel can additionally contain the information,how high the prediction reliability is with respect to the position of the pixel in object coordinates, and / or that the height map of the observed area of ​​the accumulation is determined and transmitted by means of a 3D scanner (80).

4. Method according to one of the preceding claims, characterized in that the method is set up for locating objects of the observed area suitable for removal, in which the observed area of ​​the object accumulation is arranged in an open container, and the object accumulation consists of objects which lie one above the other, wherein the identified plurality of objects determined with regard to their 3-dimensional position and orientation lies on the upper side of the observed area of ​​the object accumulation.

5. Method for controlling a movable gripping and / or transport element (50) for removing an object (10) from an object accumulation (40) present in an open container (30) consisting of a plurality of objects of at least one object type lying one above the other, wherein the method carries out the method steps contained in claim 4 and the 3- dimensional position and orientation of the objects lying on the upper side of the observed area of ​​the object accumulation is transmitted to a control element (60) of the gripping and / or transporting element, wherein by means of the control element, from the transmitted 3-dimensional position and orientation of the objects, corresponding control signals for the movement of the gripping and / or transporting element for removing the corresponding object from the container are calculated at least for a subset of the objects lying on the upper side.

6. Method according to one of claims 1 to 3, characterized in that the method is set up for the localization of objects moving in the observed area, in which at a predetermined time or several predetermined times the height map of the observed area of ​​the object accumulation is received, the plurality of objects arranged in the observed area is identified and their 3-dimensional position and orientation is determined and provided at the interface of the data processing unit.

7. Method for controlling a movable gripping and / or transporting element for manipulating an object from an object accumulation consisting of a plurality of objects of at least one object type arranged one above the other and / or next to one another, wherein the method carries out the method steps contained in claim 6 and the 3-dimensional position and orientation of the objects arranged in the observed area of ​​the object accumulation is transmitted to a control element of the gripping and / or transporting element after being made available at the interface, wherein by means of the control element, corresponding control signals for the movement of the gripping and / or transporting element for manipulating the corresponding object are calculated from the transmitted 3-dimensional position and orientation of the objects, at least for a subset of the objects.

8. Method for controlling a movable robot element to assume a predetermined position relative to an object from an object accumulation consisting of a plurality of objects of at least one object type arranged one above the other and / or next to one another, the method comprising the steps of carries out the method steps contained in claim 6 and the 3-dimensional position and orientation of the objects arranged in the observed area of ​​the object accumulation is transmitted to a control element of the robot element after being made available at the interface, wherein by means of the control element, from the transmitted 3-dimensional position and orientation of the objects, corresponding control signals for the movement of the robot element to assume the predetermined position relative to the corresponding object are calculated at least for a subset of the objects.

9. Device (20) for the automatic localization of objects (10) from an observed area of ​​an object accumulation (40) consisting of a plurality of objects of at least one object type arranged next to one another and / or one above the other, with a data processing unit (23) which is set up in such a way that it: • receives a data model of at least one object type, • receives an NN algorithm (70) trained for the specific at least one object type, wherein the NN algorithm is trained based on the data model of the at least one object type, • receives a height map (91, 191) of the observed area of ​​the object accumulation, • identifies a plurality of objects arranged on top of the observed area of ​​the object cluster and determines the 3-dimensional position and orientation of this plurality of objects in a predetermined coordinate system based on the NN algorithm trained with respect to the at least one object type using the data model and exclusively using the height map of the observed area of ​​the object cluster with respect to the observed area, and • providing the 3-dimensional position and orientation of the plurality of objects arranged on top of the object cluster in the observed area at an interface (22) of the data processing unit.

10. Device according to claim 9, characterized in that that the data processing unit is configured such that, for the purpose of identifying and determining the position and orientation of the plurality of objects arranged in the observed area of ​​the object cluster, it determines an object coordinate map (92) of the object cluster and a segmentation mask (94) of the object cluster from the height map of the object cluster by means of the NN algorithm, wherein the object coordinate map contains the position of pixels in object coordinates and the segmentation mask for each pixel contains information about the pixel's affiliation to a segment of a plurality of segments, wherein, for example, the segmentation mask for each pixel can additionally contain information about the degree of prediction certainty with regard to the position of the pixel in object coordinates, and / or that the data processing unit is configured such that it applies the segmentation mask to the object coordinate map,by means of the application, individual objects are identified as being located in the observed area of ​​the object cluster and pixels of the object coordinate map and the height map belonging to the corresponding object are determined, whereby it determines from this the 3-dimensional coordinates in a predetermined coordinate system and the 3-dimensional coordinates in the coordinate system of the respective object for each pixel belonging to an identified object, and from the coordinate pairs thus determined of all pixels of the respective identified object, the 3-dimensional position and orientation of these identified objects in the predetermined coordinate system, and / or that the device additionally comprises a 3D scanner (80) which determines the height map of the observed area of ​​the cluster and transmits it to the data processing unit.

11. Device according to one of claims 9 to 10, characterized in that the device is set up to locate objects of the observed area that are suitable for removal, wherein the observed area of ​​the object accumulation is arranged in an open container, and the object accumulation consists of objects that lie one above the other, wherein the identified plurality of objects, determined with regard to their 3-dimensional position and orientation, lies on the upper side of the observed area of ​​the object accumulation.

12. System for controlling a movable gripping and / or transport element (50) for removing an object (10) from an observed area of ​​an object accumulation (40) present in an open container (30) consisting of a plurality of objects of at least one object type lying one above the other, wherein the system comprises a device according to claim 11 and a control device (60), wherein a data processing unit (23) of the device transmits the 3-dimensional position and orientation of the objects lying on top of the observed area of ​​the object accumulation to the control device,wherein the control device determines, from the transmitted 3-dimensional position and orientation of the objects, at least for a subset of the objects lying on the upper side, corresponding control signals for the movement of the gripping and / or transport element for removing the corresponding object from the container and moves the gripping and / or transport element in accordance with the determined control signals.

13. Device according to one of claims 9 to 10, characterized in that the device is set up to localize objects moving in the observed area, wherein the device is set up to repeatedly receive the height map of the observed area of ​​the object accumulation at a predetermined time or at several predetermined times, to identify the plurality of objects arranged in the observed area and to determine their 3-dimensional position and orientation and to provide them at the interface of the data processing unit.

14. System for controlling a movable gripping and / or transporting element for manipulating an object from an object accumulation consisting of a plurality of objects of at least one object type arranged one above the other and / or next to one another, the system comprising a device according to claim 13 and a control device, a data processing unit of the device being configured to transmit the 3-dimensional position and orientation of the objects arranged in the observed area of ​​the object accumulation to the control device of the gripping and / or transporting element, the control device being configured to use the transmitted 3-dimensional position and orientation of the objects at least for a subset of the To determine corresponding control signals for the movement of the gripping and / or transport element for manipulating the corresponding object and to move the gripping and / or transport element according to the determined control signals.

15. System for controlling a movable robot element so that it assumes a predetermined position relative to an object from an object accumulation consisting of a plurality of objects of at least one object type arranged one above the other and / or next to one another, wherein the system comprises a device according to claim 13 and a control device, wherein a data processing unit of the device is configured to transmit the 3-dimensional position and orientation of the objects arranged in the observed area of ​​the object accumulation to the control device of the robot element, wherein the control device is configured toto determine, from the transmitted 3-dimensional position and orientation of the objects, at least for a subset of the objects, corresponding control signals for the movement of the robot element to assume the specified position relative to the corresponding object and to move the robot element in accordance with the determined control signals.

16. A method for machine training of a NN algorithm for identifying and localizing objects (10) from an observed area of ​​an object cluster (40) by means of a data processing unit (23), comprising the following steps: a) using a physical model to arrange a predetermined number of objects of at least one object type, each of which corresponds to the data model of this object type, in a container such that a ground-truth object cluster is created from the predetermined number of objects, b) rendering an ideal height map of the observed area of ​​the object cluster based on the ground-truth object cluster, c) rendering an ideal object coordinate map of the observed area of ​​the object cluster based on the ground-truth object cluster, d) Rendering an ideal segmentation mask of the observed area of ​​the object cluster based on the ground-truth object cluster, e) Modifying the ideal height map using a first change function to create a training height map, f) Determining a predicted object coordinate map and a predicted segmentation mask from the training height map using the current version of the NN algorithm for the observed area of ​​the object cluster, g) Determining damage using a loss function from a comparison of the predicted object coordinate map with the ideal object coordinate map and from a comparison of the predicted segmentation mask with the ideal segmentation mask, h) Modifying the NN algorithm to minimize the value of the loss function, i) Repeating steps e) to h) with different change functions,which differ from the first change function, and / or steps a) to h) with a different number of objects of the same at least one object type and / or different orientation of the container and / or different observation areas of the container until a termination criterion is reached, j) storing the last version of the trained NN algorithm in a data memory of the data processing unit., 17. Method according to claim 16 or one of claims 1, 5, 7 and 8, characterized in that the NN algorithm (70) is based on a U-Net architecture and / or that in determining the damage of the loss function, a symmetry possibly present in an object type is taken into account, wherein the damage is determined as damage of the loss function in which the object coordinate map loss function component is minimal for different positions of the respective object with regard to the symmetry, and / or that when determining the damage of the loss function, the pixels of the object boundaries of the predicted segmentation mask are weighted more heavily than the remaining pixels by multiplication with a weighting function (75a) and / or that the predicted segmentation mask additionally contains prediction certainty information for each pixel, which results from the object coordinate map loss function portion of the loss function.