Method and device for detecting objects in an environment using a multiplicity of observations

The method addresses the challenge of accurate and reliable object detection by determining state vectors and their uncertainties from spatially resolved images, resulting in improved object recognition with comprehensive uncertainty estimation.

WO2025113994A1PCT designated stage expired Publication Date: 2025-06-05TECHNISCHE UNIVERSITAT MUNCHEN
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/082111
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-01
Filing Date
2024-11-13
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing object detection methods in computer-assisted object recognition struggle with accurately and reliably detecting objects in environments while providing a comprehensive understanding of the associated uncertainties.

Method used

A computer-implemented method that determines state vectors from spatially resolved images to characterize object arrangements, calculates element-specific uncertainties, and compares these vectors to generate error-weighted state vector estimates for objects, taking into account the uncertainties.

Benefits of technology

This method enables accurate and reliable object detection while providing a realistic estimation of uncertainties, improving the understanding and accuracy of object recognition processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024082111_05062025_PF_FP_ABST
    Figure EP2024082111_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method for detecting objects in an environment, to a computer program for detecting objects in an environment, to a machine-readable storage medium for detecting objects in an environment, and to a device for detecting objects in an environment. The method according to the invention is used to detect objects in an environment using a multiplicity of observations, wherein each of the observations in each case comprises a spatially resolved recording of at least one part of the environment. State vectors are determined for each of the observations using the corresponding recording, wherein the state vectors each characterize the arrangement of an object detected in the recording in the environment, and wherein an element-specific uncertainty is also determined for each of the elements of the state vectors using the corresponding recording. The state vectors determined for different observations are compared in order to determine sets of state vectors each linked to a common object. For each of the sets of state vectors linked to a common object, an error-weighted state vector estimation is determined for the relevant object using the set of state vectors and taking into account the uncertainties of the elements of the relevant state vectors.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for detecting objects in an environment based on a plurality of observations AREA [OOOI] The invention lies in the field of computer-assisted object recognition. The invention relates to a computer-assisted method for recognizing objects in an environment based on a plurality of observations, as well as a corresponding computer program, a corresponding machine-readable storage medium, and a device for recognizing objects in an environment. BACKGROUND

[0002] The automatic computer-aided detection of objects in an environment is of great importance for a wide variety of technical applications, for example in industrial production and logistics or even autonomous driving. In recent years and decades, a large number of methods have been developed, particularly using machine learning and artificial neural networks, which enable, for example, the detection of objects in two-dimensional camera images, including videos, and / or in data from depth sensors such as lidar or radar sensors. In addition to simply detecting the presence of an object, object detection can include, for example, determining the position, extent and / or orientation of the object and / or classifying the object into a variety of object classes.

[0003] In practice, the detection of objects and the measurements or recordings used for this purpose are always subject to uncertainties. In order to increase the accuracy and reliability of object detection, these uncertainties should be minimized as much as possible. Furthermore, the precise knowledge and consideration of such uncertainties is essential for many applications and can, for example, help to avoid critical errors. However, with known methods, such uncertainties are often not or only insufficiently known and / or cannot be reproduced or explained, for example, because the behavior of the neural networks used is not or cannot be fully understood or explained and essentially represent a “black box”. OVERVIEW

[0004] It is therefore an object of the invention to provide a method for detecting objects in an environment which enables accurate and reliable detection of the objects as well as a realistic and comprehensible estimation of the associated uncertainties.

[0005] This object is achieved according to the invention by a computer-implemented method for detecting objects having the features of claim 1, a computer program having the features of claim 15, a machine-readable storage medium having the features of claim 16, and a device for detecting objects having the features of claim 17. Embodiments of the invention are specified in the dependent claims.

[0006] According to a first aspect of the present invention, a computer-implemented method for detecting objects in an environment based on a plurality of observations is provided. Each of the observations comprises a spatially resolved image of at least part of the environment. For each of the observations, state vectors are determined from the corresponding image, each of which characterizes the arrangement of an object detected in the image in the environment. For the elements of the state vectors, an element-specific uncertainty is determined from the corresponding image. The state vectors determined for different observations are compared to determine sets of state vectors, each of which is linked to a common object.For each of the sets of state vectors associated with a common object, an error-weighted state vector estimate for the object in question is determined using the set of state vectors, taking into account the uncertainties of the elements of the state vectors in question.

[0007] The method according to the invention can be carried out using the device according to the fourth aspect of the present invention. In particular, the method can be carried out entirely or partially by such a device, for example, entirely or partially automated. Alternatively or additionally, the method can also be carried out using any other suitable computer device. [ooo8] The method according to the invention is not limited to specific types of objects, but can be used to detect a variety of types of objects. The objects can be or include physical objects such as, for example, manufacturing parts, manufactured products, containers, vehicles, pedestrians, robots, industrial equipment, storage facilities, buildings and / or building parts. The objects can be or include moving and / or stationary objects. The objects can have a fixed or variable shape. The objects can be spaced apart from one another and / or in contact with one another.

[0009] For the purposes of this disclosure, "environment" refers to the three-dimensional space in which the objects are arranged. The environment may, for example, be or include an interior space such as an industrial hall or warehouse and / or an exterior space such as an outdoor storage area or a street.

[0010] The objects are recognized based on a large number of observations. In the context of the present disclosure, an “observation” can refer to any type of measurement or recording by means of which the objects in the environment can be observed, in particular such measurements or recordings which allow conclusions to be drawn about the arrangement of the objects in the environment. Each of the observations comprises a spatially resolved recording of at least one part (e.g., a section) of the environment. A spatially resolved recording can, for example, comprise a large number of measurements or recordings, each of which is linked to a specific point and / or area in the environment. The spatially resolved recording can, for example, be or comprise a two- or three-dimensional camera image.Alternatively or additionally, the spatially resolved recording may include other types of measurements or recordings, in particular depth measurements (i.e., distance measurements relative to an observation point). The observations and recordings may be numbered consecutively with the index m.

[0011] The observations (e.g. the corresponding images) can be made with corresponding observation parameters, e.g. from a specific observation point, in a specific observation direction (viewing direction), with a specific observation section, at a specific observation time and / or with specific observation settings (e.g. with a specific focus, a specific spatial resolution and / or a specific observation duration). Preferably, some or all observations differ in at least one observation parameter, e.g. in the Observation point and / or the observation direction. Based on the observation parameters, in particular the observation point and the observation direction, the observations and the quantities derived from them, in particular the state vectors, can be linked to a reference system of the environment ("world coordinates"), e.g., transformed into this.

[0012] The observations may be made as part of the method according to the invention. In particular, the method according to the invention may be used for object detection in real time, for example, by continuously making observations and iteratively improving the error-weighted state estimates based on the new observations. In other examples, the observations may also have been made in whole or in part in advance before executing the method according to the invention.

[0013] For each of the observations, state vectors bbi = (bi, b bi j2 , ...) are determined, which each characterize the arrangement, in particular the three-dimensional arrangement, of an object detected in the image in the environment (e.g., indicate and / or allow conclusions about it). In other words, the elements bi, bi, 2 , ... the state vectors are parameters that describe the arrangement of the respective object in the environment. The state vectors can, for example, characterize a position, an orientation, and / or an extent of the respective object. The state vectors can be extracted from the corresponding image using computer-aided object recognition, for example as described below. In some embodiments, the state vectors can be determined using principal component analysis (PCA).

[0014] In addition, for the elements bi, b bi, 2 , ... of the state vectors, each with an element-specific uncertainty Oi based on the corresponding image, b Oi j2 , ... are determined, which can be summarized in an uncertainty vector Qi. The element-specific uncertainty can characterize (e.g., specify and / or allow conclusions to be drawn about) a confidence interval and / or confidence level for the respective element of the respective state vector. The element-specific uncertainty can, for example, characterize a statistical error (e.g., a variance or a standard deviation), a systematic error, and / or a confidence level of a neural network for the respective element.

[0015] The element-specific uncertainties can be determined individually for the respective element or a subgroup of elements (in contrast to, for example, a “global” uncertainty, which is considered to be the same for all elements of a state vector). is assumed). Accordingly, the element-specific uncertainties can be different and / or determined differently for at least two, in some examples for all elements of the state vector in question. For example, the element-specific uncertainties can be different and / or determined differently for elements linked to an orientation of the object, for elements linked to a position of the object, and / or for elements linked to an extent of the object. The element-specific uncertainties can also be determined individually for each corresponding recording, i.e., they can be element- and recording-specific (in contrast, for example, to a constant uncertainty that is assumed to be the same for all recordings or observations, e.g., a constant uncertainty linked to a specific sensor).In addition to an element- and recording-specific contribution, the element-specific uncertainties may also contain one or more global and / or constant contributions.

[0016] The elements of the state vectors may be uncorrelated, for example, when the objects are stationary. In other examples, two or more elements of the state vectors may be correlated, for example, when some or all of the objects are moving. In these examples, in addition to the element-specific uncertainties, covariances between the elements of the state vectors can also be determined, and Oi can be expressed accordingly as a matrix.

[0017] The state vectors bb™ determined for different observations m are compared to determine sets of state vectors, each of which is linked to a common object. This can, for example, comprise determining a similarity measure and assigning the state vectors to an object / set of state vectors based on the similarity measure. The method according to the invention is not limited to a specific similarity measure, and the comparison of state vectors can be carried out using any suitable similarity measure that characterizes, in particular quantifies, the similarity between two (or more) state vectors. The similarity measure can, for example, be a scalar product between the respective state vectors, a distance between the objects and / or enclosing bodies characterized by the respective state vectors (e.g.their centers) and / or a spatial overlap between the objects and / or bounding bodies characterized by the respective state vectors. In a preferred embodiment, the similarity measure is a "bounding box disparity", for example as described in MG Adam et al., "Bounding box disparity: 3d metrics for object detection with full degree of freedom", IEEE ICIP 2022, Bordeaux, France (Oct 2022). The "bounding box disparity" can be a similarity of state vectors (e.g. a pair of state vectors) by taking into account a distance between the objects and / or bounding bodies characterized by the respective state vectors (e.g., their centers) and a spatial overlap between the objects and / or bounding bodies characterized by the respective state vectors (e.g., their Jaccard coefficient (Intersection over Union, IoU)). The bounding box disparity (bbd) can, for example, be defined as bbd = 1 - IoU + v2v, where v2v indicates a distance between the objects and / or bounding bodies characterized by the respective state vectors, preferably a shortest distance between the objects or bounding bodies. The bounding box disparity can provide a continuous similarity measure, which allows one to decide (e.g., by means of a suitably defined threshold value) whether it is the same object (i.e.,both state vectors characterize the same object). In some embodiments, the similarity measure can be determined by means of Monte Carlo sampling of different instances of the bounding boxes (e.g. based on the corresponding uncertainties) and calculating an averaged “bounding box disparity”. The state vectors can be compared directly with one another, e.g. by determining a similarity measure between two or more state vectors. Alternatively or additionally, the comparison can also be carried out indirectly, in particular by determining a similarity measure between a state vector and a state vector estimate, e.g. as described below. The sets of state vectors associated with a common object can be determined explicitly, for example as a list of state vectors {bb. -1 , bb , Alternatively or additionally, one or more of these sets of state vectors can also be determined only implicitly, for example as those state vectors which are used to determine a state vector estimate, but for example without the corresponding set of state vectors being explicitly determined and / or stored as a list.

[0018] For each of the sets of state vectors associated with a common object, an error-weighted state vector estimate bbj is determined for the respective object. This state vector estimate is determined using the respective set of state vectors, for example, by suitable averaging. The state vector estimate is determined taking into account the uncertainties of the elements of the respective state vectors in order to obtain an error-weighted state vector estimate. The state vectors bbf 1can, for example, be weighted depending on, in particular with (i.e. proportional to) the respective uncertainties ct. Preferably, the state vector estimate is determined element-wise error-weighted, for example by element-wise error-weighted averaging. An element bbj,k of the state vector estimate can be determined, for example, by determining the relevant elements bb^ of the state vectors bb™ depending on, in particular with (i.e. proportional to) the respective uncertainties o™ k Determining the error-weighted state vector estimate bbj for the object in question can allow the determination of element-specific uncertainties a J k for the state vector estimation, e.g., based on the element-specific uncertainties of the respective state vectors. The element-specific uncertainties of the state vector estimation can be summarized in an uncertainty vector (or matrix) Oj.

[0019] In a preferred embodiment, the error-weighted state vector estimate is determined using one (or more) iterative stochastic filters. The stochastic filter can iterate over the observations and can model the information available at the respective time (for the current observation) to determine the state vector estimate. The stochastic filter can be a filter that models a Markov process. The stochastic filter can, for example, be or comprise a Bayesian filter, in particular a Kalman filter. Alternatively or additionally, the stochastic filter can be or comprise a particle filter (also known as a sequential Monte Carlo method) and / or a Benes filter.

[0020] The state vectors can each characterize a bounding body for the respective object detected in the image in the environment, in particular the arrangement (e.g., position, extent, and / or orientation) of the bounding body in the environment. The bounding body can be a simple geometric body that encloses the respective object. The bounding body can, for example, be an ellipsoidal or spherical bounding body (bounding sphere) or, preferably, a cuboid-shaped bounding body (bounding box).

[0021] The state vectors can each comprise elements that indicate the position(s) of the enveloping body along one, two, or preferably three absolute spatial directions of a reference system of the environment (i.e., in a coordinate system fixed relative to the environment). The state vectors can, for example, contain a first coordinate (x) along a first absolute spatial direction, a second coordinate (y) along a second absolute spatial direction, and / or a third coordinate (z) along a third absolute spatial direction, and / or elements that can be transformed into the corresponding coordinates.

[0022] Alternatively or additionally, the state vectors may each comprise elements that specify one, two, or preferably three angles of rotation, which indicate the orientation of the enveloping body relative to the three absolute spatial directions. The state vectors may for example, contain a first angle of rotation (φ) with respect to a first axis of rotation, a second angle of rotation (φ) with respect to a second axis of rotation and / or a third angle of rotation (φ) with respect to a third axis of rotation and / or elements which can be transformed into the corresponding angles of rotation. The angles of rotation can, for example, define a rotation transformation between the reference system of the environment and a reference system of the respective enveloping body. The angles of rotation can, for example, each be an angle of rotation about one of the absolute spatial directions of the reference system of the environment and / or an angle of rotation about a relative spatial direction of the reference system of the respective enveloping body. The angles of rotation are preferably Eulerian angles.

[0023] Alternatively or additionally, the state vectors may each comprise elements that indicate the dimensions of the enveloping body along one, two, or preferably three relative spatial directions of the reference system of the respective enveloping body. The state vectors may, for example, have a first dimension (di, e.g., a length) along a first relative spatial direction, a second dimension (d 2 , e.g. a width) along a second relative spatial direction and / or a third dimension (d 3 , e.g., a height) along a third relative spatial direction and / or elements that can be transformed into the corresponding dimensions. The relative spatial directions can, for example, correspond to axes of the enveloping body, in particular to principal axes of the enveloping body. In the case of cuboid-shaped enveloping bodies, the relative spatial directions can, for example, run parallel to the cuboid edges.

[0024] Determining the state vectors based on the corresponding image can comprise determining an orientation of a surface of the object detected in the image. Preferably, the orientation of the surface of the object in question is determined at a plurality of points. The points can be distributed over the entire part of the object contained in the image (e.g., imaged and / or visible), for example in order to determine the orientations of all surfaces of the object contained in the image. Determining the orientation can, for example, comprise determining a normal vector to the surface, e.g., at the corresponding point. In one example, the orientation is determined in the form of an object-specific normal point cloud as described below. The rotation angles contained in the state vectors can be determined based on the determined surface orientation(s), for example, based on one or more modes (e.g.,B. Accumulations such as angular ranges in which there are increased surface orientations) of the angular distribution, for example one or more maxima of the angular distribution of the determined surface orientations, in particular based on a dominant mode or a dominant (e.g. highest) maximum of the angular distribution. The modes or maxima can, for example, each be linked to a specific surface of the object and thus to a specific relative spatial direction of the object's reference system.

[0025] The uncertainty(s) of one or more, preferably all, rotation angles contained in the state vectors can be determined based on a width of the angular distribution of the determined surface orientations, for example based on the width(s) of one or more modes (e.g. one or more maxima) of the angular distribution of the determined surface orientations, in particular based on the width of the dominant mode or the dominant maximum of the angular distribution. The width can, for example, characterize an angular range over which the corresponding mode or the corresponding maximum of the angular distribution extends, for example an angular range between a lower percentile (e.g. 10%) and an upper percentile (e.g. 90%) of the corresponding mode or the corresponding maximum, a standard deviation of the corresponding mode orof the corresponding maximum and / or a full width at half maximum of the corresponding mode or maximum. The width(s) of the angular distribution of the determined surface orientations can provide a good estimate of the uncertainties with which the surface orientations and thus the rotation angles can be determined.

[0026] The uncertainty(s) of one or more, preferably all, positions of the enveloping body and / or the uncertainty(s) of one or more, preferably all, dimensions of the enveloping body can each be determined taking into account an observation direction (viewing direction) of the corresponding image. The observation direction can influence how accurately the positions and / or dimensions of the enveloping body can be determined from the image. For example, a position or dimension in a direction at a larger angle, in particular perpendicular, to the observation direction can generally be determined with a lower uncertainty than a position or dimension in a direction at a smaller angle, in particular parallel to the observation direction.In a preferred embodiment, the uncertainties of the respective positions and / or the respective dimensions of the enveloping body are therefore each determined based on an angle between the observation direction of the corresponding image and one of the three relative spatial directions of the reference system of the respective enveloping body. Alternatively or additionally, the uncertainties of the respective positions and / or the respective dimensions of the enveloping body can also be determined based on an angle between the observation direction of the corresponding image. observation and one of the three absolute spatial directions of the surrounding reference system. The uncertainty in question can be determined, for example, using the scalar product between the direction of observation and the relevant relative or absolute spatial direction, e.g., it can be proportional to or equal to this scalar product.

[0027] In addition to elements that characterize the arrangement of the object in question in the environment, the state vectors can also each comprise one or more elements that characterize a classification of the object in question detected in the image into a plurality of object classes. The element(s) in question can, for example, indicate an object class from the plurality of object classes to which the object in question has been assigned, for example by means of one-hot coding. Preferably, the element(s) in question characterize probabilities with which the object in question has been assigned to one of the object classes. The classification can be determined by means of computer-aided image recognition, e.g. by means of a neural network. An element-specific uncertainty for the element(s) in question can, for example, be determined based on a confidence level output by the neural network.

[0028] The spatially resolved images can, for example, each comprise a two-dimensional camera image of at least part of the environment. The camera images can have been recorded using a camera device, in particular a camera device of a device according to the invention. The camera images can each contain a light intensity, preferably a spectrally resolved light intensity (e.g., as an RGB image), as a function of location. The camera images can each depict a specific field of view (a specific image section). In one example, the camera images are 360° images.

[0029] Alternatively or additionally, the spatially resolved images can, for example, each comprise a depth point cloud with depth information relating to at least part of the environment. The depth point clouds can have been recorded using a depth sensor, in particular a depth sensor of a device according to the invention, e.g. a 3D camera, a radar sensor or preferably a laser scanner such as a lidar sensor. The depth point clouds can each comprise a plurality of depth measurements ("points"), each of which is linked to a specific point and / or area in the environment. Each of the depth measurements can, for example, represent a distance between the observation point of the relevant image and the corresponding point or characterize an area in the surrounding area, e.g., as an absolute distance and / or as a travel time. The points of the depth point clouds can be arranged in a regular or irregular pattern.

[0030] In a preferred embodiment, the spatially resolved images each comprise a two-dimensional camera image of at least part of the environment and a depth point cloud with depth information regarding the relevant part of the environment, at least partially, preferably completely, overlapping the camera image. In one example, the spatially resolved images can comprise three-dimensional camera images with depth information, for example (two-dimensional) camera images in which some or all pixels are additionally provided with a depth or distance indication (thus forming the depth point cloud). The method according to the invention is not limited to specific types of observations or images and can be carried out with any type of observation and spatially resolved images that can be used to determine state vectors that characterize the arrangement of objects detected in the image in the environment.

[0031] For each of the observations, the method can comprise detecting objects in the corresponding two-dimensional camera image using computer-aided object recognition, e.g., using transformer-based and / or convolutional neural network-based object recognition. By detecting objects in the two-dimensional camera images, complex 3D object recognition can be avoided. Detecting the objects in the corresponding two-dimensional camera image can comprise segmenting the corresponding two-dimensional camera image into object masks, each of which is linked to one of the detected objects. Each of the object masks can specify a part or section, in particular a contiguous part or section, of the camera image that is linked to the object in question, e.g., in which the object in question is located and / or which is covered by the object in question.In some embodiments, the object masks can be filtered, for example, using a geometric and / or statistical filter, for example to filter out high-frequency components and / or edges of the object masks. Alternatively or additionally, detecting the objects can also include classifying the objects into a plurality of object classes, for example, to determine the relevant elements of the state vector.

[0032] The method may further comprise, for each of the observations, assigning points of the depth point cloud to one of the objects detected in the image. An object-specific depth point cloud can be determined for each object detected in the recording. The object-specific depth point cloud can comprise those points of the (complete) depth point cloud that are linked to the object in question, for example to a point and / or area of ​​the object. The object-specific depth point cloud is preferably determined on the basis of an object mask for the object in question, which can be determined, for example, as described above by segmenting the two-dimensional camera image. The object-specific depth point cloud can, for example, comprise all points of the (complete) depth point cloud that are located within the object mask (i.e. the object mask can be used as a clipping mask for the depth point cloud).

[0033] In some embodiments, the method may further comprise determining, for each of the observations, object-specific normal point clouds for objects detected in the recording based on the depth point cloud. The points of the object-specific normal point clouds may each characterize a normal vector of a surface of the object in question, e.g., indicate the direction of the normal vector. The object-specific normal point clouds may be determined based on the depth point cloud, e.g., by forming a discrete derivative, preferably by convolution. In one example, a normal point cloud is determined from the depth point cloud, and points of the normal point cloud are assigned to one of the objects detected in the recording, e.g., based on the object mask for the object in question, in order to determine the object-specific normal point cloud.Alternatively or additionally, the object-specific normal point cloud can be determined from the object-specific depth point cloud, for example by discrete derivation.

[0034] The state vectors and / or the element-specific uncertainties can be determined based on the object-specific depth point cloud and / or the object-specific normal point cloud. This can involve transforming the corresponding object-specific point cloud(s) into the reference system of the environment, for example, based on the observation point and the observation direction of the relevant image. In some embodiments, the corresponding object-specific point cloud(s) can be filtered, for example, using a geometric and / or statistical filter, for example to filter out high-frequency components and / or edge regions of the point cloud(s).

[0035] Elements of the state vectors that characterize the orientation of the object in question (e.g. some or all of the rotation angles cp, 0, ip) and / or the relevant element-specific uncertainties can be determined, for example, based on the angular distribution of the normal vectors characterized by the object-specific normal point cloud. The angular distribution of the normal vectors can, for example, have one or more modes (e.g. clusters such as angular ranges in which there are more normal vectors), in particular one or more maxima. These modes or maxima can each be linked to one of the relative spatial directions of the reference system of the object or the enveloping body and / or to one of the rotation angles cp, 0, ip. The rotation angles cp, 0, ip can, for example, be determined based on a center of gravity, a geometric center, a median and / or a maximum value (maximum) of one or more of these modes orMaxima (preferably at least of the dominant mode / dominant maximum of the angular distribution) can be determined. The relevant element-specific uncertainties can be determined analogously based on one or more widths of the angular distribution of the normal vectors, in particular based on one or more widths of the clusters or maxima (preferably at least based on the width of the dominant mode / dominant maximum of the angular distribution), for example as described above.

[0036] Elements of the state vectors that characterize the extent of the object in question (e.g. some or all of the dimensions di, d 2 , d 3), and / or the relevant element-specific uncertainties can be determined, for example, based on the extent of the object-specific point cloud(s). For this purpose, the point cloud(s) can be suitably transformed, for example, depending on the orientation of the object in question (e.g., the rotation angles cp, 0, ip). The point cloud(s) can, for example, be transformed into the reference system of the object or envelope in question. The dimensions di, d 2 , d 3of the object in question can be determined, for example, based on the distances between the furthest apart points and / or modes of the point cloud(s) along the corresponding spatial direction of the relative reference system, e.g. corresponding to these distances. The element-specific uncertainties in question can be determined based on the distribution of the points of the point cloud(s) along the corresponding spatial direction of the relative reference system, for example based on a width and / or gradient of the distribution, in particular based on the widths of the furthest apart modes along the corresponding spatial direction of the relative reference system. Alternatively or additionally, the uncertainties in question can be determined taking into account the direction of observation of the corresponding image, for example as described above.

[0037] Elements of the state vectors that characterize the position of the object in question (e.g., some or all of the positions x, y, z) and / or the relevant element-specific uncertainties can be determined, for example, based on a center of the object-specific point cloud(s), for example, based on a center of gravity, a geometric center, and / or a median of the object-specific point cloud(s). In a preferred embodiment, the elements of the state vectors that characterize the position of the object in question (e.g., some or all of the positions x, y, z) and / or the relevant element-specific uncertainties are determined based on a front surface of the object pointing in the direction of the observation point (opposite the observation direction), e.g., based on that surface of the object whose normal vector has the smallest angle relative to the direction from the object to the observation point.The positions x, y and / or z can be determined, for example, by adding half of the corresponding dimension di, d to the position of a corner and / or edge of the front surface. 2 or d 3 is added. The corresponding quantities can already be known from determining the object's dimensions, so no additional calculation steps are necessary. Alternatively or additionally, the positions x, y, and / or z can be determined, for example, by explicitly determining a center of the front surface.

[0038] In some embodiments, the elements of the state vectors that characterize the orientation, extent and / or position of the object in question (e.g., some or all of the rotation angles <p, 0, r , einige oder alle der Abmessungen di, d 2 , d 3and / or some or all of the positions x, y, z) and / or the relevant element-specific uncertainties can be determined using principal component analysis (PCA). However, this generally requires that at least three sides of the object in question are visible in the image.

[0039] When determining the element-specific uncertainties, one or more of the following contributions can also be taken into account: one or more uncertainties associated with the respective object mask (e.g., a confidence level of a neural network that determines the object mask); a size of the object mask, a size of the object-specific depth point cloud (e.g., a number of points) and / or a size of the object-specific normal point cloud (whereby smaller object masks and / or point clouds can, for example, be associated with a greater uncertainty); a distance between the object in question and the observation point (whereby a greater distance can, for example, be associated with a greater uncertainty); a result of filtering the object mask, the object-specific depth point cloud and / or the object-specific normal point cloud (whereby, for example, a higher proportion of filtered points indicate poorer data (e.g. a poorer mask and / or point cloud) and may therefore be associated with a greater uncertainty), an effect of one or more element-specific uncertainties relating to the elements of the state vector associated with the orientation of the object on the uncertainties of the elements associated with the position and / or extent of the object (e.g. by way of error propagation); an uncertainty (e.g. depending on a similarity measure) when comparing the state vectors determined for different observations; an uncertainty of the observation point and / or an uncertainty of the observation direction.

[0040] In some embodiments, the determined state vectors can be filtered based on the element-specific uncertainties. For example, state vectors and / or entire observations can be excluded where one or more uncertainties exceed certain thresholds and / or a certain number of uncertainties each exceed a corresponding threshold.

[0041] In a preferred embodiment, the comparison of the state vectors determined for different observations and the determination of the error-weighted state vector estimates are performed iteratively. This can be done by iterating over the multitude of observations / images, for example, by additionally considering a new ("current") observation m in each iteration step. This approach is particularly suitable for applications in which object recognition is to take place in real time.

[0042] The iteration steps can, for example, each determine a similarity measure between previous state vector estimates bb] determined in the previous iteration step m-1.” -1 and the state vectors bb / determined for the current observation m 1 . The previous state vector estimates can be stored in an object list = {bb / ^ / bb / 1 , ...} The state vectors determined for the current observation can be stored in an observation list I m= {bb™ , bb™ , ...}. There are no particular restrictions regarding the similarity measure, and a similarity measure as described above can be used, for example a distance between the objects and / or bounding bodies characterized by the respective vectors (e.g. their centers), a spatial overlap between the objects and / or bounding bodies characterized by the respective vectors, and / or preferably a bounding box disparity. The determined similarity measures can, for example, be summarized in the form of a matrix, where each row (or column) is linked to a state vector determined for the current observation m, and each Column or row with a previous state vector estimate determined in the previous iteration step mi.

[0043] The iteration steps can furthermore each assign the state vectors bb]” determined for the current observation to a corresponding state vector estimate of the previous state vector estimates bb]" -1based on the similarity measure. For example, each previous state vector estimate can be assigned the state vector determined for the current observation that has the highest similarity and has not already been assigned to another state vector estimate (e.g. because it has an even higher similarity to this one). Preferably, the state vectors determined for the current observation are only assigned to a previous state vector estimate if the relevant similarity measure exceeds a predefined similarity threshold, e.g. (depending on the definition of the similarity measure) is greater or smaller than a predefined threshold. State vectors that cannot be assigned to a previous state vector estimate in this way (e.g. because they do not have a sufficiently high similarity to any of these estimates) can be classified as new, previously unobserved objects.These state vectors can be used as a first state vector estimate for the object in question and can be converted into a current object list J. m of current state vector estimates. Accordingly, in the first iteration step, initial state vector estimates can be obtained from the state vectors determined for the objects detected in the first image.

[0044] The iteration steps can further comprise determining current state vector estimates bb]” based on a respective previous state vector estimate bb]” -1 and the state vector bb] determined for the current observation associated with the respective previous state vector estimate, taking into account the uncertainties of the elements of the respective state vector ff]” and the uncertainties of the elements of the respective previous state vector estimate ff]” -1This is preferably done using an iterative stochastic filter, e.g. as described above.

[0045] In a particularly preferred embodiment, the current state vector estimates bb]" are determined using a Kalman filter. For this purpose, for example, the uncertainties ff]" of the elements of the respective state vector and the uncertainties ff]" -1 of the elements of the respective previous state vector estimation, a Kalman gain can be determined, for example as = ff _jm-l ( ( ^am-1 + , ff —m Kj j ) 1 ■ Using the Kalman gain, the current state vector estimate can be determined by error-weighted interpolation between the previous state vector estimate and the corresponding state vector, for example as The current state vector estimate can then be added to the current object list J mThe element-specific uncertainties of the current state vector estimate can be determined from the element-specific uncertainties of the previous state vector estimate and the Kalman gain, for example as where I denotes the identity matrix.

[0046] In the case of moving objects, the movement of the objects (which can be described, for example, by additional elements of the state vectors that characterize, for example, a velocity and / or acceleration of the respective object) can also be taken into account, for example by appropriately transforming the previous state vector estimate and / or the current state vector. In this case, the elements of the state vectors can be correlated with each other, and the uncertainties of the elements of the respective state vector and the state vector estimate can be described accordingly by an uncertainty matrix instead of an uncertainty vector.

[0047] In some embodiments, at least one of the state vectors determined for the observations may be an incomplete state vector, in which one or more of the elements of the state vector are empty. This may be due, for example, to corresponding parts of the object in question not being visible in the image. For example, the object may be oriented relative to the observation direction such that only a front surface of the object is visible. In such a case, for example, a dimension and / or position of the object along the observation direction cannot be determined.

[0048] In one embodiment, the method further comprises determining an observation point and / or an observation direction for a future observation (e.g., the next observation) based on the element-specific uncertainties of the determined state vectors and / or based on the element-specific uncertainties of the error-weighted state vector estimates. The observation point or the observation direction can, for example, be chosen so that missing (empty) elements of the error-weighted state vector estimates can be determined and / or certain state vector elements can be determined again and / or with a lower uncertainty. For example, a dimension of an object in a certain direction may be unknown or known only with great uncertainty. The observation point or direction can then be chosen, for example, so that the object in question is observed perpendicular or approximately perpendicular to this direction.

[0049] The method according to the invention enables object recognition by determining corresponding state vector estimates based on a large number of observations, taking into account observation- and element-specific uncertainties. This improves the accuracy of the state vector estimates and allows for realistic estimation of element-specific uncertainties in the state vector estimates. The method according to the invention delivers explainable and comprehensible results (white box), thus facilitating a better understanding of the estimated state vectors and the associated uncertainties. The method according to the invention enables three-dimensional object recognition without requiring neural networks for three-dimensional object recognition and corresponding training data.The method according to the invention can also be used for object recognition in real time and is not limited to certain types of observations, spatially resolved recordings and methods for determining the state vectors, but can also be used, for example, to combine data from different recording devices (e.g. sensors) and / or object detectors.

[0050] According to a second aspect of the present invention, a computer program for detecting objects in an environment is provided. The computer program comprises instructions which, when executed by a processor, cause the processor to execute the method according to the first aspect of the invention according to any of the embodiments described herein. The computer program can, in particular, be stored on and / or executed by a device according to the invention.

[0051] According to a third aspect of the present invention, a machine-readable storage medium for detecting objects in an environment is provided. The storage medium comprises instructions which, when executed by a processor, cause the processor to perform the method according to the first aspect of the invention according to any of the steps described herein. The machine-readable storage medium can, for example, be intended for use with a device according to the invention and, in particular, be provided as part of such a device. The machine-readable storage medium can be a volatile or, preferably, a non-volatile storage medium.

[0052] According to a fourth aspect of the present invention, a device for detecting objects in an environment is provided. The device comprises a recording device configured to record spatially resolved images of at least part of an environment of the device. The device further comprises a control unit configured to carry out the method according to the first aspect of the invention according to any of the embodiments described herein in order to detect objects in the environment of the device based on images recorded by the recording device.

[0053] The recording device can comprise a camera device and / or a depth sensor. The camera device can be configured to record two-dimensional camera images of at least part of the environment of the device, for example as described above. The depth sensor can be configured to record depth point clouds with depth information relating to at least part of the environment, for example as described above. Preferably, the depth sensor is configured to record depth point clouds which each overlap with one of the camera images recorded by the camera device, for example are recorded in the same observation direction and / or from the same observation point.

[0054] The device may be motorized and, for example, comprise an electric motor that drives the device, allowing the device to be moved within the environment or to move independently, for example, to approach various observation points. Alternatively or additionally, the device may be portable (for example, by an operator or user) and / or configured to be mounted on a motorized drive device.

[0055] The camera device can be configured as a 360° all-round camera. The depth sensor can be configured, for example, as a radar sensor and / or laser scanner, in particular as a lidar sensor. In some embodiments, the recording device can be configured as a 3D camera device configured to record three-dimensional camera images with depth information (i.e., functions as a combined camera and depth sensor device).

[0056] The control unit may include a processor and a storage medium. The storage medium may contain instructions that can be executed by the processor to provide the functionality described herein. The storage medium may be the storage medium according to the invention. SHORT DESCRIPTION OF THE CHARACTERS

[0057] The invention is explained in more detail below using exemplary embodiments with reference to the accompanying drawings. The figures show schematically:

[0058] Fig. 1a: an exemplary environment with objects arranged therein;

[0059] Fig. lb: one of the objects from Fig. la and a shell for this object;

[0060] Fig. 2: a flowchart of a method for detecting objects according to an example;

[0061] Fig. 3: the determination of an error-weighted state vector estimate for an object according to an example;

[0062] Fig. 4: a flowchart of a method for detecting objects by iteratively determining error-weighted state vector estimates according to an example;

[0063] Fig. 5: the iterative determination of an error-weighted state vector estimate for an object according to an example;

[0064] Fig. 6: a flowchart of a method for determining state vectors for detected objects and associated element-specific uncertainties according to an example;

[0065] Fig. ad: the determination of object-specific depth and normal point clouds according to an example;

[0066] Fig. 8: a machine-readable storage medium and a computer program for detecting objects in an environment according to an example, and

[0067] Fig. 9: a device for detecting objects in an environment according to an example. DESCRIPTION OF THE EMBODIMENTS

[0068] Fig. 1a shows an exemplary environment 100 with a plurality of objects 102 arranged therein. The environment 100 may, for example, be an industrial or warehouse facility, and the objects 102 may, for example, be manufacturing parts and / or containers such as containers and / or packages. In another example, the environment 100 may be a street environment, and the objects 102 may, for example, be vehicles and / or pedestrians.

[0069] Spatially resolved images of the surrounding area 100 or a part of it can be taken iO4m-i, 104m from different observation points io6 m -i, io6 m The spatially resolved images iO4 m -i, iO4 meach represent an observation and can, for example, include a two- or three-dimensional camera image and / or a depth point cloud with depth information regarding the environment or the relevant part thereof.

[0070] In a spatially resolved image iO4 m -i, iO4 m the objects 102 contained therein can be recognized by means of computer-aided image recognition and for each of the recognized objects a state vector bbi = (bi,i, bi, 2 , ...) can be determined, which characterizes the arrangement of the respective object 102 in the environment 100. The state vector can, for example, characterize an enveloping body 108 for the respective object 102 and the arrangement of this enveloping body 108 in the environment 100.

[0071] This is shown by way of example in Fig. 1b, which shows one of the objects 102 from Fig. 1a and an associated enclosing body 108 in an enlarged view. The enclosing body 108 is cuboid-shaped in this case (i.e., a "bounding box"). The state vector can specify the positions of the enclosing body 108 (e.g., the center or a corner thereof) along three absolute spatial directions of a reference system of the environment 100, for example, the coordinates x, y, and z along the x, y, and z directions from Fig. 1a, 1b. The state vector can further specify three rotation angles α, θ, α (e.g., the three Eulerian angles), which specify the orientation of the enclosing body 108 relative to the three absolute spatial directions. In addition, the state vector can specify the dimensions dd 2 , d 3of the enveloping body 108 along three relative spatial directions of a reference system of the enveloping body 108 (which may be rotated relative to the reference system of the environment 100), for example the dimensions along the main axes and / or side edges of the enveloping body 108 as illustrated in Fig. 1b. Furthermore, the state vector may comprise one or more elements c, which indicate a classification of the respective object 102 (e.g., classification probabilities into a plurality of object classes (e.g., "Production Part A", "Production Part B", "Container", and "Package" or "Automobile", "Truck", "Motorcycle", "Bicycle", and "Pedestrian"). Accordingly, the state vector may, for example, have the 10-dimensional form bbi = (x, y, z, cp, 0, ip, di, d 2 , d 3 , c) accept.

[0072] For bounding bodies without a clear orientation (such as a bounding cuboid where the front and back are indistinguishable), a parameterization as described above can lead to an overdefinition, so that several state vectors can describe the same arrangement of the bounding body in the environment. For example, a 90 0 - Rotation around the axis with dimension d 3 linked relative spatial direction equivalent to an exchange of the dimensions di and d 2 These symmetries can be taken into account by appropriately restricting the angular ranges for the rotation angles cp, 0, ip. For example, the rotation angles can be restricted to the range from 0° to 90° (e.g., using a modulo operation), whereby when leaving this range or projecting onto this range (e.g., using a modulo operation), the dimensions dd 2 and d 3 can be swapped in pairs.

[0073] Fig. 2 shows a flowchart of an exemplary method 200 for detecting objects 102 in an environment 100 based on a plurality of observations according to the first aspect of the present invention. The method 200 is computer-implemented and can be carried out, for example, by means of a device according to the fourth aspect of the present invention, for example by the control unit 902 of the device 900 from Fig. 9. The execution of the method 200 is not limited to the order indicated by the flowchart in Fig. 2. As far as technically possible, the steps of the method 200 can be carried out in any order, in particular at least partially simultaneously. In some embodiments, the method 200 can be carried out iteratively, for example similarly to the method 400 from Fig. 4.

[0074] In step 202, state vectors are calculated for each of the observations based on the corresponding spatially resolved image iO4 m -i, 104m are determined, wherein the state vectors each characterize the arrangement of an object 102 detected in the recording in the environment 100. For example, as described above, for each of these objects 102, the positions (x, y, z), the angles of rotation (cp, 0, ip), the dimensions (di, d 2 , d 3 ) and optionally a classification c can be determined, for example by means of computer-aided image recognition, e.g. as described below with reference to Fig. 6. For the elements of the state vectors, an element-specific uncertainty is also determined based on the corresponding image 104m, iO4m-i, for example, in the above example, the element-specific uncertainties Oi = (ox, o y , o z , o <p, oe, Oip, Odi, Od 2 , Od 3 , o c), for example as described below with reference to Fig. 6.

[0075] In step 204, the state vectors determined for different observations are compared to determine sets of state vectors, each associated with a common object 102. For this purpose, the state vectors can be compared directly with each other, for example, to determine the most similar state vectors in the different observations that presumably describe the same object. Alternatively, the comparison can be performed indirectly, for example, by comparing the state vectors of one observation with state vector estimates determined from state vectors determined for other observations, for example, as described below for the method 400 of Fig. 4.

[0076] In step 206, for each of the sets of state vectors linked to a common object 102, an error-weighted state vector estimate bbj is determined for the respective object 102. The state vector estimate is determined based on the respective set of state vectors, taking into account the uncertainties of the elements of the respective state vectors. For example, an average value can be formed from the state vectors element by element (i.e., an x-average, a y-average, a z-average, and so on), wherein the elements of the state vectors are each weighted depending on (e.g., proportional to) the uncertainty of the respective element. Preferably, the error-weighted state vector estimates are determined using an iterative stochastic filter, for example, as described below for the method 400 of Fig. 4.

[0077] The determination of an error-weighted state vector estimate for an object 102 is shown as an example in Fig. 3 using the dimensions di, d 2 and d 3 In this example, two spatially resolved images are used to show iO4 m -i, 104m each state vectors bb -1 , bb (including dimensions dd 2 and d 3 ) for the object 102 (or its envelope 108). The image iO4 m -i is recorded from an observation direction that is almost perpendicular to a front surface of the object 102 or the enveloping body 108 (ie, runs almost parallel to a side edge of the enveloping body 108). One of the dimensions (dj) is therefore difficult to determine and is accordingly associated with a large uncertainty Odi, while the other two dimensions (d 2 , d 3 ) respectively can be determined with a significantly lower uncertainty. Image 104m is taken from a different observation direction, which differs significantly from the normal vectors of the surfaces of object 102 or the enveloping body 108. Accordingly, in this case, the dimension di can be determined with a significantly lower uncertainty. For example, the uncertainties of the other two dimensions can be larger than in image iO4. m -i. By element-wise error-weighted averaging, iO4 can be calculated from the two images. m -i, iO4 m An optimized state vector estimate bb™ can be determined whose uncertainty with respect to each of the dimensions (or each element) can be smaller than the corresponding uncertainties of the individual observations. In particular, the more precise determination of di based on image 104m can be given significantly greater weight than the less precise determination based on image iO4.m -i.

[0078] Fig. 4 shows a flowchart of another exemplary method 400 for detecting objects 102 in an environment 100 based on a plurality of observations according to the first aspect of the present invention. The method 400 is computer-implemented and can be carried out, for example, by means of a device according to the fourth aspect of the present invention, for example by the control unit 902 of the device 900 from Fig. 9. The execution of the method 400 is not limited to the order indicated by the flowchart in Fig. 4. As far as technically possible, the steps of the method 400 can be carried out in any order, in particular at least partially simultaneously.

[0079] Method 400 is an iterative method in which a plurality of observations are analyzed sequentially (e.g., in real time). The current observation is indexed in Fig. 4 by the counting variable m. In each iteration step, steps 402 to 408 are executed for the current observation m. The iterative determination of an error-weighted state vector estimate for an object 102 using method 400 is illustrated by way of example in Fig. 5.

[0080] In step 402, similar to step 202 of the method 200, state vectors bb are determined for the current observation m based on the corresponding recording. 1, which characterize the arrangement of the objects detected in the image 104m in the environment, as well as element-specific uncertainties for the elements of the state vectors. This is preferably done as described below with reference to Fig. 6. The state vectors determined for the current observation can be stored in an observation list I m = {bb™ , bb™ , ...}, where the second lower index indexes the object in question. [oo8i] In step 404, the state vectors bb]" determined for the current observation are taken from the observation list I m with the previous state vector estimates bb determined in the previous iteration step mi]" -1 which are in a previous object list / m-1 = {bb i 1 -1 , bb]^ -1 , ...} can be summarized. For this purpose, a distinction can be made between the elements of the lists l m and A similarity measure is determined pairwise, which quantifies the similarity of the vectors in question, for example a “bounding box disparity” as described in MG Adam et al., “Bounding box disparity: 3d metrics for object detection with full degree of freedom”, IEEE ICIP 2022, Bordeaux, France (Oct 2022).

[0082] In step 406, based on the comparison in step 404 (e.g. based on the similarity measure), the state vectors bb]" determined for the current observation are extracted from the observation list I m one of the previous state vector estimates bb]" -1 from the object list J m~ assigned. For this purpose, each previous state vector estimate can be assigned the state vector determined for the current observation which has the highest similarity and has not already been assigned to another state vector estimate (e.g. because it has an even higher similarity to this one), whereby the assignment only occurs if the relevant similarity measure exceeds a predetermined similarity threshold.

[0083] In step 408, a current object list J m determined. For this purpose, the previous state vector estimates bb]" -1 by means of an iterative stochastic filter based on the state vector bb]" assigned to the respective previous state vector estimate from the current observation list I mupdated to determine a current state vector estimate bb]. The iterative stochastic filter takes into account both the element-specific uncertainties o -1 of the previous state vector estimates bb]" -1 as well as the element-specific uncertainties <r]" des betreffenden Zustandsvektors bb]". In einer bevorzugten Ausführungsform ist das iterative stochastische Filter ein Kalman-Filter, wobei die aktuelle Zustandsvektor-Schätzung bb]" und ihre elementspezifischen Unsicherheiten o]" beispielsweise anhand der folgenden Gleichungen ermittelt werden kann: K]" = o]" -1 (O]" -1 + ct]")" 1 bb]" = bb]" -1 + ]" (bb]" - bb]" -1 )

[0084] In addition to the current state vector estimates bb]", the current object list J m furthermore, some or all of the current state vectors bb]" are included, which are in Step 406 could not be assigned to any previous state vector estimate (e.g., because they do not show a sufficiently high similarity to any of these estimates). These state vectors may represent new objects that have not yet been observed in any image.

[0085] Subsequently, the method 400 can return to step 402 for the next iteration step (with iterator m incremented by one), for example, when a new observation / recording is available. Thus, the state vector estimates can be successively improved over a large number of observations.

[0086] Fig. 6 shows a flowchart of an exemplary embodiment of the determination of the state vectors and their element-specific uncertainties in step 402 of the method 400. The execution of step 402 is also not limited to the order indicated by the flowchart in Fig. 6. As far as technically possible, the substeps of step 402 can be executed in any order, in particular at least partially simultaneously.

[0087] In the example of Fig. 6, the spatially resolved images include iO4 meach as shown by way of example in Fig. 7a, a two-dimensional camera image 700 of at least a part of the environment 100 and a depth point cloud 702 overlapping with the camera image 700 and containing depth information (e.g. distance measurements) regarding the relevant part of the environment 100. The depth point cloud 702 can, for example, have been determined using a laser scanner, e.g. a lidar sensor, or can be the depth information of a three-dimensional camera image, the two other dimensions of which form the two-dimensional camera image 700. Each of the points of the depth point cloud 702 shown as crosses in Fig. 7a can, for example, indicate a distance between the observation point of the recording 104m and the corresponding point in the environment 100. The points can be arranged in a regular pattern (for example in a three-dimensional camera image) or as in Fig.7a may be schematically indicated in an (at least partially) irregular pattern or randomly arranged.

[0088] In step 402-1, objects 102 contained in the two-dimensional camera image 700 are recognized by means of computer-aided object recognition. This can be done, for example, using transformer-based and / or convolutional neural network-based object recognition, for example using the CBNetV2 model (cf. arXiv:2i07.00420 [cs.CV]). Alternatively or additionally, one or more of the methods available at https: / / papers- withcode.com / sota / instance-segmentation-on-coco or http: / / web.archive.org / web / 2O23iii3i55ioo / https: / / paperswithcode.com / sota / instance-segmentation-on-coco can be used, e.g., EVA, FD-SwinV2-G, Mask Frozen-DETR, BEiT-3, MasK DINO, ViT-Adapter-L, SwinV2-G, Soft Teacher + Swin-L, ViT-Adapter-L, and / or Mask DINO. Preferably, as shown schematically in Fig. 7b, an object mask 704 is determined for the respective object 102 by segmenting the camera image 700, wherein the object mask 704 can, for example, contain all pixels of the camera image 700 over which the respective object 102 extends.

[0089] In step 402-2, an object-specific depth point cloud 702A is determined for each of the objects detected in step 402-1, as schematically shown in Fig. 7d, for example, by assigning the points of the depth point cloud 702 to a coder (or to a non-associated object 102) using the object masks 704. The object masks 704 can be used, for example, as clipping masks to "cut out" the object-specific depth point clouds 702A from the depth point cloud 702.

[0090] In step 402-3, an object-specific normal point cloud 706A is determined in a similar manner for each of the objects detected in step 402-1, as also schematically illustrated in Fig. 7d. For this purpose, a normal point cloud 706 can first be determined by discrete derivation of the depth point cloud 702, e.g., by convolution with a kernel of the form (-1, +1), as schematically illustrated in Fig. 7c, wherein each of the points of the normal point cloud 706 illustrated as circles in Fig. 7c can indicate a normal vector of a surface in the environment 100 (e.g., a surface of an object 102). The points of the normal point cloud 706 can then be assigned to an associated object 102 (or no associated object 102) using the object masks 704.The object-specific depth point clouds 702A and / or the object-specific normal point clouds 706A may be filtered in some embodiments, for example with a statistical and / or geometric filter, for example to filter out projection errors at the mask edges (e.g. due to errors in the high-frequency components of the masks and / or point clouds).

[0091] In step 402-4, the state vectors for the objects detected in step 402-1 are determined using the object-specific depth point clouds 702A and the object-specific normal point clouds 706A. For this purpose, the following procedure can be followed for each of these objects, for example:

[0092] First, the orientation of the object in question can be determined. This can be done, for example, based on the angular distribution of the object-specific normal point cloud 706A (e.g., in the form of histograms), for example, by determining the dominant mode of the angular distribution (e.g., the largest maximum of the angular distribution). Based on this mode (e.g., a mean normal vector of this mode), the rotation angles cp, 0, ip can be determined.

[0093] Subsequently, the object-specific depth point cloud 702A can be transformed based on the rotation angles cp, 0, ip (e.g., rotated in a direction parallel or perpendicular to the mean normal vector of the dominant mode). Based on the position distribution of the points of the (e.g., suitably transformed) depth point cloud 702A, the dimensions di, d 2 , d 3be determined, e.g. based on the modes occurring therein, which may each be linked to a side surface of the object 102, and / or based on the maximum distance between the points in a certain direction.

[0094] The positions x, y, z of the enveloping body 108 in the reference system of the environment can be determined, for example, based on a center of gravity, median, and / or geometric center of the object-specific depth point clouds 702A. Preferably, the positions x, y, z are determined based on a front surface of the object 102 pointing in the direction of the observation point (opposite the observation direction), e.g., based on that surface of the object whose normal vector has the smallest angle relative to the direction from the object to the observation point (and can be linked, for example, to the dominant mode of the angular distribution of the object-specific normal point cloud 706A). The positions x, y, and / or z can be determined, for example, by adding half of the corresponding dimension di, d to the position of a corner and / or edge of the front surface. 2 or d 3 is added.

[0095] Finally, in step 402-5, the element-specific uncertainties of the elements of the state vectors are determined. For this purpose, the following procedure can be used for each of the objects detected in step 402-1:

[0096] The uncertainties Oqj, oe, cfy of the rotation angles can be determined based on the widths (e.g., the variances and / or half-widths) of one or more modes of the angular distribution of the object-specific normal point cloud 706A, in particular based on a width of the dominant mode of the angular distribution. The uncertainties Oq,, oe, cty of the rotation angles can, for example, be set equal to the width (e.g., the variance or the half-width) of the dominant mode of the angular distribution.

[0097] The uncertainties o x , o y , a z the positions and the uncertainties Odi, Od 2 , Od 3The dimensions can be determined in each case based on an angle between the observation direction of the corresponding image and one of the three relative spatial directions (basis vectors) of the reference system of the respective enveloping body 108, for example, set equal to the amount of the scalar product between the observation direction and the relevant relative spatial direction (which can be determined, for example, based on the amounts of the elements of the product from the transposed rotation matrix of the object 102 or the enveloping body 108 (depending on the angles of rotation <p, 0, ip) und dem Positionsvektor des Objektes 102 bzw. des Hüllkörpers 108, jeweils im Bezugssystem der Aufnahme, berechnet werden kann, wobei die Mat- rixmultiplikation die einzelnen Skalarprodukte des Beobachtungsvektors mit den Basisvektoren des Objekt-Koordinatensystems liefert).

[0098] Fig. 8 shows a schematic representation of an exemplary machine-readable storage medium 800 for detecting objects in an environment according to the third aspect of the present invention. The storage medium 800 can be a non-volatile storage medium, such as a flash memory. The storage medium 800 can be used as the storage medium 906 of the device 900 described below from Fig. 9 or can serve as a data carrier, e.g., to transfer the computer program 802 stored on the storage medium 800 to the storage medium 906 of the device 900.

[0099] An exemplary computer program 802 for detecting objects in an environment according to the second aspect of the present invention is stored on storage medium 800. Computer program 802 includes commands 804-808 (e.g., instructions and / or program statements) that, when executed by a processor such as processor 904 of device 900, cause the processor to execute the inventive method for detecting objects in an environment based on a plurality of observations according to one of the embodiments described herein.

[0100] In the example of Fig. 8, the computer program 802 contains instructions 804-808 for executing the method 200 from Fig. 2, namely instructions 804 for executing step 202 (i.e., for determining state vectors and element-specific uncertainties for the elements of the state vectors using a spatially resolved image), instructions 806 for executing step 204 (i.e., for comparing the state vectors determined for different observations in order to determine sets of state vectors each linked to a common object) and instructions 808 for executing step 206 (i.e., for determining error-weighted state vector estimates using sets of state vectors linked to a common object, taking into account the Uncertainties of the elements of the respective state vectors). Alternatively or additionally, the computer program 802 may also contain instructions for executing the method 400 of Fig. 4 and / or steps 402-1 to 402-5 of Fig. 6.

[0101] Fig. 9 shows a schematic representation (not to scale) of an exemplary apparatus 900 for detecting objects in an environment according to the fourth aspect of the present invention.

[0102] The device 900 has a control unit 902, which in the example of Fig. 9 comprises a processor 904 and a storage medium 906. The processor 904 can be embodied, for example, as a central processing unit (CPU), graphics processor (Graphics Processing Unit), field programmable gate array (FPGA), application-specific integrated circuit (ASIC), and / or a combination thereof. The storage medium 906 can comprise a volatile memory (e.g., a RAM or cache memory) and / or a non-volatile memory (e.g., a flash memory). The storage medium 906 can store instructions that can be executed by the processor 904 to provide the functionality described herein. The storage medium 906 can be embodied like the storage medium 800 of Fig. 8 and / or can store the computer program 802 of Fig. 8.

[0103] The device 900 further comprises a recording device 908 configured to record spatially resolved images of at least part of the environment of the device. In the example of Fig. 9, the recording device 908 comprises a camera device 910 and a depth sensor 912. The camera device 910 is configured to record two-dimensional camera images of at least part of the environment of the device 900 and can be embodied, for example, as a CCD or CMOS camera. The depth sensor 912 is configured to record depth point clouds with depth information relating to at least part of the environment.The depth sensor 912 is arranged on the same side of the device 900 as the camera device 910 and has the same or substantially the same orientation, so that the depth point clouds recorded by the depth sensor 912 each overlap with one of the camera images recorded by the camera device 910. The depth sensor 912 can be designed, for example, as a lidar sensor. The control unit 902 is coupled to the recording device 908 and can be configured to control the recording device 908, for example, to receive a spatially resolved image (e.g., a camera image and / or a depth point cloud) from the recording device 908 or to read it out from it. and / or to control the recording device 908 to capture a spatially resolved image (e.g., a camera image and / or a depth point cloud). In some embodiments, the depth sensor 912 can be integrated into the camera device 910, and the camera device 910 can be configured, for example, as a βD camera and / or βD scanner.

[0104] The control unit 902 is configured to carry out the inventive method for detecting objects in an environment based on a plurality of observations according to any of the embodiments described herein. As a result, the control unit 902 can detect objects in the environment of the device 900 based on images captured by the recording device 908. In particular, the control unit 902 can be configured to carry out some or all of the steps of the method 200 from Fig. 2, some or all of the steps of the method 400 from Fig. 4, and / or some or all of the steps 402-1 to 402-5 from Fig. 6.

[0105] The described embodiments of the invention and the figures serve only as examples. The invention may vary in form without changing the underlying functional principle. The scope of protection of the method according to the invention, the computer program according to the invention, the storage medium according to the invention, and the device according to the invention arises solely from the following claims.

Claims

Claims 1. Computer-implemented method (200, 400) for detecting objects (102) in an environment (100) based on a plurality of observations, wherein each of the observations comprises a spatially resolved image (104m, iO4 m -i) of at least a part of the environment (100), the method (200, 400) comprising: for each of the observations, determining state vectors (bbi) from the corresponding recording (iO4 m , iO4 m -i), where the state vectors (bbi) each represent the arrangement of a signal in the recording (104m, iO4 m -i) characterize the detected object (102) in the environment (100) and wherein for the elements of the state vectors (bbi) an element-specific uncertainty is additionally determined on the basis of the corresponding recording (iO4 m , iO4 m -i) is determined; Comparing the state vectors (bbi) determined for different observations to determine sets of state vectors (bbi) each associated with a common object (102); and for each of the sets of state vectors (bbi) associated with a common object (102), determining an error-weighted state vector estimate (bbj) for the respective object (102) from the set of state vectors (bbi), taking into account the uncertainties of the elements of the respective state vectors (bbi).

2. The method (200, 400) according to claim 1, wherein the error-weighted state vector estimate (bbj) is determined by means of an iterative stochastic filter, in particular a Kalman filter and / or a particle filter.

3. Method (200, 400) according to claim 1 or 2, wherein the state vectors (bbi) each represent an enveloping body (108) for the respective object in the recording (104™, iO4 m-i) characterize the detected object (102) in the environment (100), in particular wherein the state vectors (bbi) each comprise elements which indicate: the positions (x, y, z) of the enveloping body (108) along three absolute spatial directions of a reference system of the environment (100), three angles of rotation ( <p, 0, rp), welche die Orientierung des Hüllkörpers (108) relativ zu den drei absoluten Raumrichtungen angeben, sowie die Abmessungen (d d2, d3) des Hüllkörpers (108) entlang von drei relativen Raumrichtungen eines Bezugssystems des jeweiligen Hüllkörpers (108).

4. The method (200, 400) according to claim 3, wherein: determining the state vectors (bbi) based on the corresponding image (104m, iO4m-i) comprises determining an orientation of a surface of the object (102) detected in the image (104m, iO4m-i) at a plurality of points; and the uncertainties of the three rotation angles (cp, 0, ip) are each determined based on a width of the angular distribution of the determined surface orientations.

5. Method (200, 400) according to claim 3 or 4, wherein the uncertainties of the positions (x, y, z) and / or the dimensions (d b d2, d3) of the enveloping body (108) taking into account an observation direction of the corresponding image (iO4 m , iO4 m -i) can be determined.

6. Method (200, 400) according to claim 5, wherein the uncertainties of the positions (x, y, z) and / or the dimensions (d bd2, d3) of the enveloping body (108) based on an angle between the observation direction of the corresponding image (104m, iO4 m -i) and one of the three relative spatial directions of the reference system of the respective enveloping body (108) and / or based on an angle between the observation direction of the corresponding image (104m, iO4 m -i) and one of the three absolute spatial directions of the environment's reference system.

7. Method (200, 400) according to one of the preceding claims, wherein the state vectors (bbi) each comprise one or more elements which allow a classification of the respective image in the recording (104m, iO4 m -i) characterize the recognized object (102) into a plurality of object classes.

8. Method (200, 400) according to one of the preceding claims, wherein the spatially resolved images (104m, iO4 m -i) a two-dimensional camera image (700) of at least a part of the environment (100) and a depth point cloud (702) overlapping with the camera image (700) with depth information relating to the relevant part of the environment (100).

9. The method (200, 400) according to claim 8, wherein the method (200, 400) comprises, for each of the observations, detecting objects (102) in the corresponding two-dimensional camera image (700) by means of computer-aided object recognition, wherein detecting the objects (102) in the corresponding two-dimensional camera image (700) preferably comprises segmenting the corresponding two-dimensional camera image (700) into object masks (704) each associated with one of the detected objects (102).

10. The method (200, 400) according to claim 8 or 9, wherein the method (200, 400) further comprises, for each of the observations, assigning points of the depth point cloud (702) to one of the depth data points recorded in the image (104m, iO4 m-i) detected objects in order to determine an object-specific depth point cloud (702A) for each object (102) detected in the recording (104m, iO4m-i).

11. The method (200, 400) of claim 10, wherein the method (200, 400) further comprises, for each of the observations, determining object-specific normal point clouds (706A) for the objects in the image (iO4 m , iO4 m -i) recognized objects (102) based on the depth point cloud (702), wherein the points of the object-specific normal point clouds (706A) each characterize a normal vector of a surface of the object in question.

12. The method (200, 400) according to claim 10 or 11, wherein the state vectors (bbi) and / or the element-specific uncertainties are determined based on the object-specific depth point cloud (702A) and / or the object-specific normal point cloud (706A).

13. Method (200, 400) according to one of the preceding claims, wherein the comparison of the state vectors (bbi) determined for different observations and the determination of the error-weighted state vector estimates are carried out iteratively (bbj), wherein the iteration steps each comprise: Determining a similarity measure between previous state vector estimates determined in the previous iteration step (bb™ -1 ) and the state vectors (bb™) determined for the current observation; Assigning the state vectors (bb™) determined for the current observation to 5 a corresponding state vector estimate of the previous state vector Estimates (bb™ -1 ) based on the similarity measure; and Determining current state vector estimates (bb™) based on a respective previous state vector estimate (bb™ -1 ) and the respective previous state vector estimate (bb™ -1) assigned to the current observation, taking into account the uncertainties of the elements of the respective state vector (bb™) and the uncertainties of the elements of the respective previous state vector estimate (bb™ -1 ).

14. Method (200, 400) according to one of the preceding claims, wherein at least one of the state vectors (bbi) determined for the observations is an incomplete 15 state vector in which one or more of the elements of the state vector (bbi) are empty.

15. A computer program (802) for detecting objects (102) in an environment (100), comprising instructions which, when the program (802) is executed by a processor (904), cause the processor (904) to execute the method (200, 400) according to one of the preceding claims.

16. A machine-readable storage medium (800) for detecting objects (102) in an environment (100), comprising instructions which, when executed by a processor (904), cause the processor (904) to execute the method (200, 400) according to any one of claims 1 to 14.

17. Device (900) for detecting objects (102) in an environment (100), the device (900) comprising: a recording device (908) which is configured to take spatially resolved images (iO4 m , iO4 m -i) from at least a part of an environment (100) of the device (900); and a control unit (902) which is configured to carry out the method (200, 400) according to one of claims 1 to 14, in order to use images (104m, iO4) taken by the recording device (908) m -i) to carry out a detection of objects (102) in the environment (100) of the device (900).

18. Device (900) according to claim 17, wherein the recording device (908) comprises a camera device (910) and a depth sensor (912), wherein the camera device (910) is configured to record two-dimensional camera images (700) of at least a part of the environment (100) of the device (900) and the depth sensor (912) is configured to record depth point clouds (702) with depth information relating to at least a part of the environment (100), which point clouds each overlap with one of the camera images (700) recorded by the camera device (910).

Citation Information

Patent Citations

  • Cross-modal sensor data alignment

    CN114097006A

  • Method and apparatus for sensor fusion

    US20160314097A1

  • System and method for tracking detected objects

    US20220180117A1