System and method for object recognition
The system uses a depth learning model with a DNN to generate three-dimensional point clouds and provide uncertainty information for each feature, addressing the challenges of object recognition in varying conditions, thereby improving the reliability and robustness of autonomous systems.
Patent Information
- Application Number
- CN202510048003.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-12
- Filing Date
- 2025-01-13
- Publication Date
- 2025-07-15
AI Technical Summary
The existing deep neural networks cannot effectively estimate uncertainty in object recognition, resulting in insufficient reliability and safety of autonomous driving systems in complex environments.
The uncertainty estimation of the DNN model is extended by generating a three-dimensional point cloud and using a deep neural network to calculate the output characteristics of the object, while determining the associated uncertainty information and outputting this information in the bounding box.
It improves the reliability and robustness of object recognition, and can more accurately identify and evaluate the uncertainty of objects in the autonomous driving system, ensuring the safety and reliability of the system.
Smart Images

Figure CN120318784A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to systems and methods for object recognition, in which deep learning-based uncertainty estimation using point clouds is utilized. Background Art
[0002] For autonomous driving to be practical and safe, vehicles must be able to independently recognize objects, especially traffic signs, lane markings, people, bicycles, and other vehicles. Autonomous vehicles can understand the environment to effectively avoid accidents.
[0003] Advanced driver assistance systems (ADAS) and autonomous driving (AD) require an accurate representation of the vehicle's surrounding environment, where radar sensors are the most commonly used perception sensors. By using radar, the surrounding environment can be represented as a point cloud.
[0004] US20200050191 describes a general modeling of perception uncertainty based on ground truth data and sensor data. An abstract architecture for localization uncertainty modeling is described therein.
[0005] Here, reliable perception is not trivial, especially when the system is used in various situations. It should be taken into account that: sensors may be affected, for example, by different weather conditions. It is also possible that there are situations and objects that have not been observed before. However, this complexity cannot be detected using traditional algorithms. Therefore, machine learning ML, especially deep neural networks (DNNs), is mainly used here. And these deep neural networks must be protected in their aspect. Due to the data-driven nature of neural networks, this poses challenges.
[0006] Large-scale labeled data collected in different scenarios can be used to train deep neural networks. In addition to its own internal structure, the recognition performance of deep neural networks is also strongly affected by the amount of available data, the quality of ground truth labels, and the perceived surrounding environment. Therefore, it is important to predict the uncertainty of the deep neural network, which corresponds to the current recognition quality of the neural network.
[0007] Autonomous vehicles or other devices can be equipped with a large number of different sensors ranging from cameras, radars, and lidar systems to ultrasonic sensors. These sensors continuously provide a large amount of data about their surrounding environment. However, this data in its raw form cannot be easily used to control the corresponding device. This is where perception comes into play, i.e., the interpretation of the provided sensor data. Precise and reliable surrounding environment perception is crucial here for the reliable operation of the device, especially autonomous vehicles.
[0008] In the publication “Labels are not perfect: Inferring spatial uncertainty in object detection” by D. Feng, Z. Wang, Y. Zhou, L. Rosenbaum, F. Timm, K. Dietmayer, M. Tomizuka, and W. Zhan, IEEE Transactions on Intelligent Transportation Systems, Vol. 23, No. 8, pp. 9981–9994, 2022, the imperfection of labels in real datasets was considered. An index for evaluating label uncertainty was proposed. Using the uncertainty measure in the labeled data, a DNN model that creates the uncertainty inherent in the derived labels or markings was created.
[0009] The uncertainty evaluation for a neural classifier is described in DE 102021210566 A1. Based on a trainable evaluation network, a deviation from an expected correlation (such as the gradient norm) is identified. However, this only provides an uncertainty value for each classification (for a classifier) or for each object recognition (for an object detector). In order to integrate a DNN-based object detector into a complex ADAS or AD system, an uncertainty estimate is required for each feature (position, speed, classification, etc.) of the identified object. SUMMARY OF THE INVENTION
[0010] The present invention creates a system for object recognition according to claim 1 of the patent and a method for recognizing objects in the surrounding environment according to claim 7 of the patent.
[0011] According to a first aspect, the present invention creates a system for object recognition, the system having:
[0012] A distance sensor for generating a three-dimensional point cloud of the surrounding environment, wherein each data point of the three-dimensional point cloud has spatial coordinates of a position on the surface of an object in the surrounding environment of the distance sensor;
[0013] A deep neural network for calculating output features of an object based on the three-dimensional point cloud generated by the distance sensor, wherein for each output feature, associated uncertainty information is determined during its calculation; and
[0014] An output unit for outputting the object recognized based on the calculated output features in a bounding box assigned to the object, the bounding box reproducing the determined uncertainty information for the calculated output features.
[0015] The deep neural network (DNN) used is an artificial neural network (KNN) with a large number of intermediate layers.
[0016] Deep learning (German: mehrschichtiges Lernen (multi-layer learning), tiefes Lernen (deep learning), or tiefgehendes Lernen (in-depth learning)) here represents a method of machine learning, i.e., ML, which uses an artificial neural network (KNN) with many intermediate layers (English: hidden layer) between the input layer and the output layer, and thus forms a large-scale internal structure.
[0017] According to another aspect, the present invention also creates a method for recognizing an object in the surrounding environment, including the following steps:
[0018] Generating a three-dimensional point cloud of the surrounding environment, where each data point of the three-dimensional point cloud has spatial coordinates of a position on the surface of an object in the surrounding environment;
[0019] Based on the generated three-dimensional point cloud, calculating output features of the object through a deep neural network, where associated uncertainty information is determined for each output feature during its calculation; and
[0020] Outputting the object recognized based on the calculated output features in a bounding box assigned to the object, where the bounding box reproduces the determined uncertainty information for the calculated output features.
[0021] The basic idea of the present invention is to extend the deep learning model for object recognition of the point cloud as follows: such that it provides uncertainty information for each output feature.
[0022] The uncertainty information can provide a deeper understanding of the behavior of the object detector, which constitutes an important part of data-driven engineering, especially in the release aspect of safety-critical applications.
[0023] In a possible implementation of the method for recognizing an object in the surrounding environment according to the present invention, the uncertainty information is determined during the training phase of the deep neural network (DNN) and during the execution of object recognition using the trained deep neural network (DNN).
[0024] This enables the uncertainty estimation to also be used for inference in various applications.
[0025] In the context of radar object recognition based on deep learning, the system 1 according to the invention and the method according to the invention are characterized in that they extend the DNN model by providing uncertainty information for each output feature AM.
[0026] In a possible implementation of the system for object recognition according to the invention, the output features include the 3D position of the object, the size of the object, the speed of the object, the orientation of the object, the classification of the object, and the recognition rate regarding the object.
[0027] These different output features describe various basic properties of the object, and these properties can also be evaluated in real time for different application scenarios and / or through subsequent data processing stages.
[0028] In a possible implementation of the system for object recognition according to the invention, each calculated output feature has an associated mean value and an associated standard deviation, which reproduce the determined uncertainty information for the output feature.
[0029] The mean value and the standard deviation can be effectively calculated in a short time.
[0030] In a possible implementation of the system for object recognition according to the invention, the distance sensor has a radar sensor that generates a three-dimensional radar point cloud.
[0031] The radar sensors work reliably. They are robust against fluctuations in environmental conditions and are widely used.
[0032] In a possible implementation of the system for object recognition according to the invention, the distance sensor has a LiDAR sensor that generates a three-dimensional LiDAR (Light Detection and Ranging) point cloud.
[0033] The LiDAR sensors are also widely used.
[0034] In a possible implementation of the method for recognizing objects in the surrounding environment according to the invention, the uncertainty of the ground truth labels of the training data set is calculated in the training phase for training the deep neural network.
[0035] In machine learning ML, ground truth is referred to in the context of data that allows checking the model quality.
[0036] In a possible implementation of the method for recognizing objects in the surrounding environment according to the invention, a heatmap with uncertainty information is generated in order to train the deep neural network using the heatmap loss.
[0037] The generated heatmap or attention map can also help to understand the decisions made by the deep neural network here and better identify possible errors in its actions.
[0038] In a possible implementation of the method for identifying objects in the surrounding environment according to the present invention, the generated heatmap with uncertainty information is used to perform bounding box regression, which generates a bounding box with uncertainty information.
[0039] In a possible implementation of the method for identifying objects in the surrounding environment according to the present invention, the generated bounding box with uncertainty information is used to train the deep neural network using the bounding box loss.
[0040] In a possible implementation of the method for identifying objects in the surrounding environment according to the present invention, the trained deep neural network outputs bounding boxes with associated uncertainty information generated by using bounding box regression based on the three-dimensional point cloud generated by the range sensor and based on other data.
[0041] The other data includes data regarding the sensing detection of the surrounding environment by the range sensor.
[0042] In a possible implementation of the method for identifying objects in the surrounding environment according to the present invention, the other data regarding the sensing detection of the surrounding environment by the range sensor includes, on the one hand, data regarding the quality of the used range sensor and, on the other hand, data regarding the environmental conditions present in the sensing detection of the surrounding environment.
[0043] Thereby, the reliability and robustness of object recognition are improved.
[0044] The system for object recognition according to the present invention can be used in a large number of different devices or facilities.
[0045] According to another aspect, the present invention creates a device, in particular a vehicle, in which a system for object recognition is integrated, the system having:
[0046] A range sensor for generating a three-dimensional point cloud of the surrounding environment, wherein each data point of the three-dimensional point cloud has spatial coordinates of a position on the surface of an object in the surrounding environment of the range sensor;
[0047] A deep neural network for calculating output features of an object based on the three-dimensional point cloud generated by the range sensor, wherein associated uncertainty information is determined for each output feature during its calculation; and
[0048] An output unit for outputting, in a bounding box assigned to the object, the object identified based on the calculated output features, the bounding box reproducing the determined uncertainty information for the calculated output features.
[0049] In one embodiment of the device, the device has a device for autonomous driving, in particular an autonomous household appliance.
[0050] In another embodiment of the device, the device has a monitoring device, in particular a monitoring device for monitoring a location or a building and access control.
[0051] In another embodiment of the device, the device has a traffic monitoring device for monitoring traffic participants. Description of the Drawings
[0052] Possible embodiments of the system for object recognition according to the invention and of the method according to the invention are described in more detail below with reference to the drawings. Among them:
[0053] Figure 1 A block diagram is shown for schematically representing a possible embodiment of the system according to the invention;
[0054] Figure 2 A flowchart is shown for schematically representing a possible embodiment of the method according to the invention;
[0055] Figure 3 A diagram is shown for clarifying the working mode of the system according to the invention and of the method according to the invention;
[0056] Figure 4 A schematic diagram is shown for clarifying the training process for training a deep neural network used in the system and method according to the invention;
[0057] Figure 5 A schematic diagram is shown for clarifying the operation of a deep neural network used in the system and method according to the invention;
[0058] Figures 6A - 6F A schematic diagram is shown for clarifying possible application examples of the system and method according to the invention. Detailed Description of the Invention
[0059] Figure 1 A possible embodiment of the system 1 for object recognition according to the invention is schematically shown in a block diagram.
[0060] The system 1 includes a distance sensor 2 for generating a three-dimensional point cloud PW of the surrounding environment. Each data point of the three-dimensional point cloud PW has spatial coordinates of a position on the surface of an object OBJ in the surrounding environment of the distance sensor 2.
[0061] Each point P of the point cloud PW includes the three-dimensional position (x, y, z) of the corresponding detection point and other features, such as speed, intensity, and measurement quality. The point cloud PW contains a large number of data points in three-dimensional space. Each point P corresponds to the X, Y, and Z coordinates of a position on the surface of the real object OBJ, and multiple points P together form the surface of the object OBJ. A point cloud or point cluster (English: pointcloud) forms a set of points P in a vector space, which has an unorganized spatial structure ("cloud"). The point cloud PW is described by the points P it contains.
[0062] To extract objects from the detected point cloud PW, especially from a radar point cloud, the trained deep neural network 3 of the system 1, or the deep neural network (DNN) 3, is used. The deep neural network 3 can recognize patterns in the points of the point cloud PW and represent the possible objects in the surroundings by a list of so-called bounding boxes BB with object classes.
[0063] To enable a device 7, such as a vehicle, to move safely in this surroundings, it includes a control system that can reliably recognize relevant objects OBJ in the environment of the device and take appropriate reactions respectively. The control system of the device 7 preferably includes the system 1 according to the invention for controlling the actuators of the device 7.
[0064] The system 1 according to the invention includes at least one deep neural network (DNN) 3, which is used to calculate the output feature AM of the object OBJ based on the three-dimensional point cloud PW generated by the distance sensor 2 of the system 1, wherein for each output feature AM, associated uncertainty information UINF is determined during its calculation process.
[0065] Figure 1 The shown system 1 further includes an output unit 4, which is used to output the object OBJ recognized based on the calculated output feature AM in the bounding box BB assigned to the object, and the bounding box reproduces the determined uncertainty information UINF for the calculated output feature AM. The bounding box (Bounding Box BB) is preferably a frame that can enclose the recognized object OBJ.
[0066] In a possible embodiment, the output unit 4 is a user interface UI with a display unit, which is used to show the bounding box BB enclosing the recognized object OBJ to a user, such as the driver of a bicycle or a passenger car. The bounding box BB can be displayed as a bounding box or frame (BB-UINF) with the determined uncertainty information UINF, so that the user can recognize how reliable the calculation is.
[0067] In another possible embodiment, the output unit 4 has a data interface to other data processing units, which further process the output data obtained, for example, to generate control signals for the actuators.
[0068] The uncertainty information UINF can be used as a quality metric for tracking multiple objects in subsequent modules. Bounding boxes BB with lower uncertainty are weighted higher here, indicating the confidence or certainty in correct object recognition and object tracking.
[0069] In a possible embodiment of the object recognition system 1 according to the invention, the output feature AM has the 3D position of the object, the size of the object, the speed of the object, the orientation of the object, the classification of the object, and / or the recognition rate regarding the object OBJ. Depending on the application, other output features are possible.
[0070] In a possible embodiment of the object recognition system 1 according to the invention, each calculated output feature AM has a calculated associated mean value and a calculated associated standard deviation, which reproduce the determined uncertainty information UINF for the output feature AM.
[0071] In a possible embodiment of the object recognition system 1 according to the invention, the distance sensor 2 of the system 1 has a radar sensor that generates a three-dimensional radar point cloud R-PW.
[0072] RADAR stands for "Radio Detection and Ranging" and represents an active transmission and reception method in microwaves up to the GHz range. Radar sensor systems are used for non-contact detection, tracking, and positioning of one or more objects OBJ by means of electromagnetic waves.
[0073] The radar antenna emits signals in the form of radar waves. If the radar wave hits the object OBJ, the signal changes and is reflected back to the sensor. The reflected signal arriving at the radar antenna contains information about the detected object. The received signal is processed to identify and locate the object OBJ using the data.
[0074] In another possible embodiment of the object recognition system 1 according to the invention, the distance sensor 2 of the system 1 has a LiDAR sensor that generates a three-dimensional LiDAR point cloud L-PW.
[0075] A LiDAR sensor (an abbreviation for "light detection and ranging", optical detection and ranging) is a distance measurement sensor, such as a radar or sonar. The LiDAR distance sensor 2 emits laser pulses that are reflected by the object OBJ. They can thereby sense the structure of their surroundings. They detect the reflected light energy and determine the distance to the object OBJ in order to create a 2D or 3D representation of the surroundings.
[0076] In a possible embodiment of the system 1 for object recognition according to the invention, the determination of uncertainty information is carried out during the training phase TP of the deep neural network 3 and during the object recognition carried out using the trained deep neural network 3.
[0077] Figure 2 A possible embodiment of the method for recognizing an object OBJ according to the invention is illustrated in a flowchart.
[0078] In the illustrated embodiment, the method has a plurality of main steps S1 to S3.
[0079] In a first step S1, a three-dimensional point cloud PW of the surroundings is generated, wherein each data point of the three-dimensional point cloud PW has spatial coordinates of a position on the surface of the object OBJ in the surroundings. The point cloud PW is generated by the distance sensor 2 or other sensors. The point cloud PW has, for example, a radar point cloud R-PW and a LiDAR point cloud L-PW.
[0080] In a further step S2, the deep neural network 3 calculates output features AM of the object OBJ based on the generated three-dimensional point cloud PW, wherein associated uncertainty information UINF is determined during the calculation for each output feature AM.
[0081] In a further step S3, the object identified based on the calculated output features, to which a bounding box BB of the object OBJ is assigned, is output, and the bounding box reproduces the determined uncertainty information UINF for the calculated output features AM.
[0082] In a possible embodiment of the method according to the invention, uncertainty information is determined during the training phase TP of the deep neural network (DNN) 3 and during the object recognition carried out using the trained deep neural network (DNN) 3.
[0083] In a possible embodiment of the method according to the invention, the trained deep neural network (DNN) 3 outputs bounding boxes BB-UINF each having associated uncertainty information based on the three-dimensional point cloud PW generated by the distance sensor 2 and based on other data regarding the sensing detection of the surroundings by the distance sensor 2 by using Bounding-Box-Regression.
[0084] In a possible embodiment of the method according to the invention, the other data includes data regarding the sensing detection of the surroundings by the distance sensor 2, data SQ (sensor quality data) regarding the quality of the used distance sensor 2, and / or data regarding the environmental conditions UB present in the sensing detection of the surroundings.
[0085] Various other embodiments are possible. In a possible embodiment, different detected point clouds PW, in particular the detected radar point cloud R-PW and the detected LiDAR point cloud L-PW, are processed in parallel by associated trained deep neural networks 3-R and 3-L for object recognition of an object OBJ in the surroundings.
[0086] According to another aspect, the invention also creates a device 7 in which a system 1 for object recognition is integrated.
[0087] The system 1 integrated in the device 7 includes: a distance sensor 2 for generating a three-dimensional point cloud PW of the surroundings, wherein each data point of the three-dimensional point cloud PW has spatial coordinates of a position on the surface of an object in the surroundings of the distance sensor 2; a deep neural network for calculating output features AM of an object OBJ based on the three-dimensional point cloud PW generated by the distance sensor 2, wherein associated uncertainty information UINF is determined during the calculation of each output feature AM; and an output unit 4 for outputting the object recognized based on the calculated output features AM in a bounding box BB-UINF assigned to the object, the bounding box reproducing the determined uncertainty information UINF for the calculated output features AM.
[0088] In the context of deep learning-based radar object recognition, the system 1 according to the invention and the method according to the invention are characterized in that they extend the DNN model by providing uncertainty information UINF for each output feature AM, as Figure 3 schematically shown. Object recognition determines and classifies an object or entity based on the provided sensor data.
[0089] Figure 3 For elucidating the system 1 according to the invention and the method according to the invention.
[0090] A point cloud PW with a large number of points P is sensed by the distance sensor 2 of the system 1 and fed to the trained DNN object detector 3. It calculates (PRE stands for prediction) the output feature AM of the object OBJ based on the three-dimensional point cloud PW generated by the distance sensor 2, where for each output feature AM, the associated uncertainty information UINF (Uncertainty) is determined or calculated during its calculation.
[0091] Each labeled bounding box gBB shows the recognized object. The bounding box BB describes features such as 3D position, size, speed, orientation, classification, and recognition rate. Each feature has an associated mean value and an associated standard deviation, which illustrate the uncertainty of the feature.
[0092] Figure 3 The mean value of the spatial uncertainty distribution (dashed box BB-UINF) is shown. The bounding box BB (BB-UINF) is output with the uncertainty information UINF for each feature AM.
[0093] In addition, the obtained uncertainty information UINF can be used for releasing safety-critical perception systems. A perception with a high degree of uncertainty indicates a safety-critical perception, which, for example, requires the intervention of a human operator, such as in autonomous driving.
[0094] The method and system 1 according to the present invention include: the calculation of the uncertainty distribution of the labeled features (represented, for example, by a Gaussian representation) and an adapted loss function VF (Loss) of the trainable system, which takes into account the prediction and labeled uncertainties in order to implicitly learn the uncertainty distribution for each feature.
[0095] In this way, each output feature AM recognized for each object is extended with the uncertainty information UINF, which provides basic information about the reliability of the object recognition model. Therefore, during the training process ( Figure 4 ) and the evaluation process ( Figure 5 ), the added uncertainty estimates are calculated. This enables the uncertainty estimates to also be used for inference in various applications.
[0096] A specific application is to use the uncertainty estimate to determine the variance information required for running an object tracker (such as a Kalman filter). In this case, the variance information is crucial for the construction of a reliable and safe ADAS or AD system that meets high automotive standards.
[0097] In addition, the provision of uncertainty information provides deeper insights into the behavior of object detectors, which forms an important part of data-driven engineering, especially with regard to the release of safety-critical applications. Additionally, the system 1 according to the invention can detect deterioration of object recognition or support sensor output according to specifications. It can also be used to identify incorrect sensor positions and to identify blindness.
[0098] Finally but not least, the method according to the invention can be incorporated into integrated learning methods or used for improved visualization of corresponding products in vehicle perception and vehicle human-machine interfaces (such as monitors, Carplay, etc.).
[0099] As Figure 4 and Figure 5 shown, the method according to the invention can be applied to the training and evaluation of DNN-based radar object recognition.
[0100] Figure 4 A method for determining uncertainty during the training phase TP of the DNN 3 is shown.
[0101] The training data can be classified (English "labeled") or unclassified. Depending on the situation, it can be called supervised learning (using labeled data for it) or unsupervised learning (using unclassified training data).
[0102] Supervised learning is a more common technique in practice for implementing machine learning ML. The first step in supervised machine learning is to collect, classify, or label the training data. For this purpose, a series of features corresponding to the test data are set.
[0103] For example, if one wants to implement an algorithm for handwriting recognition, images of handwritten characters and the correct assignment to numbers or letters are required. These data sets are also called ground truth.
[0104] Supervised learning is based on the so-called ground truth. There is training data in which the input parameters and the results are known. In deep learning, classification or regression is learned based on unstructured data. Here, the (deep) neural network 3 is used. In order for the deep neural network 3 to learn, there must be so-called ground truth for a sufficient amount of data, i.e., knowledge for correct classification.
[0105] In machine learning ML, ground truth is addressed in the context of data that allows for the inspection of the quality of a model. This means, for example, having data where it is known how the model should be evaluated against it, because, for example, humans have previously evaluated it manually (and with a high degree of certainty).
[0106] The quality of the available training data is a decisive factor in the development of an object recognition system based on artificial intelligence.
[0107] In supervised learning, a machine learning algorithm is presented with a data set in which the target variables are known. The algorithm learns to interpret the relationships and dependencies in the data of these target variables. After training, the quality of the predictions is evaluated in order to then apply the learned patterns to unknown data and create forecasts and predictions.
[0108] In a possible embodiment of the method according to the invention for recognizing objects in the surroundings, in a training phase TP for training a deep neural network (DNN) 3, first the uncertainty (GTL-UINF) of the ground truth labels GTL of the training data set used is calculated by the computing unit UINF-BER. After heatmap generalization (Heatmap-Generalisierung) HM-GEN, the ground truth labels with the uncertainty GTL-UINF are used to determine the heatmap loss HM-VF.
[0109] In a possible embodiment of the method according to the invention for recognizing objects in the surroundings, a heatmap HM-UINF with uncertainty information is also generated, and the uncertainty information is used to train the deep neural network (DNN) 3 using the heatmap loss (HM-VF).
[0110] The training is preferably accomplished using the backpropagation (BP) algorithm.
[0111] In a possible embodiment of the method according to the invention for recognizing an object OBJ in the surroundings, the generated heatmap HM-UINF with uncertainty information is used to perform bounding box regression BB-REG, which generates a bounding box BB-UINF with uncertainty information.
[0112] In a possible embodiment of the method according to the invention for recognizing an object OBJ in the surroundings, the generated bounding box BB-UINF with uncertainty information is used to train the deep neural network (DNN) 3 using the bounding box loss BB-VF.
[0113] First, the uncertainty, i.e., the uncertainty of the ground truth label GTL, is calculated based on the point cloud PW with features such as the distance to the ego vehicle, intensity, radial velocity, etc., based on the sensor quality SQ of the spacing sensor 2 in terms of point resolution and noise, based on the environmental conditions UB, such as weather and visibility conditions, and based on other features, such as the spacing to the ego vehicle, the number of points within the ground truth label, etc.
[0114] The uncertainty of the ground truth label GTL can be determined by the point distribution within the label.
[0115] Each bounding box label can be defined as y, and all points within the bounding box BB are defined as x1:K, where K represents the number of points within the object. For simplicity of calculation, the points x1:K are projected onto the bounding box boundary. The variable v represents the points projected onto the boundary line of the bounding box BB.
[0116] Given the points x1:K, the uncertainty distribution p(y|x1:K) of the bounding box label y can be estimated using Bayes' rule:
[0117] p(y|x1:K) ∝ p(x1:K|y,v)p(y) (1)
[0118] where p(x1:K|y,v) is the observation probability, which is correlated with the noise of the spacing sensor, and
[0119] p(y) is the prior uncertainty distribution of the label y.
[0120] The solution of the uncertainty distribution p(y|x1:K) can be represented by the covariance matrix Σ(y), as illustrated in Equation (2):
[0121]
[0122] where Σ0(y) represents the prioritized covariance matrix, which is defined as the variance of all object truth bounding boxes in the dataset,
[0123] σ2 represents the sensor observation noise described in Equation (4),
[0124] Φk,m represents the point registration probability described in Equation (3), and where
[0125] J(v*) represents the transformation matrix.
[0126] To solve for the point registration probability φk,m, the points xk within the bounding box BB are projected onto the boundary of the bounding box BB.
[0127] Assumption: If the point \(x_k\) can be projected onto the \(M\) nearest points on the boundary, then \(v_{k,m}\in V\) represents the \(m\)-th point on the boundary of the bounding box \(BB\) that is closest to the point \(x_k\).
[0128] Therefore, the point registration probability \(\varphi_{k,m}\) can be determined in Equation (3) as follows:
[0129]
[0130] where the form of the given sensor observation noise \(\sigma^2\) is:
[0131]
[0132] In fact, the sensor observation noise can also be determined empirically based on the prior knowledge of the distance sensor 2.
[0133] With the aid of Equations (2) and (3), the uncertainty of the ground truth label \(GTL\) in the dataset is determined or calculated (see also Figure 4 ).
[0134] After calculating the uncertainty of the ground truth label \(GTL\), a heatmap \(HM\) with uncertainty information \(HM - UINF\) is created.
[0135] This heatmap \(HM\) is used to train the DNN network 3 using the heatmap loss \(HM - VF\) (Losses (Loss)) of the heatmap \(HM\) (e.g., focal loss). By using the heatmap \(HM - UINF\) with uncertainty, bounding box regression \(BB - REG\) is performed to obtain a bounding box \(BB\) with uncertainty \(BB - UINF\), that is, the geometric description of the bounding box \(BB\) is carried out using the mean and standard deviation of each individual parameter.
[0136] To calculate the error, a so-called loss function \(VF\) (English Loss (Loss)) is used in the neural network, and different methods are available for this loss function, such as cross-entropy, mean squared error, or Poisson.
[0137] The heatmap \(HM - UINF\) with uncertainty and the bounding box \(BB - UINF\) with uncertainty are trained using the heatmap loss \(HM - VF\) (Loss) of the heatmap \(HM\) and the bounding box loss \(BB - VF\) of the bounding box \(BB\) respectively, and the losses take into account the uncertainty as (as Figure 4 shown).
[0138] The heatmap \(HM\) or the attention map helps to understand the decision-making of the neural network and identify errors in its actions.
[0139] The present invention also supports the determination of uncertainty during evaluation and online testing. For example, the trained DNN 3 can be used in a test vehicle such that the current recognition or detection results with associated uncertainty can be displayed on a monitor.
[0140] As Figure 5 shown, the trained DNN network 3 can output a bounding box BB-UINF with associated uncertainty, in such a way that a deep neural network (DNN) 3 and a bounding box regression BB-REG are used, which are trained in the above training pipeline.
[0141] Figure 5 An exemplary block diagram of the system 1 in evaluating DNN-AUSW according to the present invention is shown. By using the proposed method with uncertainty information, DNN-based radar perception can output a bounding box BB-UINF with uncertainty information.
[0142] The architectures shown and described above represent specific embodiments or implementations that can be modified in various aspects.
[0143] The general basic architecture can be described as follows:
[0144] To train the network using the uncertainty information UINF, the first method is: based on the quality of the predicted bounding box BB, freeze the object recognition part of the network and apply the loss only to the uncertainty prediction. Another method is to train the object recognition and uncertainty together such that unreliable recognition causes less loss when there is an actual error. The input to the network is a set of unordered points P, which form a point cloud PW, and the point cloud is provided to a LiDAR, radar, or other sensor.
[0145] The uncertainty-based heatmap loss HM-VF(Loss) and the bounding box loss BB-VF can be replaced by other loss functions VF according to the needs of the application. The heatmap HM can be used per capita, per category, or as a general feature map. By adapting the network, the generalization HM-GEN of the heatmap HM and the loss HM-VF of the heatmap HM can be removed based on the uncertainty.
[0146] The deep neural network 3 preferably also learns only from the labels with the associated calculated uncertainty.
[0147] System 1 can be trained using any type of DNN that takes into account uncertainty information UINF. System 1 can also use all types of DNN perception methods. A large number of neural networks can be used as the backbone and the head, including convolutional neural networks, Transformer architectures, etc. Other output formats besides the oriented bounding box BB are also possible, such as semantic segmentation, occupancy grid, or visibility map.
[0148] Figure 6 shows different application scenarios in which System 1 described here can be used.
[0149] System 1 can be used in an automated assembly system, for example, to identify components as objects OBJ on a conveyor belt and their orientations to determine the grasping points, as Figure 6A shown.
[0150] System 1 can be used in every autonomous driving device 7, especially household appliances (vacuum cleaners, lawn mowers), for example, to identify objects OBJ (obstacles) or occupancy, as Figure 6B shown.
[0151] System 1 can be used for automatic access control, for example, for person identification and discrimination for automatic door opening, as Figure 6C schematically shown.
[0152] System 1 can be used for monitoring a location or a building, for example, for the identification, inspection, and classification of dangerous goods, as Figure 6D schematically shown.
[0153] System 1 can be used for traffic monitoring using a fixed radar sensor 2, as Figure 6E shown in.
[0154] System 1 can be used to identify and classify traffic participants in an auxiliary system of a bicycle or other two-wheeled vehicles (motorcycles, mopeds, etc.), as Figure 6F shown. Other application areas are radar sensors that use object information, such as fixed radar sensors for traffic monitoring or radar sensors for bicycles.
[0155] The developed uncertainty model can be used in a test vehicle with a radar sensor. The identified objects with uncertainty are preferably represented or displayed as bounding boxes BB with variances. System 1 can be used in auxiliary systems and autonomous driving functions. In addition, System 1 can be used in the field of robotics (e.g., obstacle identification in an automatic lawn mower).
Claims
1. A system (1) for object recognition, the system having: A distance sensor (2) for generating a three-dimensional point cloud (PW) of the surrounding environment, wherein each data point of the three-dimensional point cloud has spatial coordinates of a position on the surface of an object (OBJ) in the surrounding environment of the distance sensor (2); A deep neural network (3) for calculating output features (AM) of an object (OBJ) based on the three-dimensional point cloud (PW) generated by the distance sensor (2), wherein associated uncertainty information (UINF) is determined during the calculation process for each output feature (AM); and An output unit (4) for outputting the object (OBJ) recognized based on the calculated output features (AM) in a bounding box (BB-UINF) assigned to the object or in another output format, the bounding box or the other output format reproducing the determined uncertainty information (UINF) for the calculated output features (AM).
2. The system for object recognition according to claim 1, wherein, The output features (AM) include the 3D position of the object, the size of the object, the speed of the object, the orientation of the object, the classification of the object, and the recognition rate regarding the object.
3. The system for object recognition according to claim 1 or 2, wherein each calculated output feature (AM) has an associated mean value and an associated standard deviation, which reproduce the determined uncertainty information (UINF) for the output feature (AM).
4. The system for object recognition according to any one of the preceding claims 1 to 3, wherein, The distance sensor (2) has a radar sensor for generating a three-dimensional radar point cloud.
5. The system for object recognition according to any one of the preceding claims 1 to 4, wherein, The distance sensor (2) has a LiDAR sensor for generating a three-dimensional LiDAR point cloud.
6. The system for object recognition according to any one of the preceding claims 1 to 5, wherein, The determination of the uncertainty information is performed during the training phase (TP) of the deep neural network (3) and during object recognition performed using the trained deep neural network (3).
7. A method for recognizing an object in a surrounding environment, the method comprising the following steps: Generating (S1) a three-dimensional point cloud (PW) of the surrounding environment, wherein each data point of the three-dimensional point cloud (PW) has spatial coordinates of a position on the surface of an object (OBJ) in the surrounding environment; Calculating (S2) output features (AM) of an object (BJ) based on the generated three-dimensional point cloud (PW) by a deep neural network (3), wherein associated uncertainty information (UINF) is determined during the calculation process for each output feature (AM); and Outputting (S3) the object (OBJ) recognized based on the calculated output features (AM) in a bounding box (BB-UINF) assigned to the object or in another output format, wherein the bounding box or the other output format reproduces the determined uncertainty information (UINF) for the calculated output features (AM).
8. The method for recognizing an object in a surrounding environment according to claim 7, wherein the uncertainty information is determined during the training phase (TP) of the deep neural network (3) and during object recognition performed using the trained deep neural network (3).
9. The method for identifying an object in the surrounding environment according to claim 8, wherein during a training phase (TP) for training the deep neural network (3), the uncertainty of the ground truth label (GTL) of the training data set is calculated.
10. The method for identifying an object in the surrounding environment according to the claim, wherein a heat map with uncertainty information (HM-UINF) is generated, which is used to train the deep neural network (3) using a heat map loss (HM-VF).
11. The method for identifying an object in the surrounding environment according to claim 10, wherein the generated heat map with uncertainty information (HM-UINF) is used to perform bounding box regression (BB-REG), which generates a bounding box with uncertainty information (BB-UINF).
12. The method for identifying an object in the surrounding environment according to claim 11, wherein the generated bounding box with uncertainty information (BB-UINF) is used to train the deep neural network (3) using a bounding box loss (BB-VF).
13. The method for identifying an object in the surrounding environment according to any one of the preceding claims 7 to 12, wherein the trained deep neural network (3) outputs a bounding box (BB-UINF) with associated uncertainty information respectively based on a three-dimensional point cloud (PW) generated by a distance sensor (2) and based on other data regarding the sensing detection of the surrounding environment by the distance sensor (2) by using bounding box regression.
14. The method for identifying an object in the surrounding environment according to claim 13, wherein the other data regarding the sensing detection of the surrounding environment by the distance sensor (2) includes data regarding the quality of the used distance sensor (2) (SQ) and / or data regarding environmental conditions (UB) present in the sensing detection of the surrounding environment.
15. A device (7) having integrated therein a system (1) for object recognition according to any one of claims 1 to 6, wherein the device (7) comprises: An automatically driving device, in particular an automatically driving household appliance or an automatically driving vehicle; A monitoring device, in particular a monitoring device for monitoring a location or a building and for access control; and a traffic monitoring device, in particular a traffic monitoring device for monitoring traffic participants.
Citation Information
Patent Citations
Quantitative assessment of the uncertainty of a classifier's statements based on measurement data and several processed products of the same.
DE102021210566A1
Perception uncertainty modeling from actual perception systems for autonomous driving
US20200050191A1