Method and apparatus for detecting an object

Through the radar point cloud detection method of the two-level neural network architecture, the problems of large computing workload and poor scalability are solved, and efficient and accurate multi-object detection is achieved.

CN114387203BActive Publication Date: 2025-07-11APTIV TECHNOLOGIES AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111223865.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-19
Filing Date
2021-10-18
Publication Date
2025-07-11
Estimated Expiration
2041-10-18

AI Technical Summary

Technical Problem

The existing object detection method based on radar point cloud has a large amount of calculation work, which is difficult to scale to multiple object classes and there is a problem of inconsistent bounding box estimation.

Method used

Using a two-level neural network architecture, the geometric shape spatial conditions of the object are first estimated through semantic segmentation and rough approximation, and then the data point subset is selected by refining the approximation and confidence scores, and finally the precise geometry and confidence of the object are estimated.

Benefits of technology

It reduces the computational workload, can effectively scale to multiple object classes, and improves the accuracy and efficiency of object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114387203B_ABST
    Figure CN114387203B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and apparatus for detecting an object. A method for detecting an object by using a radar sensor and by using an apparatus configured to build a neural network is provided. A plurality of raw radar data points are captured. At least one object class including a predetermined object type and a geometry for surrounding the object is defined. Via a first stage of the neural network, semantic segmentation of the data points with respect to the object class and the background is performed, and for each data point, a rough approximation of the spatial condition of the geometry is estimated. Based on the rough approximation, via a second stage of the neural network, a subset of the data points is selected based on the semantic segmentation, and for each data point of the subset, a refined approximation of the spatial condition of the geometry and a confidence score of the refined approximation are estimated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method and apparatus for detecting an object based on radar point cloud. Background Art

[0002] Modern vehicles typically include cameras, sensors, and other devices for monitoring the vehicle environment. For example, reliable perception of the vehicle environment is very important for safety functions and driver assistance systems installed in the vehicle. In the field of autonomous driving, reliable perception of the vehicle environment is a crucial issue. For example, objects such as other vehicles or pedestrians must be detected and tracked to achieve correct performance of autonomous driving.

[0003] In many cases, only noisy data from sensors (such as radar and / or LIDAR sensors and cameras) can be used to detect and track objects. For example, data from a radar sensor can be provided as sparse data forming a so-called point cloud, which includes data points caused by radar reflections belonging to certain objects (such as other vehicles) or belonging to background reflections. Therefore, it may be challenging to correctly detect different objects based only on radar data.

[0004] Therefore, machine learning methods for detecting objects based on radar point clouds have been proposed. For example, Danzer etal.:"2D car detection in radar data with PointNets",2019IEEE IntelligentTransport Systems Conference,Auckland,New Zealand,New Zealand,Oct 27-30,2019 describes a method that combines object classification with object bounding box estimation based on a neural network architecture called PointNet (see also: Charles,R.Qi et al.:"PointNet:Deep Learning on Point Sets for3D Classification and Segmentation".2017IEEE Conference on Computer Visionand Pattern Recognition(CVPR),July 21-26,2017,Honolulu,HI,USA). This basic neural network architecture shares fully connected layers among multiple points and performs so-called "max pooling" among point features to obtain a global signature of the entire point cloud. In the above method, the object detector applied to the original radar point cloud uses this basic neural network architecture multiple times in a hierarchical manner to perform classification, that is, only the points belonging to a certain object are extracted, and the bounding box that completely encloses the detected object is estimated, thus defining its spatial conditions.

[0005] However, this method requires a large amount of computational work because it is based on so-called "patch proposals" for individual data points. Using "anchor boxes" for bounding box estimation increases the computational workload. In addition, although existing object detection methods can be extended to multiple object classes, such as different types of vehicles, pedestrians, etc., extending to multiple object classes will further increase the computational workload because, for example, copies of the entire network will be required. During classification, redundant operations may also be performed, which may lead to inconsistent determination of object bounding boxes.

[0006] Therefore, there is a need for methods and devices for detecting objects based on radar point clouds that allow for the detection of multiple object classes and require low computational workload. Summary of the Invention

[0007] In one aspect, the present disclosure relates to a computer-implemented method for detecting objects by using a radar sensor and by using a device configured to establish a neural network. According to the method, an original radar point cloud including a plurality of data points is captured via the radar sensor, and at least one object class is defined, wherein each object class includes a predefined object type and a geometry configured to enclose the object of the predefined object type. Via a first stage of the neural network, semantic segmentation is performed on the plurality of data points with respect to the at least one object class and with respect to the background, and for each data point, a rough approximation of the spatial condition of the geometry of the at least one object class is estimated. Based on the rough approximation of the spatial condition of the geometry, via a second stage of the neural network, a subset of the plurality of data points is selected based on the semantic segmentation, and for each data point of the subset, a refined approximation of the spatial condition of the geometry of the object class, and a confidence score of the refined approximation of the spatial condition of the geometry are estimated.

[0008] As input, the method uses the original radar point cloud, which can be based on radar signals reflected by objects in the radar sensor environment and acquired by the sensor.

[0009] If the radar sensor and the device for establishing the neural network are installed in a vehicle, the at least one object class can relate to vehicles and / or pedestrians. For example, in such a case, the corresponding object types can be "vehicle" and / or "pedestrian". Examples of suitable geometries for the at least one object class can be a bounding box, a sphere, or a cylinder that completely encloses the object. For such geometries, the spatial condition can be defined at least by the central position of such a geometry (e.g., a sphere). Additionally, if desired, the spatial condition can include further parameters for defining, for example, the size and orientation of the geometric object.

[0010] To perform semantic segmentation, the layers of the neural network reflect the probability that each data point of the point cloud belongs to the at least one object class or belongs to the background. Thus, semantic segmentation determines whether a data point belongs to the at least one object class. Therefore, via semantic segmentation, data points belonging to the at least one object class can be selected within the second stage of the neural network.

[0011] Furthermore, if the method is to be implemented in a vehicle, two or more object classes can be defined, i.e., different types of vehicles as object types and / or pedestrians as additional object types. Since the network structure includes two stages of the neural network, the method can be extended to multiple object classes without significantly increasing the required computational effort. This is partly due to the two steps of selection of the subset of the original data points and approximation of the spatial condition of the geometry.

[0012] Due to the two levels of the method (i.e., the neural network), redundant operations can also be avoided, which again reduces the computational workload. In addition, the two levels of the neural network can be trained separately, which reduces the requirements for applying the method. Although the method is described as including a neural network with two levels, in fact, the neural network can include two or more separate neural networks connected to each other.

[0013] Since the method obtains a refined approximation of the spatial conditions of the geometric shape, the object to be detected is located within the geometric shape. For example, it can be assumed that a vehicle is located within the corresponding bounding box for which the spatial conditions have been estimated by the method. That is, object detection is performed by estimating, for example, the position and orientation of the geometric shape surrounding the object. In addition, the confidence score defines the reliability of the refined approximation, which can include a final estimate of the position and orientation of the geometric shape. For example, the confidence score is important for further applications within a vehicle that rely on the output of the method (i.e., the refined approximation of the spatial conditions of the geometric shape).

[0014] The method may include one or more of the following features:

[0015] The spatial conditions may include the center of the geometric shape, and estimating the rough approximation and the refined approximation of the spatial conditions may include estimating, for each data point, a corresponding vector that moves the corresponding data point towards the center of the geometric shape. The spatial conditions may further include the size of the geometric shape, and estimating the refined approximation of the spatial conditions may include estimating the size of the geometric shape. In addition, the spatial conditions may further include the orientation of the geometric shape, and estimating the refined approximation of the spatial conditions may include estimating the orientation of the geometric shape. Estimating the orientation of the geometric shape may be decomposed into a step of regressing the angular information of the geometric shape and a step of classifying the angular information.

[0016] The first level of the neural network may include corresponding output paths for semantic segmentation and for at least one parameter of the spatial conditions, and the second level of the neural network may further include corresponding refined output paths for the confidence score and for at least one parameter of the spatial conditions. Each output path of the neural network can be trained based on a separate objective function.

[0017] The execution of semantic segmentation can be trained by defining the ground truth of the geometric shape and by determining whether the corresponding data points are within the ground truth of the geometric shape. The first level of the neural network can be deactivated while the second level of the neural network is being trained.

[0018] The geometry of at least one object class may include a bounding box, and a refined approximation of the estimated spatial condition may include estimating the dimensions and orientation of the bounding box. Additionally, at least two different object classes may be defined, and the at least two different object classes may include different object types and different geometries.

[0019] A refined approximation of the spatial condition of the geometry can be estimated by performing farthest point sampling only on a selected subset of the plurality of data points. Estimating a refined approximation of the spatial condition of the geometry may include grouping the data points of the selected subset to train a second level of a neural network with respect to the local structure.

[0020] According to an embodiment, the spatial condition may include the center of the geometry, and estimating a rough approximation and a refined approximation of the spatial condition may include estimating a corresponding vector for each data point that moves the corresponding data point towards the center of the geometry. The shift vector may also be regarded as the offset vector of the corresponding data point. The final estimate of the center of the geometry can be determined by calculating the average of all the refined approximations of the center of the geometry. Since the localization of the center of the geometry can be carried out in two steps, a rough and a refined approximation, and since background points are discarded before the refined approximation, the computational effort required to estimate the center of the geometry can be reduced.

[0021] Furthermore, the spatial condition may also include the dimensions of the geometry, and estimating a refined approximation of the spatial condition may include estimating the dimensions of the geometry. The dimensions of the geometry may include, for example, the length, width, and height of the bounding box of the object, the radius of the sphere enclosing the object, or the radius and height of the cylinder used to enclose the object. Alternatively or additionally, the spatial condition may further include the orientation of the geometry, and estimating a refined approximation of the spatial condition may include estimating the orientation of the geometry. The orientation of the geometry may include, for example, the yaw angle, roll angle, and pitch angle of the bounding box. If the radar sensor provides only two-dimensional data, the orientation of the geometry may include only the yaw angle (e.g., for the base region of the bounding box). If the spatial condition of the geometry can be defined by its center, dimensions, and orientation, the method provides a direct way to localize the geometry and detect the corresponding object.

[0022] Additionally, estimating the orientation of the geometry can be decomposed into a step of regressing the angular information of the geometry and a step of classifying the angular information. Due to the periodicity of the angle with respect to π or 2π, it may be difficult for a neural network to regress the angle. Therefore, the reliability of the regression regarding the orientation of the geometry can be improved by decomposing this task into a regression step and a classification step. For example, the neural network can regress the absolute values of the sine and cosine of the final yaw angle of the bounding box. For classification, the neural network can determine the signs (i.e., ±1) of the sine and cosine of the final yaw angle.

[0023] According to a further embodiment, the first stage of the neural network may include respective output paths for semantic segmentation and for at least one parameter for spatial conditions, and the second stage of the neural network may include respective output paths for confidence scores and for a refinement of at least one parameter for spatial conditions. The respective output paths of the neural network may be trained based on separate objective functions. Due to the separate objective functions, the total workload for training the neural network can be reduced. In addition, the reliability of the final estimate or final prediction of the neural network can be improved by adapting the respective objective functions to the tasks leading to the respective output paths.

[0024] The execution of semantic segmentation can be trained by defining the ground truth of the geometry and by determining whether the respective data points are within the ground truth of the geometry. Thus, semantic segmentation can be performed based on a direct prerequisite, namely, the localization of known objects that define the ground truth of the geometry. In addition, the transition region can be defined as being close to the surface of, for example, the bounding box that defines the ground truth, i.e., inside the bounding box but close to its boundary. This can avoid the estimation of the center of the geometry that is far from the true center defined by the ground truth of the bounding box.

[0025] According to a further embodiment, while training the second stage of the neural network, the first stage of the neural network can be deactivated. In other words, the two stages of the neural network can be trained independently of each other. Thus, the workload for training the neural network can be reduced again.

[0026] In addition, the geometry of at least one object class may include a bounding box, and the refined approximation for estimating the spatial conditions may include estimating the dimensions and orientation of the bounding box. The bounding box can be, for example, a direct representation of the geometry for detecting a vehicle as an additional object.

[0027] According to a further embodiment, two different object classes can be defined, and at least two different object classes may include different object types and different geometries. Due to the two-stage architecture of the neural network, the method can be extended in a straightforward manner to at least two different object classes without significantly increasing the computational workload. For detecting objects in a vehicle environment, the two different geometries can be a bounding box representing the vehicle and a cylinder representing a pedestrian.

[0028] The refined approximation of the spatial conditions of the geometry can be estimated by performing farthest point sampling only on a selected subset of the plurality of data points. Alternatively or additionally, estimating the refined approximation may include grouping the data points of the selected subset to train the second stage of the neural network with respect to the local structure. Since these steps are based only on the selected subset of data points, the workload for the refined approximation can be reduced again.

[0029] In another aspect, the present disclosure relates to a device for detecting an object. The device includes a radar sensor that obtains radar signals reflected by an object in the environment of the radar sensor and provides an original radar point cloud based on the radar signals, wherein the original radar point cloud includes a plurality of data points. The device further includes a module configured to establish a neural network including a first stage and a second stage and define at least one object class, each object class including a predefined object type and a geometry configured to enclose an object of the predefined object type. The first stage of the neural network is configured to perform semantic segmentation on the plurality of data points with respect to the at least one object class and with respect to the background, and estimate a rough approximation of the spatial condition of the geometry of each object class for each data point. Based on the rough approximation of the spatial condition of the geometry, the second stage of the neural network is configured to select a subset of the plurality of data points based on the semantic segmentation, and estimate a refined approximation of the spatial condition of the geometry of the object class for each data point in the subset and estimate a confidence score of the refined approximation of the spatial condition of the geometry.

[0030] As used herein, the term "module" may refer to any of the following, a part of any of the following, or include any of the following: an application specific integrated circuit (ASIC); an electronic circuit; a combinational logic circuit; a field programmable gate array (FPGA); a processor (shared, dedicated, or group) that executes code; other suitable components that provide the function; or a combination of some or all of the above, such as in a system on a chip. The term "module" may include a memory (shared, dedicated, or group) that stores code executed by the processor.

[0031] In summary, the device according to the present disclosure includes a radar sensor and a module for performing the steps described above for the corresponding method. Therefore, the benefits, advantages, and disclosures of the above method are also effective for the device according to the present disclosure.

[0032] In another aspect, the present disclosure relates to a computer system configured to perform some or all of the steps of the computer-implemented method described herein.

[0033] The computer system may include a processing unit, at least one memory unit, and at least one non-transitory data storage. The non-transitory data storage and / or the memory unit may include a computer program for instructing the computer to perform some or all of the steps or aspects of the computer-implemented method described herein.

[0034] In another aspect, the present disclosure relates to a non-transitory computer-readable medium that includes instructions for performing some or all of the steps or aspects of the computer-implemented methods described herein. The computer-readable medium can be configured as: an optical medium such as a compact disc (CD) or a digital versatile disc (DVD); a magnetic medium such as a hard disk drive (HDD); a solid state drive (SSD); a read-only memory (ROM) such as a flash memory; and so on. Additionally, the computer-readable medium can be configured as a data storage accessible via a data connection such as an Internet connection. The computer-readable medium can be, for example, an online data repository or a cloud storage.

[0035] The present disclosure also relates to a computer program that instructs a computer to perform some or all of the steps or aspects of the computer-implemented methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The exemplary embodiments and functions of the present disclosure are described herein in conjunction with the following drawings shown schematically:

[0037] Figure 1a 、 Figure 1b and Figure 1c depicts a schematic diagram of an apparatus for detecting an object by performing a method according to the present disclosure,

[0038] Figure 2 is an illustrative diagram of an offset vector applied to a data point,

[0039] Figure 3 depicts Figure 1a 、 Figure 1b and Figure 1c the internal mechanism of a part of the apparatus in

[0040] Figure 4 depicts the masking step applied to a data point in the second level as shown in Figure 1a 、 Figure 1b and Figure 1c shown,

[0041] Figure 5 and Figure 6 depicts an example of the spatial conditions for determining a bounding box based on original data points by the apparatus as shown in Figure 1a 、 Figure 1b and Figure 1c shown, and

[0042] Figure 7 depicts a graph of the precision and recall determined by detecting two different object types by the apparatus according to the present disclosure. DETAILED DESCRIPTION

[0043] Figure 1a 、Figure 1b and Figure 1c depicts a schematic view of a device 11 for detecting an object. The device 11 includes a radar sensor 13 and a module 17. The radar sensor 13 provides a raw radar point cloud including a plurality of data points 15. The module 17 is configured to establish a neural network including a first stage 19 and a second stage 21. The plurality of data points 15 from the raw radar point cloud are used as inputs to the module 17, that is, inputs to the first stage 19 and the second stage 21 of the neural network.

[0044] The radar sensor 13 acquires radar signals reflected by an object in the environment of the radar sensor 13. For example, the device 11 is mounted in a host vehicle 20 (see Figure 4 ). In this case, the radar signals are reflected by other vehicles located within the bounding boxes 61 and other objects outside these bounding boxes 61 (see Figure 4 's upper figure). Thus, the data point cloud is represented by a plurality of two-dimensional data points, and the x and y coordinates of these data points are known relative to the coordinate system of the host vehicle 20, that is, relative to the radar sensor 13 mounted in the host vehicle 20.

[0045] Figure 1a provides a high-level overview of the device 11, while Figure 1b shows the internal structure of the first stage 19 of the neural network established by the module 17, Figure 1c shows the internal structure of the second stage 21 of the neural network established by the module 17. The first stage 19 of the neural network receives the plurality of data points 15 as inputs and outputs a semantic segmentation 23 of the plurality of data points 15 and an offset 25 of each data point 15. The second stage 21 receives the plurality of data points 15 from the radar sensor 13 and the semantic segmentation 23 and the offset 25 from the first stage 19 as inputs and outputs parameters 27 of the bounding boxes 61 (see Figures 2 to 6 ). The parameters 27 define the spatial conditions of the bounding boxes 61 and thus define the spatial conditions of the objects to be detected surrounded by these bounding boxes 61.

[0046] The method is described below as including a neural network with two stages. However, the neural network can also be considered to include two or more separate neural networks connected to each other.

[0047] Figure 1bdepicts the internal structure of the first stage 19 of the neural network established by module 17. The first stage 19 includes a preprocessing module 29 that encodes the radar information (such as x and y coordinates, Doppler velocity, and radar cross-section) of each data point among the plurality of data points 15 to provide deeper feature vectors. The encoding makes the plurality of data points 15 used as the input of the first stage 19 a better representation for the tasks of the first stage 19, that is, performing semantic segmentation 23 and determining the offset 25.

[0048] The output of the preprocessing module 29 is received by another neural network structure 31 called "PointNet++", which is described in detail in Charles, R. Qi et al.: "PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space", arXiv:1706.02413v1 [cs.CV] 7 Jun 2017. This neural network structure 31 includes a so-called set-abstraction layer and a fully connected layer shared by all data points 15, which are described in detail below in the context of the second stage 21. The first stage 19 of the neural network also includes a first output path for performing semantic segmentation 23 and a second output path for determining the offset vector 25 of each data point among the plurality of data points 15.

[0049] For semantic segmentation 23, object classes are defined, where each object class includes a predetermined object type and a geometry configured to enclose the object of the predetermined object type. If the device 11 is installed in the host vehicle 20 (see Figure 4 ), then there are two object classes, where the first object class includes the object type "vehicle", and for this object type, the bounding box 61 (see Figure 4 ) is used as the geometry, while the second object class includes the object type "pedestrian", and for this object type, a cylinder is used as the geometry. When performing semantic segmentation 23, for each data point 15, a probability of belonging to one of the predetermined object classes (i.e., vehicle or pedestrian) or belonging to the background is established. Specifically, the weights of the layer 33 of the neural network for semantic segmentation 23 will reflect the probability that each data point 15 belongs to one of the object classes or belongs to the background, and the layer 33 will classify all data points 15 accordingly. The data points 15 that are not classified as the background will be used as the center proposal of the bounding box within the second stage 21 of the neural network.

[0050] To learn the task of semantic segmentation 23, the focal classification loss is used as the objective function. Additionally, ground truth bounding boxes are defined, and those points from multiple data points 15 that are within the corresponding ground truth bounding boxes are labeled as positive. Further, a transition region can be defined, where points within the ground truth bounding box but close to its boundary are marked but not labeled as positive.

[0051] Data points 15 that are classified as belonging to one of the object classes and not classified as background will be used as proposals for the centers of the bounding boxes in the second stage 21 of the neural network. If the transition region is used when learning the semantic segmentation task, center proposals that are far from the true center of the bounding box are avoided.

[0052] In the second output path of the first stage 19, for each data point 15, a corresponding offset vector 25 that moves the corresponding data point 15 towards the center of the object or its bounding box is determined. This is shown in Figure 2 where five data points 15 are shown together with their corresponding offset vectors 25. Additionally, the bounding box 61 is shown having a true center 63. Each data point 15 is shifted closer to the center 63 via its corresponding offset vector 25.

[0053] Specifically, the x and y components of the offset vector are regressed by the neural network. The network predicts values in the range from -1 to +1, which are multiplied by a suitable scalar. The scalar is chosen to be large enough so as not to limit the output representation of the network, i.e., the offset vector 25. However, by choosing a suitable scalar, data points 15 can be prevented from being shifted towards the center of a bounding box belonging to another object. To learn the task of determining the offset vector 25, the smooth L1 regression loss is used as the objective function.

[0054] Thus, the two output paths that produce the semantic segmentation 23 and the offset vector 25 are independently optimized by their corresponding objective functions. However, the common part of the PointNet++ 31 of the neural network is optimized or trained by using the objective functions of these two output paths.

[0055] The internal structure of the second stage 21 of the neural network is as shown in Figure 1c The second stage 21 receives the data points 15 as well as the semantic segmentation 23 and the offset vector 25 output by the first stage 19 of the neural network as inputs. In the shifting step 35, each data point 15 is shifted by its corresponding offset vector 25 so as to end up with the shifted data points 37. The shifted data points 37 are the inputs to the additional layers 39, 47 of the second stage 21 of the neural network.

[0056] Figure 5 and Figure 6An example of the shifting step 35 applied to the data point 15 is shown. For each data point 15 forming the original radar point cloud provided by the radar sensor 13, the corresponding offset vector 25 moves the data point 15 in the direction of the assumed center of the corresponding bounding box 61. That is, based on the original data point 15 and based on the offset vector 25 which is the result of the first stage 19 of the neural network, the shifted data point 37 is generated as the shifted input of the layer of the second stage 21. Thus, the shifted data point 37 can be regarded as the first approximation for estimating the true center of the bounding box 61. The first approximation represented by the shifted data point 37 will be refined by the second stage 21.

[0057] To perform this refinement, the second stage 21 includes a set abstraction layer (SA layer) 39, which is an essential element of the PointNet++ architecture. However, this set abstraction layer is enhanced by a masking or selection step 41 applied to the shifted data point 37. To perform this masking or selection 41, the result of the semantic segmentation 23 output by the first stage 19 of the neural network is used.

[0058] Figure 4 The effect of the masking or selection step 41 is shown. The original radar point cloud including the data point 15 is depicted before (upper part) and after (lower part) the masking or selection step 41. Note, Figure 4 the upper part of Figure 5 and Figure 6 shows the entire point cloud before applying the shifting step 35 (see also

[0059] ). Before the masking 41, the point cloud includes all data points, that is, the data points 65 classified as background and the data points 67 classified as vehicles and the data points 68 classified as pedestrians. After the masking or selection step 41, only the data points 67, 68 classified as belonging to vehicles or pedestrians are considered.

[0060] After performing the masking or selection step 41 using the shifted data point 37, a so-called farthest point sampling (FPS) 43 is applied to the remaining data points. The farthest point sampling means selecting a subset of m points (n > m) from n points such that each point from the subset is the farthest point with respect to all the proceeding points within the subset. Due to the farthest point sampling 43, at least one data point is sampled for each object to be detected. Figure 3An example of the grouping step 45 (layer 1) is depicted on the left. The shifted data points 37 are grouped or clustered using corresponding circles 66 having a certain predetermined radius around the sampled data points 64. The output of the grouping step 45 and of the entire set abstraction layer 39 is a list of sampled points with deep feature vectors containing object information, rather than the abstract local encodings used in the comparative methods of the related art. Finally, the output of the set abstraction layer 39 is processed by another shared fully-connected layer 47 of the neural network in order to regress the parameters 27 of the bounding boxes 61, that is, in order to estimate the spatial conditions of these bounding boxes. This is shown on the left in Figure 3 which the result of "learning" from the local structure (i.e., from the surrounding data points of the respective sampled data points) via another layer 2 is shown.

[0061] The second stage 21 of the neural network includes four output paths, each output path corresponding to one of the parameters 27 of the bounding box. Specifically, the four output paths correspond to the refined offset vector 49, the bounding box size (BB size) 51, the bounding box yaw angle (BB yaw) 53, and the bounding box confidence score (BB score) 55.

[0062] The refined offset vector 49 is determined in a very similar way to the offset vector 25 in the first stage 19. That is, the refined vector 49 is determined which moves the shifted and selected data points closer to the true center of the bounding box 61. However, the scalar used is half of the scalar of the first stage. In order to learn the task of determining the refined offset vector 49, the smooth L1 regression loss is again used as the objective function.

[0063] The bounding box size 51 refers to the width and length of the considered bounding box. The fully-connected layer 47 of the network predicts a value x in the range from -1 to +1 in order to calculate, for example, the width w of the bounding box as follows:

[0064] w = s·e x (1)

[0065] where x is the predicted value of the network and s is another scalar. Taking a vehicle as an example, the scalar s can be set to 6m, in which case the neural network is able to predict the width (and in the same way, the length) of the vehicle (as the object to be detected) in the range from 2.2m to 16m.

[0066] In order to learn the task of determining the bounding box size 51, the objective function is defined as:

[0067]

[0068] where w pred is the predicted width of the bounding box, w gtis the width of the ground truth bounding box. With such an objective function, overestimation of the bounding box size 51 is avoided.

[0069] The third output path of the second stage 21 refers to the orientation of the bounding box, which is the bounding box yaw angle 53 in the case of two-dimensional raw radar input data. Due to the periodicity with respect to π or 2π, angle regression is a difficult task for a neural network, so the task of determining the bounding box orientation is decomposed into a regression step and a classification step. For the regression step, the neural network regresses the absolute values of the sine and cosine of the yaw angle to be determined. The classification step determines the signs of the sine and cosine of this yaw angle, i.e., ±1. To learn the task of determining the bounding box yaw angle 53, the mean squared error loss (MSE loss) is used as the objective function for learning the regression task, and the binary cross-entropy loss is used as the objective function for the classification task.

[0070] Since the loss objectives must be combined for the two tasks, the classification is dynamically weighted, where the dynamic weight is calculated as:

[0071] weight = 2·[1 - max(sin 2 θ gt , cos 2 θ gt )](3)

[0072] where θ gt is the ground truth yaw angle. For all m (positive integer or negative integer or zero), the weight function has a minimum value of 0 at m·π / 2 and a maximum value of 1 at (2m + 1)·π / 4. With this weighting, when the ground truth yaw angle is near the classification boundary, the classification has less influence on the overall task of determining the orientation. In this way, even for incorrect sign estimations (due to classification) that do not lead to significant errors, the neural network can perform the correct regression of the sine and cosine.

[0073] The fourth output path of the second stage 21 refers to the bounding box confidence score 55. It is important to provide a confidence score for network predictions if non-maximum suppression is applied in a post-processing step (not shown in the figure). Such post-processing may be required if several overlapping bounding box estimates have to be merged into a single bounding box estimate.

[0074] To estimate the bounding box confidence score 55, a ground truth bounding box needs to be defined, and for all bounding boxes for which the bounding box parameters 27 have been determined, the so-called Intersection over Union score (IoU score) is calculated. The IoU score compares the predicted bounding box with the ground truth bounding box and is defined as the ratio of the intersection or overlap of the predicted bounding box and the ground truth bounding box to their union. If the IoU score is higher than a given threshold, e.g., 0.35, the detection of the bounding box is considered positive. To avoid confusion in the network, bounding boxes with an IoU score slightly lower than this threshold (e.g., in the range of 0.2 to 0.35) are masked out.

[0075] To estimate the bounding box confidence score 55 for misaligned bounding boxes, an empirical approximation is used. First, the centers of the predicted bounding box and the ground truth bounding box are determined. Around these centers, bounding boxes with an estimated size or estimated dimensions are formed, and both bounding boxes are rotated by the ground truth yaw angle. Then, the intersection of these aligned boxes is calculated. The value of the intersection is weighted by the absolute cosine value of the angular error and rescaled to fit the range from 0.5 to 1. Then, the final bounding box confidence score or IoU score is calculated by dividing the approximate intersection by the union area (i.e., the sum of the areas of the two boxes, the ground truth bounding box and the predicted bounding box, minus the approximate intersection).

[0076] To learn the task of determining the bounding box confidence score, the focal segmentation loss is again used as the objective function. Additionally, the four output paths of the second stage 21 are independently trained with respect to their objective functions, and the common part of the second stage 21 (i.e., the set abstraction layer 39) is trained by using all the objective functions. Additionally, when training the second stage 21, the first stage 19 is deactivated or "frozen".

[0077] In Figure 5 and Figure 6 the effect of the second stage 21 is shown by the refined offset vector 49 and the bounding box 61. Specifically, the data point 15 that is shifted by the shift step 35 to end at the shifted data point 37 is moved closer to the center of the bounding box 61 by the refined offset vector 49 represented by the inner arrow or the second arrow in Figure 5 and Figure 6 After applying the refined offset vector 49, the final center proposal 69 is found at the end of the "second" arrow 49. For the bounding box 61, the second stage 21 also determines the bounding box dimensions 51 (i.e., its length and width) and the bounding box yaw angle 53 in order to define the orientation of the bounding box. Thus, the spatial condition of the bounding box 61 is estimated by estimating its center, its dimensions, and its yaw angle.

[0078] Some data points 15 are classified as background points 65 in the first stage 19 of the neural network. For these background points 65, due to the masking step 41 based on semantic segmentation 23, the second stage 21 of the neural network does not estimate the refined offset vectors 49. In other words, the background points 65 are sorted out in the second stage 21.

[0079] To test the apparatus 11 and method according to the present disclosure, two object classes are defined. The first object class includes "vehicle" as the object type and includes the above-mentioned bounding box as the geometry surrounding the object. The second object class includes "pedestrian" as the object type and includes a cylinder as the geometry. For the first object class (vehicle), the bounding box parameters 27 are estimated via the second stage 21 of the neural network. For the second object class (pedestrian), only the refined offset vectors 49 and the bounding box confidence scores 55 are estimated. For the cylinder as the geometry, it is assumed that the fixed diameter is 0.5 meters, and the yaw angle does not need to be determined for two-dimensional data.

[0080] Furthermore, the network is tested with two different setup configurations. For the first setup configuration, the data points 15 from a single radar frame (1 frame) are used as the input for estimating the bounding box parameters 27. For the second configuration, the data points 15 from four radar frames (4 frames) are used, and these data points are compensated to account for the ego-motion of the host vehicle 20. For the 4-frame setup, four measurements of the scene (i.e., measurements of the environment of the host vehicle 20) are fused, and an additional layer is provided to the neural network.

[0081] In Figure 7 , the results of the tests are shown, i.e., the precision on the y-axis versus the recall on the x-axis. Precision and recall are typical quality factors of a neural network. Precision describes the fraction of actual correct positive identifications, while recall describes the fraction of actual positive identifications that are correctly identified. Furthermore, recall is a measure of the stability of the neural network. The precision and recall of an ideal neural network would be in the range of 1, i.e., Figure 7 the data point in the upper right corner. Furthermore, the 4-frame setup contains a higher number of floating point operations (FLOPS) than the 1-frame setup, as shown in Figure 7 the lower right corner of

[0082] The cross labeled 73 is the result of the 1-frame setup for the vehicle object class, while the dot labeled 75 is the result of the 4-frame setup for the same object class (i.e., "vehicle"). Furthermore, the cross labeled 77 is the result of the 1-frame setup and the pedestrian object class, while the dot labeled 79 represents the result of this pedestrian object class based on the 4-frame setup. The results show that the neural network is able to detect vehicles and pedestrians in a manner suitable for both setups.

[0083] In summary, the 4-frame setting has better performance than the 1-frame setting. However, the results shown as Figure 7 below indicate that the apparatus 11 and method according to the present disclosure allow a reasonable trade-off between performance and cost.

[0084] List of Reference Numerals

[0085] 11 Apparatus for detecting an object

[0086] 13 Radar sensor

[0087] 15 Data points

[0088] 17 Module

[0089] 19 First stage

[0090] 20 Host vehicle

[0091] 21 Second stage

[0092] 23 Semantic segmentation

[0093] 25 Offset vector

[0094] 27 Bounding box parameters

[0095] 29 Preprocessing module

[0096] 31 PointNet++

[0097] 33 Fully connected layer

[0098] 35 Shifting step 37 Shifted data points

[0099] 39 Set abstraction layer

[0100] 41 Masking or selection step 43 Furthest point sampling

[0101] 45 Grouping and clustering step

[0102] 47 Fully connected layer

[0103] 49 Refined offset vector

[0104] 51 Bounding box size

[0105] 53 Bounding box yaw angle

[0106] 55 Bounding box confidence score

[0107] 61 Bounding box

[0108] 63 Bounding box center

[0109] 64 Sampled data points

[0110] Data points classified as background

[0111] 66 Circle

[0112] 67 Data points classified as "vehicle"

[0113] 68 Data points classified as "pedestrian"

[0114] 69 Final center suggestion

[0115] 73 Test results for "vehicle" with 1-frame setting

[0116] 75 Test results for "vehicle" with 4-frame setting

[0117] 77 Test results for "pedestrian" with 1-frame setting

[0118] 79 Test results for "pedestrian" with 4-frame setting

Claims

1. A computer-implemented method for detecting an object by using a radar sensor (13) and by using a device (11, 17) configured to establish a neural network, the method comprising the steps of: Capturing, via the radar sensor (13), an original radar point cloud including a plurality of data points (15), Defining at least one object class, each object class including a predetermined object type and a geometry configured to enclose an object of the predetermined object type, i) Via a first stage (19) of the neural network: Performing, for the plurality of data points (15), semantic segmentation (23) with respect to the at least one object class and the background, and For each data point (15), estimating a first approximation (25) of a spatial condition of the geometry of the at least one object class, and ii) Based on the first approximation (25) of the spatial condition of the geometry, via a second stage (21) of the neural network: Selecting a subset of the plurality of data points (15) based on the semantic segmentation (23), and For each data point (15) of the subset: Estimating a second approximation (49, 51, 53) of the spatial condition of the geometry of the object class, the second approximation refining the first approximation, and Estimating a confidence score (55) of the second approximation of the spatial condition of the geometry, wherein The spatial condition of the geometry includes a center (63) of the geometry, and The steps of estimating the first approximation (25) and the second approximation (49, 51, 53) of the spatial condition of the geometry include: for each data point, estimating a corresponding vector that moves the data point towards the center (63) of the geometry.

2. The method according to claim 1, wherein The spatial condition further includes a dimension (51) of the geometry, and The step of estimating a second approximation (49, 51, 53) of the spatial condition comprises: Estimating the dimension (51) of the geometry.

3. The method according to claim 2, wherein The spatial condition further includes an orientation (53) of the geometry, and The step of estimating a second approximation (49, 51, 53) of the spatial condition comprises: Estimating the orientation (53) of the geometry.

4. The method according to claim 3, wherein The step of estimating the orientation (53) of the geometry is decomposed into a step of regressing angle information of the geometry and a step of classifying the angle information.

5. The method according to claim 1, wherein The first stage (19) of the neural network includes respective output paths for the semantic segmentation (23) and for at least one parameter (27) of the spatial condition, The second stage (21) of the neural network includes respective output paths for the confidence score (55) and for the at least one parameter (27) of the spatial condition that are refined, and Each output path of the neural network is trained based on a separate objective function.

6. The method according to claim 1, wherein The step of performing semantic segmentation (23) is trained by: defining the ground truth of the geometric shape and determining whether the corresponding data points (15) are within the ground truth of the geometric shape.

7. The method according to claim 1, wherein while training the second stage (21) of the neural network, the first stage (19) of the neural network is deactivated.

8. The method according to claim 1, wherein the geometric shape of the at least one object class includes a bounding box (61), and The step of estimating a second approximation (49, 51, 53) of the spatial condition comprises: the dimensions (51) and orientation (53) of the bounding box (61) are estimated.

9. The method according to claim 1, wherein at least two different object classes are defined, and the at least two different object classes include different object types and different geometric shapes.

10. The method according to claim 1, wherein the step of estimating a second approximation (49, 51, 53) of the spatial condition of the geometric shape is performed by: performing farthest point sampling only on a selected subset of the plurality of data points (15).

11. The method according to claim 1, wherein The step of estimating a second approximation (49, 51, 53) of the spatial condition of the geometric shape comprises: the data points (15) of the selected subset are grouped to train the second stage (21) of the neural network with respect to the local structure.

12. A device (11) for detecting an object, the device comprising: a radar sensor (13) that acquires radar signals reflected by objects in the environment of the radar sensor (13) and provides a raw radar point cloud based on the radar signals, wherein the raw radar point cloud includes a plurality of data points (15), a module (17) configured to establish a neural network including a first stage (19) and a second stage (21) and define at least one object class, each object class including a predetermined object type and a geometric shape configured to surround the object of the predetermined object type, wherein, i) the first stage (19) of the neural network is configured to perform the following operations: perform semantic segmentation (23) on the plurality of data points (15) with respect to the at least one object class and the background, and for each data point (15), estimate a first approximation (25) of the spatial condition of the geometric shape of each object class, and ii) the second stage (21) of the neural network is configured to perform the following operations based on the first approximation (25) of the spatial condition of the geometric shape: select a subset of the plurality of data points (15) based on the semantic segmentation (23), and for each data point (15) of the subset: estimate a second approximation (49, 51, 53) of the spatial condition of the geometric shape of the object class, the second approximation refining the first approximation, and estimate a confidence score (55) of the second approximation of the spatial condition of the geometric shape, wherein, the spatial condition of the geometric shape includes the center (63) of the geometric shape, and When estimating the first approximation and the second approximation of the spatial conditions of the geometry of the object class, the neural network is configured to estimate, for each data point, a corresponding vector that moves the data point towards the center (63) of the geometry.

13. A computer system configured to perform a computer-implemented method according to any one of claims 1 to 11.

14. A non-transitory computer-readable medium comprising instructions for performing a computer-implemented method according to any one of claims 1 to 11.