METHOD, PROCESSOR CIRCUIT AND COMPUTER-READABLE STORAGE MEDIUM FOR PERFORMING TRAFFIC OBJECT DETECTION IN A MOTOR VEHICLE
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- CARIAD SE
- Filing Date
- 2023-05-16
- Publication Date
- 2026-05-21
AI Technical Summary
Existing pedestrian detection systems in vehicles fail to recognize pedestrians that are only partially visible due to occlusion, requiring additional computing power that is typically not available in vehicle control units, leading to undetected pedestrians.
A method using statistical distribution models to analyze feature vectors of image data, reducing dimensions, and comparing them with predefined models to identify partially visible pedestrians, without requiring specialized classifiers or additional computing power.
Efficiently detects partially visible pedestrians with minimal computational effort, ensuring reliable pedestrian detection for automated driving systems.
Description
[0001] The invention relates to a method for operating a traffic object detection system in a control unit of a motor vehicle. The invention also relates to a processor circuit for carrying out the method and a computer-readable storage medium for enabling the processor circuit to carry out the method.
[0002] Traffic object detection uses image data from a camera image or image dataset—that is, a representation of the surroundings—to determine, using at least one machine learning model (ML model), whether and where a pedestrian is depicted in the respective camera image. From this, the relative position of the pedestrian to the vehicle can be determined by converting the sensor coordinate system of the surroundings sensor into an absolute coordinate system of the vehicle. This information can then be signaled to an automated driving function of the vehicle, which can subsequently calculate a driving trajectory for the vehicle to pass the pedestrian without collision.The automated driving function can be, for example, a driver assistance function (such as lane keeping assistance and / or parking assistance) and / or an autonomous driving function (autopilot) that can plan a driving trajectory for automated, collision-free driving of the motor vehicle when it is known where, for example, pedestrians are located in the surroundings.
[0003] The relevant state of the art in this area is known, for example, from the following scientific publications: Shifeng Zhang, Longyin Wen, Xiao Bian, Zhen Lei, and Stan Z. Li, "Occlusion-aware R-CNN: Detecting Pedestrians in a Crowd", European Conference on Compuer Vision - ECCV 2018; Wanli Ouyang and Xiaogang Wang, "A Discriminative Deep Model for Pedestrian Detection with Occlusion Handling", IEEE, 2012; Chen Ning, Li Menglu, Yuan Hao, Su Xueping, Li Yunhong, "Survey of pedestrian detection with occlusion", Complex & Intelligent Systems, 2021; Alonso I.P., Llorca D.F., Sotelo M.Ä., Bergasa L.M., de Toro P.R., Nuevo J., Ocaña M., Garrido M.A.G., "Combination of Feature Extraction Methods for SVM Pedestrian Detection", Transactions on Intelligent Transportation Systems - IEEE 2007; He Y., Zhu C., Yin X.-C, "Mutual-Supervised Feature Modulation Network for Occluded Pedestrian Detection", 25th International Conference on Pattern Recognition (ICPR), 2021; Zhou J., Hoang J., "Real Time Robust Human Detection and Tracking System", Computer Society Conference on Computer Vision and Pattern Recognition, IEEE, 2005 CAO GONG et al., "Solving Occlusion Problem in Pedestrian Detection by Constructing Discriminative Part Layers", 2017 IEEE Winter Conference on Applications of Computer Vision (WACV), IEEE, 24. März 2017, Seiten 91-99; Yanqiu Xiao et al., "Deep learning for occluded and multi-scale pedestrian detection: A review", IET Image Processing, IET, UK, Bd. 15, Nr. 2, 14. Dezember 2020, Seiten 286-301. .
[0004] An important consideration when detecting pedestrians based on images of the environment is that pedestrians are not always fully visible. Occlusion can occur, for example, when several pedestrians are standing next to each other and / or a pedestrian is standing behind an object, such as a traffic signpost. Pedestrian detection based on at least one machine learning model can fail in such cases, as described in the aforementioned publications (occlusion problem).The additional operation of further ML models trained to detect partially obscured pedestrians, of whom only a part is visible, is avoided in connection with pedestrian detection in a motor vehicle because a correspondingly additional computing power would be necessary, which is usually not available in a processor circuit of a motor vehicle's control unit.
[0005] The invention is based on the objective of efficiently recognizing, from an image data set, whether it also depicts a traffic object that is only partially visible (such as, in particular, a pedestrian that is only partially visible).
[0006] The problem is solved by the subject matter of the independent patent claims. Advantageous further developments of the invention are described by the dependent patent claims, the following description, and the figures.
[0007] As one solution, the invention comprises a method for operating or performing pedestrian detection in a processor circuit of a motor vehicle. Such a processor circuit can be formed by a control unit or a group of several control units of the motor vehicle.
[0008] Pedestrian detection is based, in a manner known per se, on the reception of at least one image data set (i.e., a camera image or a corresponding image sequence) from at least one environmental sensor of the vehicle. Such a single image data set describes a specific representation, image, or single frame of the vehicle's surroundings. The at least one environmental sensor can, in a manner known per se, incorporate one or more environmental cameras for this purpose. Alternatively, the image data set can, for example, be a single frame from a video data stream.
[0009] The method assumes that, in a manner known from the prior art, feature data of image features are extracted from the image data of the image dataset using at least one machine learning model (ML model) and a feature extraction unit of the at least one ML model. The individual pixels or image points of the image are thus grouped or examined to determine whether they represent at least one predetermined image feature, such as a structure, texture, or pattern. Edge detection is one example of such feature extraction.
[0010] Furthermore, the method is based on the prior art, which recognizes that bounding boxes or boundary contours can be identified in the respective image or representation (i.e., the image data of the image dataset). These bounding boxes frame or delimit an image area that potentially contains an image of a passerby. Such bounding boxes can be determined based on the described image features and / or by means of a preceding part of at least one machine learning model. Such a bounding box represents a "hypothesis" or pre-detection based on the fact that image features have been detected or are present in the image area, which together indicate a probability that a passerby is depicted, or that at least one object (of a still unknown object class) is present, above a predefined threshold.Within each bounding box, a classifier unit of at least one machine learning model uses the image features from that bounding box to detect or classify a fully or predominantly depicted passerby (if a person or passerby in general is actually depicted in the bounding box). If the bounding box contains image features that represent a fully visible passerby, the classifier unit responds with a detection result indicating the presence or recognition of such a passerby (otherwise, preferably not).
[0011] A classifier unit has a detection threshold above which it will detect a pedestrian within a bounding box, even if the pedestrian is not fully, but at least predominantly, for example, more than 70 percent, or generally more than a predefined minimum percentage. This minimum percentage can range from 65 to 90 percent, to give examples. Therefore, it is sufficient for the classifier unit if a pedestrian is visible to a corresponding predominant degree.
[0012] The detection result of the classifier unit then identifies or indicates the respective bounding box for which the classifier unit has detected or classified a depicted pedestrian. Each bounding box can be identified by a detection signal. Thus, all those bounding boxes for which the classifier unit has detected the presence or image of a pedestrian are known. The detection signal can, for example, specify an ID of the bounding box and / or coordinates, to name just a few examples.
[0013] As previously explained, such a classifier unit has a detection threshold. This means that if a pedestrian is only partially depicted in a bounding box—specifically, if less of the pedestrian is visible than the detection threshold of the classifier unit requires—then this pedestrian will not be detected by the classifier unit. This results in the previously described effect of occlusion, meaning the pedestrian remains undetected by the classifier unit.
[0014] In order to provide an additional check for the presence of the occlusion problem, which also detects such passers-by that have remained undetected by the actual classifier unit, the invention now provides that for some or all of the bounding boxes and / or for additionally formed bounding boxes (as will be explained further below), the image features contained or delimited therein may indicate an overlooked passer-by.
[0015] This does not require a specialized classifier unit of a machine learning model, such as an artificial neural network. Instead, the image features are grouped into a feature vector for efficient analysis. The image features are thus arranged into a vector or represented by it. This feature vector represents a point in a feature space. The number of dimensions in this feature space corresponds, in a known manner, to the number of vector components of the feature vector. Statistical distribution models are defined within this feature space. These models define, or specify, for each point in the feature space whether, or with what probability, this point (i.e., the feature vector describing the point) represents a portion of a passerby and / or a person's body—in other words, a previously undetected passerby.
[0016] In other words, several statistical distribution models are defined, each of which models a statistical distribution of image features of a partially visible passerby and / or of a person's body visible only to a certain extent. The statistical distribution divides the feature space in which the described image features are grouped into a respective feature vector. "Partial area" here means that the distribution models are based on feature vectors that represent a non-predominant depiction of the passerby, for example, only a single body part or an occlusion of more than one such "occlusion component," which can range from 25 percent to 80 percent.For example, when it comes to person detection, one distribution model might represent only a foot, and / or another distribution model might represent only a hand, and / or another distribution model might represent only a person's head. A distribution model could also represent, for example, a person's torso or a pair of legs—that is, more than just one limb. In the case of a motor vehicle, for instance, a section might represent only a wheel arch with a wheel, or a side view of only a trunk.
[0017] A distance measure, used to determine the distance between a feature vector and a given distribution model, can be based, for example, on a probability value that the statistical distribution model can output for that feature vector. The distance measure can also be a binary measure indicating whether or not a feature is associated with a feature. For example, a Gaussian kernel distribution (for probability values) or a Support Vector Machine (SVM) can be used as the statistical distribution model. Alternatively, a classifier can be used as the statistical distribution model, which, for example, can be based on an artificial neural network. Using a correspondingly short feature vector results in minimal computational effort.
[0018] Thus, for each feature vector of image features within a bounding box and for each distribution model, there is a distance value that indicates how similar the feature vector is to the distribution model or to that part of a passerby or body part of a person that is represented by the respective distribution model. The distance value is compared to a predefined threshold or trigger value, and if, according to the comparison, the distance value to one of the distribution models is less than the threshold, a signal is triggered indicating that a passerby remained undetected by the classifier unit and has now been detected based on the distribution model.If the distance value is smaller than the threshold value, then there is a correspondingly high degree of similarity or belonging of the feature vector to the distribution model, meaning that the feature vector represents with a correspondingly high probability a sub-area of a passerby or body area of a person and / or a human body.
[0019] The feature vector determined for a given bounding box can be a concatenation or a listing of all image features from the bounding box. However, it is preferred to reduce the length of the feature vector, since a bounding box can contain a large number of image features, particularly more than 100 or more than 1000. Accordingly, the invention provides that the feature vector is generated by combining the image features into a temporary vector, i.e., a vector that encompasses all image features. This temporary vector is then reduced to the feature vector by means of a dimension-reducing mapping. The feature vector thus has fewer vector components than the temporary vector.For example, one of the following methods known from the prior art can be used as a suitable dimension-reducing mapping: Multidimensional scaling (MDS), greedy-forward selection, or correlation-based feature selection.
[0020] The invention offers the advantage that, with relatively little computational effort, namely calculating a feature vector and comparing it with several distribution models, it can be determined whether the classifier unit of at least one ML model of a passerby in an image data set has overlooked a feature because only a partial area is visible or depicted, which in particular was smaller than the detection threshold of the classifier unit.
[0021] The invention can also be applied to road users other than pedestrians, i.e., to motor vehicles and / or cyclists, i.e., to road user detection. The invention can also be applied to other stationary traffic infrastructure objects, i.e., to traffic signs, road markings, and traffic signals (traffic lights), i.e., to traffic infrastructure detection. In general, the invention can be applied to traffic objects, i.e., road users and traffic infrastructure objects, i.e., to traffic object detection.
[0022] Therefore, when the term "pedestrian" is used here, it could also refer to a road user of another type (motor vehicle and / or cyclist) and / or to a piece of traffic infrastructure. In general, it doesn't necessarily have to be a pedestrian, but could refer to any traffic object in general.
[0023] The invention also includes further developments that result in additional advantages.
[0024] The bounding boxes described, which are generated as hypotheses, suggestions, or input data for the classifier unit, can be very numerous when using a prior art algorithm; for example, there can be more than 100 or even more than 1000 bounding boxes per image or image dataset. To avoid having to calculate a feature vector for all bounding boxes and compare it with statistical distribution models, a further development proposes excluding those bounding boxes for which the classifier unit already recognizes that they depict a pedestrian. Thus, the detection result or the recognition result of the classifier unit is excluded because no pedestrian could have been overlooked there.Furthermore, bounding boxes are excluded if they are found to overlap with a bounding box representing a pedestrian, as determined by the classifier unit, by more than a predefined minimum area. This minimum area can range from 80% to 99%. This is based on the understanding that prior art algorithms used to determine bounding boxes can output or signal multiple, offset bounding boxes for a single pedestrian. All of these bounding boxes can be excluded if the classifier unit signals that a pedestrian is represented within one of them.
[0025] To prevent the classifier unit from signaling a single pedestrian while overlooking a second, partially obscured pedestrian, a subtraction function can be used. This subtracts or removes from a bounding box the portion of its area that belongs to a bounding box with a detected pedestrian. In other words, any overlap between a bounding box with a pedestrian and another bounding box (for which no pedestrian has yet been detected) is removed. The non-overlapping area is then described by at least one additional bounding box. Defining multiple additional bounding boxes may be necessary if they must have a predefined basic shape, such as a rectangle.Thus, even behind a recognized or detected pedestrian, the visible portion of another pedestrian who would otherwise be hidden can be detected using distribution models.
[0026] A particularly preferred approach, according to a further development, involves the dimension-reducing mapping comprising a transformation of the temporary vector using principal component analysis. This results in so-called principal components as the transformed vector. For the dimensionality reduction, only a predetermined subset of the vector components, i.e., fewer than all vector components, is then used for the feature vector. In particular, the first N principal components are used (N being an integer). Thus, for example, the number N of vector components of the feature vector can be reduced to less than 200, and in particular, less than 100.
[0027] As previously explained, the classifier unit is expected to have, or inherently possesses, a pedestrian detection threshold. This threshold specifies the percentage or proportion of a pedestrian that must be visible within the bounding box for the classifier unit to detect them. Each distribution model is preferably designed to model a sub-area that lies below the detection threshold. Thus, the distribution model can identify a representation of a pedestrian's sub-area as belonging to a pedestrian for whom the classifier unit cannot perform a detection due to its detection threshold. The respective distribution model can be designed with a corresponding feature vector of representations of pedestrian sub-areas.This can be done in a known manner based on histograms of a corresponding number of feature vectors from training datasets, as will be described below.
[0028] To obtain a suitable sampling point or location for generating or extracting feature data from image features, a further development proposes using a convolutional neural network (CNN) as the feature extraction unit. Such a feature extraction unit can detect related pixels within an image or within pixels of an environment that collectively represent a specific pattern or structure, such as edges and / or regions of a particular color or color pattern, and / or basic shapes like angles, corners, or pairs of eyes, to name just a few examples. The CNN contains corresponding filters that, through correlation, can locate the pattern or structure within the image.
[0029] A deep artificial neural network (DNN) can be used as a classifier unit following such a feature extraction unit, as is provided for in a further development. Such a DNN is also referred to as a fully connected neural network (FCNN). It assigns a detection class or a detection result to the respective extracted feature data, for example, the statement of whether, or with what percentage probability, the group of feature data in a bounding box represents a pedestrian.
[0030] The training of convolutional neural networks and / or deep artificial neural networks can be performed using an algorithm known from the prior art, such as the backpropagation algorithm. Reference is made to the prior art publications described at the beginning.
[0031] As feature data representing or describing the extracted image features, the activation values of artificial neurons in at least one network layer of the feature extraction unit are used or determined according to a further development. Such an activation value is, in a manner known per se (depending on the specific scientific terminology), the input or output value of an activation function (for example, a sigmoid function) of the respective artificial neuron. At an output layer or in the last or subsequent layers of a feature extraction unit, such an activation value already represents a complete image feature, for example, a monochrome area or an edge, to name just a few examples.In particular, feature data can be accessed or read from an area that, in a machine learning model consisting of CNN and DNN, is referred to as the "bottleneck", i.e., the transition area or interface between these two networks.
[0032] If statistical distribution models for a bounding box detect that it contains an undetected pedestrian (i.e., a pedestrian only partially depicted), a predetermined safety measure is triggered in the vehicle if the undetected pedestrian is signaled. In other words, the signaling of the undetected pedestrian can be linked to or coupled with such a safety measure. This safety measure could, for example, involve the automated driving function, which receives the pedestrian detection result, reducing the vehicle's speed and / or adjusting or checking a planned trajectory to see if it points towards the undetected pedestrian, and if so, rerouting the vehicle around the undetected pedestrian.
[0033] The statistical distribution models described above are necessary for the procedure to be carried out. These can be generated, as previously described, based on histograms of feature vectors that depict individuals only partially or within a subset of the image. To generate such histograms systematically, a further development proposes that, to generate the distribution models from training datasets (i.e., images of environments with pedestrians), bounding boxes of fully depicted individuals are divided into subsets, and the image features contained in each subset are then grouped into corresponding training feature vectors. The previously described dimension-reducing mapping, such as PCA, can also be used in this process.In particular, the training feature vectors are generated in the same way as the feature vectors already described, as used in the operation of the person detection system, to ensure that identical feature vectors are produced. The training feature vectors result in point clouds or a point cloud within the described feature space. The training feature vectors, or rather their points in the feature space, are then grouped into clusters using a clustering algorithm. For example, the K-means algorithm can be used for this purpose. Alternatively, clusters can be determined using a Support Vector Machine (SVM). Each cluster represents one of the described statistical distribution models.If, during person detection, it is determined that a feature vector in a bounding box has a distance or distance to one of the clusters or a cluster center that is less than the specified threshold, then this feature vector is assigned to the cluster, thus determining that the feature vector, and therefore the bounding box, represents a part of a person. Otherwise, the feature vector represents an undetected passerby. For example, a single leg, a pair of legs, or a head can each be represented by a cluster. Therefore, detection, or at least an indication of an undetected person, is also possible.
[0034] To carry out the method, the invention also includes the described processor circuit for a motor vehicle. The processor circuit can be implemented by a control unit or a group of several control units in the motor vehicle. The processor circuit is configured to execute an embodiment of the method according to the invention. For this purpose, the processor circuit can comprise at least one microprocessor and / or at least one microcontroller and / or at least one FPGA (Field Programmable Gate Array) and / or at least one DSP (Digital Signal Processor). Furthermore, the processor circuit can comprise program code configured to execute the embodiment of the method according to the invention when carried out by the processor circuit. The program code can be stored in a data memory of the processor circuit.
[0035] For use cases or application situations that may arise during the procedure and are not explicitly described here, it may be provided that, according to the procedure, an error message and / or a request for user feedback is issued and / or a default setting and / or a predetermined initial state is set.
[0036] To enable a conventional motor vehicle processor circuit to execute the method, the invention also provides a computer-readable storage medium containing program instructions that, when executed by a motor vehicle processor circuit, cause it to perform an embodiment of the inventive method for pedestrian detection in the motor vehicle. A further claimed computer-readable storage medium contains program instructions that, when executed by a computer—such as one located at a manufacturer of control units for motor vehicles and / or in a laboratory or workshop, or such a computer implemented as a backend—cause this computer to perform the described determination of training feature vectors and clusters therefrom according to the described method.Such a computer-readable storage medium can therefore be used to generate the statistical distribution models in a laboratory or at a manufacturer's facility, which can then be used in a processor circuit of the motor vehicle to carry out the described procedure.
[0037] The invention also includes a motor vehicle in which an embodiment of the described processor circuit is coupled to at least one environmental vector of the motor vehicle for receiving image data sets and to an automated driving function for providing detection results. The motor vehicle according to the invention is preferably configured as a motor vehicle, in particular as a passenger car or truck, or as a passenger bus or motorcycle.
[0038] As a further solution, the invention also includes a computer-readable storage medium comprising program code which, when executed by a computer or a computer network, causes it to execute an embodiment of the method according to the invention. The storage medium can, for example, be provided at least partially as a non-volatile data storage medium (e.g., as flash memory and / or as an SSD - solid state drive) and / or at least partially as a volatile data storage medium (e.g., as RAM - random access memory). The storage medium can be implemented in the processor circuit within its data storage. However, the storage medium can also be operated, for example, as a so-called app store server on the internet. The computer or computer network can provide a processor circuit with at least one microprocessor. The program code can be binary code or assembly language and / or source code of a programming language (e.g.,C) and / or be provided as a program script (e.g. Python).
[0039] The following are exemplary embodiments of the invention described. This is illustrated by: Fig. 1 a schematic representation of an embodiment of the motor vehicle according to the invention; Fig. 2 a sketch to illustrate an embodiment of the method according to the invention.
[0040] In the figures, identical reference symbols denote functionally equivalent elements.
[0041] Fig. 1 Figure 11 shows a motor vehicle 10, which may be a car, in particular a passenger car or truck. The motor vehicle 10 may have an automated driving function 11, by which an actuator 12 of the motor vehicle 10 can be controlled automatically or without driver intervention by means of a control signal 13. The actuator 12 may be designed for lateral control (steering) and / or longitudinal control (acceleration and braking) of the motor vehicle 10. By means of the control signal 13, the driving function 11 can thus guide the motor vehicle 10 along a driving trajectory 14 by controlling the actuator 12. The driving function 11 can calculate this trajectory in order to guide the motor vehicle 10 through an environment 15, for example, road traffic or a road network, without collisions.To plan the driving trajectory 14, pedestrian detection 16 can be placed upstream of the driving function 11, which can be carried out by a machine learning model ML.
[0042] The pedestrian detection 16 can receive image data sets 19 from at least one environmental sensor 17 of the motor vehicle 10, for example a camera, whose detection range 18 can be directed into the environment 15, each of which can depict the environment 15 with potentially visible persons or pedestrians 20, 21 therein.
[0043] The ML model can include a feature extraction unit 22, which may, for example, be based on a convolutional neural network (CNN). Using the feature extraction unit 22, image features can be extracted from the image datasets 19, as is generally known for convolutional neural networks or for computer vision processing of image datasets 19. Additionally or alternatively, bounding boxes 23 can be determined from the image datasets 19. These bounding boxes delineate or frame regions or image areas of the images according to the image datasets 19 in which, according to the feature-extracted image features 24, a person could be located as a passerby 20, 21. These are so-called hypotheses or suggestions. The image data of the individual bounding boxes 23 can be provided to a classifier unit 25, which may, for example, be based on a fully connected neural network (FCNN).The classifier unit 25 can, based on the image features 24 from the individual bounding boxes 23, generate a recognition result 26 in a manner known per se, which indicates which of the bounding boxes 23 actually contains a person as a passerby 20, 21.
[0044] For the purposes of further explanation, it is assumed that the classifier unit 25 can detect the fully visible pedestrian 20, who is completely visible or depicted in the respective image data set 19. In contrast, pedestrian 21 is only partially depicted or only a portion of pedestrian 21 is shown and was overlooked or not detected by the classifier unit 25 in this example. In this embodiment, a pedestrian is a person.
[0045] The detection result or recognition result 26 of the detection can be signaled to the driving function 11. Based on the position of the detected pedestrian 20 in the image according to image data set 19, the driving function 11 can determine a relative position of the pedestrian 20 in the environment 15 with respect to the motor vehicle and then plan the driving trajectory 14.
[0046] To check whether the classifier unit 25 has missed a pedestrian, for example pedestrian 21, in addition to those bounding boxes 23 that do not frame or cover any of the detected pedestrians 20, 21, a respective feature vector 31 can be determined using a dimension-reducing mapping 30. For this purpose, the image features 24 of the respective bounding box 23 can, for example, be combined into a preliminary vector 33, which can be reduced in dimension or length using mapping 30 to generate the feature vector 31. A principal component analysis (PCA) can, for example, be used as mapping 30.
[0047] The feature vector 31 can be compared with statistical distribution models 35 in a distance calculation 34, or a membership value or a value for an occurrence probability can be determined for the feature vector 31 according to the respective statistical distribution model 35. From this, a respective distance value 36 of the feature vector 31 with respect to the statistical distribution model 35 can be determined. The smallest distance value 36 can be used further, since in particular only the smallest distance value 36 needs to be used for the subsequent procedural steps. For example, the reciprocal of a probability value that yields an occurrence probability of the feature vector 31 according to the respective statistical distribution model 35 can be used as the distance value 36.A Gaussian kernel distribution function, for example, can be used as a statistical distribution model 35, which can specify a probability of occurrence for a feature vector 31. If a support vector machine (SVM) is used as the statistical distribution model 35, the detection result is a binary distance value (belong or not belong).
[0048] In a threshold comparison 37, the distance value 36 can be compared with a threshold value 38, which is indicated here by the Greek letter β. If the distance value is greater than the threshold value 38 (symbolized by a plus symbol), then there is no match between the tested bounding box according to feature vector 31 and any of the statistical distribution models 35, and thus the bounding box does not represent a sub-area of a passerby (symbolized by an "OK" checkmark).
[0049] If, on the other hand, in the threshold comparison 37 the distance value 36 is smaller than the threshold value 38 (symbolized by a minus symbol), a signal 39 can indicate that an undetected or overlooked pedestrian is present in the image data set 19 underlying the bounding boxes. The signal 39 can, for example, trigger the described safety measure 40; that is, the driving function 11 can, for example, reduce the speed of the vehicle 10 contrary to the previously planned trajectory 14 or ensure that the speed is kept below a specified maximum speed.
[0050] Fig. 2 illustrates by way of example how the distribution models 35 can be formed and how the undetected passerby 21 can be determined using the distribution models 35 based on the feature vector 31.
[0051] Fig. 2 to this end (with reference to Fig. 1 ) shows how training datasets 60 ( Fig. 1 ) can be analyzed in the same way using the feature extraction unit 22 as already described. The training datasets 60 can be labeled in the known manner, i.e., it can be known which bounding box 23 represents or depicts a fully visible person 61 as a passerby. Such a bounding box 23 can be divided or decomposed into sub-areas 62, i.e., the image features within the bounding box 23 can be assigned to the respective sub-area 62. It is known from the prior art that an image feature 24 is also assigned a location or position within an image in an image dataset 19.
[0052] The image features 24 of each sub-area 62 can now be mapped or transformed into training feature vectors 63 in the manner described, using the dimension-reducing mapping. The training feature vectors 63 can be grouped together using a clustering algorithm 64. This results in several clusters 66 in a feature space 65, each of which represents a distribution model 35. The feature space 65 is represented here, for the sake of simplicity, as a two-dimensional space (plane). For individual areas or points in the feature space 65, the respective distribution model 35 can indicate whether and / or with what probability such a point in the feature space 65 belongs to, or is caused or generated by, the distribution model 35.
[0053] For example, if the bounding box 23 of pedestrian 21 is transferred or transformed into the feature vector 31, this feature vector 31 represents a point 67 in the feature space 65, which has a distance value 36 to the distribution model 35 that is smaller than the threshold value 38. Accordingly, the point 67 can be determined as belonging to the distribution model 35, i.e., its cluster 66. Thus, the threshold comparison 37 shows that an undetected pedestrian 21 is present, and therefore the signaling 39 must be triggered or started.
[0054] This provides an overall reliability metric to find the hidden data points based on their distance from the center of mass or geometric center of the clusters from the training data.
[0055] Based on this, a system installed in the vehicle is created that uses a reference cluster model, which was created in the backend based on the training data set used for the optimization of the DNN (generally the classifier unit), to check whether the currently processed image contains obscured pedestrians.
[0056] In every phase of perception and decision-making where DNNs make predictions, the reliability of such predictions is crucial for the automated driving system as a whole, since the system's upcoming decisions can be influenced by these reliability values. In other words, if a prediction from a subsystem, such as perception, proves unreliable, the system must make alternative decisions instead of the uncertain ones, as otherwise the safety of passengers or other road users could be jeopardized. In the case of this patent, this algorithm checks the possibility that a hidden pedestrian might be overlooked by the DNN.
[0057] Current state-of-the-art methods rely on computationally intensive approaches that are not always easily applicable given the limited resources of operational equipment such as vehicles. Furthermore, their predictions are misleading when dealing with adverse disturbances, where they still exhibit a high degree of confidence in the incorrectly predicted data points. Our method is based on lightweight statistical models that require only a fraction of the computing power needed by the main DNN, thus preserving the efficiency of the overall recognition system. Because the method is based on statistical analysis techniques, system engineers can also define decision boundaries to rely only on a specific range of valid recognitions and consider the rest unreliable. This is particularly helpful in establishing reliable and safe decision ranges within which the DNN can perform well.
[0058] Furthermore, our method can also be used for the security argumentation of perceptual DNNs, whereby the DNNs are evaluated on a variety of input data points at distances from their true class cluster centers, and evidence can be generated based on this.
[0059] An overview of a particularly preferred embodiment of the method is as follows. The figures show that the method decides whether a prediction by the automated driving system in the vehicle is reliable or not. The decision as to whether a prediction by the automated driving system in the vehicle is reliable or not is made in the following six steps. Steps S1 to S3 are performed in the backend during the system design phase, while steps S4 to S6 are performed in the vehicle during runtime, requiring minimal computing power. The steps, in particular, are as follows: (1) [In the backend], the activations of one or more CNN layers (or a similar encoder) are extracted for the entire training dataset for the fully visible pedestrians. These are then divided into several random splits, each covering a subset of the pedestrians. Subsequently, each split is simplified using linear dimensionality reduction methods such as principal component analysis (PCA), resulting in a simplified feature space that can be divided into different classes. Due to its linearity, this model is very small and lightweight, requiring minimal computational power. Consequently, at the end of this step, many statistical distributions are extracted on different subdivisions of the pedestrians, which are used as a reference for the next steps. (2) [In the backend] Based on the results from (1), a model is developed that forms a cluster for each split.This is then used to estimate a probability score that defines the probability that a data point belongs to a cluster. This probability value is calculated in such a way that the vehicle requires as little computing power as possible. (3) [In the backend], if required for the clustering method, a probability threshold is defined for each cluster, representing the minimum probability that a data point belongs to that cluster or not. (4) [In the vehicle] With conventional 2D object recognition methods, thousands of 2D "proposals" are generated, which are suppressed in the later stages of recognition, and only a few of which lead to the final result. Our algorithm uses these many 2D proposals to extract the corresponding filter activations from the CNN layer mentioned above.These suggestions are then narrowed down to exclude those that overlap with a pedestrian already detected by the main DNN. (The goal is to find the missed obscured pedestrians.) Finally, they are transferred to the new space using the same PCA model as in (1) and compared with the clusters formed in (2). (5) [In the vehicle] The probability value of the results from (4) is estimated with respect to the split clusters, and the nearest cluster is returned. (6) [In the vehicle] If the probability value estimated in step S5 is less than the threshold calculated in step S3, then the final prediction is considered to be "a likely obscured pedestrian missed by the main pedestrian detector."
[0060] Overall, the examples show how, for an automated driving function, an additional check for overlooked or undetected pedestrians can be provided during pedestrian detection, which can be based on clusters (distribution models) in a reduced feature space and can therefore be carried out with low computational effort.
Claims
1. A method for operating traffic object detection (16) in a processor circuit of a motor vehicle (10), wherein at least one image data set (19) describing relevant imaging (30) of an environment (15) of the motor vehicle (10) is received from at least one environment sensor (17), and by means of at least one machine learning model (ML model), on the basis of the relevant image data set (19), • determining bounding boxes (23) containing potential images of traffic objects (20, 21) in the imaging (30) and • extracting feature data of image features (24) from image data of the image data set (19) by means of a feature extraction unit (22) of the at least one ML model; • detecting a traffic object (20, 21) that is completely depicted or is depicted by more than a predetermined minimum fraction, the minimum fraction being in a range of 65 to 90 percent, within the respective bounding box (23) on the basis of the image features (24) contained therein by means of a classifier unit (25) of the at least one ML model, and the relevant bounding box (23) depicting a traffic object (20, 21) is identified by a detection signal as a result of the detection, the detection signal indicating an ID of the bounding box (23) and / or coordinates, characterized in that, • for further bounding boxes (23) formed by subtraction, the image features (24) contained therein are combined to form a feature vector (31) in each case, and, during the subtraction, that area portion is subtracted or removed from a bounding box (23) which belongs to a bounding box (23) containing a traffic object (20, 21) identified by means of the classifier unit (25), and • a relevant distance value (36) of the feature vector (31) is determined to form a plurality of statistical distribution models (35), each of which models a statistical distribution of such image features (24) of only one relevant portion (62) of a traffic object (20, 21) and / or a body of a person, and in this case, "portion" (62) means that the distribution models (35) are based on those feature vectors which represent nonpredominant imaging of the traffic object (20, 21) or a pedestrian, i.e. represent only a single body part or obscuring of the traffic object (20, 21) or pedestrian by more than such an "occlusion fraction", which may be in a range of 25 percent to 80 percent, and • comparing the distance value (36) with a predetermined threshold value (38) and, • if, according to the comparison, the distance value (36) for one of the statistical distribution models (35) is less than the threshold value (38), it is indicated that a traffic object (20, 21) has gone undetected by the classifier unit (25); if the distance value is namely less than the threshold value, there is an accordingly high level of similarity or affiliation of the feature vector with the distribution model (35), i.e. the feature vector represents, with an accordingly high probability, a portion (2) of a traffic object (20, 21) or bodily region of a person and / or a human body, • the feature vector (31) being formed by the image features (24) being combined to form a temporary vector and the temporary vector being reduced to form the feature vector (31) by means of dimension-reducing imaging (30).
2. The method according to claim 1, wherein, before determining the distance value (36), those bounding boxes (23) are excluded for which it is identified that they overlap in area with a bounding box (23) for which it is indicated by the classifier unit (25) that it depicts a traffic object (20, 21) by more than a predetermined minimum fraction, so as not to calculate a feature vector (31) for all the bounding boxes (23) and not to have to compare said feature vector with the statistical distribution models, and / or wherein the subtraction comprises at least one of the further bounding boxes (23) being formed in that, for one of the bounding boxes (23) which has an overlap with a bounding box (23) for which it is indicated by the classifier unit (25) that it depicts a traffic object (20, 21), the nonoverlapping part is described as at least one further bounding box (23), in order to prevent the classifier unit (25) indicating a single pedestrian while a second, partially obscured pedestrian behind them is overlooked.
3. The method according to any one of the preceding claims, wherein the dimension-reducing imaging (30) involves a transformation of the temporary vector by means of a principal component analysis, and only a predetermined partial number of the vector components from the transformed vector are used for the feature vector (31).
4. The method according to any one of the preceding claims, wherein each of the distribution models models a portion (62) that is below a detection threshold of the classifier unit (25).
5. The method according to any one of the preceding claims, wherein a convolutional network (CNN) is used as the feature extraction unit (22) and / or a deep artificial neural network (DNN) is used as the classifier unit (25).
6. The method according to any one of the preceding claims, wherein activation values of artificial neurons of at least one network layer of the feature extraction unit (22) are determined as feature data from the feature extraction unit (22).
7. The method according to any one of the preceding claims, wherein if an undetected traffic object (20, 21) is indicated, a predetermined safety measure (40) is triggered in the motor vehicle (10).
8. The method according to any one of the preceding claims, wherein for generating the distribution models (35) from training data sets (60), bounding boxes (23) of completely depicted traffic objects (20, 21) are decomposed into the portions (62) and the image features (24) contained in the relevant portion (62) are combined to form respective training feature vectors (63), and the determined training feature vectors (63) are divided into clusters (66) by means of a cluster algorithm, wherein each cluster (66) constitutes one of the statistical distribution models (35).
9. A processor circuit for a motor vehicle (10), wherein the processor circuit is configured to carry out a method according to any one of claims 1 to 7.
10. A computer-readable storage medium containing program instructions which, when executed by a processor circuit, cause the processor circuit to carry out a method according to any one of claims 1 to 7 or, when executed by a computer, cause the computer to carry out a method according to claim 8.