Method, device and machine readable storage medium for determining a quality level of a data set of a sensor
By training deep neural networks to evaluate the quality level of sensor datasets, the problem of sensor modal quality assessment is solved, thereby improving the accuracy and safety of autonomous driving systems.
Patent Information
- Application Number
- CN202011108465.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-16
- Filing Date
- 2020-10-16
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2040-10-16
AI Technical Summary
Existing technologies struggle to effectively assess the quality of each modality in sensor data fusion, making it difficult to balance false positive and false negative rates, which affects the safety and accuracy of autonomous driving systems.
By employing machine learning models, particularly deep neural networks, and training and fusing datasets from multiple sensors, the quality level of each sensor's dataset is evaluated by leveraging the differences in data features across different modalities, thereby optimizing sensor data fusion and environmental perception.
It improves the accuracy and reliability of sensor data fusion, reduces false positive and false negative rates, and enhances the safety and recognition capabilities of autonomous driving systems.
Smart Images

Figure CN112668602B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to a method for training a machine learning model for determining a quality level of a data set of a sensor and to a method for determining a quality level of a data set of a sensor. The invention further relates to a device, a computer program and a machine readable storage medium. BACKGROUND
[0002] The realization of a perception or rather representation (English: Perception) of the surroundings is of great significance for the realization of at least partially automated driving and is also increasingly important for driver assistance systems. Here, the surroundings are detected by means of sensors and, for example, recognized by means of pattern recognition methods.
[0003] The data of the sensors are thus converted into a symbolic description of the surroundings-related aspects. This symbolic description of the surroundings then forms the basis for implementing actions in the surroundings described in this way or for the application or use purposes of, for example, at least partially automated mobile platforms. A typical example of a symbolic description of the surroundings is the description of static and dynamic objects by means of attributes which characterize, for example, the position, shape, size and / or speed of the respective object. The objects can relate, for example, to obstacles which have to be avoided in order to collide.
[0004] The environmental perception is usually based on a combination or fusion of data of a plurality of sensors of different modalities. For example, different sensor types such as optical cameras, radar systems, lidar systems and ultrasonic sensors can be combined into a multi-modal sensor group.
[0005] Here, the fields of view of some different sensors of different modalities can overlap, so that within the field of view relevant to the application there is always sensor data of at least two different modalities. This has the advantage that deficiencies of a single sensor of one modality can in principle be compensated by other sensors of another modality. For example, in a combination of optical cameras and radar systems, the limited visibility of the optical cameras in fog can be mitigated by perceiving objects by means of radar measurements. SUMMARY
[0006] In order to mitigate deficiencies of sensors in the context of sensor data fusion, different approaches can be applied. For example, in certain applications, it can be sufficient to optimize only the false positive rate of object detection or only the false negative rate of object detection.
[0007] For example, if—for at least partially automated robots—the most important criterion is avoiding collisions in at least partially automated mode, and manual control can be performed after the robot stops and detects a false positive (“Ghost”), then the modules for sensor data fusion and environment awareness (MSFU) can be optimized to avoid false negative detections. For simplicity, in such optimization, if at least one of the sensor modalities generating data senses an object, that object is assumed to be real.
[0008] In this scenario, if one sensor fails to detect an object (a false negative), while the other sensor modality remains unaffected by this erroneous detection, the deficiency of the first sensor is successfully mitigated by the other sensor modality. However, if one of the sensor modalities is affected by false positives (“Ghosts”), errors will occur in the method because this situation cannot be mitigated by using other modalities in this optimization. This is a drawback of existing solutions.
[0009] Similarly, in Automatic Emergency Breaking (AEB), for example, it makes sense to optimize the system to avoid false positives in such a way that the module for sensor data fusion and environment awareness (MSFU) only considers an object real when it is perceived by all sensor modalities within the object's field of view. In this case, existing systems also have drawbacks because the optimization for avoiding false positives results in a higher false negative recognition rate compared to individual sensor modalities.
[0010] The feasible solutions described in the existing system for mitigating sensor limitations do not use any information to indicate which sensor mode is more or less reliable. This at least partially explains the reasons for the aforementioned shortcomings.
[0011] Continuous and automated evaluation of sensor quality is a problem that is difficult to solve using traditional methods, as multiple different conditions can affect quality in each mode. This problem has not been solved or is only insufficient in existing methods for representing the surrounding environment.
[0012] This invention discloses a method for training a machine learning model for determining the quality level of a sensor dataset, a method for fusing sensor datasets, a method for determining the quality level of a sensor dataset, and a method, an apparatus, a computer program, and a machine-readable storage medium according to features of the invention. Advantageous configurations are preferred embodiments and the subject matter described below.
[0013] This invention is based on the knowledge that a machine learning model (e.g., a neural network) can be trained to determine the quality of a dataset from one of multiple sensors in a specific surrounding environment of the sensor. The machine learning model makes the determination using multiple datasets generated under different conditions in different surrounding environments by multiple sensors of different modalities.
[0014] According to one aspect, a method is proposed for training a machine learning model to determine the quality level of at least one dataset for each of a plurality of sensors, wherein the plurality of sensors are configured to generate a representation of the surrounding environment.
[0015] Here, in one step, multiple datasets are provided for each of the multiple sensors, corresponding to multiple representations of the surrounding environment.
[0016] In another step, property data for multiple ground-truth objects representing the surrounding environment are provided.
[0017] In another step, the quality level of the corresponding dataset for each of the multiple sensors is determined using a metric (Metrik), wherein the metric compares at least one variable determined using the corresponding dataset with at least one attribute data of at least one reference truth object representing the surrounding environment.
[0018] In another step, a machine learning model is trained using multiple (especially new) datasets from each of the multiple sensors and their respective determined quality levels to determine the quality level of at least one dataset from each of the multiple sensors.
[0019] According to one approach, a machine learning model is trained by taking a dataset of multiple datasets from each of multiple sensors as input signals and providing the corresponding quality level as the target output to the machine learning model in multiple training steps, so as to enable the machine learning model to match training samples with the datasets and the corresponding quality levels.
[0020] After this training, which is performed in multiple iterations, ends, the machine learning model is set up to provide output variables when a new dataset from each sensor is fed as an input signal to the machine learning model, so as to output a prediction of the quality level of each individual dataset for each of the multiple sensors.
[0021] The method steps are shown in the order described in the overall description of the invention to make the method easy to understand. However, those skilled in the art will recognize that many method steps can be traversed in different orders and still result in the same outcome. In this sense, the order of the method steps can be changed accordingly and is therefore also disclosed.
[0022] Here, attributes are used to describe static and dynamic objects representing the surrounding environment, wherein the attributes, for example, relate to the spatial location of the object, and the attributes describe the shape, size, or speed of the corresponding object.
[0023] The dataset can, for example, originate from sensors of different modalities, which are individually or collectively configured to generate a representation of the surrounding environment. Examples of different modalities, or different sensor types, could be optical cameras, radar systems, lidar systems, or ultrasonic sensors. Additionally, the sensors can include all other types of sensors that represent the desired representation of the surrounding environment or its state—e.g., temperature sensors and / or microphones—that are capable of generating a dataset of noise about the surrounding environment accordingly. Multiple such sensors can then be combined into a multimodal sensor group.
[0024] Multiple sensors can include a small number of sensors—such as two or three sensors—or 100 sensors, or even more sensors.
[0025] Here, the term "benchmark truth object" or "benchmark truth surroundings representation" should be understood as these descriptions of the object or surroundings representing a reference that describes the reality of the surroundings with sufficient accuracy for the corresponding purpose. In other words, the benchmark truth object or surroundings is an observed or measured object or surroundings that has been objectively analyzed.
[0026] Therefore, the training of machine learning models is carried out using labeled training samples. The labeled samples (quality standard samples) used to train the machine learning model are determined by another labeled sample for which the attributes of the benchmark ground truth objects representing the surrounding environment (environment samples) have been determined and stored for multiple datasets of multiple sensors.
[0027] Because the quality level can be determined from environmental samples using metrics, this quality level can be used as the training target for a machine learning model. Therefore, this involves a monitored training process.
[0028] Multiple datasets from each of a plurality of sensors are fed as input signals to a machine learning model, which may involve, for example, a deep neural network. These sensor datasets can be considered to represent the raw measurement data of all sensors in a multi-mode sensor group, since these datasets have not yet been transformed into descriptions of the object. The output variables of the machine learning model represent an evaluation of the quality of each sensor's representation of its corresponding surrounding environment.
[0029] The advantage of using a dataset of multiple sensors (not yet converted into objects representing the surrounding environment) to train a machine learning model is that it can be trained using the publicly available functional correlations of the generated sensor data, and the metrics used to determine quality can, on the one hand, use the properties of the benchmark ground truth objects representing the surrounding environment as reference variables, and on the other hand, be matched for the corresponding modes of the sensors.
[0030] Because all sensor data in a multi-mode group is processed by the same model, in order to determine the quality criteria for a specific sensor, the model uses information not only from that sensor but also from all other sensors. This increases the reliability of the quality criteria. As a simple example, determining the quality criteria for a lidar system could include images from a camera from which weather conditions could be identified.
[0031] The specific examples below further illustrate the advantages of using multimodal data. In this respect, it is essentially possible to distinguish between two different sets of situations.
[0032] In the first group, it is helpful to include environmental conditions or surrounding environmental conditions that can only be evaluated by other modes in order to evaluate the quality of a determined mode.
[0033] In the second group of cases, either the surrounding environment is irrelevant, or the surrounding environmental conditions cannot be detected using any modality. In this second group of cases, the output of the learned model (e.g., a DNN) is based on a comparison of the data of the modality to be evaluated with data from at least one other modality, because the model learns the feature differences between modalities from the training data.
[0034] Therefore, the first group includes external sensor conditions that negatively impact the determined mode, but are best, or even only, determined by another mode. As mentioned above, an example of this is weather conditions, which can be determined from camera images. A concrete example is that weather conditions such as snow, rain, or fog reduce the effective performance of lidar sensors due to laser scattering. In extreme cases, such as heavy snow, rain, or fog, laser reflection from objects can no longer be observed, making lidar data analysis unsuitable for identifying these sensor deficiencies. Even when laser reflection exists under unfavorable weather conditions for lidar, the measurement quality is negatively affected, and lidar data analysis cannot draw reliable conclusions about data quality because reflection intensity depends on other factors (such as the reflectivity of the object). On the other hand, weather conditions can be well identified using machine learning models based on camera images.
[0035] Another reason for variations in the quality of sensor measurement data is the influence of other objects in the surrounding environment of the object being measured. For example, if the resolution of a radar sensor is insufficient to separate two objects that are less than a certain distance apart, a trained model can identify this by incorporating information from other modalities (such as lidar or cameras) and thus reduce the quality of the corresponding radar data.
[0036] Another concrete example of errors in radar sensors due to the measurement principle is the potential for interference when radar waves of the same or similar frequencies are simultaneously received from radar sensors installed in other vehicles—that is, when radar waves transmitted by sensors installed in one's own vehicle and reflected by an object overlap. However, if the resulting error occurs frequently enough in the training data, it can be learned using machine learning methods. In this case, the identification is based on a comparison of radar measurements with measurements from at least one other modality using a trained model (e.g., a DNN), because inconsistencies between data from different modalities can be observed in this situation.
[0037] For each sensor in the sensor group, this quality level can, for example, have a value between 0 and 1, where a value of 1 corresponds to the sensor’s best performance—that is, there is no sensor deficiency, while a value of 0 is output if the sensor does not output reliable information.
[0038] Since the training of a machine learning model largely depends on the chosen machine learning model, this section will use a deep neural network as an example to illustrate this training process.
[0039] Neural networks provide a framework for many different algorithms used in machine learning, collaboration, and handling complex data inputs. These neural networks learn to perform tasks based on examples, typically without requiring programming with task-specific rules.
[0040] This type of neural network is based on a collection of connected units or nodes called artificial neurons. Each connection enables the transmission of signals from one artificial neuron to another. The receiving artificial neuron can process the received signal and then activate other artificial neurons connected to it.
[0041] In a typical implementation of a neural network, the signals at the connections of an artificial neuron are real numbers, and the output of the artificial neuron is calculated as a nonlinear function of the sum of its inputs. The connections of an artificial neuron typically have weights that correspond to the progress of learning. These weights increase or decrease the strength of the signal at the connection. An artificial neuron may have a threshold such that it outputs a signal only when the total signal exceeds that threshold.
[0042] Typically, multiple artificial neurons are combined in layers. Different layers may perform different types of transformations on their inputs. The signal may propagate from the first layer (input layer) to the last layer (output layer) after multiple traversals of these layers.
[0043] Here, the sensor dataset is fed as an input signal to the deep neural network, and the quality level is compared with the output variable of the neural network as the target variable.
[0044] At the start of training, such a neural network receives, for example, a random initial weight for each connection between two neurons. Input data is then fed to the network, and each neuron weights the input signal with its own weights, continuing to feed the results to neurons in the next layer. The overall result is then provided in the output layer. The magnitude of the error, and the proportion of that error for each neuron, can be calculated, and the weights of each neuron are then changed in the direction that minimizes the error. Further iterations are then performed iteratively, remeasuring the error and rematching the weights until the error falls below a pre-given limit.
[0045] According to one perspective, the machine learning model has a neural network or a recurrent neural network, and the corresponding output variable is a vector having at least one vector element that describes a quality level.
[0046] Machine learning methods can involve models such as deep neural networks (DNNs) with multiple hidden layers.
[0047] Other machine learning models that can be used in this method include linear discriminant analysis (LDA) as a trained method for feature transformation. In addition, support vector machines (SVM), Gaussian mixture models (GMM), or hidden Markov models (HMM) can be used as trained classification methods.
[0048] When using feedforward neural networks with datasets from multi-mode sensors, datasets with a defined time step are typically processed independently of previous datasets (“single-frame processing”).
[0049] By using recurrent neural networks (RNNs), it is also possible to incorporate temporal correlations between sensor datasets. Furthermore, the structure of an RNN can incorporate long short-term memory (LSTM) units to improve information incorporation over longer time periods, allowing that information to better influence the current quality criterion estimation.
[0050] Neural network architectures can be adapted to different formats (e.g., data structure, data volume, and data density) for different sensor modalities. For example, classic grid-based convolutional neural networks (CNNs) can be advantageously used to process sensors whose data at each time step (e.g., camera image data) can be represented as one or more two-dimensional matrices. Architectures such as PointNet or PointNet++ can be advantageously used to process other sensors (e.g., LiDAR) that provide point clouds at each time step; these architectures are point-based and can directly process the raw point cloud as input compared to grid-based methods.
[0051] According to another proposal, the neural network has multiple layers, with datasets from at least two sensors being input independently to different parts of the input layer of the neural network, and the fusion of the datasets from the at least two different sensors being performed in another layer within the neural network.
[0052] Depending on the sensor types used by at least two of the multiple sensors, the sensor datasets can be independently input into different parts of the input layer of the neural network, and then fusion can be performed within the neural network (“intermediate fusion”). Therefore, if the neural network has multiple layers, the fusion of the corresponding datasets from at least two of the multiple sensors can be performed in different layers of the neural network.
[0053] When fusing datasets from at least two different sensors, especially when fusing information from at least two of multiple sensors.
[0054] In particular, by fusing sensor datasets, information from different sensors can be fused before being input into the neural network (“early fusion”).
[0055] Alternatively, this fusion can also be achieved by having multiple neural networks process the data independently, thus identifying the data independently, and then fusing the network outputs (i.e., object representations) (“post-fusion”).
[0056] In other words, different methods can be considered to fuse data from sensors with different modalities in order to achieve multimodal data recognition. Therefore, the fusion of different sensor modalities can be achieved through early, intermediate, or late fusion. In early fusion, multimodal data from the sensors is typically fused after appropriate transformations and / or data selection (e.g., for calibrating the field of view), before the data is fed into the input layer of the neural network. In late fusion, each modality is processed independently of the others by a DNN suitable for that modality, and each produces the desired output. All outputs are then fused. In both early and late fusion, the fusion occurs outside the DNN architecture; in the case of early fusion, it occurs before processing by the DNN, and in the case of late fusion, it occurs after processing by the DNN.
[0057] Unlike the intermediate fusion, here the input of the multimodal data and the processing in some "lower" (i.e., near the input) layers are initially separate from each other. Then, fusion is performed within the DNN, and the fused and processed data is further processed in other "upper" (i.e. near the output) layers.
[0058] The possible fusion architectures that can be found in the literature are MV3D[1], AVOD[2], Frustum PointNet[3] and Ensemble Proposals[4].
[0059] [1] "Multi-view 3d object detection network for autonomous driving", X. Chen et al., Computer Vision and Pattern Recognition (CVPR), IEEE Conference 2017, IEEE, 2017, pp. 6526–6534.
[0060] [2] "Joint 3d proposal generation and object detection from view aggregation", J.Ku et al., arXiv: 1712.02294, 2017.
[0061] [3] Frustum pointnets for 3D object detection from rgb-d data, CRQi, W. Liu, C. Wu, H. Su and LJ Guibas, Computer Vision and Pattern Recognition (CVPR), IEEE Conference 2018, IEEE, 2018.
[0062] [4] "Robust detection of non-motorized road users using deep learning on optical and lidar data", T. Kim and J. Ghosh, Intelligent Transportation Systems (ITSC), 2016, pp. 271–276.
[0063] In accordance with the standard, the architecture mentioned herein outputs the classification and size estimation of objects in the field of view. For the invention proposed herein, for example, the architecture and operation of fusion using DNN proposed in [1, 2, 3, 4] can be used, where the output layer is matched accordingly, and regression of the quality criterion is computed using one or more (e.g., fully connected) layers.
[0064] According to one aspect, the following operation is proposed: fusing relevant information from at least two of multiple sensors by adding, averaging, or cascading the data from datasets of different sensors and / or the properties of objects representing the environment as determined by the relevant datasets of different sensors.
[0065] This form of fusion of datasets from different sensors can be used in conjunction with the earlier fusion methods described above.
[0066] Especially in the later stages of fusion, the fusion of output datasets from multiple modality-specific deep neural networks (DNNs) can also be achieved through addition, averaging, or cascading.
[0067] By integrating different architectures and operations, one or more neural networks can be matched with the corresponding data structures of different sensor modalities, so as to achieve optimal data processing through neural networks.
[0068] A neural network can be constructed for quality standard regression such that a single output layer neuron is assigned to each scalar output quality value.
[0069] Alternatively, regression can be achieved by representing each scalar quality criterion using multiple neurons in the output layer. For this purpose, a suitable partitioning of the value range of the quality criterion is chosen.
[0070] For example, in the case where the quality standard ranges from 0 to 1, the first partition among the ten partitions can correspond to the range from 0 to 0.1, the second partition can correspond to the range from 0.1 to 0.2, and so on. Then, the output quality for the regression of each individual scalar quality standard g is realized by one output neuron for each partition [a; b]. Here, if the belonging partition is completely lower than the quality value to be represented (b ≤ g), the expected output value of the corresponding neuron is equal to 1. If the partition belonging to the neuron is completely higher than the quality standard (g ≤ a), the expected output is equal to 0. If the quality standard lies within the partition of the neuron (a < g < b), the value ((g - a) / (b - a)) should be adopted.
[0071] According to another aspect, at least one of the plurality of sensors generates the following data sets: These data sets have a time series of the data of the at least one sensor.
[0072] Since a typical sensor can provide a new value, for example, every 10 ms, and the data sequence can also be processed in the corresponding structure of the neural network, the time correlation can affect the determination of the quality of the corresponding sensor. For example, every 100 consecutive data sets can be provided as an input signal for the neural network, respectively.
[0073] According to one aspect, at least two of the plurality of sensors have different modalities, and the metric for determining the quality level depends on the corresponding modalities of the sensors.
[0074] By having the metric for determining the quality level depend on the corresponding modalities of the sensors, it is possible to achieve a specific coordination of the corresponding determination of the quality with the corresponding sensor types. This results in a higher accuracy in the description of the quality level.
[0075] According to another aspect, at least one variable determined by means of the corresponding data sets is compared with at least one attribute data of at least one belonging reference truth object of the surrounding environment representation by means of the metric.
[0076] To calculate the quality, the corresponding metric can particularly determine the distance. Here, according to the metric, a smaller distance should result in a better quality level, or vice versa. To determine such a metric, attributes specific to the sensor modality should be considered.
[0077] In one step of the comparison, at least one variable of the corresponding data set of the sensor is associated with at least one reference truth object of the surrounding environment representation belonging thereto.
[0078] In another step of the comparison, quantitative consistency between at least one variable and at least one attribute of a benchmark truth object representing the surrounding environment is determined using a metric that depends on the sensor modality, in order to determine the quality level of the corresponding dataset.
[0079] Here, the association between at least one variable of the corresponding dataset of the sensor and at least one benchmark truth object of the associated environment is established using other metrics (e.g., a distance measurement of the variables in the corresponding dataset of the benchmark truth object). In other words, this means that a specific variable from the measured dataset is associated with an object whose metric (e.g., a distance metric) is minimized. Additionally, a threshold is defined for the metric, according to which the value must be below or above for the association to be valid. Otherwise, the measurement will be discarded as noise (“clutter”).
[0080] With appropriate normalization, for example, taking into account the distribution of the spacing criteria defined modally on the sample (where the distribution also depends on the associated threshold, for example), a quality criterion that is typically between 0 and 1 can be derived.
[0081] Typically, the measurement points of a LiDAR sensor's point cloud are located on the surface of an object. Therefore, in this case, it is meaningful to define the shortest Euclidean distance from the 3D measurement points to the surface of the 3D object model (e.g., the 3D bounding box) as a distance metric associated with each LiDAR measurement point for each object. The distance metric for the object can then be calculated overall as the average interval of all associated points.
[0082] When using spacing metrics to calculate the quality standard of radar positions, the process is essentially similar to that of lidar. However, it should be noted in this case that radar positions are typically not measured on the surface of the object. Therefore, possible spacing metrics can be defined by first estimating or modeling the distribution of radar positions on the object—ideally depending on relevant parameters (e.g., the position of the object being measured relative to the sensor, the spacing to the sensor, and the type of object). This distribution estimation can be performed based on a suitable sample.
[0083] Since radar position measurements also include velocity estimates compared to lidar point measurements, these measured attributes can be incorporated into the calculation of metrics, which involves not only correlation but also the calculation of quality standards.
[0084] Especially in lidar systems (and radar systems), it is meaningful to additionally include the number of points being measured. If the number of measured points is less than expected when considering the distance to the object, the quality standard will deteriorate. A point-count-based metric can be combined with the aforementioned distance metrics.
[0085] According to one aspect, the metrics are based on the object detection rate (e.g., false positive rate and / or false negative rate) and / or combination rate (e.g., mean average precision) and / or F1 metric of at least one of a plurality of sensors.
[0086] If object estimates exist in each modality in addition to the raw measurement data from the sensors, other metrics can be used, such as algorithms for object detection and tracking, to obtain the object estimates from the raw measurement data.
[0087] In this context, determining the association between the modal object and the baseline truth object enables the computation of the following metrics:
[0088] The false positive rate (also known as "fallout") indicates the proportion of objects detected by a sensor that do not correspond to the true reference object. A value of 0 indicates that none of the detected objects are false positives; a value of 1 indicates that none of the detected objects correspond to the true reference object.
[0089] Conversely, the false negative rate (also known as the "miss rate") indicates the proportion of baseline ground truth objects that were not detected by the sensor out of the total number of actually existing objects. If no such object exists, the false negative rate is 0; if the sensor does not detect any actually existing objects, the false negative rate is 1.
[0090] The F1 score is the harmonic mean of precision and recall, so it takes into account not only false positives but also false negatives.
[0091] Precision indicates which part of the objects detected by the sensor are correct (precision = 1 - false positive rate), while recall (also known as sensitivity) indicates the proportion of actual objects that are correctly detected (recall = 1 - false negative rate).
[0092] The mean precision (mAP) also combines precision and recall into a unique value, but additionally, it should be considered that precision and recall are correlated. If recall deteriorates, the object detector can be optimized for precision, or vice versa. mAP can only be determined if confidence values exist for each object identified by the detector. In this case, if the confidence values change, the average precision is calculated for all recall values from 0 to 1.
[0093] One approach proposes determining the association by using the probability that the dataset belongs to a certain object in the surrounding environment.
[0094] A method is proposed for determining the corresponding quality level of each dataset of at least two of a plurality of sensors using a machine learning model trained according to the above method, wherein the plurality of sensors are configured to generate a representation of the surrounding environment; the datasets of the at least two sensors are provided as input signals to the trained machine learning model; and the machine learning model provides output variables for outputting the quality level.
[0095] Here, the term “dataset” includes not only a single data point from a sensor (e.g., if the sensor outputs only one value), but also a range of values, such as two-dimensional or multi-dimensional (e.g., the range of values of a sensor that records digital images of the surrounding environment), and all other data quantities that represent typical output variables for the corresponding sensor type.
[0096] The output variable of a machine learning model depends on the machine learning model used. For neural networks, the output variable can be, in particular, the quality level.
[0097] With the help of a machine learning model trained according to the method shown, it is possible to provide a quality level for each dataset of one of multiple sensors in a simple way.
[0098] A method is proposed for fusing datasets from a plurality of sensors, wherein the plurality of sensors are configured to generate a representation of the surrounding environment, and wherein, when fusing datasets from a plurality of sensors, the overall result of environment recognition is influenced by datasets weighted by corresponding determined quality levels.
[0099] Fusion can be used not only for the identification of information or datasets but also for the fusion of information or datasets. In this sense, a method for fusing datasets from multiple sensors can be understood as a method for identifying and fusing datasets from multiple sensors.
[0100] As mentioned above, it is advantageous to set current quality standards for each sensor modality as inputs for modules used in sensor data fusion and environment perception (MSFU), in addition to the sensor data itself.
[0101] MSFU can also be implemented using machine learning methods, but it can also be implemented using conventional algorithms (through "manual engineering"), such as Kalman filtering of sensor data. Combining ML-based reliability determination of sensors with a separate module for generating environmental performance has several advantages compared to systems implemented solely using ML methods.
[0102] For example, one advantage is that, in the case of non-ML-based implementations of sensor data fusion and environmental perception, the overall system's requirements for computing power and storage capacity are lower or can be better reduced from the outset, making it easier to implement the entire system on computer hardware suitable for mobile use in robots or vehicles.
[0103] In this context, ML-based reliability determination can operate at a lower update rate and / or with only a portion of multi-mode sensor data, while modules for sensor data fusion and environmental perception can operate at a higher update rate and / or use all available sensor data.
[0104] In this way, on the one hand, high accuracy and low latency in environmental perception can be achieved, and on the other hand, the automated assessment of sensor reliability can take advantage of the advantages of machine learning methods to consider the complex correlations between multi-mode sensor data under multiple conditions, which is impossible or unreliable with non-ML-based methods, or impossible for a similar number of different conditions (e.g., detectable on training samples).
[0105] Implementing the MSFU independently of the module used to determine sensor quality criteria can also have the advantage of enabling a smaller and / or more cost-effective overall system implementation. For example, in applications requiring redundant design due to reliability requirements, redundant design of the MSFU (rather than the module used for quality criterion determination) is sufficient.
[0106] In this scenario, MSFU can be implemented such that it continues processing multi-mode sensor data even if the module used to determine the quality criteria fails, even though the sensor quality criteria no longer exist. This could be achieved, for example, by not using weighting in the fusion process or by using fixed preset parameters. Therefore, the overall reliability of the MSFU output becomes lower, and there is no overall quality. However, if this is properly considered in the functionality, system safety can still be guaranteed, ensuring that the non-redundant design of the modules is sufficient to determine the sensor quality criteria.
[0107] An example of “appropriate consideration in functionality” could be: initiating a handover to a human driver in the event of a failure in a module whose quality standards for sensors are defined, so that information about the environment, which is less reliable, must be used only for a limited time period.
[0108] Modules used for sensor data fusion and environmental perception could, for example, utilize an assessment of the current reliability of each individual sensor output by modules used for quality criteria determination, in order to improve the accuracy and reliability of environmental perception. This increases the safety of autonomous driving systems.
[0109] A concrete example of how reliability assessment can be used is to incorporate weighting into the fusion of sensor data, where data from more reliable sensors at a given time are weighted more heavily and thus have a stronger impact on the overall outcome of environmental identification compared to data from sensors assessed as less reliable.
[0110] Compared to the conventional system described above, this invention achieves the following improvements: it provides additional information for sensor data fusion based on the quality assessment of each sensor. Therefore, it mitigates shortcomings in a context-specific manner and achieves better identification accuracy and reliability under different usage conditions, environmental conditions, and consequently, different sensor conditions.
[0111] One example is that a vehicle's autonomous driving capabilities are matched to an overall assessment, and therefore to the current, estimated accuracy and reliability of environmental perception. A concrete example is that, if overall reliability or accuracy is assessed as poor (e.g., if the quality assessments of all sensors are below a defined threshold), the speed of the autonomous vehicle is reduced for an extended period until overall reliability or accuracy improves again. Another concrete example is requesting handover to a human driver—i.e., terminating autonomous driving—when overall reliability is assessed as no longer sufficient.
[0112] Based on one approach, the quality standards of each sensor are combined into an overall evaluation.
[0113] A comprehensive assessment of the controls for mobile platforms that are at least partially automated can additionally improve the security of the functions.
[0114] A method is proposed in which a control signal for operating at least a partially automated vehicle is provided based on the value of at least one quality level of at least one of a plurality of sensors; and / or a warning signal for warning vehicle occupants is provided based on the value of at least one quality level of at least one of a plurality of sensors.
[0115] The term "based on" should be understood broadly as referring to providing control signals based on the value of at least one quality level. It should be understood that considering the use of the value of at least one quality level for any determination or calculation of the control signal does not preclude considering the use of other input variables in such determination. Similarly, this also applies to warning signals.
[0116] A mobile platform can be understood as a mobile, at least partially automated system and / or a driver assistance system for a vehicle. An example could be a vehicle that is at least partially automated or has a driver assistance system. This means that, in this context, a system that is at least partially automated includes a mobile platform in terms of at least partially automated functionality, but a mobile platform also includes vehicles and other mobile machines that include driver assistance systems. Other examples of mobile platforms could be driver assistance systems with multiple sensors, mobile multi-sensor robots (e.g., robotic vacuum cleaners or lawnmowers), multi-sensor monitoring systems, manufacturing machines, personal assistants, or access control systems. Each of these systems can be fully or partially automated.
[0117] An apparatus is described, configured to perform the method as described above. With such an apparatus, the method can be easily integrated into different systems.
[0118] A computer program is described, comprising instructions that, when executed by a computer, cause the computer to perform one of the methods described above. This computer program enables the use of the described methods in various systems.
[0119] This describes a machine-readable storage medium on which the aforementioned computer program is stored. Attached Figure Description
[0120] refer to Figure 1 and Figure 2 Embodiments of the invention are shown and further described below. The accompanying drawings illustrate:
[0121] Figure 1 This shows a flowchart for training a machine learning model;
[0122] Figure 2 A method for fusing datasets from multiple sensors and manipulating actuators is shown. Detailed Implementation
[0123] Figure 1 A flowchart of a method 100 for training a machine learning model to determine the quality level of a dataset 23 for each of a plurality of sensors 21 is shown. In step S0, a dataset 23 (raw sensor measurement data) is pre-generated by each of the plurality of sensors 21, wherein the plurality of sensors 21 are configured to generate raw measurement data from which a representation of the surrounding environment can be determined.
[0124] Multiple sensors 21 can be, for example, a vehicle equipped with multiple sensors 21 or a fleet of such vehicles.
[0125] In step S1, multiple datasets 23 corresponding to multiple ambient environment representations for each of the multiple multi-mode sensors 21 are provided. That is, multiple datasets 23 have been generated by the multiple multi-mode sensors 21 corresponding to multiple ambient environment representations, and in this step, said multiple datasets are provided to obtain so-called unlabeled samples of the datasets 23 for each of the multiple multi-mode sensors 21.
[0126] In another step S2, attribute data of a benchmark ground truth object representing multiple environmental representations is provided. This can be done, for example, by manual labeling or by processing the dataset 23 of each of the multiple multimode sensors 21 as a whole to obtain so-called labeled samples of the dataset 23 of each of the multiple multimode sensors 21. The attribute data respectively represent symbolic descriptions of the corresponding surrounding environment representations.
[0127] In another step S3, the quality levels 26a, 26b, 26c, and 26d of the corresponding dataset 23 for each of the multiple sensors 21 are determined using a metric corresponding to one of the methods described above. Here, the quality levels 26a, 26b, 26c, and 26d can vary according to the corresponding dataset 23, such that the final quality levels 26a, 26b, 26c, and 26d of each sensor in the time series of dataset 23 are time-dependent, because the quality level characterizes the quality of the sensor data of the corresponding sensor at a specific moment.
[0128] In another step S4, a machine learning model is trained or optimized using multiple datasets 23 for each of the multiple sensors 21 and the respective assigned quality levels 26a, 26b, 26c, 26d for each of the corresponding datasets 23 of each sensor 21.
[0129] Figure 2 The diagram illustrates a process for generating a dataset 23 using multiple multi-mode sensors 21 until the possible functions and controls of an actuator 31 for at least partial automation of a mobile platform are realized, or for partial autonomous driving functions of a driver assistance system in a vehicle, or for fully autonomous driving functions in a vehicle are realized.
[0130] Here, multiple sensors 21 are configured to generate a dataset 23 for representing the surrounding environment. These sensors 21 may have different modes, i.e., different sensor types 21 (e.g., optical cameras, radar systems, lidar systems, or ultrasonic sensors), and may be combined into a multi-mode sensor group.
[0131] These datasets 23 represent the input data of module 25, which uses a machine learning model trained according to the methods described above—for example, a deep neural network (DNN) trained in this way—to determine the current quality level 26a, 26b, 26c, 26d for each of the multiple multi-mode sensors 21 based on the currently existing datasets 23.
[0132] In addition to dataset 23, these quality levels 26a, 26b, 26c, and 26d are also provided as input to module 27 for sensor data fusion and environmental perception (MSFU).
[0133] MSFU module 27 uses these quality levels 26a, 26b, 26c, 26d of the datasets 23 of multiple multi-mode sensors 21 in the scope of identification and fusion, for example by weighting the representation of the surrounding environment determined by the dataset 23 of the corresponding dataset in the fusion scope, so as to ensure the accuracy and / or reliability of the representation of the surrounding environment 28a generated by MSFU module 27 even if there are errors in the dataset 23 (e.g., due to insufficient sensors).
[0134] Additionally, an overall quality 28b can be determined, which describes the reliability of the generated representation of the surrounding environment by taking into account all quality levels 26a, 26b, 26c, 26d of the dataset 23 from multiple multi-mode sensors 21.
[0135] The surrounding environment 28a is represented and the overall quality 28b is provided to the input of module 29 in order to realize the function and control of transmission mechanism 31.
[0136] Compared to conventional methods, module 29 is able to achieve the function with better performance and higher reliability, partly because the reliability of the surrounding environment representation 28a is improved by means of the method shown herein, and partly because the overall quality 28b is additionally provided as supplementary information.
[0137] The overall quality 28b can, for example, be used to design at least partially automated driving functions to match the current quality levels 26a, 26b, 26c, 26d of the dataset 23 of multiple multi-mode sensors 21—that is, to pre-determine a more cautious driving approach (e.g., with a lower available speed) when the overall quality 28b represented by the surrounding environment is poor.
[0138] The overall quality 28b can also be used to determine when autonomous or partially autonomous driving capabilities are no longer available and therefore must be handed over to a human driver, for example, by comparing the overall quality 28b to a threshold.
Claims
1. A method (100) for training a machine learning model to determine the quality level (26a, 26b, 26c, 26d) of at least one dataset (23) for each of a plurality of sensors (21), wherein, The plurality of sensors (21) are configured to generate a representation of the surrounding environment, and the method comprises the following steps: Provide multiple datasets (23) (S1) for each of the multiple sensors (21) and corresponding multiple representations of the surrounding environment. Provide attribute data of the baseline truth objects representing the plurality of surrounding environments (S2); The quality level (26a, 26b, 26c, 26d) of the corresponding dataset (23) for each of the plurality of sensors (21) is determined using a metric (S3), wherein the metric compares at least one variable determined using the corresponding dataset with at least one attribute data of at least one reference truth object to which the surrounding environment belongs; and The machine learning model is trained (S4) using the multiple datasets (23) of each of the multiple sensors (21) and the assigned determined quality levels (26a, 26b, 26c, 26d) to determine the quality level (26a, 26b, 26c, 26d) of at least one dataset (23) of each of the multiple sensors (21).
2. The method (100) according to claim 1, wherein, The machine learning model has a neural network or a recurrent neural network, and the corresponding output variable is a vector having at least one vector element describing the quality level (26a, 26b, 26c, 26d).
3. The method (100) according to claim 2, wherein, The neural network has multiple layers, and the datasets (23) of at least two sensors (21) are independently input into different parts of the input layer of the neural network, and the fusion of the datasets (23) of the at least two different sensors (21) is carried out in another layer within the neural network.
4. The method (100) according to any one of the preceding claims, wherein, The following operation is performed: information from at least two of the plurality of sensors (21) is fused by adding, averaging, or cascading the data sets (23) of the different sensors (21) and / or the attributes of objects representing the environment as determined by the corresponding data sets (23) of the different sensors (21).
5. The method (100) according to any one of claims 1 to 3, wherein, At least one of the plurality of sensors (21) generates the following dataset (23): the dataset has a time series of data from the at least one sensor (21).
6. The method (100) according to any one of claims 1 to 3, wherein, At least two of the plurality of sensors (21) have different modes, and the metric used to determine the quality grades (26a, 26b, 26c, 26d) depends on the corresponding modes of the sensors (21).
7. The method (100) according to any one of claims 1 to 3, wherein, The following comparison is performed using a metric: a comparison is made using at least one variable determined by the corresponding dataset (23) with at least one attribute data of at least one benchmark truth object to which the surrounding environment belongs, the comparison comprising the following steps: Associate at least one variable of the corresponding dataset (23) of the sensor (21) with at least one object representing the surrounding environment of the reference truth; The quantitative consistency of the at least one variable with at least one attribute of the benchmark truth object represented by the surrounding environment is determined by using a metric that depends on the modality of the sensor (21) in order to determine the quality level (26a, 26b, 26c, 26d) of the corresponding dataset (23).
8. The method (100) according to any one of claims 1 to 3, wherein, The metric is based on the object detection rate and / or combination rate of at least one of the plurality of sensors (21), the object detection rate being, for example, the false positive rate and / or the false negative rate, and the combination rate being, for example, the mean average accuracy and / or the F1 criterion.
9. The method (100) according to claim 7, wherein, The association is determined by the probability that the dataset (23) belongs to an object in the surrounding environment.
10. A method for determining corresponding quality grades (26a, 26b, 26c, 26d) for each dataset (23) of at least two of a plurality of sensors (21), said method being performed by means of a machine learning model trained using the method according to any one of claims 1 to 9, wherein, The plurality of sensors (21) are configured to generate a representation of the surrounding environment; the dataset (23) of at least two sensors (21) is provided as an input signal to a trained machine learning model; the machine learning model provides output variables for outputting the quality levels (26a, 26b, 26c, 26d).
11. The method according to claim 10, wherein, To provide control signals for operating at least a partially automated vehicle based on the value of at least one quality level (26a, 26b, 26c, 26d) of at least one of the plurality of sensors (21); and / or to provide warning signals for warning vehicle occupants based on the value of at least one quality level (26a, 26b, 26c, 26d) of at least one of the plurality of sensors (21).
12. A method for fusing a dataset (23) of a plurality of sensors from multiple sensors (21), wherein, The plurality of sensors (21) are configured to generate a representation of the surrounding environment, wherein, in the fusion of the datasets (23) of the plurality of sensors (21), the datasets (23) weighted by means of the quality levels (26a, 26b, 26c, 26d) determined according to claim 10 influence the overall result of the environment identification.
13. An apparatus configured to perform the method according to any one of claims 1 to 12.
14. A computer program product comprising instructions that, when implemented by a computer, cause the computer to perform the method according to any one of claims 1 to 12.
15. A machine-readable storage medium on which a computer program product according to claim 14 is stored.
Citation Information
Patent Citations
Multi-sensor mobile robot slam mapping method and system for complex environment
CN109059927A
Method, device, mobile application device, computer program for machine learning for a vehicle sensor system
DE102017213692A1