Method for estimating data completeness
The method addresses the lack of objective criteria in data completeness by partitioning datasets into scenario types and using probability distributions to ensure complete data collection, enhancing safety and reducing costs in autonomous systems.
Patent Information
- Application Number
- JP2025067215
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-17
- Filing Date
- 2025-04-16
- Publication Date
- 2025-11-05
AI Technical Summary
Existing methods for determining data completeness in autonomous systems lack objective criteria, leading to incomplete data collection and potential safety risks, with subjective expert knowledge being the primary means of assessment.
A method for estimating data completeness by partitioning datasets into scenario types, applying occurrence probability distributions to determine scenario type and scenario completeness criteria, and terminating data detection when these criteria are met.
Ensures objective and generalizable data completeness, reducing the risk of incomplete data sets in autonomous systems, thereby enhancing safety and reducing unnecessary data collection costs.
Smart Images

Figure 2025165895000002 
Figure 2025165895000003 
Figure 2025165895000004
Abstract
Description
[Background technology]
[0001] For example, in the robotics and automotive industries, there is an increasing focus on highly automated or autonomous systems. A particular challenge in developing such systems is that autonomous systems must function safely and reliably under a wide variety of different circumstances. It is estimated that in modern autonomous pilot systems, beyond automation level 3, maintaining the system can become so complex that validation accounts for 80% of the resource costs, while development accounts for only 20%. To adequately identify the requirements for the system and its components and enable subsequent testing, it is necessary to capture as many representative scenarios (situations) in which the autonomous system must operate. For this purpose, datasets can be used to verify and validate the autonomous system. For example, datasets can also serve to train and / or validate machine learning models.
[0002] As the automation level (SAE level) increases, the set of scenarios to be considered becomes very large and, as a consequence of the open context, these scenarios may change over time. In particular, data is captured and automatically evaluated to ensure that all relevant scenarios are taken into account.
[0003] In the context of data detection systems (e.g., continuous data), the question must be answered as to when the detected data is complete so that data detection can be terminated. Typically, in the prior art, completeness is measured against a reference derived from an already existing dataset or equivalent model. However, existing methods do not provide a measure of the required degree of similarity. Furthermore, problems arise in applying such methods when a reference dataset is not available. In order to prove the safety of autonomous systems with a reference dataset, this dataset must take into account the characteristics of each system that could potentially lead to unsafe situations.
[0004] For this reason, in the prior art, subjective criteria in the form of expert knowledge are mainly applied in such systems. As a result of this subjectivity, the detected data may be incomplete or incomplete, which may lead to excessive costs on the one hand and unpredictable safety risks on the other hand. Specifically, this means that further data may be collected even though all the data is already available, or that risks may be overlooked and not considered. Thus, there is a need for a method for determining the completeness of data. Summary of the Invention
[0005] A first general aspect of the present disclosure relates to a method for estimating data completeness, including receiving a dataset resulting from a plurality of observations from data detection, partitioning data of the dataset into one or more scenarios, grouping the one or more scenarios into a plurality of scenario types, the plurality of scenario types including previously observed scenario types of a universe of scenario types, estimating whether previously observed scenario types satisfy a criterion for scenario type completeness in terms of the universe of scenario types, determining whether one or more scenarios assigned to a scenario type satisfy a criterion for scenario completeness of a corresponding scenario type based on one or more occurrence probability distributions, and outputting a command to terminate data detection if the criterion for scenario type completeness and the criterion for scenario completeness are satisfied.
[0006] A second general aspect of the present disclosure relates to a computer system designed to implement the method for data integrity estimation according to the first general aspect (or an embodiment thereof).
[0007] A third general aspect of the present disclosure relates to a computer program designed to implement the method for data integrity estimation according to the first general aspect (or an embodiment thereof).
[0008] A fourth general aspect of the present disclosure relates to a computer-readable medium or signal storing and / or embodying a computer program according to the third general aspect (or an embodiment thereof).
[0009] The method proposed in the present disclosure according to a first general aspect (or an embodiment thereof) serves to provide a method for estimating the completeness of data using specific criteria. One advantage of the technique of the present disclosure can be seen in that the criteria for the completeness of a dataset are objective and generalizable. By way of example, it can prevent too little data from being collected or a particular scenario type or scenarios within a scenario type from being unobserved.
[0010] In this way, risks that would arise if an autonomous system were tested, validated, or verified with an incomplete data set can be reduced, thereby, for example, increasing safety for vehicle occupants of an autonomous vehicle, for example, increasing safety for humans interacting with a robot, and, for example, preventing the collection of too much or too long data, for example, preventing the collection of data in an attempt to observe scenarios that are completely unobservable in the autonomous system's operating environment, thereby providing cost benefits for data discovery.
[0011] The disclosed technique may further have the advantage of determining a stopping criterion for discovering scenarios on the one hand, and for detecting data for the scenarios on the other hand. That is, the advantage lies in being able to improve data quality on a scenario-by-scenario-type or scenario-by-scenario basis. Furthermore, the method may serve to determine an estimate of the probability distribution of the data and to determine the associated inaccuracy. This may be advantageous, for example, for improving the data detection process.
[0012] Yet another advantage is that the detected data assessed as complete by the method can serve to test, validate, and / or verify vehicle, robot, building automation, power tool automation, and / or household appliance automation functions. Illustratively, the data can serve to train, test, and / or validate machine learning models. Yet another advantage is that machine learning models can be utilized to control and / or regulate vehicle, robot, building automation, power tool automation, and / or household appliance automation functions.
[0013] In this disclosure, several terms are used as follows:
[0014] A "scenario" may, for example, generally represent a grouping or classification of specific situations, conditions, or environments to which an autonomous system or technology should or can react. Scenarios may be used, for example, to structure and organize a large number of potential usage scenarios so that they can be systematically analyzed, developed, or tested. Examples of scenarios are driving on the autobahn at 130 km / h, driving in a 30 km zone, or driving on a country road during heavy snowfall.
[0015] Scenario types allow scenarios to be grouped based on criteria. Since each system may have different processing stages, there can be a wide variety of criteria for defining a scenario type. Two examples: Accordingly, the criterion "brake activation" would result in all scenarios requiring the host vehicle to brake being placed into one category. Alternatively, the criterion "obstacle detection" would result in all scenarios involving pedestrians, vehicles, or objects in the path of the host vehicle being aggregated into one category. A scenario category can contain various scenarios that can be similarly described by characteristics or parameters. For example, within a scenario category, a first parameter value of a parameter can define a first scenario, and a second parameter value of the same parameter can define a second scenario.
[0016] A "vehicle" may be any device that carries passengers and / or cargo. A vehicle may be a motor vehicle (e.g., a car or truck) or a tracked vehicle. A vehicle may be a motorized two-wheeled or three-wheeled vehicle. However, water-traveling and flying devices may also be vehicles. A vehicle may be operated or assisted at least partially autonomously. [Brief explanation of the drawings]
[0017] [Figure 1-A] 1 illustrates a schematic diagram of an exemplary method for estimating data completeness. [Figure 1-B] 1 illustrates a schematic diagram of an exemplary method for estimating data completeness. [Figure 1-C] 1 illustrates a schematic diagram of an exemplary method for estimating data completeness. [Figure 2] 1 illustrates a schematic diagram of an exemplary architecture for implementing a method for data completeness estimation. [Figure 3] 1 shows a schematic overview of multiple scenario types and the universe of scenario types. DETAILED DESCRIPTION OF THE INVENTION
[0018] 1A to 1C are flowcharts illustrating possible steps of a method 100 for estimating data completeness.
[0019] A method 100 for estimating data completeness includes receiving 110 a dataset 10 resulting from multiple observations from data detection, partitioning 120 data from the dataset 10 into one or more scenarios, and grouping 130 the one or more scenarios into multiple scenario types 20a, where the multiple scenario types 20a include previously observed scenario types from the universe of scenario types 20. The method 100 further includes estimating 140 whether previously observed scenario types 20a satisfy a criterion for scenario type completeness from the universe of scenario types 20, determining 150 whether one or more scenarios assigned to a scenario type satisfy the criterion for scenario completeness of the corresponding scenario type based on one or more occurrence probability distributions, and outputting a command to terminate data detection if the criterion for scenario type completeness and the criterion for scenario completeness are satisfied 160. Illustratively, the method 100 can include integrating the criterion for scenario type completeness and the criterion for scenario completeness into an overall criterion for data completeness. Illustratively, issuing a command to terminate data detection 160 may be performed if the overall criteria are met.
[0020] Illustratively, the data in the dataset 10 can be assigned to scenario types based on a discrete model 30, e.g., as shown in FIG. 2 . For example, the discrete model can include a histogram, a Markov chain, a support vector machine (SVM), a k-nearest neighbor (k-NN), or a logistic regression. Illustratively, segmenting 120 the data in the dataset 10 includes forming a histogram reflecting scenario types 20a. Illustratively, the universe of scenario types 20 can include all, or nearly all, of the scenario types that could potentially occur during operation of the autonomous system. FIG. 3 shows a schematic diagram of multiple scenario types 20a and the universe of scenario types 20. Illustratively, the universe of scenario types 20 can include scenario types that could realistically occur, i.e., with a certain probability. Illustratively, scenario types 20a that have already been observed can form a first subset of the universe of scenario types 20, and scenario types 20b that have not yet been observed can form a second subset of the universe of scenario types 20. Illustratively, determining 150 the type of scenario being observed and / or the completeness of the scenario can include determining the probability of occurrence of previously unobserved data. Illustratively, the probability of occurrence can be determined by an occurrence probability distribution. Illustratively, such probability of occurrence is estimated by an estimation method.
[0021] In one embodiment, the criteria for scenario type completeness may be based on the probability τ that each scenario type in the universe of scenario types 20 has been observed at least once. Illustratively, the criteria for scenario type completeness may include that the probability that each scenario type in the universe of scenario types 20 has been observed once is greater than a threshold. Illustratively, the probability that each scenario type in the universe of scenario types 20 has been observed at least once is determined based on a solution to the "Coupon Collector's Problem."
[0022] For example, a random variable X can be defined that represents the number of observations required before each scenario type in the universe of scenario types 20 is observed at least once. By way of further example, the probability P(X≦B) that each scenario type in the universe of scenario types 20 is observed at least once after B observations can be determined. For example, for the probability that each scenario type in the universe of scenario types 20 is observed once,
number
[0023] For application of these illustrative methods to estimate the observation time until a previously unknown scenario type is observed, the discrete model underlying the analytical or numerical method is correspondingly extended so that the previously unknown scenario type is modeled with some degree of probability, for example, by expert estimation or through more advanced estimation methods (e.g., nonlinear regression).
[0024] To address the question of how many scenario types to extend a discrete model with, the total number of scenario types can be estimated. For example, the estimation of the number of scenario types can be performed using discrete estimation. For example, Chao's method (Chao, Anne, and Shen-Ming Lee, "Estimating the number of classes via sample coverage," Journal of the American Statistical Association, Vol. 87, No. 417 (1992): pp. 210-217) can be applied, which uses sampling selection to estimate all scenario types. To do this, the number and frequency of observed scenario types are modeled. The ratio of scenario types observed once to those observed twice can be used, for example, to estimate scenario types that have not yet been observed. For example, the appropriateness of the number of scenario types in the entire set can be determined based on the total number of scenario types.
[0025] Illustratively, each scenario type of the plurality of scenario types 20a may include one or more parameters. Illustratively, the received data set may include parameter values for one or more parameters. Illustratively, a scenario of a scenario type may be characterized by at least one parameter value. For example, within a scenario type, a first parameter value of a parameter may define a first scenario, and a second parameter value of the same parameter may define a second scenario.
[0026] Illustratively, grouping 130 one or more scenarios into multiple scenario types 20a may be performed based on common indicators of one or more scenarios. For example, the indicators may include criteria such as those described above. For example, the indicators may include a "Brake Activation" criterion and / or an "Obstacle Detection" criterion.
[0027] Illustratively, each scenario type of the plurality of scenario types 20a may have one or more characteristics. Illustratively, the dataset 10 may include attributes of one or more characteristics. Illustratively, a scenario of a scenario type may be characterized by at least one attribute. Illustratively, the one or more characteristics may be derived from a common indicator.
[0028] Illustratively, scenarios assigned to a scenario type may occur with a probability distribution within that scenario type, where the probability of occurrence may depend on parameter values of one or more parameters or on values of one or more attributes.
[0029] The method 100 may include determining one or more occurrence probability distributions for the attribute and / or parameter values. Illustratively, the attribute or parameter values may be described by one or more occurrence probability distributions. Illustratively, criteria for the occurrence probability distributions may be optionally defined. Illustratively, the criteria may be evaluated depending on the application.
[0030] Illustratively, determining 150 can be based on a probability of occurrence of at least one parameter value and / or attribute. Illustratively, the probability of occurrence can be determined by an occurrence probability distribution of the corresponding parameter and / or characteristic. Illustratively, method 100 can include aggregating one or more occurrence probability distributions of all parameters and / or characteristics of the scenario into a multi-dimensional occurrence probability distribution.
[0031] Illustratively, estimating 140 can be performed by continuous estimation 50. Illustratively, the occurrence probability distribution can be determined based on previously received parameter values and / or attributes. Illustratively, the occurrence probability distribution can provide information about the occurrence probability of previously unobserved parameter values and / or attributes. Illustratively, the occurrence probability of received data can include or represent the occurrence probability of parameter values and / or attributes. For example, a scenario type can define relevant parameters or characteristics of the scenario. Illustratively, the method can include selecting relevant parameters or characteristics from the dataset 10 for scenarios of the scenario type. For illustrative purposes, the scenario "Following journey without critical situations" can be described. For this scenario type, the scenario can be described by relevant parameters such as speed, relative speed to the leading vehicle, and / or distance to the leading vehicle. Illustratively, a space of possible parameter values for each parameter can be defined. For example, the space of possible parameter values for the parameter "distance" can include 0.5 m, 0.3 m, 0.2 m, 1.2 m, etc. Illustratively, the space of possible parameter values may be subdivided into discrete portions. For example, the discretization may be based on the physical resolution limit of the detection system utilized for data detection. For example, for the parameter "relative distance," a resolution of 10 cm may be assumed if this corresponds to the resolution limit of radar. Illustratively, the method 100 may include determining, for each scenario type of the plurality of scenario types 20a, one or more occurrence probability distributions of parameter values.
[0032] In one embodiment, the criteria for scenario completeness may include the occurrence probability of at least one attribute and / or parameter value exceeding a threshold. One or more occurrence probability distributions for parameter values and / or attributes may determine the occurrence probability of previously unobserved data. Illustratively, the occurrence probability of previously unobserved data, parameter values, or attributes may be lower than a threshold within a scenario type. Illustratively, it may be concluded from this that the corresponding scenario type can be classified as fully observed.
[0033] By way of example, the criteria for scenario completeness may be met if, for each of the scenario types 20a that have already been observed, the probability of occurrence for previously unobserved data is lower than a threshold. As explained above, relevant parameters and / or characteristics may be determined for a scenario type. By way of example, the criteria for scenario completeness may be met if, for each of the scenario types 20a that have already been observed, the probability of occurrence for previously unobserved data or parameter values for relevant parameters or relevant characteristic attributes is lower than a threshold. This may be desirable to define a stopping criterion so that data detection does not continue even though it is no longer likely that the missing parameter values and / or attributes will still be observed.
[0034] By way of example, a criterion for scenario completeness may be met if a weighted function of the probability of occurrence of a scenario in a scenario type is below a threshold value.
[0035] In one embodiment, the one or more occurrence probability distributions may be determined by a non-parametric method, which may include kernel density estimation.
[0036] In one embodiment, the method 100 can include calculating a difference between one or more occurrence probability distributions for previously observed data and one or more newly modeled occurrence probability distributions based on the newly received data. Illustratively, the criteria for scenario completeness can include a difference greater than a threshold. For example, calculating the difference can be based on at least one of the Kolmogorov-Smirnov distance (discrete form), the Kullback-Leibler information (discrete form), the Anderson-Darling test, or a chi-square test.
[0037] Illustratively, the occurrence probability distributions may be determined at different times during data detection, each of which is different from the others. For example, in the case of kernel density estimation, a narrow bandwidth may be applied that is below the physical resolution limit for the parameter in question. Illustratively, the method 100 may include comparing the determined occurrence probability distributions. Illustratively, a criterion for scenario completeness may include a difference between the determined occurrence probability distributions exceeding a threshold.
[0038] In one embodiment, if the criteria for scenario type completeness and the criteria for scenario completeness are not met, method 100 can continue. Illustratively, method 100 can further include determining 152 a separation between the occurrence probability distribution of the newly received data and the determined one or more occurrence probability distributions to determine a first validity of the determined one or more occurrence probability distributions. Illustratively, the criteria for data completeness can be based on the first validity. Illustratively, validity can be met if the separation between the occurrence probability distribution of the newly received data and one or more occurrence probability distributions of parameter values of one or more parameters and / or attributes of one or more characteristics is less than a separation threshold. Illustratively, the criteria for scenario completeness can be met if the occurrence probability for previously unobserved data within a scenario type is less than a threshold and the separation between the occurrence probability distribution of the newly received data and one or more occurrence probability distributions is less than a separation threshold. Illustratively, if validity is not given, another approach for determining one or more occurrence probability distributions can be applied. For example, if no validity is given, kernel density estimation can be performed using a different bandwidth.
[0039] In one embodiment, the method 100 may include performing 151 a hypothesis test utilizing the dataset 10 to determine a second plausibility of the determined one or more occurrence probability distributions. Illustratively, a criterion for scenario completeness may be based on the second plausibility. Illustratively, the hypothesis test may represent a fit of the data to the determined one or more occurrence probability distributions.
[0040] By way of example, the method can be applied to combine the primary plausibility, secondary plausibility, and numerical plausibility and therefrom determine an overall plausibility for the data set. By way of example, a data set can be assessed as complete given the plausibility for the estimates of unobserved scenario types and the plausibility for the scenarios of all observed scenario types.
[0041] In some embodiments, the method can include utilizing the dataset for validation and / or verification of a vehicle function, a robot function, a building automation function, a power tool automation function, and / or a household appliance automation function. For example, the method can provide criteria for classifying a dataset 10 on which validation and / or verification has been performed as complete.
[0042] In one embodiment, method 100 can include utilizing dataset 10 for training, testing, and / or validating a machine learning model. Illustratively, the method can include utilizing the machine learning model to control and / or regulate a vehicle function, a robot function, a building automation function, a power tool automation function, and / or a household appliance automation function. Illustratively, method 100 can include importing the machine learning model into a computer system of a vehicle, a robot, a building, a power tool, a household appliance, a machine tool, a personal assistant, an access control system, and / or a medical appliance.
[0043] Illustratively, the robot may include a vehicle (either autonomously or assisted driving). For example, the vehicle function may be a function for autonomous and / or assisted driving. In some examples, the machine learning model and / or the method 100 for performing data integrity estimation may be designed to run on a computer system of a vehicle (e.g., an autonomous, highly automated, or assisted driving vehicle), i.e., may be implemented, for example, on a computer. For example, the computer system may be implemented locally on the vehicle or (at least partially) on a backend communicatively coupled to the vehicle. For example, the computer system may include a controller capable of executing the method 100 and / or the machine learning model. In some examples, the vehicle may include a computer system having a communication interface that allows communication with the backend. For example, the method 100 for performing data integrity estimation and / or the machine learning model may be executed on such a backend. Illustratively, the controller function of the vehicle may include the method 100 for performing integrity estimation and / or the machine learning model or may have access to it when executed remotely, for example, on a cloud. Illustratively, the dataset 10 may validate and / or verify the vehicle's controller function.
[0044] In another example, as alluded to above, the method 100 for estimating data integrity and / or the machine learning model may be designed to run on a robot. Illustratively, the dataset 10 may be designed for validation and / or verification of a robot function. Illustratively, the machine learning model may be designed for controlling and / or monitoring a robot function (particularly for controlling and / or monitoring a motor function of the robot). In some examples, the method 100 and / or the machine learning model may run on a computer system of the robot. For example, the computer system may be implemented locally on the robot or may be implemented (at least in part) in a backend communicatively connected to the robot.
[0045] Illustratively, the data integrity estimation method 100 and / or the machine learning model may be designed to run in a building. Illustratively, the dataset 10 serves to validate and / or verify a building function (particularly for controlling a building automation function). Illustratively, the machine learning model serves to control and / or monitor a building function (particularly for controlling a building automation function). For example, the building function may be a function for adjusting room temperature, humidity, and / or safety equipment. In some examples, the method 100 may be designed to run the machine learning model on a computer system inside the building. For example, the computer system may be implemented locally to the building or may be implemented (at least partially) in a backend communicatively connected to the building. For example, the computer system may include a control system or building automation controller, on which the data integrity estimation method 100 and / or the machine learning model may be executed. Illustratively, the building may be equipped with a computer system having a communication interface that allows communication with an external backend. For example, the data integrity estimation method 100 and / or the machine learning model may be executed on such a backend. By way of example, the input data and / or operational data for the machine learning model can be based on information such as room temperature, lighting, or the presence of people. In some cases, the input data and / or operational data for the machine learning model can include relative temperature differences, lighting intensity, or distance to specific locations or objects within the building. By way of example, the information can originate from a network, such as sensor data or settings of other buildings or building components. Such information can be provided by communication between buildings or building parts or via an external backend.
[0046] In another example, the method 100 and / or the machine learning model may be designed to run on a power tool. Illustratively, the dataset 10 may be designed to validate and / or verify power tool functionality (particularly the work functions of the power tool). Illustratively, the machine learning model may be designed to control and / or monitor power tool functionality (particularly the work functions of the power tool). In some examples, the method 100 and / or the machine learning model may run on a computer system of the power tool. For example, the computer system may be implemented locally on the power tool or may be implemented (at least in part) in a backend communicatively connected to the power tool.
[0047] In another example, the method 100 for estimating data integrity and / or the machine learning model may be designed to run on a household appliance. Illustratively, the dataset 10 may be designed for validating and / or verifying a household appliance function (particularly for controlling and / or monitoring a working function of the household appliance). Illustratively, the machine learning model may be designed for controlling and / or monitoring a household appliance function (particularly for controlling and / or monitoring a working function of the household appliance). In some examples, the method 100 and / or the machine learning model may run on a computer system of the household appliance. For example, the computer system may be implemented locally to the household appliance or may be implemented (at least in part) in a backend communicatively connected to the household appliance.
[0048] In another example, the dataset 10 may serve to validate and / or verify the functionality of a machine tool, a personal assistant, an access control system, and / or a medical device. By way of example, the machine learning model may be designed to run on a machine tool, a personal assistant, an access control system, and / or a medical device, e.g., for medical imaging, and / or may be accessible over a network.
[0049] By way of example, the dataset 10 may include audio data and / or video data. By way of example, data detection may be performed during vehicle driving operations. By way of example, the dataset 10 may include sensor signals, in particular sensor signals from sensors for measuring speed, distance, angular velocity, current intensity, voltage, temperature, pressure, air pressure, weight, deformation, liquid or gas flow, or gas composition. By way of example, the dataset 10 serves for object recognition. By way of example, the dataset, or the sensor data contained therein, may be subjected to semantic segmentation, for example, with respect to traffic signs, road surfaces, pedestrians, vehicles, electric wires, water pipes, gas pipes, biochemical reactions, or natural obstacles, for example in route planning.
[0050] In some embodiments, one or more method steps of method 100 may be computer-implemented.
[0051] Further, a computer system designed to perform the method 100 for estimating data integrity is disclosed. The computer system may include at least one processor and / or at least one working memory. The computer system may further include a (non-volatile) memory. Illustratively, all steps of the method 100 may be performed by the computer system. In some examples, individual steps of the method 100 may be performed by the computer system. Optionally, results of individual method steps that are not performed by the computer system may be received by the computer system.
[0052] Further disclosed is a computer program designed to perform the data integrity estimation method 100. The computer program may exist, for example, in an interpretable or compiled form. The computer program may, for example, be (partially) loaded into a computer's RAM for execution as a bit or byte sequence.
[0053] Further disclosed is a computer readable medium or signal storing and / or containing the computer program or at least a part thereof. The medium may include for example RAM, ROM, EPROM, HDD, SDD, ... on which the signal is stored.
Claims
1. A method (100) for estimating the completeness of data, comprising: The method comprises: receiving (110) a dataset (10) resulting from a plurality of observations from data detection; Partitioning (120) the data of the dataset (10) into one or more scenarios; grouping (130) the one or more scenarios into a plurality of scenario types (20a), the plurality of scenario types (20a) including previously observed scenario types of the universe of scenario types (20); estimating (140) whether the previously observed scenario types (20a) satisfy criteria for scenario type completeness in terms of the universe of scenario types (20); determining (150) whether the one or more scenarios assigned to a scenario type satisfy criteria for scenario completeness of the corresponding scenario type based on one or more occurrence probability distributions; outputting (160) a command to terminate the data detection if the criteria for the scenario type completeness and the criteria for the scenario completeness are met.
2. 2. The method (100) of claim 1, wherein the criterion for the scenario type completeness is based on the probability that each scenario type in a universe of scenario types (20) has been observed at least once.
3. 3. The method (100) of claim 1 or 2, wherein the estimating (140) includes estimating (141) the number of observations still required until each scenario type in the universal set of scenario types (20) has been observed at least once with at least some probability.
4. The method (100) comprises:
10. The method (100) of any one of the preceding claims, further comprising estimating a total number of scenario types (20a) based on the plurality of scenario types (20a) to determine the numerical relevance.
5. 10. The method (100) of claim 1, wherein the grouping (130) of the one or more scenarios into a plurality of scenario types (20a) is performed based on common indicators of the one or more scenarios.
6. 10. The method (100) of claim 1, wherein each scenario type of the plurality of scenario types (20a) comprises one or more parameters, the dataset (10) comprises parameter values for the one or more parameters, and a scenario of a scenario type is characterized by at least one parameter value.
7. 10. A method (100) according to any one of the preceding claims, wherein each scenario type of the plurality of scenario types (20a) has one or more characteristics, the dataset (10) includes attributes of the one or more characteristics, and scenarios of a scenario type are characterized by at least one attribute.
8. 8. The method (100) according to claim 6 or 7, wherein the one or more occurrence probabilities for the attribute or parameter values comprise determined occurrence probabilities, and optionally criteria for the occurrence probability distribution are defined and evaluated depending on the application case.
9. The method (100) of claim 8, wherein the criteria for the scenario completeness include a probability of occurrence for at least one attribute and / or parameter value exceeding a threshold.
10. If the criteria for the scenario type completeness and the criteria for the scenario completeness are not met, the method continues, the method comprising:
10. The method (100) of claim 8 or 9, further comprising: determining (152) a degree of separation between an occurrence probability distribution of newly received data and the determined occurrence probability distribution(s) to determine a first validity of the determined occurrence probability distribution(s), wherein the criterion for scenario completeness is further based on the first validity.
11. The method (100) comprises:
10. The method (100) of any one of the preceding claims, further comprising: performing a hypothesis test (151) using the dataset to determine a second plausibility of the determined one or more occurrence probability distributions, wherein the criterion for scenario completeness is further based on the second plausibility.
12. The method (100) comprises: aggregating the first validity, the second validity, and the numerical validity to determine an overall validity for the data set; and / or 10. The method (100) of any one of the preceding claims, further comprising integrating the criteria for the scenario completeness and the criteria for the scenario type completeness to determine an overall completeness for the data set.
13. The method (100) comprises:
10. The method (100) of any one of the preceding claims, comprising utilizing the dataset (10) for validation and / or verification of a vehicle function, a robot function, a building automation function, a power tool automation function, and / or a household appliance automation function.
14. The method (100) comprises: using the dataset (10) for training, testing, and / or validating a machine learning model; and 10. The method (100) of any one of the preceding claims, comprising utilizing the machine learning model to control and / or regulate a vehicle function, a robot function, a building automation function, a power tool automation function, and / or a household appliance automation function.
15. A computer system designed to implement a method (100) for estimating data integrity according to any one of the preceding claims 1 to 14.
16. 15. A computer program comprising instructions which, when executed by a computer system, direct the computer program to perform a method (100) for estimating data integrity according to any one of the preceding claims 1 to 14.
17. 17. A computer readable medium or signal storing and / or embodying a computer program according to claim 16.