Data Completeness Estimation Through Scenario-Class Coverage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for data completeness in autonomous systems rely on subjective criteria, leading to incomplete or excessive data collection, which poses safety risks and inefficiencies.
Innovation Solution
A method for estimating data completeness by dividing data into scenario classes, using probability distributions to determine if observed scenario classes meet predefined criteria, and issuing a command to stop data collection when these criteria are met, ensuring objective and generalized data quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If subjective criteria based on expert knowledge are used to determine data completeness, then data collection can be initiated, but the data may be either overly complete or incomplete leading to excessive effort or safety risks
Solution Approach 1:
The patent transforms the subjective assessment of data completeness into an objective parameter-based evaluation system. It defines specific completeness criteria including scenario class coverage (percentage of total scenario classes observed) and scenario coverage within classes (using probability distributions and confidence intervals). This parameter transformation resolves the contradiction by providing measurable thresholds that eliminate expert subjectivity while maintaining manageable complexity through automated calculation.
Solution Approach 2:
The patent replaces the mechanical/expert judgment system with an automated computational system. Instead of relying on expert knowledge and subjective assessment, the system uses automatic scenario classification, probability distribution analysis, and statistical testing to determine data completeness. This substitution eliminates the reliability-complexity contradiction by providing consistent, objective measurements without the burden of expert involvement.
2Reliability
If data collection continues until all scenarios are observed, then completeness increases, but time and resources are wasted on unobservable or already sufficient scenarios
Solution Approach 1:
The patent applies preliminary action by establishing completeness criteria and probability thresholds before data collection begins. The system pre-defines scenario classes, confidence intervals, and stopping rules based on statistical theory. During data collection, the system continuously monitors whether predefined criteria are met, allowing early termination when sufficiency is achieved. This prevents wasted time on unobservable scenarios while ensuring reliability through pre-planned statistical safeguards.
Solution Approach 2:
The patent implements continuous feedback mechanisms during data collection. The system monitors the observed scenario classes and their distribution, comparing against the predefined completeness criteria. When the feedback indicates that confidence intervals are satisfied and scenario class coverage thresholds are met, the system triggers a stopping command. This feedback loop resolves the contradiction by dynamically adjusting data collection duration based on actual progress toward completeness.
3Adaptability or versatility
If no reference dataset is available for comparison, then data acquisition can start without bias, but existing completeness assessment methods cannot be applied
Solution Approach 1:
The patent applies segmentation by dividing the data collection problem into distinct scenario classes based on the autonomous system's functional requirements and operating conditions. Instead of requiring a reference dataset, the system segments scenarios into meaningful categories (e.g., weather conditions, traffic situations, road types) and establishes completeness criteria for each segment. This segmentation enables precision measurement of completeness without external references, resolving the contradiction between adaptability and measurement precision.
Data Source
Figure 1A~1B
Figure 1C~2
Figure 3
AI summary
A general aspect of the present disclosure relates to a method for estimating data completeness. The method comprises receiving a data set as a result of a plurality of observations from a data acquisition, dividing the data of the data set into one or more scenarios, grouping 130 the one or more scenarios into a plurality of scenario classes, wherein the plurality of scenario classes comprises the previously observed scenario classes of a total set of scenario classes, estimating whether the previously observed scenario classes meet a criterion for scenario class completeness with respect to the total set of scenario classes, determining, based on one or more probability distributions, whether the one or more scenarios assigned to a scenario class meet a criterion for scenario completeness of the corresponding scenario class,and issuing a command to stop data collection when the scenario class completeness criterion and the scenario completeness criterion are met.