Semantic Domain Model for Trigger Event Coverage in ML Data Sets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data sets for training deep neural networks (DNNs) often lack comprehensive coverage of real-world scenarios, leading to biased decisions and vulnerabilities in safety-critical applications, such as autonomous driving and cancer detection, which can result in erroneous outputs and safety failures.
Innovation Solution
A method that utilizes a semantic domain model to evaluate the coverage of trigger events in a data set, identifying and mitigating recurring errors by generating synthetic data to improve coverage, thereby certifying the data set for safety-critical applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing data sets are used for training DNNs, then the training process can proceed with available data, but the data set coverage of real-world scenarios is insufficient leading to biased decisions and safety failures
Solution Approach 1:
The patent applies preliminary action by performing coverage evaluation of trigger events before the DNN is deployed for safety-critical applications. The method evaluates whether the training data set adequately covers potential failure scenarios (trigger events) that could cause erroneous outputs. By conducting this evaluation in advance, the patent identifies gaps in data coverage before they lead to safety failures, allowing for proactive improvement of the data set rather than reactive fixes after deployment.
2Reliability
If data set coverage is increased to include all trigger events, then reliability improves, but the complexity of evaluating and ensuring complete coverage increases
Solution Approach 1:
The patent applies parameter changes by transforming the complex evaluation problem into a more manageable form through parameterization. Specifically, the method represents trigger events and data coverage using structured parameters within a semantic domain model. This parameterization allows for systematic evaluation of coverage by comparing specific parameters (such as presence/absence of trigger event characteristics in the data) rather than requiring complex qualitative assessment of entire data sets, thereby reducing evaluation complexity while maintaining reliability.
3Adaptability or versatility
If synthetic data is generated to improve coverage, then data set completeness improves, but the time and computational resources required for data processing increase
Solution Approach 1:
The patent applies copying by generating synthetic data that replicates the characteristics of missing or underrepresented trigger events in the original data set. Instead of collecting additional real-world data (which would be time-consuming), the method creates artificial copies of desired data patterns through synthetic data generation. These synthetic samples are designed to match the statistical and semantic properties of real trigger events, thereby improving coverage efficiency by producing usable training data much faster than traditional data collection methods.
Data Source
AI summary
A method of evaluating a data set with respect to a coverage of trigger events, which can produce erroneous outputs when processed by a machine learning system. The method includes: providing a semantic domain model as well as a data set;validating the machine learning system on at least a part of the data set, wherein for recurring incorrect outputs of the machine learning system with the same objects, these objects are identified as trigger events; determining a coverage of the trigger events by the data set depending on the semantic domain model.


