Semantic Domain Model for Trigger Event Coverage in ML Data Sets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data sets for training deep neural networks (DNNs) often lack comprehensive coverage of real-world scenarios, leading to biased decisions and vulnerabilities in safety-critical applications, such as autonomous driving and cancer detection, which can result in erroneous outputs and safety failures.

Innovation Solution

A method that utilizes a semantic domain model to evaluate the coverage of trigger events in a data set, identifying and mitigating recurring errors by generating synthetic data to improve coverage, thereby certifying the data set for safety-critical applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing data sets are used for training DNNs, then the training process can proceed with available data, but the data set coverage of real-world scenarios is insufficient leading to biased decisions and safety failures

Engineering Contradiction:
Improvesafety of DNN in critical applicationsVSAvoidcoverage of real-world scenarios
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent applies preliminary action by performing coverage evaluation of trigger events before the DNN is deployed for safety-critical applications. The method evaluates whether the training data set adequately covers potential failure scenarios (trigger events) that could cause erroneous outputs. By conducting this evaluation in advance, the patent identifies gaps in data coverage before they lead to safety failures, allowing for proactive improvement of the data set rather than reactive fixes after deployment.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If data set coverage is increased to include all trigger events, then reliability improves, but the complexity of evaluating and ensuring complete coverage increases

Engineering Contradiction:
Improvecoverage of trigger eventsVSAvoidevaluation process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by transforming the complex evaluation problem into a more manageable form through parameterization. Specifically, the method represents trigger events and data coverage using structured parameters within a semantic domain model. This parameterization allows for systematic evaluation of coverage by comparing specific parameters (such as presence/absence of trigger event characteristics in the data) rather than requiring complex qualitative assessment of entire data sets, thereby reducing evaluation complexity while maintaining reliability.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If synthetic data is generated to improve coverage, then data set completeness improves, but the time and computational resources required for data processing increase

Engineering Contradiction:
Improvedata set coverageVSAvoidtime for data generation and retraining
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent applies copying by generating synthetic data that replicates the characteristics of missing or underrepresented trigger events in the original data set. Instead of collecting additional real-world data (which would be time-consuming), the method creates artificial copies of desired data patterns through synthetic data generation. These synthetic samples are designed to match the statistical and semantic properties of real trigger events, thereby improving coverage efficiency by producing usable training data much faster than traditional data collection methods.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20230186051A1Method and device for determining a coverage of a data set for a machine learning system with respect to trigger events
Publication Date: 2023.06.15 ROBERT BOSCH GMBH
  • US20230186051A1 patent drawing
  • US20230186051A1 patent drawing
  • US20230186051A1 patent drawing

AI summary

A method of evaluating a data set with respect to a coverage of trigger events, which can produce erroneous outputs when processed by a machine learning system. The method includes: providing a semantic domain model as well as a data set;validating the machine learning system on at least a part of the data set, wherein for recurring incorrect outputs of the machine learning system with the same objects, these objects are identified as trigger events; determining a coverage of the trigger events by the data set depending on the semantic domain model.