AI Service Evaluation Using Held-Out Data Taxonomy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

AI services often perform poorly when there is a mismatch between the training data and the customer's data, leading to inconsistent results across different data sets and potential biases in evaluation metrics.

Innovation Solution

A method is developed to partition data into in-domain and out-of-domain data, using a taxonomy to define held-out data for evaluation, which is excluded from training, and determining performance metrics and guarantees using bootstrap validation processing, providing confidence intervals based on these metrics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If AI services are trained on customer data, then the service can be customized for specific customers, but the service may not perform well when there is a mismatch between training data and customer data

Engineering Contradiction:
Improvecustomization capabilityVSAvoidperformance consistency
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent applies preliminary action by pre-partitioning data into in-domain and out-of-domain portions before training begins. This allows the system to prepare evaluation datasets in advance that represent different customer scenarios, ensuring that performance can be reliably assessed across diverse conditions before deployment. The taxonomy-building process also occurs preliminarily to structure the evaluation framework.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the training and evaluation process by dividing data into distinct in-domain and out-of-domain portions. This segmentation allows separate evaluation of performance on familiar versus novel data types, enabling the system to demonstrate both customization capability and generalization reliability through structured performance breakdowns.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If evaluation metrics are computed on validation data, then performance can be measured, but the validation data may be contaminated and metrics may be biased

Engineering Contradiction:
Improveperformance measurementVSAvoidevaluation objectivity
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent extracts held-out data from the training process entirely, creating a separate evaluation dataset that is never used for training. This extraction ensures complete independence between training and evaluation, eliminating contamination and bias. The held-out data is set aside from the beginning and used solely for objective performance measurement.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces held-out data as an intermediary between training and final evaluation. This intermediary dataset acts as a buffer that prevents direct contamination while still providing realistic performance measurement, as it mirrors the structure and characteristics of actual customer data without being part of the training process.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If a taxonomy of domains and sub-domains is built for evaluation, then performance can be evaluated across different settings, but the complexity of the evaluation process increases

Engineering Contradiction:
Improveevaluation coverageVSAvoidevaluation process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the evaluation process into hierarchical levels (domains and sub-domains) based on the taxonomy structure. This segmentation allows systematic evaluation across multiple settings without requiring a completely complex monolithic approach. Each taxonomic level can be evaluated independently and aggregated.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a taxonomic dimension to the evaluation process, organizing performance metrics along hierarchical axes of domain and sub-domain. This dimensional organization structures the complexity in a manageable way, allowing comprehensive coverage across different settings while maintaining clear organizational boundaries that simplify the overall process.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11829496B2Workflow for evaluating quality of artificial intelligence (AI) services using held-out data
Publication Date: 2023.11.28 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11829496B2 patent drawing
  • US11829496B2 patent drawing
  • US11829496B2 patent drawing

AI summary

One embodiment provides for a method for evaluation of an artificial intelligence (AI) service, the method includes partitioning, by a processor, data into in-domain data and out-of-domain data. The processor defines held-out data from both of the in-domain data and the out-of-domain data for evaluation by each of domain and sub-domain based on building a taxonomy of both domains and sub-domains for the AI service. The held-out data is excluded from training data used for training the AI service. The processor further determines distribution underlying performance metrics for the held-out data using bootstrap validation processing. The processor also determines performance guarantees for multiple settings conditioned on multiple characteristics of an application scenario for the held-out data of the taxonomy based on the underlying performance metrics.