AI Model Evaluation via Perturbation Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional evaluation techniques for artificial intelligence systems, particularly large language models, struggle to quantify the robustness and consistency of generated responses, leading to challenges in institutional adoption due to uncertainties and limited insights into the chain of thought and robustness of answers.

Innovation Solution

A method involving automated crowdsourcing of question perturbations using rephrasing and response models to generate lexical variants of inquiries, cluster responses based on shared characteristics, and compute metrics for evaluating the robustness and consistency of generated responses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional evaluation techniques are used for artificial intelligence systems, then the implementation is simple, but the ability to quantify robustness and consistency of generated responses is insufficient

Engineering Contradiction:
Improvequantification of robustness and consistencyVSAvoidevaluation system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The evaluation system is segmented into multiple independent AI agents, each responsible for specific evaluation tasks such as generating perturbations, assessing responses, and computing metrics. This segmentation enables precise quantification of robustness and consistency while maintaining manageable system complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Independent AI agents serve as intermediaries between the original AI system and the evaluation process. These agents automatically generate perturbations, evaluate responses, and compute metrics without requiring direct human intervention, thereby achieving precise measurement while keeping the system automated and scalable.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If automated crowdsourcing of question perturbations is implemented, then the evaluation of robustness and consistency is improved, but the system complexity increases

Engineering Contradiction:
Improverobustness and consistency evaluationVSAvoidautomated crowdsourcing system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The independent AI agents are designed with multi-functionality, capable of performing multiple tasks including generating perturbations, evaluating responses, and computing metrics. This universality reduces the need for separate specialized components, thereby improving evaluation reliability while controlling system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system automatically varies parameters such as perturbation magnitude, question formulations, and evaluation metrics to assess robustness and consistency across different conditions. This parameter-based approach enables comprehensive evaluation without requiring complex manual intervention for each test scenario.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If multiple metrics are computed for each block of responses, then the insights into model confidence and hallucinations are enhanced, but the computational overhead increases

Engineering Contradiction:
Improveinsights into chain of thought and robustnessVSAvoidcomputational resource consumption
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The system computes multiple metrics including accuracy, robustness, plurality voting, agreement, and reliability metrics for each block of responses. This comprehensive metric computation provides detailed insights into model confidence and hallucinations, with the computational overhead justified by the significant information gain about model behavior.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250209276A1Method and system for evaluating artificial intelligence models via perturbations
Publication Date: 2025.06.26 JPMORGAN CHASE BANK NA
  • US20250209276A1 patent drawing
  • US20250209276A1 patent drawing
  • US20250209276A1 patent drawing

AI summary

A method for facilitating automated model evaluation based on question perturbations is disclosed. The method includes receiving, via an application programming interface, inputs that include an inquiry in a natural language format; generating, via a rephrasing model, questions based on the inquiry, each of the questions corresponding to a lexical variant of the inquiry; determine, via response models, an initial response for each of the questions and the inquiry; clustering the initial response for each of the questions and the inquiry into blocks based on shared characteristics; and computing metrics for each of the blocks.