Generative AI System Evaluation via Quantitative Metrics and Blueprints

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing generative AI systems face challenges in efficiently developing, assessing, and monitoring their performance, particularly in terms of accuracy, relevance, and preventing biases and hallucinations.

Innovation Solution

The development of a method that constructs multiple generative AI systems using modeling blueprints, evaluates them using a set of quantitative metrics, and provides recommendations for their use based on these evaluations, while also incorporating features like retrieval-augmented generation and monitoring models to improve performance and prevent undesirable outputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple generative AI systems are constructed and evaluated using quantitative metrics, then the accuracy and reliability of system selection is improved, but the complexity of the development process increases

Engineering Contradiction:
Improvesystem selection reliabilityVSAvoiddevelopment process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the generative AI system construction into multiple independent instances, each evaluated separately using standardized quantitative metrics. This allows systematic comparison while maintaining manageable complexity through modular evaluation procedures.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs multiple quantitative metrics to evaluate different aspects of generative AI systems (accuracy, relevance, bias, hallucination prevention). By changing evaluation parameters systematically, the patent achieves reliable system selection without overwhelming complexity.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If retrieval-augmented generation is incorporated to improve accuracy and reduce hallucinations, then the quality of generated content is improved, but the computational resources and processing time increase

Engineering Contradiction:
Improvecontent generation accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent incorporates retrieval-augmented generation that performs preliminary information retrieval and verification before content generation. This preliminary action improves accuracy and reduces hallucinations by ensuring the generative model works with verified information, though it increases computational resources and processing time.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If monitoring models are used to detect biases and hallucinations in real-time, then the quality control of generated content is improved, but the system complexity and processing overhead increase

Engineering Contradiction:
Improvecontent quality controlVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements monitoring models that provide real-time feedback on generated content for detecting biases and hallucinations. This feedback mechanism improves content quality control by identifying issues during generation, but increases system complexity and processing overhead through additional monitoring layers.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250190459A1Systems and methods for development, assessment, and/or monitoring of a generative ai system
Publication Date: 2025.06.12 DATAROBOT INC
  • US20250190459A1 patent drawing
  • US20250190459A1 patent drawing
  • US20250190459A1 patent drawing

AI summary

A method for developing a generative AI system may include constructing a plurality of generative AI systems, wherein constructing the generative AI systems includes executing at least one modeling blueprint; providing a plurality of queries to each of the generative AI systems, the queries being part of an evaluation dataset; during processing of the queries by each generative AI system, monitoring values of one or more quantitative metrics; providing, for display by a user device, data indicating the values of the quantitative metrics for each generative AI system; and providing, for display by the user device, a recommendation regarding use or non-use of at least one generative AI system included in the plurality of generative AI systems.