LLM Output Evaluation via Criteria and Converter Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large language models often generate unexpected, unreasonable, or incorrect responses, known as 'hallucinations,' which can lead to errors when used by other algorithms. It is desirable to mitigate or eliminate such responses or recognize them immediately to prevent further errors.

Innovation Solution

A method involving a machine learning model ensemble, comprising a criteria model and a converter model, to evaluate the output of a primary large language model. The criteria model compares each sentence of the output to a reference source, generating a data structure with evaluations and reasons for consistency or inconsistency. The converter model then converts this data structure into a vector storing consistency scores, allowing for the generation of a metric indicating the overall consistency of the output with the reference source.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a primary large language model generates output automatically, then productivity is improved, but reliability deteriorates due to hallucinations and unexpected responses

Engineering Contradiction:
Improveoutput generation speedVSAvoidresponse accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces a criteria model as an intermediary between the primary large language model and the final output. This criteria model evaluates whether the generated output is reasonable and consistent with the input, acting as a mediator to filter out hallucinations while preserving the automated generation capability. The intermediary layer enables both high productivity and improved reliability by separating the generation function from the validation function.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If manual evaluation of large language model output is performed, then reliability is improved, but productivity deteriorates due to time-consuming human review

Engineering Contradiction:
Improveoutput accuracyVSAvoidevaluation speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements self-service evaluation where the large language model's output is automatically evaluated by the criteria model without requiring human intervention. The system serves itself by using the input and output to generate evaluation criteria and assess consistency automatically. This eliminates the need for time-consuming manual review while maintaining reliable evaluation through automated reasoning.

Inventive Principle:
Principle #25Self-service

3Reliability

If automated evaluation using a criteria model is implemented, then reliability is improved by reducing hallucinations, but device complexity increases due to additional models and data structures

Engineering Contradiction:
Improveresponse consistencyVSAvoidsystem architecture
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the evaluation process into distinct functional components: the criteria model that generates evaluation criteria, the evaluation mechanism that assesses consistency, and the feedback loop that improves future generations. By dividing the complex evaluation task into modular segments, the system achieves high reliability through systematic assessment while managing complexity through clear separation of concerns and reusable components.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250139374A1Automated evaluation of large language models
Publication Date: 2025.05.01 INTUIT INC
  • US20250139374A1 patent drawing
  • US20250139374A1 patent drawing
  • US20250139374A1 patent drawing

AI summary

Providing an output of a primary large language model to a criteria model including a second large language model. The criteria model compares each of the sentences to a reference source and generates a first data structure including a first vector. The first vector stores, for each of the sentences, a corresponding evaluation of a given sentence as being consistent or inconsistent with the reference source, and a corresponding reason for the corresponding evaluation of the given sentence. The first data structure is provided to a converter model including a third large language model. The converter model converts the first data structure to a second data structure. The second data structure includes a second vector storing scores indicating a corresponding consistency value for each of the sentences. A metric, indicating an overall consistency of the output with respect to the reference source, is generated from the second data structure.