LLM Auto-Evaluation Using Application-Specific Metric Prompts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing evaluation metrics for large language models (LLMs) are limited by the need for costly reference outputs, inability to capture detailed and subjective criteria, and reliance on pre-trained NLP models, making them inaccurate and non-scalable for specific applications.

Innovation Solution

Develops application-specific metrics with human-aligned guidelines, converted into metric prompts, using LLMs like GPT-4 for evaluation, and employs few-shot in-context learning to improve accuracy and consistency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing evaluation metrics use pre-trained NLP models and reference outputs, then evaluation can be automated, but the accuracy and scalability are limited for specific applications

Engineering Contradiction:
Improveevaluation accuracyVSAvoidapplication-specific adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent changes the fundamental parameters of evaluation by transitioning from fixed pre-trained NLP models to flexible LLM-based evaluators that can be prompted with application-specific criteria. This allows the evaluation system to adapt to different applications while maintaining automated operation, resolving the contradiction between measurement precision and adaptability.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the evaluation process into distinct components: generating reference outputs separately, creating evaluation prompts with specific criteria, and using LLMs to score against these criteria. This segmentation allows each component to be optimized independently, improving both accuracy and application-specific adaptability.

Inventive Principle:
Principle #1Segmentation

2Reliability

If detailed and subjective evaluation criteria are implemented, then evaluation comprehensiveness improves, but cost and complexity increase

Engineering Contradiction:
Improveevaluation comprehensivenessVSAvoidevaluation system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces LLMs as intermediaries that bridge detailed subjective criteria and automated evaluation. The LLMs process complex evaluation prompts containing multiple criteria and generate consistent scores, maintaining comprehensiveness while reducing the need for manual evaluation and simplifying the overall system architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates structured evaluation prompts that copy and formalize human evaluation criteria into reusable templates. These prompts encapsulate detailed subjective criteria in a standardized format that LLMs can process consistently, improving comprehensiveness without proportionally increasing system complexity.

Inventive Principle:
Principle #26Copying

3Measurement precision

If reference outputs are required for evaluation, then objective scoring is possible, but cost and scalability are reduced

Engineering Contradiction:
Improvescoring objectivityVSAvoidevaluation scalability
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent implements self-service evaluation where the system generates its own reference outputs and evaluation criteria without requiring external human annotators or pre-prepared reference data. The LLM evaluators autonomously process inputs and generate scores based on programmed criteria, maintaining objectivity while dramatically improving scalability and reducing costs.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20260023929A1Application Specific Auto-evaluation for Large Language Models (LLMs)
Publication Date: 2026.01.22 ORACLE INT CORP
  • US20260023929A1 patent drawing
  • US20260023929A1 patent drawing
  • US20260023929A1 patent drawing

AI summary

In one embodiment, a non-transitory computer-readable media stores instructions executable by processors for generating a prompt configured for eliciting outputs from large language models (LLMs) based on information associated with a task, inputting the prompt to a first LLM configured to output a response based on processing the prompt, determining metrics for evaluating the first LLM based on the task, wherein each of the metrics is associated with a scoring guideline, generating metric prompts based on the respective metrics and the scoring guidelines associated with the respective metrics, inputting the response and the metric prompts to second LLMs configured to output scores corresponding to the respective metrics based on processing the response and the metric prompts, and generating an analysis report based on the metrics and their corresponding scores.