LLM Auto-Evaluation Using Application-Specific Metric Prompts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing evaluation metrics for large language models (LLMs) are limited by the need for costly reference outputs, inability to capture detailed and subjective criteria, and reliance on pre-trained NLP models, making them inaccurate and non-scalable for specific applications.
Innovation Solution
Develops application-specific metrics with human-aligned guidelines, converted into metric prompts, using LLMs like GPT-4 for evaluation, and employs few-shot in-context learning to improve accuracy and consistency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing evaluation metrics use pre-trained NLP models and reference outputs, then evaluation can be automated, but the accuracy and scalability are limited for specific applications
Solution Approach 1:
The patent changes the fundamental parameters of evaluation by transitioning from fixed pre-trained NLP models to flexible LLM-based evaluators that can be prompted with application-specific criteria. This allows the evaluation system to adapt to different applications while maintaining automated operation, resolving the contradiction between measurement precision and adaptability.
Solution Approach 2:
The patent segments the evaluation process into distinct components: generating reference outputs separately, creating evaluation prompts with specific criteria, and using LLMs to score against these criteria. This segmentation allows each component to be optimized independently, improving both accuracy and application-specific adaptability.
2Reliability
If detailed and subjective evaluation criteria are implemented, then evaluation comprehensiveness improves, but cost and complexity increase
Solution Approach 1:
The patent introduces LLMs as intermediaries that bridge detailed subjective criteria and automated evaluation. The LLMs process complex evaluation prompts containing multiple criteria and generate consistent scores, maintaining comprehensiveness while reducing the need for manual evaluation and simplifying the overall system architecture.
Solution Approach 2:
The patent creates structured evaluation prompts that copy and formalize human evaluation criteria into reusable templates. These prompts encapsulate detailed subjective criteria in a standardized format that LLMs can process consistently, improving comprehensiveness without proportionally increasing system complexity.
3Measurement precision
If reference outputs are required for evaluation, then objective scoring is possible, but cost and scalability are reduced
Solution Approach 1:
The patent implements self-service evaluation where the system generates its own reference outputs and evaluation criteria without requiring external human annotators or pre-prepared reference data. The LLM evaluators autonomously process inputs and generate scores based on programmed criteria, maintaining objectivity while dramatically improving scalability and reducing costs.
Data Source
AI summary
In one embodiment, a non-transitory computer-readable media stores instructions executable by processors for generating a prompt configured for eliciting outputs from large language models (LLMs) based on information associated with a task, inputting the prompt to a first LLM configured to output a response based on processing the prompt, determining metrics for evaluating the first LLM based on the task, wherein each of the metrics is associated with a scoring guideline, generating metric prompts based on the respective metrics and the scoring guidelines associated with the respective metrics, inputting the response and the metric prompts to second LLMs configured to output scores corresponding to the respective metrics based on processing the response and the metric prompts, and generating an analysis report based on the metrics and their corresponding scores.


