Large Language Model Evaluation via Coarse and Fine-Grained Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for evaluating large language models face challenges in accurately assessing responsiveness, particularly when there is a conflict between user input and system setting information, and struggle to identify issues like system customization jailbreaking and leakage, as well as comprehensively evaluating model responses.
Innovation Solution
A method and apparatus for evaluating large language models by using a preset evaluation rule to assess response information from multiple models, combining coarse-grained and fine-grained evaluation dimensions to determine responsiveness, ensuring accurate and comprehensive evaluation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single evaluation method is used to assess large language models, then the evaluation process is simple, but the accuracy and comprehensiveness of responsiveness assessment deteriorates
Solution Approach 1:
The evaluation process is segmented into two distinct stages: coarse-grained evaluation (first evaluation information) that screens for obvious issues, and fine-grained evaluation (second evaluation information) that assesses multiple dimensions including responsiveness, accuracy, and completeness. This segmentation allows comprehensive assessment while managing complexity through structured progression.
Solution Approach 2:
The patent introduces multiple evaluation dimensions (responsiveness, accuracy, completeness) and a two-stage evaluation structure, transforming a single-dimensional simple evaluation into a multi-dimensional comprehensive assessment system that captures the complexity of model performance.
2Measurement precision
If multiple evaluation dimensions are used to comprehensively assess model responses, then the evaluation accuracy improves, but the evaluation complexity increases
Solution Approach 1:
Multiple evaluation dimensions are segmented into two stages: the first stage evaluates basic correctness and safety, while the second stage evaluates detailed dimensions including responsiveness, accuracy, and completeness. This segmentation manages complexity by organizing multiple dimensions into a structured two-stage process.
Solution Approach 2:
The coarse-grained evaluation is performed as a preliminary action before the fine-grained evaluation. This preliminary assessment filters out obviously incorrect or unsafe responses, allowing the system to focus computational resources on detailed evaluation only when necessary, thereby managing overall complexity.
3Reliability
If detailed evaluation of multiple dimensions is performed, then the identification of system customization jailbreaking and leakage improves, but the evaluation time increases
Solution Approach 1:
The coarse-grained evaluation serves as a preliminary filtering stage that quickly identifies obviously incorrect, unsafe, or problematic responses including potential jailbreaking attempts. This preliminary action reduces the number of cases requiring time-consuming fine-grained evaluation, thereby reducing overall evaluation time while maintaining reliable issue identification.
Solution Approach 2:
The evaluation process is segmented into fast coarse-grained assessment and detailed fine-grained analysis. This segmentation allows the system to efficiently handle the majority of cases with quick assessments while applying comprehensive multi-dimensional evaluation only when the coarse stage identifies cases requiring detailed analysis, thus balancing reliability and time efficiency.
Data Source
AI summary
A method for evaluating a large model, an electronic device and a computer readable storage medium are provided, which relate to a field of artificial intelligence technology, and in particular to fields of large models technology and deep learning technology. The method includes: evaluating a response information of each of M large language models for an input instruction based on a preset evaluation rule, so as to obtain a first evaluation information for each response information, where M is a positive integer greater than 1; evaluating, in response to the first evaluation information for the M large language models being consistent with each other, each response information in a plurality of evaluation dimensions, so as to obtain a second evaluation information for each response information; and determining an evaluation result representing a responsiveness of each large language model, according to the second evaluation information for each response information.


