Generative Model Evaluation via Multi-Dimensional Quality Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current evaluation methods for generative models, such as automatic and manual evaluation, fail to accurately assess the quality of generated content, particularly in aspects like logic performance and information quality, and lack comparability between different models.

Innovation Solution

A method and apparatus for model evaluation that involves providing inputs to a generative model, labeling outputs across multiple quality evaluation dimensions, and determining an overall quality score based on predefined quality levels and corresponding scores, allowing for objective and comparable assessment of model performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automatic evaluation methods are used, then evaluation efficiency is improved, but evaluation accuracy deteriorates

Engineering Contradiction:
Improveevaluation efficiencyVSAvoidevaluation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The evaluation system segments the evaluation process into multiple independent quality evaluation dimensions (e.g., fluency, coherence, relevance, creativity). Each dimension is evaluated separately using automated methods, and then the results are aggregated to produce an overall evaluation score. This segmentation allows automated evaluation to maintain high efficiency while improving accuracy by focusing on specific aspects rather than attempting to evaluate everything at once.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If manual evaluation methods are used, then evaluation accuracy is improved, but evaluation time deteriorates

Engineering Contradiction:
Improveevaluation accuracyVSAvoidevaluation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system merges automated evaluation and manual evaluation into a unified framework. Automated evaluation performs initial screening and scoring across multiple dimensions, while human evaluators review and adjust scores for complex or ambiguous cases. This combination allows the system to maintain high accuracy through human judgment while reducing overall evaluation time by leveraging automated processing for routine assessments.

Inventive Principle:
Principle #5Merging (Combining)

3Adaptability or versatility

If evaluation criteria are not standardized, then flexibility is improved, but comparability between models deteriorates

Engineering Contradiction:
Improveevaluation flexibilityVSAvoidmodel comparability
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The evaluation system establishes a universal set of quality evaluation dimensions and corresponding scoring criteria that can be applied across different generative models and task types. The framework includes standardized definitions for dimensions such as fluency, coherence, relevance, and creativity, along with consistent scoring scales. This universality enables fair comparison between models while maintaining flexibility to adapt to specific application domains through configurable weightings and dimension selections.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240412049A1Method, device and storage medium for model evaluation
Publication Date: 2024.12.12 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US20240412049A1 patent drawing
  • US20240412049A1 patent drawing
  • US20240412049A1 patent drawing

AI summary

In embodiments of the present disclosure, a solution for model evaluation is provided. The method comprises: providing inputs in an input set to a first generative model, to obtain a first output set output by a first generative model, wherein the first output set comprises a plurality of outputs corresponding to the plurality of inputs; obtaining first labelling information corresponding to a plurality of outputs in the first output set, the first labelling information indicating a quality level of each output marked in a plurality of quality levels divided in each quality evaluation dimension in the plurality of quality evaluation dimensions; and determining a first overall quality score of the first generative model at least based on the first labelling information of the outputs and respective quality scores corresponding to the plurality of quality levels divided in the plurality of quality evaluation dimensions.