Generative evaluation system and method based on large language model

By constructing a diverse training dataset and a policy gradient optimization training and evaluation model, the problems of low evaluation accuracy and insufficient generalization ability in existing technologies are solved, and a more efficient and explainable generative evaluation effect is achieved.

CN120670846AInactive Publication Date: 2025-09-19SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510773933.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing generative evaluation methods based on large language models suffer from low accuracy, poor versatility, lack of explainability in the evaluation process, and insufficient generalization ability.

Method used

A generative evaluation system based on a large language model is designed, including a data construction module, a model training module, and an evaluation benchmark module. By screening and reconstructing public datasets, rejecting sampling, and classifying and synthesizing multiple data types, the evaluation model is trained using critical thinking and policy gradient loss optimization. The consensus of multiple existing models is combined as the benchmark truth value to calculate the score of the initial evaluation model.

Benefits of technology

It improves the accuracy and interpretability of evaluation, enhances the generalization ability of the model, improves the robustness of evaluation and training efficiency, and can adapt to diverse evaluation scenarios and stylized evaluation requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670846A_ABST
    Figure CN120670846A_ABST
Patent Text Reader

Abstract

The invention relates to a generative evaluation system and method based on a large language model.The system comprises an input unit and an evaluation unit, a pre-trained evaluation model is arranged in the evaluation unit and used for evaluating to-be-evaluated data received by the input unit and outputting a result, and the method comprises the steps that a public data set is collected, performing screening reconstruction processing, sampling rejection processing and classified synthesis processing on related data in the public data set to construct a training data set; training the large language model by using the training data set, and training to obtain an evaluation initial model in combination with SFT loss and strategy gradient loss; performing score screening on the evaluation initial model, and determining an optimal evaluation initial model as a trained evaluation model; and inputting the current to-be-evaluated data into the trained evaluation model, and outputting to obtain a corresponding generative evaluation result. Compared with the prior art, the accuracy, interpretability and generalization of evaluation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a generative evaluation system and method based on a large language model. Background Art

[0002] With the rapid development of NLG (Natural Language Generation) technology, it is becoming increasingly important to establish reliable evaluation methods to accurately measure the quality of generated content. Traditional NLG evaluation metrics, such as BLEU, ROUGE, and TER, focus primarily on surface-level differences in text, are often insufficient in assessing semantics, and can lead to misleading research conclusions. In addition, other methods that use neural embeddings to calculate scores, although they assess aspects such as semantic equivalence and fluency, have limited flexibility and scope. In addition, these traditional methods often differ significantly from human judgment and lack interpretability of scores.

[0003] Existing research, taking into account the powerful generative capabilities of Large Language Models (LLMs), has designed an LLM as an evaluator for NLG tasks. This approach, known as LLM-as-judge, fine-tunes the LLM to evaluate model outputs and render judgments. This approach not only provides reward scores but also analyzes the reasoning behind the decisions. Unlike traditional reward models that only output a single reward value, LLMs can provide more refined feedback by explaining the judgment logic.

[0004] However, current evaluation methods that rely on large language models still have some flaws. Traditional reward models can only be evaluated by directly scoring the model's responses, lacking interpretability and being "black box" evaluations, leaving users unaware of the specific evaluation process and reasons. Existing generative judge models often have limited generalization capabilities and can only be applied to fixed dialogue templates, failing to adapt to a wide range of downstream tasks. For example, the CompassJudger-1 model enhances its judgment capabilities through SFT (Supervised Fine-Tuning) data, but does not address scoring bias. The GRPO model uses group relative strategy optimization, but requires online response generation, which is computationally expensive. The Con-J model uses a bootstrapped optimization method called DPO (Direct Preference Optimization), but lacks interpretability for complex tasks.

[0005] In summary, existing evaluation models are mostly trained for specific prompts or datasets, resulting in poor generalization and insufficient world knowledge, leading to inaccurate evaluation of knowledge-intensive queries. Furthermore, existing evaluation model training lacks unified supervisory signals and optimization objectives, their evaluation benchmarks cover limited scenarios, and their ground-truth accuracy is insufficient. These factors lead to low accuracy and poor generalizability in practical generative evaluation. Summary of the Invention

[0006] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a generative evaluation system and method based on a large language model, which can improve the accuracy, interpretability and generalization of the evaluation.

[0007] The objectives of the present invention can be achieved through the following technical solutions: a generative evaluation system based on a large language model, comprising an input unit and an evaluation unit, wherein the input unit is used to receive data to be evaluated input by a user; the evaluation unit is provided with a pre-trained evaluation model, which is used to output an evaluation result corresponding to the data to be evaluated; the evaluation unit includes a data construction module, a model training module and an evaluation benchmark module, wherein the data construction module is used to reconstruct the training data;

[0008] The model training module trains the large language model using the training data to obtain an initial evaluation model;

[0009] The evaluation benchmark module is used to screen the initial evaluation model to obtain the optimal initial evaluation model, which is a pre-trained evaluation model.

[0010] Furthermore, the data construction module includes a data sorting submodule and a data synthesis submodule, wherein the data sorting submodule is used to perform screening, reconstruction and rejection sampling processing on the collected public data sets;

[0011] The data synthesis submodule is used to classify and synthesize the knowledge-based and chat-based data output by the data sorting submodule.

[0012] Furthermore, the large language model used in the model training module specifically adopts a transformer architecture.

[0013] Furthermore, the evaluation benchmark module includes an evaluation data set and a comparison submodule, and the comparison submodule is used to compare and score the output of the initial evaluation model with the standard answer in the evaluation data set, and calculate the score result of the initial evaluation model.

[0014] A generative evaluation method based on a large language model includes the following steps:

[0015] S1. Collect public data sets and perform screening and reconstruction, rejection sampling, and classification and synthesis on relevant data in the public data sets to construct a training data set;

[0016] S2. Use the training dataset to train the large language model, combining SFT loss and policy gradient loss to train and obtain the initial evaluation model;

[0017] S3. Scoring and screening the initial evaluation model to determine the optimal initial evaluation model as the trained evaluation model;

[0018] S4. Input the current data to be evaluated into the trained evaluation model and output the corresponding generative evaluation results.

[0019] Furthermore, the public data set in step S1 includes public evaluation data, public reward data and general instruction data.

[0020] Furthermore, the specific process of step S1 includes:

[0021] S11. Perform time screening on public evaluation data to filter out outdated data, and introduce a thought chain mechanism to reconstruct outdated data to obtain public evaluation data with enhanced diversity;

[0022] S12. Perform rejection sampling processing on the public reward data to obtain optimized public reward data;

[0023] S13. Classify and synthesize the diversity-enhanced public evaluation data and the optimized public reward data to obtain knowledge-based data and chat-based data;

[0024] S14. Construct a training dataset, including diversity-enhanced public evaluation data, optimized public reward data, knowledge-based data, chat-based data, and general instruction data.

[0025] Furthermore, the training process in step S2 includes: using critical thinking, designing a special thinking chain prompt template for judgment, and breaking down the judgment task;

[0026] Design the evaluation reward, establish the objective function of maximizing the expected reward, combine the SFT loss and policy gradient loss to calculate the total loss of the large language model output, and optimize the training to obtain the initial evaluation model.

[0027] Furthermore, the step S2 specifically decomposes the evaluation task into:

[0028] 1. User demand analysis: Analyze the specific requirements of user instructions;

[0029] 2. Model A / B Strengths: Evaluate the strengths of Model A and Model B;

[0030] 3. Model A / B weaknesses: Point out the weaknesses of Model A and Model B;

[0031] 4. Reasoning: Make inferences based on the above analysis;

[0032] 5. Prediction: Output the final prediction result;

[0033] The large language model output in step S2 includes the analysis process and the final answer, and SFT loss calculation is performed for the analysis process and policy gradient loss calculation is performed for the final answer.

[0034] Furthermore, the process of scoring the initial evaluation model in step S3 includes:

[0035] Use multiple existing models to generate responses corresponding to real user queries and construct an evaluation dataset;

[0036] Based on the evaluation dataset, the consensus of judges from multiple existing models is integrated to obtain the Mixture of Judges (MoJ) consensus as the benchmark truth value.

[0037] Using sample-level accuracy and model-level ranking consistency as evaluation indicators, the output of the initial evaluation model is compared with the benchmark truth value to obtain the score of the initial evaluation model.

[0038] Compared with the prior art, the present invention has the following advantages:

[0039] The evaluation unit designed in this invention includes a data construction module, a model training module, and an evaluation benchmark module. The data construction module processes public data to reconstruct training data. The model training module uses the training data to train a large language model to obtain an initial evaluation model. Furthermore, the evaluation benchmark module filters the initial evaluation model to obtain the optimal initial evaluation model. The data construction module filters, reconstructs, and rejects sampling the collected public data sets, and performs classification and synthesis of knowledge-based and chat-based data, implementing a judgment-oriented thought chain data generation scheme. This scheme can construct diverse, high-quality training data, which is beneficial for improving the reliability of subsequent large language model training and, in turn, the generalization performance of the evaluation.

[0040] When optimizing and training a large language model, the present invention first uses critical thinking to design a special thought chain prompt template for judgment, decomposes the judgment task into multiple key processes, and designs judgment rewards. Maximizing the expected reward is used as the optimization goal, and the analysis process and final answer output by the large language model are respectively subjected to SFT loss and policy gradient loss calculations. A boundary policy gradient optimization scheme combined with rejection sampling is implemented, which can perform verifiable reward-supervised training with boundary policy gradient loss for the large language model, greatly improving the training efficiency and robustness of the evaluation model, and making the evaluation results well interpretable.

[0041] This method scores the trained initial evaluation model. It uses multiple existing models to construct an evaluation dataset and integrates their consensus, using the Mixed Judges (MoJ) consensus as the ground truth. It also considers both sample-level accuracy and model-level ranking consistency as evaluation metrics. This ensures that the optimal evaluation model is selected, thereby improving evaluation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 Schematic diagram of the method flow of the present invention;

[0043] Figure 2 This is a schematic diagram of the architecture of the data construction module in the embodiment;

[0044] Figure 3 Schematic diagram of the training process of the large language model in the embodiment;

[0045] Figure 4 Schematic diagram of the data types included in the evaluation data set constructed in the embodiment. DETAILED DESCRIPTION

[0046] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0047] Example

[0048] To address the shortcomings of existing generative evaluation models, this proposal designs a unified judgment model training paradigm to improve judgment accuracy and generalization ability. It also guides intrinsic critical reasoning by supervising judgment tasks with verifiable rewards. In addition, it constructs a comprehensive and reliable judgment model evaluation benchmark, JudgeBenchV2, in order to achieve comparable judgment performance between small judgment models (e.g., 7B parameters) and large models (e.g., 235B parameters).

[0049] This solution provides a generative evaluation system based on a large language model, comprising an input unit and an evaluation unit. The input unit is configured to receive data to be evaluated input by a user. The evaluation unit is provided with a pre-trained evaluation model for outputting evaluation results corresponding to the data to be evaluated. The evaluation unit comprises a data construction module, a model training module, and an evaluation benchmark module. The data construction module is configured to reconstruct training data. Specifically, the data construction module comprises a data organization submodule and a data synthesis submodule. The data organization submodule is configured to perform screening, reconstruction, and rejection sampling on the collected public data set.

[0050] The data synthesis submodule is used to classify and synthesize the knowledge-based and chat-based data output by the data sorting submodule;

[0051] The model training module uses the training data to train the large language model to obtain an initial evaluation model. In this embodiment, the large language model used by the model training module specifically adopts the transformer architecture;

[0052] The evaluation benchmark module is used to screen the initial evaluation model and obtain the optimal initial evaluation model, that is, the pre-trained evaluation model. Specifically, the evaluation benchmark module includes an evaluation data set and a comparison submodule. The comparison submodule is used to compare and score the output of the initial evaluation model with the standard answer in the evaluation data set, and calculate the score result of the initial evaluation model.

[0053] Based on the above system, a generative evaluation method based on a large language model is implemented, such as Figure 1 As shown, the following steps are included:

[0054] S1. Collect public datasets (including public judgment data, public reward data, and general instruction data), and perform screening and reconstruction, rejection sampling, and classification and synthesis on the relevant data in the public dataset to construct a training dataset;

[0055] Specifically, the public evaluation data is screened by time to filter out outdated data, and a thinking chain mechanism is introduced to reconstruct outdated data to obtain public evaluation data with enhanced diversity.

[0056] Perform rejection sampling on public reward data to obtain optimized public reward data;

[0057] Then, we classify and synthesize the public judgment data with enhanced diversity and the optimized public reward data to obtain knowledge-based data and chat-based data.

[0058] Finally, a training dataset is constructed, which includes public judgment data with enhanced diversity, optimized public reward data, knowledge-based data, chat-based data, and general instruction data.

[0059] S2. Use the training dataset to train the large language model, combining SFT loss and policy gradient loss to train and obtain the initial evaluation model;

[0060] The training process includes: using critical thinking, designing a special thinking chain prompt template for judgment, and breaking down the judgment task into the following parts:

[0061] 1. User demand analysis: Analyze the specific requirements of user instructions;

[0062] 2. Model A / B Strengths: Evaluate the strengths of Model A and Model B;

[0063] 3. Model A / B weaknesses: Point out the weaknesses of Model A and Model B;

[0064] 4. Reasoning: Make inferences based on the above analysis;

[0065] 5. Prediction: Output the final prediction result;

[0066] Design evaluation rewards and establish an objective function that maximizes expected reward. Combine SFT loss and policy gradient loss to calculate the total loss of the large language model output (the large language model output includes the analysis process and the final answer, with SFT loss calculated for the analysis process and policy gradient loss calculated for the final answer). Optimize training to obtain the initial evaluation model.

[0067] S3. Scoring and screening the initial evaluation model to determine the optimal initial evaluation model as the trained evaluation model;

[0068] The process of scoring the initial model includes:

[0069] Use multiple existing models to generate responses corresponding to real user queries and construct an evaluation dataset;

[0070] Based on the evaluation dataset, the consensus of judges from multiple existing models is integrated to obtain the Mixture of Judges (MoJ) consensus as the benchmark truth value.

[0071] Using sample-level accuracy and model-level ranking consistency as evaluation metrics, the output of the initial model is compared with the benchmark truth value to obtain the score of the initial model.

[0072] S4. Input the current data to be evaluated into the trained evaluation model and output the corresponding generative evaluation results.

[0073] This embodiment applies the above solution to build a data construction module (such as Figure 2As shown in the figure, the collected public datasets are input into the data construction module. First, the data is sorted. For the public judgment data, the latest large model is used to reconstruct outdated judgments (outdated data is filtered out by time, that is, the publication date of the dataset, and this outdated data is reconstructed. During the reconstruction, a thinking chain mechanism is introduced to make the data more credible, and the correctly reconstructed data is filtered out based on the standard answer) to verify the accuracy. For the public reward data, the judgment annotations are generated through rejection sampling.

[0074] After that, data synthesis and classification are performed to obtain knowledge-based datasets (generating detailed judgments based on standardized benchmarks) and chat-based datasets (generating style-sensitive judgment data).

[0075] Finally, the training data is constructed, including diversity-enhanced public evaluation data, public reward data after rejection sampling, knowledge-based and chat-based synthetic data, and general instruction data.

[0076] It should be noted that in practical applications, different large models can also be used to generate evaluation data, and different sampling strategies can be used to construct training data.

[0077] In this embodiment, when performing model training, Figure 3 As shown, during the SFT training process of a large language model, the input is a standard question-and-answer pair: a user question and the model's answer. The answer is divided into two parts: the analysis process and the final answer. The model's output is compared with the standard answer and a loss is calculated. This is then fed back to the model to complete a full round of training. By calculating the loss for the analysis process and the final answer separately and then weighting the loss for the final answer, the model can output a more accurate final answer.

[0078] Among them, the total loss of the model includes SFT loss (for the analysis process) and policy gradient loss (for the final answer):

[0079]

[0080] This embodiment adopts critical thinking, designs a special thinking chain prompt template for evaluation, and decomposes the evaluation task into multiple key steps;

[0081] We then designed a judgement reward, defining a binary reward based on the match between the prediction and the ground truth. The reward function is: r(x,y) = 1 when y_kx = y*_kx; otherwise, it is 0. The reward function depends only on the answer position. The gradient loss is simplified, with kx representing the predicted answer position, y_kx representing the predicted answer, and y*_kx representing the standard answer.

[0082] Then, we use the policy gradient optimization method to establish the objective function of maximizing the expected reward and adopt the boundary policy gradient loss. Specifically, we use the policy gradient theorem to derive the policy gradient of the autoregressive model:

[0083]

[0084] Where D is the training data distribution, x is the training sample, π is the policy model, y is the reasoning path, and r is the reward function;

[0085] In addition, combined with rejection sampling, we generate diverse response candidates and filter out samples that do not conform to the benchmark truth value. By rejecting the sample, for N data, we sample M samples for each data, and then calculate the policy gradient loss for the answer prediction position:

[0086]

[0087] Among them, y <kx Represents all tokens inferred before position kx.

[0088] It should be noted that in practical applications, DPO loss or temperature loss can be used instead of boundary policy gradient loss, and different loss function combinations can be optimized for specific tasks.

[0089] After training and obtaining the initial evaluation model, this embodiment continues to perform score calculation for the initial evaluation model. First, an evaluation data set (such as Figure 4 Specifically, this method collects real user queries, classifies them by scenario and difficulty, and uses multiple existing high-performance models to generate responses. A hybrid judge (MoJ) approach is then used to integrate the judgment consensus of multiple existing high-performance models as the benchmark truth value. Sample-level accuracy and model-level ranking consistency are used as evaluation indicators to calculate the score of the initial model (i.e., the indicator of JudgebenchV2 in this embodiment):

[0090]

[0091] Among them, the first item is the accuracy at the sample level, which indicates the matching degree between the prediction of the model to be evaluated and the existing model at the sample level; the second item is the difference in the prediction ranking after normalization, r i,m represents the ranking of the model to be evaluated m under the score of the existing model i; the third item is the difference in prediction scores, s i,m It represents the score of the model to be evaluated s under the score of the existing model i.

[0092] It should be noted that in practical applications, different model combinations can be used as hybrid evaluators, and the weights of evaluation indicators can be adjusted to adapt to different scenarios.

[0093] To verify the effectiveness of this solution, the following experiments were conducted in this embodiment:

[0094] 1. Judgement Ability Test: This solution ranks high on multiple judgement-related benchmarks.

[0095] 2. General Ability Test:

[0096] On MMLU Pro, the 7B model achieved 52.55% accuracy;

[0097] On ArenaHard, the 7B model achieved 53.49% accuracy;

[0098] 3. Ablation Experiment

[0099] Boundary policy gradient loss brings 2.21% performance improvement;

[0100] Rejecting sampling data significantly improves judgement consistency;

[0101] 4. Judging style test:

[0102] Maintain consistent performance under stylized judgment prompts;

[0103] Outperforms RISE-32B by 10.67% on the Chat Hard subset.

[0104] Experimental results show that this scheme can effectively improve evaluation performance (the 7B parameter model achieves 60.52% accuracy on JudgeBenchV2, exceeding the average performance of similar 7B evaluation models by 10.5%), enhance generalization capabilities (excellent performance on common benchmarks such as MMLUPro and GPQA Diamond, and adapt to diverse evaluation scenarios and stylized evaluation requirements), improve evaluation reliability (JudgeBenchV2 covers 10 scenarios and 10,000 questions, and the mixed judge consensus reduces the bias of single model evaluation), and optimize training efficiency (boundary policy gradient loss brings a 2.21% performance improvement, and the rejection sampling strategy enhances model robustness). This scheme can be applied in practice and can be used for iterative optimization of LLM training; it can serve as a basic framework for multimodal evaluation; it can be extended to interactive evaluation scenarios; and it can also be applied to model security alignment evaluation.

Claims

1. A generative evaluation system based on a large language model, characterized by: The system comprises an input unit and an evaluation unit, wherein the input unit is used to receive the data to be evaluated input by the user; the evaluation unit is provided with a pre-trained evaluation model for outputting the evaluation results corresponding to the data to be evaluated; the evaluation unit comprises a data construction module, a model training module and an evaluation benchmark module; the data construction module is used to reconstruct the training data; The model training module trains the large language model using the training data to obtain an initial evaluation model; The evaluation benchmark module is used to screen the initial evaluation model to obtain the optimal initial evaluation model, which is a pre-trained evaluation model.

2. A generative evaluation system based on a large language model according to claim 1, characterized in that: The data construction module includes a data sorting submodule and a data synthesis submodule. The data sorting submodule is used to perform screening, reconstruction and rejection sampling on the collected public data sets. The data synthesis submodule is used to classify and synthesize the knowledge-based and chat-based data output by the data sorting submodule.

3. A generative evaluation system based on a large language model according to claim 1, characterized in that: The large language model used in the model training module specifically adopts the transformer architecture.

4. A generative evaluation system based on a large language model according to claim 1, characterized in that: The evaluation benchmark module includes an evaluation data set and a comparison submodule, wherein the comparison submodule is used to compare and score the output of the initial evaluation model with the standard answer in the evaluation data set, and calculate the score result of the initial evaluation model.

5. A generative evaluation method based on a large language model, characterized in that The following steps are involved: S1. Collect public data sets and perform screening and reconstruction, rejection sampling, and classification and synthesis on relevant data in the public data sets to construct a training data set; S2. Use the training dataset to train the large language model, combining SFT loss and policy gradient loss to train and obtain the initial evaluation model; S3. Scoring and screening the initial evaluation model to determine the optimal initial evaluation model as the trained evaluation model; S4. Input the current data to be evaluated into the trained evaluation model and output the corresponding generative evaluation results.

6. A generative evaluation method based on a large language model according to claim 5, characterized in that: The public data set in step S1 includes public evaluation data, public reward data and general instruction data.

7. A generative evaluation method based on a large language model according to claim 6, characterized in that: The specific process of step S1 includes: S11. Perform time screening on public evaluation data to filter out outdated data, and introduce a thought chain mechanism to reconstruct outdated data to obtain public evaluation data with enhanced diversity; S12. Perform rejection sampling processing on the public reward data to obtain optimized public reward data; S13. Classify and synthesize the diversity-enhanced public evaluation data and the optimized public reward data to obtain knowledge-based data and chat-based data; S14. Construct a training dataset, including diversity-enhanced public evaluation data, optimized public reward data, knowledge-based data, chat-based data, and general instruction data.

8. A generative evaluation method based on a large language model according to claim 5, characterized in that: The training process in step S2 includes: using critical thinking, designing a special thinking chain prompt template for judgment, and breaking down the judgment task; Design the evaluation reward, establish the objective function of maximizing the expected reward, combine the SFT loss and policy gradient loss to calculate the total loss of the large language model output, and optimize the training to obtain the initial evaluation model.

9. A generative evaluation method based on a large language model according to claim 8, characterized in that: The step S2 specifically decomposes the evaluation task into:

1. User demand analysis: Analyze the specific requirements of user instructions; 2. Model A / B Strengths: Evaluate the strengths of Model A and Model B; 3. Model A / B weaknesses: Point out the weaknesses of Model A and Model B; 4. Reasoning: Make inferences based on the above analysis; 5. Prediction: Output the final prediction result; The large language model output in step S2 includes the analysis process and the final answer, and SFT loss calculation is performed for the analysis process and policy gradient loss calculation is performed for the final answer.

10. A generative evaluation method based on a large language model according to claim 5, characterized in that: The process of scoring the initial model in step S3 includes: Use multiple existing models to generate responses corresponding to real user queries and construct an evaluation dataset; Based on the evaluation dataset, the judge consensus of multiple existing models is integrated to obtain the mixed judge MoJ consensus as the benchmark truth value; Using sample-level accuracy and model-level ranking consistency as evaluation indicators, the output of the initial evaluation model is compared with the benchmark truth value to obtain the score of the initial evaluation model.