A large language model evaluation method and system based on SimPO

CN121524565BActive Publication Date: 2026-09-11HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511643610.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-09-11
Estimated Expiration
2045-11-11

AI Technical Summary

Technical Problem

在需要综合多个指标时,权重系数(如参数量和性能分的权重)的设置往往依赖于专家经验,缺乏客观、数据驱动的确定方法

Benefits of technology

[0024]本发明公开了一种基于SimPO的大语言模型评估方法及系统,将大语言模型的动态生成表现与静态基础属性进行科学、有机的结合,构建一个全面、可解释且由数据驱动的评估框架,以实现对大型语言模型更精确、更公平的量化与排序。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121524565B_ABST
    Figure CN121524565B_ABST
Patent Text Reader

Abstract

The application discloses a large language model evaluation method and system based on SimPO. The method comprises the following steps: S1, a unified evaluation framework of mixed performance score and basic performance score is constructed to obtain an initial evaluation model; S2, an axiomatic preference dataset is used, and based on a pairwise comparison loss function, parameters in the unified evaluation framework are learned and adjusted to obtain an evaluation model; and S3, the evaluation model is used to evaluate the large language model. The method combines the dynamic generation performance of the large language model with the static basic attributes in a scientific and organic manner, constructs a comprehensive, interpretable and data-driven evaluation framework, and realizes more accurate and fair quantification and sorting of the large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, specifically relating to a method and system for evaluating large language models based on SimPO. Background Technology

[0002] Existing large language model evaluation methods have problems in the field of natural language processing, such as single evaluation dimensions, disconnect between basic attributes and actual performance, strong subjectivity in weight setting, and limitations of static benchmarks.

[0003] Traditional evaluation metrics, such as ROUGE and BLEU based on word overlap, or cosine similarity based on semantic embedding, measure the quality of generated text from a single dimension and cannot comprehensively reflect the model's overall capabilities. For example, high semantic similarity may be accompanied by low generation confidence. While benchmarks like MMLU and C-Eval are authoritative, their scores are static and cannot dynamically reflect differences in generation probabilities across specific tasks, nor can they facilitate fine-grained, interpretable comparisons between models of different scales or quantization accuracies. Existing methods typically evaluate a model's fundamental attributes, such as parameter size and quantization accuracy, separately from its performance on specific tasks. There is a lack of a unified framework to quantify how the performance advantage of a model with stronger fundamental attributes should be reflected. When multiple metrics need to be integrated, the setting of weighting coefficients (such as the weights of parameter size and performance scores) often relies on expert experience, lacking objective, data-driven methods for determination.

[0004] Therefore, there is an urgent need for a new comprehensive evaluation paradigm that can overcome the above-mentioned shortcomings, combine the basic attributes of the model with the dynamically generated performance, and determine the weights through objective methods. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present invention aims to provide a large language model evaluation method based on SimPO, which can scientifically and organically combine the dynamic generation performance of the model with its static basic attributes, thereby achieving more accurate and fair quantification and ranking of large language models.

[0006] The second objective of this invention is to provide a system for implementing the method.

[0007] This invention provides a method for evaluating large language models based on SimPO, comprising the following steps:

[0008] S1. Construct a unified evaluation framework for hybrid performance scores and basic performance scores to obtain an initial evaluation model;

[0009] S2. Using an axiomatic preference dataset and based on a pairwise comparison loss function, the parameters in the unified evaluation framework are learned and adjusted to obtain the evaluation model;

[0010] S3. Use the evaluation model to evaluate the large language model.

[0011] In step S1, the unified evaluation framework specifically involves defining a comprehensive evaluation total score. Use the following formula to calculate: ;in, For mixed performance scores; Scores for basic attributes; These are the weighting coefficients for the preferences to be learned.

[0012] The hybrid performance score is an average measure of the large language model across the entire test dataset. For a dataset containing N test samples, the hybrid performance score is expressed using the following formula: ;in, The performance score for the i-th sample is calculated using the following formula: ;in, The semantic similarity score for the i-th sample; The probabilistic cognitive bias score for the i-th sample; The weights are adjusted globally to be learned;

[0013] The probabilistic cognitive bias score is mapped to using the tanh function. The interval ensures the stability and boundedness of the correction; through The probability adjustment weights are dynamically linked to semantic similarity.

[0014] The semantic similarity score is used to measure the closeness between the answer generated by the large language model and the standard reference answer from a semantic level. Specifically, based on the ranking on the MTEB Embedding Leaderboard, a sentence embedding model is selected, and the answer generated by the large language model and the standard reference answer are mapped to high-dimensional vectors respectively. Then, the cosine similarity between the two vectors is calculated.

[0015] The probabilistic cognitive bias score is used to quantify the cognitive bias of the large language model in terms of confidence between its generated answers and the standard reference answers. Specifically, it is calculated using the following formula: ;in, Adjust the hyperparameters for the preset sensitivity; Answers generated for large language models; This is the standard reference answer; Let y be the probability that the model's answer is y given xi. Let i be the i-th sample.

[0016] The basic performance score The model is based on a saturation exponential function, expressed by the following formula: ;in, The weight of the parameter size relative to the base score; The weight of the base score for quantization bit width; For parameter scale; This is for quantization bit width.

[0017] In step S2, based on the common knowledge of the scale axiom and the precision axiom, a set of preference data pairs is constructed to obtain an axiomatic preference dataset as the training set;

[0018] Adopting the idea of ​​the Bradley-Terry model, a pairwise comparison loss function is defined, expressed by the following formula: ;in, For the set of parameters to be learned ; is the Sigmoid function; N is the total number of preference pairs in the dataset; w is the winning model in the preference pair; l is the losing model in the preference pair;

[0019] The gradient descent optimization algorithm is used to iteratively adjust the parameter values ​​in the set of parameters to be learned until the loss function converges.

[0020] The present invention also provides a system for implementing the SimPO-based large language model evaluation method, including a model building module, a model training module, and a large language model evaluation module;

[0021] The model building module constructs a unified evaluation framework for hybrid performance scores and basic performance scores, obtains an initial evaluation model, and uploads the data to the model training module;

[0022] The model training module uses the axiomatic preference dataset and the pairwise comparison loss function to learn and adjust the parameters in the unified evaluation framework based on the received data to obtain the evaluation model, and then uploads the data to the large language model evaluation module.

[0023] The large language model evaluation module evaluates the large language model based on the received data and using the evaluation model.

[0024] This invention discloses a method and system for evaluating large language models based on SimPO, which scientifically and organically combines the dynamic generation performance of large language models with static basic attributes to construct a comprehensive, interpretable, and data-driven evaluation framework, so as to achieve more accurate and fair quantification and ranking of large language models. Attached Figure Description

[0025] Figure 1This is a schematic flowchart of the method of the present invention;

[0026] Figure 2 This is a schematic diagram of the structure of the method of the present invention. Detailed Implementation

[0027] This invention provides a method for evaluating large language models based on SimPO, the flowchart of which is shown below. Figure 1 As shown, it includes the following steps:

[0028] S1. Construct a unified evaluation framework for hybrid performance scores and basic performance scores to obtain an initial evaluation model;

[0029] In step S1, the unified evaluation framework specifically involves defining a comprehensive evaluation total score. Use the following formula to calculate: ;in, For mixed performance scores; Scores for basic attributes; These are the weighting coefficients for the preferences to be learned.

[0030] The hybrid performance score is an average measure of the large language model across the entire test dataset. For a dataset containing N test samples, the hybrid performance score is expressed using the following formula: ;in, The performance score for the i-th sample is calculated using the following formula: ;in, The semantic similarity score for the i-th sample; The probabilistic cognitive bias score for the i-th sample; The weights are adjusted globally to be learned;

[0031] The probabilistic cognitive bias score is mapped to using the tanh function. The interval ensures the stability and boundedness of the correction; through The probability adjustment weights are dynamically linked to semantic similarity.

[0032] The semantic similarity score is used to measure the closeness between the answer generated by the large language model and the standard reference answer from a semantic level. Specifically, based on the ranking on the MTEB Embedding Leaderboard, a sentence embedding model is selected, and the answer generated by the large language model and the standard reference answer are mapped to high-dimensional vectors respectively. Then, the cosine similarity between the two vectors is calculated.

[0033] The probabilistic cognitive bias score is used to quantify the cognitive bias of the large language model in terms of confidence between its generated answers and the standard reference answers. Specifically, it is calculated using the following formula: ;in, Adjust the hyperparameters for the preset sensitivity; Answers generated for large language models; This is the standard reference answer; Let y be the probability that the model's answer is y given xi. Let i be the i-th sample.

[0034] The basic performance score The model is based on a saturation exponential function, expressed by the following formula: ;in, The weight of the parameter size relative to the base score; The weight of the base score for quantization bit width; For parameter scale; This is for quantization bit width.

[0035] S2. Using an axiomatic preference dataset and based on a pairwise comparison loss function, the parameters in the unified evaluation framework are learned and adjusted to obtain the evaluation model;

[0036] In step S2, based on the common knowledge of the scale axiom and the precision axiom, a set of preference data pairs is constructed to obtain an axiomatic preference dataset as the training set;

[0037] Adopting the idea of ​​the Bradley-Terry model, a pairwise comparison loss function is defined, expressed by the following formula: ;in, For the set of parameters to be learned ; is the Sigmoid function; N is the total number of preference pairs in the dataset; w is the winning model in the preference pair; l is the losing model in the preference pair;

[0038] The gradient descent optimization algorithm is used to iteratively adjust the parameter values ​​in the set of parameters to be learned until the loss function converges.

[0039] S3. Use the evaluation model to evaluate the large language model.

[0040] This invention also provides a system for implementing the SimPO-based large language model evaluation method, the structural diagram of which is shown below. Figure 2 As shown, it includes a model building module, a model training module, and a large language model evaluation module;

[0041] The model building module constructs a unified evaluation framework for hybrid performance scores and basic performance scores, obtains an initial evaluation model, and uploads the data to the model training module;

[0042] The model training module uses the axiomatic preference dataset and the pairwise comparison loss function to learn and adjust the parameters in the unified evaluation framework based on the received data to obtain the evaluation model, and then uploads the data to the large language model evaluation module.

[0043] The large language model evaluation module evaluates the large language model based on the received data and using the evaluation model.

[0044] The hardware and software platform for implementing the method in this application is as follows: multiple GPUs with high-performance computing capabilities are prepared as the computing platform, the operating system is Ubuntu 20.04, and the core development is based on Python 3.10+.

[0045] The training dataset for this application is derived from a Chinese question-and-answer dataset covering different fields and containing high-quality reference answers. It is preferably a subset of Align-Bench and Gaokao-Bench (subjective questions only). Each sample contains a "prompt" (question) and a "reference_answer" (standard answer).

[0046] Write an automated script to read the model_pool.json file and generate a preference dataset, preference_pairs.json, based on preset axioms.

[0047] 1) Scale Axiom: In the same model family and with the same number of quantization bits, the model with more parameters is the winner.

[0048] Example generated: {"winner": "Qwen1.5-7B-Chat", "loser": "Qwen1.5-1.8B-Chat","type": "scale_axiom"}

[0049] 2) Precision Axiom: For the same model (same family and params), the model with higher quantization bits is the winner.

[0050] Example of generation: {"winner": "Mistral-7B-Instruct-v0.2", "loser": "TheBloke / Mistral-7B-Instruct-v0.2-GPTQ", "type": "precision_axiom"}.

Claims

1. A method for evaluating large language models based on SimPO, characterized in that, Includes the following steps: S1. Construct a unified evaluation framework for hybrid performance scores and basic performance scores to obtain an initial evaluation model; S2. Using an axiomatic preference dataset and based on a pairwise comparison loss function, the parameters in the unified evaluation framework are learned and adjusted to obtain the evaluation model; S3. Evaluate the large language model using an evaluation model; In step S1, the unified evaluation framework specifically involves defining a comprehensive evaluation total score. Use the following formula to calculate: ;in, Score for mixed performance; Scores for basic attributes; The weighting coefficients for the preferences to be learned; The hybrid performance score is an average measure of the large language model across the entire test dataset. For a dataset containing N test samples, the hybrid performance score is expressed using the following formula: ;in, The performance score for the i-th sample is calculated using the following formula: ;in, The semantic similarity score for the i-th sample; The probabilistic cognitive bias score for the i-th sample; The weights are adjusted globally to be learned; The basic performance score The model is based on a saturation exponential function, expressed by the following formula: ;in, The weight of the parameter size relative to the base score; The weight of the base score for quantization bit width; For parameter scale; This is for quantization bit width.

2. The large language model evaluation method based on SimPO according to claim 1, characterized in that, The semantic similarity score is used to measure the closeness between the answer generated by the large language model and the standard reference answer from a semantic level. Specifically, based on the ranking on the MTEB Embedding Leaderboard, a sentence embedding model is selected, and the answer generated by the large language model and the standard reference answer are mapped to high-dimensional vectors respectively. Then, the cosine similarity between the two vectors is calculated. The probabilistic cognitive bias score is used to quantify the cognitive bias of the large language model in terms of confidence between its generated answers and the standard reference answers. Specifically, it is calculated using the following formula: ;in, Adjust the hyperparameters for the preset sensitivity; Answers generated for large language models; This is the standard reference answer; Let be the probability that the model's answer is y given xi. Let i be the i-th sample.

3. The large language model evaluation method based on SimPO according to claim 1, characterized in that, In step S2, based on the common knowledge of the scale axiom and the precision axiom, a set of preference data pairs is constructed to obtain an axiomatic preference dataset as the training set; Adopting the idea of ​​the Bradley-Terry model, a pairwise comparison loss function is defined, expressed by the following formula: ;in, For the set of parameters to be learned ; is the Sigmoid function; N is the total number of preference pairs in the dataset; w is the winning model in the preference pair; l is the losing model in the preference pair; The gradient descent optimization algorithm is used to iteratively adjust the parameter values ​​in the set of parameters to be learned until the loss function converges.

4. A system for implementing the SimPO-based large language model evaluation method according to any one of claims 1 to 3, characterized in that, It includes a model building module, a model training module, and a large language model evaluation module; The model building module constructs a unified evaluation framework for hybrid performance scores and basic performance scores, obtains an initial evaluation model, and uploads the data to the model training module; The model training module uses the axiomatic preference dataset and the pairwise comparison loss function to learn and adjust the parameters in the unified evaluation framework based on the received data to obtain the evaluation model, and then uploads the data to the large language model evaluation module. The large language model evaluation module evaluates the large language model based on the received data and using the evaluation model.

Citation Information

Patent Citations

  • System and method for automatically evaluating generation capability of large language model

    CN117521676A

  • Rational thinking method

    CN118821941A