Quantitative Evaluation Method for Generative Tasks and Related Devices
By obtaining paired evaluation data sets, building a task quantitative evaluation framework and performing variational inference training, the subjective bias and suggestive heterogeneity problems in generative task evaluation are solved, and reliable and fair quantitative evaluation results are achieved, and the accuracy and consistency of evaluation results are improved.
Patent Information
- Application Number
- CN202510423253.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-07
AI Technical Summary
There are subjective biases, suggestive heterogeneity and unscientific summary of evaluation results in generative task evaluation. It is difficult for existing methods to effectively integrate model capabilities, prompt characteristics and evaluator bias, resulting in insufficient consistency and accuracy of evaluation results.
By obtaining paired evaluation data sets, a task quantitative evaluation framework is built, a quantitative evaluation model is constructed based on project reaction theory, and variational inference training is carried out to generate evaluation results, and modeling is integrated with modeling, prompt quality, and evaluator bias as potential variables.
A reliable, fair and systematic quantitative evaluation effect was achieved, which eliminated subjective bias, improved the reliability and impartiality of the evaluation results, and ensured the fairness and transparency of the evaluation process.
Smart Images

Figure CN119961129B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a quantitative evaluation method for generative tasks and related devices. Background Art
[0002] Currently, the evaluation of generative tasks (such as natural language generation, video generation, etc.) faces three major challenges: First, the evaluation results are often affected by the subjective biases of evaluators. Individual preferences, cognitive differences, and psychological tendencies of different evaluators may lead to significant evaluation biases, thus affecting the consistency and objectivity of evaluation results. Second, the heterogeneity of prompts is another important challenge. The quality and difficulty of prompts directly determine the effectiveness and accuracy of the content generated by the model. Low-quality or ambiguous prompts may cause the content generated by the model to deviate from expectations, thus affecting the fairness of the evaluation. Finally, the process of aggregating evaluation results also presents a high degree of complexity. Existing methods usually lack a unified quantitative method that fully considers evaluator biases, prompt quality, and evaluation scale differences, resulting in insufficient uncertainty and accuracy of evaluation results.
[0003] To address these challenges, existing technical methods usually adopt rule-based scoring strategies. Such methods consider factors such as prompt quality, model answers, and evaluator consistency to a certain extent. Specifically, for different prompts, the evaluation method assigns weights according to their quality or importance; for differences between evaluators, weights are often adjusted by considering professionalism and scoring consistency. The methods for determining weights may include simple averaging, fuzzy mathematical models, or multi-criteria decision-making methods, etc.
[0004] Although existing methods attempt to alleviate the above problems through a weighting mechanism, they still have significant limitations. First, these methods fail to effectively integrate different evaluation interference factors and lack a systematic method to simultaneously consider multiple key factors such as model capabilities, prompt characteristics, and evaluator biases. Therefore, it is difficult to meet the requirements of complex evaluation scenarios. In addition, existing methods have deficiencies in theoretical basis and empirical verification and fail to provide statistical significance inferences for evaluation results. Therefore, the reliability of evaluation results is questioned, and the differences in model scores may only stem from sampling errors rather than real performance differences. Summary of the Invention
[0005] The purpose of the embodiments of this application is to propose a quantitative evaluation method, device, computer device, and storage medium for generative tasks to solve the problems of subjective bias, prompt heterogeneity, and unscientific aggregation of evaluation results in generative task evaluation, and achieve a reliable, fair, and systematic quantitative evaluation effect.
[0006] To solve the above technical problems, the embodiments of this application provide a quantitative evaluation method for generative tasks, and adopt the following technical solutions:
[0007] Obtain a paired evaluation data set;
[0008] Build a task quantitative evaluation framework, and construct a quantitative evaluation model based on the task quantitative evaluation framework;
[0009] Perform variational inference training on the constructed quantitative evaluation model to obtain a trained quantitative evaluation model;
[0010] Input the paired evaluation data set into the trained quantitative evaluation model to generate an evaluation result.
[0011] Further, the step of obtaining the paired evaluation data set specifically includes:
[0012] Obtain a preset question bank and evaluation objects;
[0013] Based on the evaluation objects, sample prompt information from the preset question bank through a randomization strategy to generate a prompt subset of the evaluation objects;
[0014] Perform generation processing on the prompt subset through a preset model to obtain answer information;
[0015] Perform paired comparison processing on the answer information to obtain the paired evaluation data set.
[0016] Further, the step of building the task quantitative evaluation framework specifically includes:
[0017] Perform quantitative evaluation on all the evaluation objects in the paired evaluation data set through a preset item response model to obtain a target evaluation probability;
[0018] Perform latent variable modeling processing on the target evaluation probability based on item response theory to obtain the task quantitative evaluation framework.
[0019] Further, the step of constructing a quantitative evaluation model based on the task quantitative evaluation framework specifically includes:
[0020] Extract the latent variable parameters in the task quantitative evaluation framework;
[0021] Import the latent variable parameters into a general model, and train the general model using an evaluation training data set to obtain a constructed quantitative evaluation model.
[0022] Further, the step of performing variational inference training on the constructed quantitative evaluation model to obtain a trained quantitative evaluation model specifically includes:
[0023] Obtain the initial value of the variational distribution parameter;
[0024] Perform variational distribution processing on the latent variable parameters based on the initial values of the variational distribution parameters to obtain an initial latent variable distribution;
[0025] Perform maximum variational lower bound processing on the initial latent variable distribution, iteratively converge the variational distribution parameters, and obtain the optimal latent variable estimate;
[0026] Import the optimal latent variable estimate into the constructed quantization evaluation model for iterative training to obtain the trained quantization evaluation model.
[0027] Further, the step of iteratively converging the variational distribution parameters to obtain the optimal latent variable estimate specifically includes:
[0028] Iterate the variational distribution parameters by the coordinate ascent method;
[0029] Calculate the optimal latent variable estimate based on the iteratively converged variational distribution parameters.
[0030] Further, the evaluation results include task ranking, task evaluation score, quantization quality evaluation, and reliability metric. The step of inputting the paired evaluation data set into the trained quantization evaluation model to generate evaluation results specifically includes:
[0031] Input the paired evaluation data set into the trained quantization evaluation model to obtain target latent variable values, where the target latent variable values include target latent ability estimate, target discrimination estimate, target effectiveness estimate, and target deviation estimate;
[0032] Generate the task ranking and the task evaluation score based on the target latent ability estimate;
[0033] Generate the quantization quality evaluation based on the target discrimination estimate and the target effectiveness estimate;
[0034] Generate the reliability metric based on the target deviation estimate.
[0035] To solve the above technical problems, an embodiment of the present application also provides a quantization evaluation device, which adopts the following technical solutions:
[0036] An acquisition module, configured to acquire a paired evaluation data set;
[0037] A construction module, configured to construct a task quantization evaluation framework and construct a quantization evaluation model based on the task quantization evaluation framework;
[0038] A training module, configured to perform variational inference training on the constructed quantization evaluation model to obtain a trained quantization evaluation model;
[0039] A generation module, configured to input the evaluation data set into the trained quantization evaluation model to generate an evaluation result.
[0040] To solve the above technical problems, an embodiment of the present application further provides a computer device, which adopts the following technical solutions:
[0041] A computer device includes a memory and a processor. Computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of the quantization evaluation method are implemented.
[0042] To solve the above technical problems, an embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solutions:
[0043] A computer-readable storage medium, characterized in that computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the steps of the quantization evaluation method are implemented.
[0044] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:
[0045] In the embodiments of the present application, by obtaining a paired evaluation data set, building a task quantization evaluation framework, constructing a quantization evaluation model based on the task quantization evaluation framework, performing variational inference training on the constructed quantization evaluation model to obtain a trained quantization evaluation model, inputting the paired evaluation data set into the trained quantization evaluation model to generate an evaluation result, and by building a new task quantization evaluation framework and modeling model capabilities, prompt quality, and evaluator bias as latent variables, the problems of subjective bias, prompt heterogeneity, and unscientific evaluation result aggregation in generative task evaluation are solved, and a reliable, fair, and systematic quantization evaluation effect is achieved. Description of the Drawings
[0046] To more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 is an exemplary architecture diagram to which the present application can be applied;
[0048] Figure 2 A flowchart of an embodiment of the quantization evaluation method according to the present application;
[0049] Figure 3It is a schematic structural diagram of an embodiment of a quantization evaluation device according to the present application;
[0050] Figure 4 It is a schematic structural diagram of an embodiment of a computer device according to the present application. Detailed implementation manners
[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.
[0052] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0053] To enable those skilled in the art of this technology to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings.
[0054] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. The terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0055] Users can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.
[0056] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop 1011, the tablet computer 1012, or the mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, a desktop computer, and the like.
[0057] The server 103 can be a server that provides various services, such as a background server that supports the pages displayed on the terminal device 101.
[0058] It should be noted that the quantization evaluation method provided by the embodiments of the present application is generally executed by the server / terminal device. Correspondingly, the quantization evaluation device is generally set in the server / terminal device.
[0059] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in
[0060] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 2 Continuing to refer to
[0061] Step S201, obtain a paired evaluation data set.
[0062] In this embodiment, the above quantization evaluation method can be deployed in a quantization evaluation platform. The above quantization evaluation platform can be constructed by a server or a server cluster. The above server or server cluster can be any electronic device with functions such as data processing, data transmission, and data storage. The electronic device (such as Figure 1 the server / terminal device shown) on which the quantization evaluation method runs can receive the paired evaluation data set through a wired connection method or a wireless connection method. It should be noted that the above wireless connection method can include but is not limited to 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other currently known or future-developed wireless connection methods.
[0063] In this embodiment, the paired evaluation dataset may be a set of data obtained by pairwise comparison of responses generated by an evaluator based on preset prompt information and a generation model. The paired evaluation dataset can be used to quantitatively evaluate the performance of different models or generative tasks.
[0064] Specifically, the paired evaluation dataset may consist of multiple task responses generated for the same prompt, with each evaluation object (such as an NLG model) generating a response. The evaluator compares each pair of generated responses according to certain quality criteria and selects the better one. The result of each comparison (i.e., the response selected by the evaluator) is recorded as a piece of data, forming the paired evaluation dataset.
[0065] The paired evaluation dataset not only includes the generated content of the evaluation object but also the results of the evaluator's comparison and selection, which are used for subsequent quantitative evaluation of model training and analysis.
[0066] Specifically, regarding the specific implementation process of obtaining the paired evaluation dataset, this application will further describe the details in subsequent specific embodiments and will not elaborate too much here.
[0067] Step S202: Build a task quantitative evaluation framework and construct a quantitative evaluation model based on the task quantitative evaluation framework.
[0068] In this embodiment, the task quantitative evaluation framework may be a model architecture constructed based on the item response theory (IRT), which can be used to quantitatively evaluate the performance of different evaluation objects in generative tasks. The task quantitative evaluation framework includes variables such as evaluating their latent ability, the discrimination of prompts, effectiveness, and the subjective bias value of evaluators.
[0069] The constructed quantitative evaluation model may be a preliminary model designed according to task requirements and the evaluation framework. Based on the task quantitative evaluation framework, the constructed quantitative evaluation model systematically models the characteristics of three dimensions: model ability, prompt characteristics, and evaluator bias using latent variable statistical methods. Compared with traditional evaluation models, the quantitative evaluation model built by this task quantitative evaluation framework can significantly improve the fairness, reliability, and statistical quantification level of the evaluation process.
[0070] The constructed quantitative evaluation model includes basic quantitative evaluation capabilities and outputs evaluation results. Specifically, the constructed quantitative evaluation model is an untrained initial model that needs to be further optimized through subsequent variational inference and training processes to provide accurate evaluation results in actual tasks.
[0071] Specifically, regarding the above-mentioned task quantification evaluation framework and the specific implementation process of constructing a quantification evaluation model based on the task quantification evaluation framework, this application will further describe the details in subsequent specific embodiments and will not elaborate too much here.
[0072] Step S203: Perform variational inference training on the constructed quantification evaluation model to obtain a trained quantification evaluation model.
[0073] In this embodiment, the above-mentioned trained quantification evaluation model can be a model based on the Item Response Theory (IRT) trained by the variational inference method. This model has been preliminarily trained with a historical evaluation dataset and can automatically learn and optimize the estimated values of latent variables (such as model ability, hint quality, and rater bias). After training, the trained quantification evaluation model can accurately generate evaluation results when new evaluation data is input, providing accurate and objective evaluation results for complex generative tasks.
[0074] Specifically, regarding the above-mentioned specific implementation process of performing variational inference training on the constructed quantification evaluation model to obtain a trained quantification evaluation model, this application will further describe the details in subsequent specific embodiments and will not elaborate too much here.
[0075] Step S204: Input the paired evaluation dataset into the trained quantification evaluation model to generate evaluation results.
[0076] In this embodiment, the above-mentioned evaluation results are the outputs generated by the trained quantification evaluation model. Specifically, the above-mentioned evaluation results can include theta parameters, gamma parameters, lambda parameters, and beta parameters.
[0077] Among them, the theta parameter can be used to represent the latent ability of the model, that is, the performance on all (non-individual) tasks. The gamma parameter can be used to represent the discrimination or difficulty of the task. The lambda parameter can be used to represent the quality or effectiveness of the task. The beta parameter can be used to represent the reliability or bias of the rater.
[0078] Specifically, regarding the above-mentioned specific implementation process of inputting the paired evaluation dataset into the trained quantification evaluation model to generate evaluation results, this application will further describe the details in subsequent specific embodiments and will not elaborate too much here.
[0079] This application obtains a paired evaluation dataset, constructs a task quantitative evaluation framework, builds a quantitative evaluation model based on the task quantitative evaluation framework, performs variational inference training on the constructed quantitative evaluation model to obtain a trained quantitative evaluation model, inputs the paired evaluation dataset into the trained quantitative evaluation model to generate an evaluation result, and constructs a new task quantitative evaluation framework, modeling model ability, prompt quality, and evaluator bias as latent variables, solving the problems of subjective bias, prompt heterogeneity, and unscientific evaluation result aggregation in generative task evaluation, and achieving a reliable, fair, and systematic quantitative evaluation effect.
[0080] In some alternative implementation manners, step S201 includes the following steps:
[0081] Obtain a preset question bank and an evaluation object;
[0082] Based on the evaluation object, sample prompt information from the preset question bank through a randomization strategy to generate a prompt subset of the evaluation object;
[0083] Perform text generation processing on the prompt subset through a preset model to obtain response information;
[0084] Perform paired comparison processing on the response information to obtain a paired evaluation dataset.
[0085] In this embodiment, the preset question bank may be a set containing multiple tasks or questions, and the above tasks can cover different evaluation scenarios and difficulty levels to ensure a comprehensive evaluation of the evaluation object.
[0086] The above evaluation object may be a text generation model (such as text generation, dialogue generation models, etc., models that can interact with various available prompt words). The above evaluation object generates corresponding content according to the prompts in the preset question bank, and its output result is evaluated for quality. When there are multiple evaluation objects, one evaluation object corresponds to one text generation model.
[0087] In a possible embodiment, prepare a preset question bank containing prompts ; meanwhile, there are evaluation objects to be evaluated , and human evaluators participate in the evaluation . Specifically, first, for each evaluation object, several prompts are sampled from the preset question bank to form a prompt subset of the evaluation object , and it is required that the preset model generates corresponding response content for each prompt (i.e., the above text generation process). The above preset model can be a generative large language model or a generative dialogue robot and other generative task models. One preset model corresponds to one evaluation object. To save evaluation time and ensure unbiased results, the sampling and distribution of prompts adopt a randomization strategy. Therefore, the evaluation object does not need to answer all prompts. In addition, when the number of prompts and the number of evaluation objects are small, it can also be required that all models generate response content for all prompts. The evaluation system will pair up the responses generated for the same prompt to form a paired response dataset.
[0088] Subsequently, the evaluator conducts paired evaluations on the content generated by the evaluation object (i.e., the above paired comparison process). Each evaluator will be randomly assigned some or all of the paired responses, compare the content, and mark the better response to generate the above paired evaluation dataset. More specifically, to save evaluation time and ensure unbiased results, this step adopts a randomization strategy, that is, not all evaluators participate in all paired comparison tasks, and each pair of comparisons may be completed by different evaluators.
[0089] This application conducts evaluations by adopting the method of paired comparison, which can effectively reduce evaluation bias, improve the reliability and fairness of evaluation results. The randomization strategy ensures the fairness of prompt distribution and evaluator participation, avoiding the influence of subjective bias. At the same time, paired comparison reduces the workload of each evaluator, improves the evaluation efficiency, ensures the diversity and representativeness of data, and provides high-quality evaluation data for the subsequent training of the IRT model.
[0090] In some alternative implementation manners, step S202 includes the following steps:
[0091] Quantitatively evaluate all evaluation objects in the paired evaluation dataset through a preset item response model to obtain a target evaluation probability;
[0092] Based on item response theory, perform latent variable modeling on the target evaluation probability to obtain a task quantitative evaluation framework.
[0093] In this embodiment, the above target evaluation probability may refer to the probability of the evaluation result calculated through an item response model (Item Response Model, IRM) based on the generated content of each evaluation object (such as an NLG model) and a given prompt. This probability reflects the performance of the evaluation object in a specific task. Through the IRM model, the response ability of each evaluation object to the task can be quantified, thereby obtaining a "probability score" (i.e., the target evaluation probability), and this score represents the possibility of the content generated by the model in a specific task.
[0094] In this embodiment, the above latent variable modeling process can be a further analysis of the target evaluation probability, using the latent variables in the IRT model for modeling. The above latent variables refer to factors that cannot be directly observed during the evaluation process, but they have an important impact on the evaluation results. Specifically, latent variable modeling is to estimate the latent variable parameters of each evaluation object under different conditions by constructing a latent variable model, that is, the above latent ability parameters, the discrimination, effectiveness of effective evaluation prompts, and the subjective deviation value of the evaluator, etc., so as to obtain a more refined task quantitative evaluation framework.
[0095] In one possible embodiment, the IRT model is used to model the probability that an evaluator selects one evaluation object to generate better content than another in a pairwise comparison. For a given prompt , from the evaluation object , the paired responses formed by the answers, and the evaluator the binary result obtained by manually evaluating it " " or " ". Among them, " " represents that the evaluator believes that the evaluation object performs better than another evaluation object in the text generation task ; on the contrary, " " represents that the evaluator believes that the evaluation object performs better than another evaluation object in the text generation task .
[0096] In one possible above, the above latent variable modeling process can be as follows:
[0097] Express the probability of " " as:
[0098]
[0099] Among them, specifically, the meaning of the above latent variable parameters is as follows:
[0100] and (the above latent ability parameters): respectively represent the latent abilities of the evaluation objects and . The latent ability represents the overall performance of the evaluation object in all prompts and evaluations. A higher value indicates a better latent ability. In addition, is the ability gap between the two evaluation objects in the pairwise comparison. The larger its absolute value, the stronger the confidence of the model in the ability ranking of the two evaluation objects.
[0101] (Discrimination of the above effective evaluation prompts): Prompt 's discrimination to capture its sensitivity to the differences in the abilities of the evaluation objects. High-discrimination prompts can more effectively distinguish the abilities of the evaluation objects. For example, difficult prompts can better distinguish the abilities of the evaluation objects than easy prompts. In addition, negative discrimination may indicate an incorrect reference answer (if any).
[0102] (The above effectiveness): Prompt 's effectiveness represents the quality of the prompt. For example, if the prompt is only "Describe a book" without providing any context information (such as the book title or characters), it will be difficult for all language models to answer accurately regardless of their potential abilities. Such low-quality prompts will lead to the invalidation of the evaluation results, making the probability of a binary choice approach 50%.
[0103] (Subjective bias value of the above evaluator): Evaluator 's bias represents the tendency to always prefer a certain option without considering the ability differences. For example, some evaluators may tend to choose the first option. A higher positive bias indicates that the evaluator has a stronger preference for the second option, which requires a greater ability gap between the evaluation objects of the first option than the second option to overcome the bias.
[0104] In some alternative implementation manners, step S202 further includes the following steps:
[0105] Extract the latent variable parameters within the task quantitative evaluation framework;
[0106] Import the latent variable parameters into the general model and use the evaluation training data set to train the general model to obtain the constructed quantitative evaluation model.
[0107] In this embodiment, the above-mentioned evaluation training dataset may refer to a structured dataset constructed based on a pairwise evaluation process and used to train a quantization evaluation model. Its core consists of a triple of "prompt - generated content by the evaluation object - result selected by the evaluator". This dataset forms large-scale pairwise comparison results by pairwise pairing the response content generated by multiple evaluation objects for different prompts and having human evaluators label their preferences for each pair of responses (i.e., select which one is better). The above-mentioned evaluation training dataset not only contains the observed results (the choices of the evaluators), but also indirectly reflects hidden potential factors, such as the performance capabilities of the evaluation objects, the distinctiveness and effectiveness of the prompts, and the subjective biases of the evaluators. By inputting this evaluation training dataset, the general model can combine the latent variable parameters in the task quantization evaluation framework to perform supervised or approximate inference training, and finally form a quantization evaluation model with actual evaluation capabilities.
[0108] In this embodiment, first, relevant latent variable parameters are extracted from the task quantization evaluation framework, such as latent ability parameters, the distinctiveness and effectiveness of effective evaluation prompts, and the subjective bias values of the evaluators. The above-mentioned latent variable parameters can be represented by preset initial values. Then, these latent variable parameters are imported into a general model, and the general model is trained using the evaluation training dataset so that the model can learn how to accurately evaluate tasks based on these latent variables. After training, a constructed quantization evaluation model is obtained, and this model can perform effective quantization evaluation in new evaluation tasks.
[0109] In some alternative implementation manners, step S203 includes the following steps:
[0110] Obtain the initial value of the variational distribution parameter;
[0111] Based on the initial value of the variational distribution parameter, perform variational distribution processing on the latent variable parameter to obtain the initial latent variable distribution;
[0112] Perform maximization of the variational lower bound processing on the initial latent variable distribution, and iteratively converge the variational distribution parameter to obtain the optimal latent variable estimate value;
[0113] Import the optimal latent variable estimate value into the constructed quantization evaluation model for iterative training to obtain a trained quantization evaluation model.
[0114] In the step of "iteratively converge the variational distribution parameter to obtain the optimal latent variable estimate value", it further includes:
[0115] Iteratively vary the variational distribution parameter by the coordinate ascent method;
[0116] Based on the iteratively converged variational distribution parameter, calculate the optimal latent variable estimate value.
[0117] In this embodiment, the above variational distribution parameters may refer to the parameters used to approximate the true distribution of latent variables during the variational inference process. Generally, the variational distribution is assumed to be a certain standard probability distribution, such as a normal distribution, and its parameters include the mean and variance. These parameters control the shape of the variational distribution. By optimizing these parameters, the variational distribution gradually approaches the true posterior distribution. Specifically, the variational distribution parameters are updated through methods such as coordinate ascent during the training process, so as to provide the best estimate of the latent variables and help improve the accuracy and generalization ability of the model.
[0118] Specifically, the initial values of the above variational distribution parameters can be randomly generated. The initial values of the variational distribution parameters provide a starting point for the variational distribution. During the subsequent optimization process, the variational distribution parameters will be adjusted through the training data and gradually converge to the optimal values that can approximate the true posterior distribution.
[0119] In this embodiment, the process of performing variational distribution processing on the latent variable parameters is to estimate the initial latent variable distribution by applying the initial mean and variance parameters to a preset variational distribution model (usually a normal distribution). At this stage, the initial mean and variance determine the preliminary estimate of the latent variables, reflecting the preliminary guess of the latent variables.
[0120] In a possible embodiment, given the above evaluation results and the IRM model, the variational inference method can be used to approximately infer the latent variables ( , where ) in the IRM model and obtain its (approximate) posterior distribution. Specifically:
[0121] The first step: Introduce the variational distribution: To approximate the true posterior distribution , where represents all the observed pairwise comparison results, introduce the variational distribution . Based on the mean-field approximation method, the variational distribution can be factorized into an independent form:
[0122]
[0123] where all the latent variable parameters are represented by . That is, the variational distribution can be further expressed as . Specifically, the variational distribution of all latent variable parameters follows a preset variational distribution family (such as a normal distribution or a log-normal distribution), and the parameters of the distribution (such as the mean and variance) will be determined through the optimization process in the following steps 2 and 3.
[0124] The second step: Define the optimization objective
[0125] By maximizing the Evidence Lower Bound (ELBO), the true posterior distribution can be indirectly minimized and the variational distribution The Kullback-Leibler (K-L) divergence between them, that is, by optimizing the parameters of the variational distribution, an approximate posterior distribution is found. Based on the factorized form of the above variational distribution, the ELBO for subsequent optimization can be expressed as follows:
[0126]
[0127] Step 3: Optimize the parameters of the variational distribution
[0128] The coordinate ascent method is used to optimize the ELBO. In each iteration, with other parameter distributions fixed , update until the preset convergence condition is met. Finally, is the approximate posterior distribution of the latent variable .
[0129] Furthermore, the estimated value of the latent variable can be obtained from the relevant statistics (such as the expected value) of its approximate posterior distribution, denoted as , where .
[0130] Subsequently, these estimated values of the latent variables are substituted into the original model to evaluate the prediction ability of the dichotomous model and verify its reliability in practical applications. Taking accuracy as an example, accuracy is the proportion of the number of correctly predicted pairwise comparisons to the total number of pairwise comparisons in all pairwise comparisons by the Item Response Model. The larger the value, the better the model performance
[0131] In some alternative implementation manners, step S204 includes the following steps:
[0132] Input the pairwise evaluation data set into the trained quantization evaluation model to obtain target latent variable values, where the target latent variable values include target latent ability estimates, target discrimination estimates, target effectiveness estimates, and target bias estimates;
[0133] Generate a task ranking and a task evaluation score based on the target latent ability estimate;
[0134] Generate a quantization quality evaluation based on the target discrimination estimate and the target effectiveness estimate;
[0135] Generate a reliability metric based on the target bias estimate.
[0136] In this embodiment, the above-mentioned target potential ability estimation value can be the comprehensive performance ability of an evaluation model or task under a given prompt, indicating the quality of the evaluation object's performance in a specific task; the above-mentioned target discrimination estimation value can be an index measuring the ability of the prompt to distinguish the abilities of the evaluation object, reflecting whether the prompt can effectively distinguish the ability differences of different models or tasks; the target effectiveness estimation value can be the quality of evaluating the prompt in generating effective and accurate content, indicating whether the prompt provides sufficient information to help the model generate high-quality answers; the above-mentioned target bias estimation value can be the systematic preference or tendency that may exist in the evaluation process by the evaluator, indicating that the evaluator, regardless of the actual performance of the evaluation object, tends to certain specific options or answers.
[0137] In a possible embodiment, first, the paired evaluation data set is input into the trained quantization evaluation model, and the model generates corresponding target latent variable values according to the input evaluation data. These latent variable values include: the target potential ability estimation value (measuring the overall performance of the model or task), the target discrimination estimation value (reflecting the effect of the prompt on distinguishing task difficulty and model ability), the target effectiveness estimation value (indicating the ability of the prompt to generate effective content), and the target bias estimation value (indicating the systematic bias of the evaluator). Then, based on the target potential ability estimation value, a task ranking and a task evaluation score are generated, and the tasks or models are sorted by comparing their abilities; based on the target discrimination estimation value and the target effectiveness estimation value, a quantization quality evaluation is generated to evaluate the task quality and the effectiveness of the prompt; finally, based on the target bias estimation value, a reliability metric is generated to evaluate the influence of bias in the evaluation process and provide a reliability reference for the final decision.
[0138] By inputting the paired evaluation data set into the trained quantization evaluation model and generating target latent variable values, this application can accurately quantify the performance of tasks or models, including task ranking, evaluation score, quality evaluation, and reliability metric. This process effectively eliminates subjective bias and provides a more fair and scientific evaluation result. At the same time, the quantization quality evaluation and reliability metric help to judge the quality of the task and the stability of the evaluation, providing strong data support for decision-making and ensuring the fairness and transparency of the evaluation process.
[0139] The embodiments of this application can obtain and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use the knowledge to obtain the best results.
[0140] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, generative task processing technology, and machine learning / deep learning.
[0141] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0142] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and they can be executed in other orders. Moreover, at least some of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. Their execution order does not necessarily have to be sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0143] Further reference Figure 3 to Figure 2 As an implementation of the method shown above, an embodiment of a quantization evaluation device is provided in this application. This device embodiment corresponds to the method embodiment shown in Figure 2 and this device can be specifically applied to various electronic devices.
[0144] As Figure 3 shown, the quantization evaluation device 300 described in this embodiment includes: an acquisition module 301, a construction module 302, a training module 303, and a generation module 304, where:
[0145] The acquisition module 301 is used to acquire paired evaluation data sets;
[0146] The construction module 302 is used to construct a task quantization evaluation framework and build a quantization evaluation model based on the task quantization evaluation framework;
[0147] A training module 303, configured to perform variational inference training on the constructed quantization evaluation model to obtain a trained quantization evaluation model;
[0148] A generation module 304, configured to input the evaluation data set into the trained quantization evaluation model to generate an evaluation result.
[0149] Among them, the obtaining module 301 includes:
[0150] A first obtaining sub-module, configured to obtain a preset question bank and an evaluation object;
[0151] A random sampling sub-module, configured to sample prompt information from the preset question bank based on the evaluation object through a randomization strategy to generate a prompt subset of the evaluation object;
[0152] A text generation sub-module, configured to perform text generation processing on the prompt subset through a preset model to obtain answer information;
[0153] A pairwise comparison sub-module, configured to perform pairwise comparison processing on the answer information to obtain the pairwise evaluation data set.
[0154] Among them, the building module 302 includes:
[0155] A quantization evaluation sub-module, configured to perform quantization evaluation on all the evaluation objects in the pairwise evaluation data set through a preset item response model to obtain a target evaluation probability;
[0156] A modeling processing sub-module, configured to perform latent variable modeling processing on the target evaluation probability based on item response theory to obtain the task quantization evaluation framework.
[0157] Among them, in the training module 303, it includes:
[0158] An extraction sub-module, configured to extract latent variable parameters in the task quantization evaluation framework;
[0159] An import sub-module, configured to import the latent variable parameters into a general model and use an evaluation training data set to train the general model to obtain a constructed quantization evaluation model.
[0160] Among them, the training module 303 includes:
[0161] A second obtaining sub-module, configured to obtain an initial value of variational distribution parameters;
[0162] A variational distribution sub-module, configured to perform variational distribution processing on the latent variable parameters based on the initial value of the variational distribution parameters to obtain an initial latent variable distribution;
[0163] An iteration sub-module, configured to perform a maximizing variational lower bound process on the initial latent variable distribution, iteratively converge the variational distribution parameters, and obtain the optimal latent variable estimate;
[0164] A training sub-module, configured to import the optimal latent variable estimate into the constructed quantization evaluation model for iterative training to obtain the trained quantization evaluation model.
[0165] Wherein, the iteration sub-module specifically includes:
[0166] A coordinate ascent unit, configured to iteratively update the variational distribution parameters by the coordinate ascent method;
[0167] A calculation unit, configured to calculate the optimal latent variable estimate based on the iteratively converged variational distribution parameters;
[0168] Wherein, the generation module 304 includes:
[0169] An input sub-module, configured to input the paired evaluation data set into the trained quantization evaluation model to obtain target latent variable values, where the target latent variable values include a target latent ability estimate, a target discrimination estimate, a target effectiveness estimate, and a target bias estimate;
[0170] A first generation sub-module, configured to generate the task ranking and the task evaluation score based on the target latent ability estimate;
[0171] A second generation sub-module, configured to generate the quantization quality evaluation based on the target discrimination estimate and the target effectiveness estimate;
[0172] A third generation sub-module, configured to generate the reliability metric based on the target bias estimate.
[0173] In this embodiment, by obtaining a paired evaluation data set, building a task quantization evaluation framework, constructing a quantization evaluation model based on the task quantization evaluation framework, performing variational inference training on the constructed quantization evaluation model to obtain a trained quantization evaluation model, inputting the paired evaluation data set into the trained quantization evaluation model to generate evaluation results, and by building a new task quantization evaluation framework, modeling model ability, prompt quality, and rater bias as latent variables, the problems of subjective bias, prompt heterogeneity, and unscientific aggregation of evaluation results in generative task evaluation are solved, and a reliable, fair, and systematic quantization evaluation effect is achieved.
[0174] In this embodiment, the operations respectively performed by the above units or modules correspond one-to-one to the steps of the quantization evaluation method in the above implementation manner, and will not be elaborated herein.
[0175] To solve the above technical problems, an embodiment of the present application further provides a computer device. Specifically, please refer to Figure 4 , Figure 4 , which is a basic structural block diagram of the computer device in this embodiment.
[0176] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are communicatively connected to each other through a system bus. It should be noted that only the computer device 4 with components 41 - 43 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0177] The computer device can be a desktop computer, a notebook, a palm computer, a cloud server, and other computing devices. The computer device can perform human-computer interaction with users through a keyboard, a mouse, a remote control, a touchpad, a voice control device, or other means.
[0178] The memory 41 at least includes one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a FlashCard, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and the external storage device of the computer device 4. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions of the quantitative evaluation method. In addition, the memory 41 may also be used to temporarily store various types of data that have been output or will be output.
[0179] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run the computer-readable instructions stored in the memory 41 or process data, such as running the computer-readable instructions of the quantitative evaluation method.
[0180] The network interface 43 may include a wireless network interface or a wired network interface, and the network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.
[0181] This embodiment provides a computer device. By obtaining a paired evaluation data set, building a task quantitative evaluation framework, constructing a quantitative evaluation model based on the task quantitative evaluation framework, performing variational inference training on the constructed quantitative evaluation model to obtain a trained quantitative evaluation model, inputting the paired evaluation data set into the trained quantitative evaluation model to generate an evaluation result, and building a new task quantitative evaluation framework to model model ability, prompt quality, and evaluator bias as latent variables, it solves the problems of subjective bias, prompt heterogeneity, and unscientific evaluation result aggregation in generative task evaluation, and achieves a reliable, fair, and systematic quantitative evaluation effect.
[0182] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions executable by at least one processor to cause the at least one processor to execute the steps of the quantization evaluation method as described above.
[0183] This embodiment provides a computer-readable storage medium. By obtaining a paired evaluation data set, building a task quantization evaluation framework, constructing a quantization evaluation model based on the task quantization evaluation framework, performing variational inference training on the constructed quantization evaluation model to obtain a trained quantization evaluation model, inputting the paired evaluation data set into the trained quantization evaluation model to generate an evaluation result, and by building a new task quantization evaluation framework and modeling model ability, prompt quality, and evaluator bias as latent variables, the problems of subjective bias, prompt heterogeneity, and unscientific evaluation result aggregation in generative task evaluation are solved, and a reliable, fair, and systematic quantization evaluation effect is achieved.
[0184] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions to cause a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0185] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The accompanying drawings show preferred embodiments of the present application, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields shall be similarly within the scope of patent protection of the present application.
Claims
1. A quantitative evaluation method for generative tasks, characterized in that It includes the following steps: Obtain a paired evaluation data set; Build a task quantization evaluation framework and construct a quantization evaluation model based on the task quantization evaluation framework; Perform variational inference training on the constructed quantization evaluation model to obtain a trained quantization evaluation model; Input the paired evaluation data set into the trained quantization evaluation model to generate an evaluation result; The step of obtaining the paired evaluation data set specifically includes: Obtain a preset question bank and evaluation objects; Based on the evaluation objects, sample prompt information from the preset question bank through a randomization strategy to generate a prompt subset of the evaluation objects; Perform text generation processing on the prompt subset through a preset model to obtain answer information; Perform paired comparison processing on the answer information to obtain the paired evaluation data set; The step of building the task quantization evaluation framework specifically includes: Perform quantization evaluation on all the evaluation objects in the paired evaluation data set through a preset item response model to obtain a target evaluation probability; Perform latent variable modeling processing on the target evaluation probability based on item response theory to obtain the task quantization evaluation framework; The latent variable modeling processing is expressed as: Among them, the evaluator believes that the evaluated object performs better than another evaluated object in the text generation task ; The said and the said respectively represent the potential capabilities of the evaluation object and ; The said is the ability gap between two said evaluation objects and ; The used for prompting discrimination, capturing its sensitivity to the differences in the abilities of the evaluation objects; The used for prompting validity, indicating the quality of the prompt; The used to represent the appraiser deviation.
2. The quantization evaluation method for generative tasks according to claim 1, characterized in that The step of constructing a quantization evaluation model based on the task quantization evaluation framework specifically includes: Extract the latent variable parameters in the task quantization evaluation framework; Import the latent variable parameters into a general model and use an evaluation training data set to train the general model to obtain a constructed quantization evaluation model.
3. The quantization evaluation method for generative tasks according to claim 2, wherein The step of performing variational inference training on the constructed quantization evaluation model to obtain a trained quantization evaluation model specifically includes: Obtain the initial value of the variational distribution parameters; Based on the initial value of the variational distribution parameters, perform variational distribution processing on the latent variable parameters to obtain an initial latent variable distribution; Perform maximum variational lower bound processing on the initial latent variable distribution, iterate and converge the variational distribution parameters to obtain an optimal latent variable estimate; Import the optimal latent variable estimate into the constructed quantization evaluation model for iterative training to obtain the trained quantization evaluation model.
4. The quantization evaluation method for generative tasks according to claim 3, wherein The step of iteratively converging the variational distribution parameters to obtain the optimal latent variable estimate specifically further includes: Iterate the variational distribution parameters by the coordinate ascent method; Calculate the optimal latent variable estimate based on the iteratively converged variational distribution parameters.
5. The quantization evaluation method for generative tasks according to claim 1, wherein The evaluation result includes task ranking, task evaluation score, quantization quality evaluation, and reliability metric. The step of inputting the paired evaluation data set into the trained quantization evaluation model to generate an evaluation result specifically includes: Input the paired evaluation data set into the trained quantization evaluation model to obtain a target latent variable value, where the target latent variable value includes a target latent ability estimate, a target discrimination estimate, a target validity estimate, and a target bias estimate; Generate the task ranking and the task evaluation score based on the target latent ability estimate; Generate the quantization quality evaluation based on the target discrimination estimate and the target validity estimate; Generate the reliability metric based on the target bias estimate.
6. A quantization evaluation device, characterized in that, When the quantization evaluation device executes, it implements the steps of the quantization evaluation method for the generative task described in any one of claims 1 to 5, including: An acquisition module, configured to acquire a paired evaluation data set; A construction module, configured to construct a task quantization evaluation framework and build a quantization evaluation model based on the task quantization evaluation framework; A training module, configured to perform variational inference training on the constructed quantization evaluation model to obtain a trained quantization evaluation model; A generation module, configured to input the evaluation data set into the trained quantization evaluation model to generate an evaluation result.
7. A computer device, characterized in that, It includes a memory and a processor. Computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, it implements the steps of the quantization evaluation method for the generative task described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by the processor, it implements the steps of the quantization evaluation method for the generative task described in any one of claims 1 to 5.
Citation Information
Patent Citations
High-speed multimode optical module performance optimization method based on variational auto-encoder
CN117354652A
Assessment method and device of large language model system and related equipment
CN119179631A