Method, device, storage medium and equipment for quality evaluation of generative large model

By generating text levels and converting them into scores, the problem of inaccurate quality assessment of generative large models is solved, achieving stable and comprehensive multi-dimensional assessment and improving the accuracy and efficiency of the assessment.

CN119005132BActive Publication Date: 2025-11-18GUANGDONG UCAP INTERNET INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410989589.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2025-11-18
Estimated Expiration
2044-07-23

AI Technical Summary

Technical Problem

Existing generative large model quality assessment methods lack a holistic semantic understanding of the assessment answers, leading to inaccurate assessments.

Method used

By obtaining triples and using a pre-trained quality assessment model, the prompts for each assessment dimension are processed to generate text levels and convert them into scores. Finally, a comprehensive quality score is calculated, and the Log-Softmax function is used to improve training efficiency and accuracy.

Benefits of technology

It enables scoring from multiple evaluation dimensions, improving the stability and accuracy of generative large model quality evaluation, simplifying the training process, and supporting dataset expansion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119005132B_ABST
    Figure CN119005132B_ABST
Patent Text Reader

Abstract

The application discloses a quality evaluation method and device of a generative large model, a storage medium and equipment, and belongs to the technical field of deep learning. A triple is acquired, the triple comprising a question, a standard answer and an answer to be evaluated generated by a generative large model; for each evaluation dimension in n evaluation dimensions, a prompt word is generated for the triple according to a prompt word template corresponding to the evaluation dimension, and n is greater than or equal to 2; a quality evaluation model is used to process the prompt word corresponding to each evaluation dimension respectively, a text level corresponding to each evaluation dimension is generated, the text level is converted into a corresponding score; n scores are comprehensively calculated to obtain a quality score of the answer to be evaluated, and the quality score is used to reflect the quality of the answer generated by the generative large model. The application can score from multiple evaluation dimensions, so that the final quality evaluation is more stable and comprehensive, and the accuracy of the quality evaluation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and in particular to a method, apparatus, storage medium and device for quality assessment of generative large models. Background Technology

[0002] With the rapid development of natural language processing technology, large language models are becoming increasingly prevalent. Large language models are generative models trained on massive datasets, possessing powerful capabilities including recognition, summarization, translation, prediction, and text generation. Generative large models have demonstrated outstanding performance on numerous tasks. However, in practical applications, evaluating the quality of generative large models is crucial.

[0003] In related technologies, computer devices can calculate evaluation metrics based on the standard answer and the answer to be evaluated, such as the Bilingual Evaluation Understudy (BLEU) metric and the Recall-Oriented Understudy for Gisting Evaluation (ROUGE) metric, and then calculate a score based on the evaluation metrics. However, this evaluation lacks an understanding of the overall semantics of the answer to be evaluated, resulting in inaccurate evaluation. Summary of the Invention

[0004] This application provides a method, apparatus, storage medium, and device for quality assessment of generative large models, to address the problem that traditional assessment metrics, when used to calculate scores, lack a comprehensive understanding of the overall semantics of the assessed answers, leading to inaccurate assessments. The technical solution is as follows:

[0005] According to a first aspect of this application, a method for quality assessment of generative large models is provided, the method comprising:

[0006] Obtain triples, wherein the triples include a question, a standard answer to the question, and an answer to be evaluated generated by a generative large model based on the question;

[0007] For each of the n evaluation dimensions, a prompt word is generated for the triple based on the prompt word template corresponding to the evaluation dimension, where n≥2;

[0008] The pre-trained quality assessment model is used to process the prompt words corresponding to each assessment dimension to generate a text level for each assessment dimension. The text level is then converted into a corresponding score. The text level is one of multiple levels that are divided into regions from good to bad and described in words.

[0009] The quality score of the answer to be evaluated is obtained by comprehensively calculating the n scores. The quality score is used to reflect the quality of the answer generated by the generative large model.

[0010] In one possible implementation, the step of processing the prompt words corresponding to each evaluation dimension using a pre-trained quality assessment model includes:

[0011] The quality assessment model is used to obtain the assessment rules for the corresponding assessment dimensions from each prompt word;

[0012] The quality assessment model is used to process the triples corresponding to the assessment dimensions according to the assessment rules.

[0013] In one possible implementation, converting the text level into a corresponding score includes:

[0014] The quality assessment model is used to obtain a preset mapping relationship, which includes the correspondence between each text level and each score;

[0015] The quality assessment model is used to find the score corresponding to the generated text level in the mapping relationship.

[0016] In one possible implementation, the method further includes:

[0017] Obtain a training dataset, wherein each training sample in the training dataset includes a triplet and n actual scores, wherein the triplet includes a question, the standard answer to the question, and the answer to be evaluated, and the n actual scores correspond to n actual text levels;

[0018] For each of the n evaluation dimensions, a prompt word is generated for the triple based on the prompt word template corresponding to the evaluation dimension;

[0019] Create a quality assessment model;

[0020] For each training sample, the quality assessment model is used to process the n prompt words corresponding to the training sample to obtain n predicted scores;

[0021] The loss is calculated for the n actual scores and n predicted scores using a preset loss function, and the quality assessment model is fine-tuned and trained based on the calculation results.

[0022] In one possible implementation, the process of using the quality assessment model to process the n prompt words corresponding to the training samples to obtain n predicted scores includes:

[0023] The quality assessment model is used to process the n prompt words corresponding to the training samples to obtain n predicted text levels corresponding to n assessment dimensions;

[0024] The n predicted text levels are converted into corresponding n predicted scores.

[0025] In one possible implementation, the loss function is Log-Softmax.

[0026] According to a second aspect of this application, a quality assessment apparatus for generative large models is provided, the apparatus comprising:

[0027] The acquisition module is used to acquire triples, wherein the triples include a question, a standard answer to the question, and an answer to be evaluated generated by a generative large model based on the question;

[0028] The generation module is used to generate a prompt word for the triplet based on the prompt word template corresponding to the evaluation dimension for each of the n evaluation dimensions, where n≥2;

[0029] The scoring module is used to process the prompt words corresponding to each evaluation dimension using a pre-trained quality assessment model, generate a text level for each evaluation dimension, and convert the text level into a corresponding score. The text level is one of multiple levels that are divided into regions from good to poor and described in words.

[0030] The evaluation module is used to perform comprehensive calculations on n scores to obtain a quality score for the answer to be evaluated. The quality score is used to reflect the quality of the answer generated by the generative large model.

[0031] In one possible implementation, the acquisition module is further configured to acquire a training dataset, wherein each training sample in the training dataset includes a triplet and n actual scores, wherein the triplet includes a question, a standard answer to the question, and an answer to be evaluated, and the n actual scores correspond to n actual text levels;

[0032] The generation module is further configured to generate a prompt word for the triplet based on the prompt word template corresponding to the evaluation dimension for each of the n evaluation dimensions;

[0033] Create a module for creating quality assessment models;

[0034] The scoring module is used to process the n prompt words corresponding to each training sample using the quality assessment model to obtain n predicted scores.

[0035] The training module is used to calculate the loss of the n actual scores and n predicted scores using a preset loss function, and to fine-tune the quality assessment model based on the calculation results.

[0036] According to a third aspect of this application, a computer-readable storage medium is provided, wherein at least one instruction is stored therein, the at least one instruction being loaded and executed by a processor to implement the quality assessment method for generative large models as described above.

[0037] According to a fourth aspect of this application, a computer device is provided, the computer device including the above-described quality assessment apparatus for generative large models.

[0038] The beneficial effects of the technical solution provided in this application include at least the following:

[0039] Fine-tuning is performed on top of a pre-trained large oracle model. Since the quality assessment model is a large language model, fine-tuning it fully utilizes the capabilities of the large oracle model, simplifying the training process and improving efficiency. Furthermore, when fine-tuning the quality assessment model using a labeled training dataset, the training dataset can be expanded according to user needs, thus enabling the expansion of the quality assessment model.

[0040] The quality assessment model can generate n scores for n assessment dimensions, and then perform comprehensive calculations on the n scores to obtain the quality score. This enables scoring from multiple assessment dimensions, making the final quality assessment more stable and comprehensive, and improving the accuracy of the quality assessment. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a flowchart illustrating a training method for a quality assessment model according to an embodiment of this application;

[0043] Figure 2 This is a flowchart of a quality assessment method for generative large models provided in one embodiment of this application;

[0044] Figure 3 This is a flowchart of a quality assessment method for generative large models provided in one embodiment of this application;

[0045] Figure 4This is a structural block diagram of a generative large model quality assessment device provided in one embodiment of this application. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0047] like Figure 1 The diagram illustrates a flowchart of a generative large model quality assessment method according to an embodiment of this application. This generative large model quality assessment method can be applied to computer devices. The generative large model quality assessment method may include:

[0048] Step 101: Obtain the training dataset. Each training sample in the training dataset includes a triplet and n actual scores. The triplet includes the question, the standard answer to the question, and the answer to be evaluated, and the n actual scores correspond to n actual text levels.

[0049] Generative large models are models used to answer input questions. To evaluate the quality of a generative large model, we can use the answer generated by the model as the answer to be evaluated, forming a triplet consisting of the question, the standard answer, and the answer to be evaluated, and then use a quality assessment model to score the quality of the answer to be evaluated. If the quality score of the answer to be evaluated is high, the corresponding generative large model is considered to be of high quality; if the quality score is low, the corresponding generative large model is considered to be of low quality. In this way, users can choose the generative large model that best suits their needs based on the quality of various generative large models.

[0050] The triples in the training samples can include any question. For example, the question might be "What is the relationship between Xiaoming and Xiaogang?", and the standard answer might be "Xiaoming and Xiaogang are the chairmen of two well-known domestic game companies." The answer to be evaluated in a generative large model might be "Xiaoming and Xiaogang are cousins."

[0051] The training samples include triples and n labeled actual scores. Here, n is the number of evaluation dimensions, and n ≥ 2. The content and number of evaluation dimensions can be set and adjusted according to actual business needs. For example, n = 2, with 2 evaluation dimensions including sentence fluency and whether the sentence contains factual errors; or n = 5, with 5 evaluation dimensions including sentence fluency, whether the sentence contains factual errors, whether the sentence is complete, whether the sentence is timely, and whether the sentence is irrelevant to the question.

[0052] Computer equipment can be configured with multiple scores for each evaluation dimension, and the number of scores can be set according to business needs. Through extensive experimentation, we found that the quality assessment model is not particularly accurate in judging numbers; it may not be able to distinguish the meaning of specific numbers 1, 2, 3, 4, and 5. This leads to some arbitrariness in the final scoring stage, such as awarding 5 points to a low-quality statement and 1 point to a high-quality statement. To address this issue, this application maps specific scores to textual levels during the training of the quality assessment model. That is, one evaluation dimension corresponds to m scores, and m scores correspond to m textual levels. For example, if m is 5, 1 point corresponds to poor, 2 points to very poor, 3 points to average, 4 points to good, and 5 points to excellent. The quality assessment model trained in this way does not provide specific numbers when scoring, but rather provides corresponding textual levels, and then converts the textual levels into corresponding scores. This effectively avoids the problem of numerical illusion in the quality assessment model, making the scoring more accurate.

[0053] For each triplet, the computer device can assign a text grade and a score to the answer to be evaluated from each evaluation dimension. Thus, a triplet can be labeled with n text grades and n actual scores. Taking the question "What is the relationship between Xiaoming and Xiaogang?" as an example, with the standard answer "Xiaoming and Xiaogang are the chairmen of two well-known domestic game companies," and the answer to be evaluated being "Xiaoming and Xiaogang are cousins," then on the evaluation dimension of "sentence fluency," the text grade can be marked as excellent, and the actual score is 5 points; on the evaluation dimension of "whether there are factual errors in the statement," the text grade can be marked as poor, and the actual score is 1 point.

[0054] In this embodiment, assuming n is 5, a training sample may include triples and the following:

[0055] 1. Assessment Dimension 1, Text Level 1, Actual Score 1;

[0056] 2. Assessment Dimension 2, Text Level 2, Actual Score 2;

[0057] 3. Evaluation dimensions: 3; Text level: 3; Actual score: 3;

[0058] 4. Assessment dimensions: 4; Text level: 4; Actual score: 4.

[0059] 5. Evaluation dimensions: 5; text level: 5; actual score: 5.

[0060] Step 102: For each of the n evaluation dimensions, generate a prompt word for the triple based on the prompt word template corresponding to the evaluation dimension.

[0061] Each evaluation dimension corresponds to a prompt word template, indicating which dimension the quality assessment model should score the answer to be evaluated. When there are n evaluation dimensions, the computer device can generate n prompt words.

[0062] Taking the assessment dimension of "whether there are factual errors in the statement" as an example, the generated prompt could be: Based on the question and standard answer 1, evaluate the quality score of the answer to be evaluated, 2, from the perspective of "whether there are factual errors in the statement." Each item's score should be an integer from 0 to 5 points. The scoring criteria are referenced in: {Scoring Criteria for Factual Errors}. Below are the question, standard answer 1, and answer to be evaluated, 2. Question: {question_i}\nStandard Answer 1: {answer_base_i}\nAnswer to be evaluated: {answer_llm_i}.\nPlease provide the score in the format: x points.

[0063] Step 103: Create a quality assessment model.

[0064] The quality assessment model is a large language model that has been pre-trained; pre-training is unsupervised training.

[0065] Step 104: For each training sample, use the quality assessment model to process the n prompt words corresponding to the training sample to obtain n prediction scores.

[0066] Each training sample corresponds to n prompt words. The quality assessment model generates a predicted score for each prompt word, and finally obtains n predicted scores.

[0067] Specifically, processing the n prompt words corresponding to the training samples using a quality assessment model to obtain n predicted scores can include: processing the n prompt words corresponding to the training samples using a quality assessment model to obtain n predicted text levels corresponding to n assessment dimensions; and converting the n predicted text levels into corresponding n predicted scores.

[0068] For example, if the predicted text generated by the quality assessment model is rated as excellent in the "sentence fluency" evaluation dimension, then its corresponding prediction score is determined to be 5 points.

[0069] Step 105: Calculate the loss for n actual scores and n predicted scores using a preset loss function, and fine-tune the quality assessment model based on the calculation results.

[0070] The raw output of large language models (typically unprocessed scores or log odds) needs to be transformed into an interpretable probability distribution. This transformation is usually achieved using the softmax loss function, which converts any real-valued vector into a probability distribution such that each element's value is between (0,1) and the sum of all elements is 1. However, when dealing with large-scale models and high-dimensional outputs, directly using softmax can lead to numerical instability and computational inefficiency. To address this issue, we set the loss function for the quality assessment model to Log-Softmax. The Log-Softmax function provides an efficient method for directly calculating the logarithm of probabilities, thereby enhancing the numerical stability and computational efficiency of large language models when handling large-scale data.

[0071] The Log-Softmax function is defined as follows:

[0072]

[0073] Among them, z i It is the i-th element of the input vector, and K is the dimension of the vector.

[0074] like Figure 2 The diagram illustrates a flowchart of a generative large model quality assessment method according to an embodiment of this application. This generative large model quality assessment method can be applied to computer devices. The generative large model quality assessment method may include:

[0075] Step 201: Obtain the triplet, which includes the question, the standard answer to the question, and the answer to be evaluated generated by the generative big model based on the question.

[0076] The triple can include any question. For example, the question is "What is the relationship between Xiaoming and Xiaogang?", the standard answer is "Xiaoming and Xiaogang are the chairmen of two well-known domestic game companies.", and the answer to be evaluated in a generative large model is "Xiaoming and Xiaogang are cousins."

[0077] Step 202: For each of the n evaluation dimensions, generate a prompt word for the triple based on the prompt word template corresponding to that evaluation dimension, where n≥2.

[0078] Each evaluation dimension corresponds to a prompt word template, indicating which dimension the quality assessment model should score the answer to be evaluated. When there are n evaluation dimensions, the computer device can generate n prompt words.

[0079] Taking the assessment dimension of "whether there are factual errors in the statement" as an example, the generated prompt could be: Based on the question and standard answer 1, evaluate the quality score of the answer to be evaluated, 2, from the perspective of "whether there are factual errors in the statement." Each item's score should be an integer from 0 to 5 points. The scoring criteria are referenced in: {Scoring Criteria for Factual Errors}. Below are the question, standard answer 1, and answer to be evaluated, 2. Question: {question_i}\nStandard Answer 1: {answer_base_i}\nAnswer to be evaluated: {answer_llm_i}.\nPlease provide the score in the format: x points.

[0080] Step 203: Use the pre-trained quality assessment model to process the prompt words corresponding to each assessment dimension, generate a text level for each assessment dimension, and convert the text level into a corresponding score. The text level is one of multiple levels that are divided into regions from good to bad and described in words.

[0081] The quality assessment model in this embodiment can be adopted. Figure 1 The method shown is used for training.

[0082] For example, if the predicted text generated by the quality assessment model is rated as excellent in the "sentence fluency" evaluation dimension, then its corresponding prediction score is determined to be 5 points.

[0083] Step 204: Perform a comprehensive calculation on the n scores to obtain the quality score of the answer to be evaluated. This quality score is used to reflect the quality of the answer generated by the generative large model.

[0084] The comprehensive calculation can be a weighted calculation or other calculations; no restrictions are made here.

[0085] In this embodiment, the quality score is positively correlated with the quality of the generative large model; that is, the higher the quality score, the better the quality of the generative large model; the lower the quality score, the worse the quality of the generative large model.

[0086] In summary, the generative large model quality assessment method provided in this application involves fine-tuning based on a pre-trained large oracle model. Since the quality assessment model is a large language model, fine-tuning it fully utilizes the capabilities of the large oracle model, simplifies the training process, and improves training efficiency. Furthermore, when fine-tuning the quality assessment model using a labeled training dataset, the training dataset can be expanded according to user needs, thereby enabling the expansion of the quality assessment model.

[0087] The quality assessment model can generate n scores for n assessment dimensions, and then perform comprehensive calculations on the n scores to obtain the quality score. This enables scoring from multiple assessment dimensions, making the final quality assessment more stable and comprehensive, and improving the accuracy of the quality assessment.

[0088] like Figure 3 The diagram illustrates a flowchart of a generative large model quality assessment method according to an embodiment of this application. This generative large model quality assessment method can be applied to a computer device. The generative large model quality assessment method may include:

[0089] Step 301: Obtain the triplet, which includes the question, the standard answer to the question, and the answer to be evaluated generated by the generative big model based on the question.

[0090] The triple can include any question. For example, the question is "What is the relationship between Xiaoming and Xiaogang?", the standard answer is "Xiaoming and Xiaogang are the chairmen of two well-known domestic game companies.", and the answer to be evaluated in a generative large model is "Xiaoming and Xiaogang are cousins."

[0091] Step 302: For each of the n evaluation dimensions, generate a prompt word for the triple based on the prompt word template corresponding to that evaluation dimension, where n≥2.

[0092] Each evaluation dimension corresponds to a prompt word template, indicating which dimension the quality assessment model should score the answer to be evaluated. When there are n evaluation dimensions, the computer device can generate n prompt words.

[0093] Taking the assessment dimension of "whether there are factual errors in the statement" as an example, the generated prompt could be: Based on the question and standard answer 1, evaluate the quality score of the answer to be evaluated, 2, from the perspective of "whether there are factual errors in the statement." Each item's score should be an integer from 0 to 5 points. The scoring criteria are referenced in: {Scoring Criteria for Factual Errors}. Below are the question, standard answer 1, and answer to be evaluated, 2. Question: {question_i}\nStandard Answer 1: {answer_base_i}\nAnswer to be evaluated: {answer_llm_i}.\nPlease provide the score in the format: x points.

[0094] Step 303: Use the quality assessment model to obtain the assessment rules for the corresponding assessment dimensions from each prompt word; use the quality assessment model to process the triples corresponding to the assessment dimensions according to the assessment rules to generate a text level for each assessment dimension, and convert the text level into a corresponding score. The text level is one of multiple levels that are divided into regions from good to bad and described in words.

[0095] One evaluation dimension corresponds to m scores, and m scores correspond to m text levels. For example, if m is 5, 1 point corresponds to poor, 2 points to poor, 3 points to average, 4 points to good, and 5 points to excellent.

[0096] The evaluation rules are the scoring criteria references mentioned in step 302. The quality assessment model can process triples according to the scoring criteria corresponding to each evaluation dimension to obtain n text levels corresponding to n evaluation criteria. Then, the quality assessment model needs to convert the text levels into scores.

[0097] Specifically, converting text levels into corresponding scores can include: using a quality assessment model to obtain a preset mapping relationship, which includes the correspondence between each text level and each score; and using the quality assessment model to find the score corresponding to the generated text level in the mapping relationship.

[0098] For example, if the predicted text generated by the quality assessment model is rated as excellent in the "sentence fluency" evaluation dimension, then its corresponding prediction score is determined to be 5 points.

[0099] Step 304: Perform a comprehensive calculation on the n scores to obtain the quality score of the answer to be evaluated. This quality score is used to reflect the quality of the answer generated by the generative large model.

[0100] The comprehensive calculation can be a weighted calculation or other calculations; no restrictions are made here.

[0101] In this embodiment, the quality score is positively correlated with the quality of the generative large model; that is, the higher the quality score, the better the quality of the generative large model; the lower the quality score, the worse the quality of the generative large model.

[0102] In summary, the generative large model quality assessment method provided in this application involves fine-tuning based on a pre-trained large oracle model. Since the quality assessment model is a large language model, fine-tuning it fully utilizes the capabilities of the large oracle model, simplifies the training process, and improves training efficiency. Furthermore, when fine-tuning the quality assessment model using a labeled training dataset, the training dataset can be expanded according to user needs, thereby enabling the expansion of the quality assessment model.

[0103] The quality assessment model can generate n scores for n assessment dimensions, and then perform comprehensive calculations on the n scores to obtain the quality score. This enables scoring from multiple assessment dimensions, making the final quality assessment more stable and comprehensive, and improving the accuracy of the quality assessment.

[0104] like Figure 4The diagram illustrates a structural block diagram of a generative large model quality assessment apparatus according to an embodiment of this application. This generative large model quality assessment apparatus can be applied to a computer device. The generative large model quality assessment apparatus may include:

[0105] Module 410 is used to obtain triples, which include a question, the standard answer to the question, and the answer to be evaluated generated by the generative big model based on the question.

[0106] The generation module 420 is used to generate a prompt word for each triplet based on the prompt word template corresponding to the evaluation dimension for each of the n evaluation dimensions, where n≥2;

[0107] The scoring module 430 is used to process the prompt words corresponding to each evaluation dimension using a pre-trained quality assessment model, generate a text level corresponding to each evaluation dimension, and convert the text level into a corresponding score. The text level is one of multiple levels that are divided into regions from good to bad and described in words.

[0108] Evaluation module 440 is used to perform comprehensive calculations on n scores to obtain a quality score for the answer to be evaluated. The quality score is used to reflect the quality of the answer generated by the generative large model.

[0109] In an optional embodiment, the scoring module 430 is further configured to:

[0110] The evaluation rules for the corresponding evaluation dimensions are obtained from each prompt word using a quality assessment model.

[0111] The quality assessment model is used to process the triples corresponding to the assessment dimensions according to the assessment rules.

[0112] In an optional embodiment, the scoring module 430 is further configured to:

[0113] The quality assessment model is used to obtain a pre-defined mapping relationship, which includes the correspondence between each text level and each score.

[0114] The quality assessment model is used to find the score corresponding to the generated text level in the mapping relationship.

[0115] In an optional embodiment, the acquisition module 410 is further configured to acquire a training dataset, wherein each training sample in the training dataset includes a triplet and n actual scores, wherein the triplet includes a question, a standard answer to the question, and an answer to be evaluated, and the n actual scores correspond to n actual text levels;

[0116] The generation module 420 is also used to generate a prompt word for the triplet based on the prompt word template corresponding to the evaluation dimension for each of the n evaluation dimensions.

[0117] Create a module for creating quality assessment models;

[0118] The scoring module 430 is used to process the n prompt words corresponding to each training sample using a quality assessment model to obtain n predicted scores.

[0119] The training module is used to calculate the loss of n actual scores and n predicted scores using a preset loss function, and to fine-tune the quality assessment model based on the calculation results.

[0120] In an optional embodiment, the scoring module 430 is further configured to:

[0121] The quality assessment model is used to process the n prompt words corresponding to the training samples to obtain n predicted text levels corresponding to n assessment dimensions;

[0122] Convert the n predicted text levels into the corresponding n predicted scores.

[0123] In an optional embodiment, the loss function is Log-Softmax.

[0124] In summary, the generative large model quality assessment device provided in this application provides fine-tuning based on a pre-trained large oracle model. Since the quality assessment model is a large language model, fine-tuning it fully utilizes the capabilities of the large oracle model, simplifies the training process, and improves training efficiency. Furthermore, when fine-tuning the quality assessment model using a labeled training dataset, the training dataset can be expanded according to user needs, thereby enabling the expansion of the quality assessment model.

[0125] The quality assessment model can generate n scores for n assessment dimensions, and then perform comprehensive calculations on the n scores to obtain the quality score. This enables scoring from multiple assessment dimensions, making the final quality assessment more stable and comprehensive, and improving the accuracy of the quality assessment.

[0126] One embodiment of this application provides a computer-readable storage medium storing at least one instruction that is loaded and executed by a processor to implement the quality assessment method for generative large models as described above.

[0127] One embodiment of this application provides a computer device that includes the above-described quality assessment apparatus for arbitrary generative large models.

[0128] It should be noted that the generative large model quality assessment device provided in the above embodiments is only illustrated by the division of the above functional modules when performing quality assessment of generative large models. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the generative large model quality assessment device can be divided into different functional modules to complete all or part of the functions described above. In addition, the generative large model quality assessment device and the generative large model quality assessment method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0129] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0130] The above description is not intended to limit the embodiments of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of the embodiments of this application.

Claims

1. A method for quality assessment of generative large models, characterized in that, The method includes: Obtain triples, wherein the triples include a question, a standard answer to the question, and an answer to be evaluated generated by a generative large model based on the question; For each of the n evaluation dimensions, a prompt word is generated for the triple based on the prompt word template corresponding to the evaluation dimension, where n≥2; The pre-trained quality assessment model is used to process the prompt words corresponding to each assessment dimension to generate a text level for each assessment dimension. The text level is then converted into a corresponding score. The text level is one of multiple levels that are divided into regions from good to bad and described in words. A quality score is obtained by comprehensively calculating the n scores, and the quality score is used to reflect the quality of the answer generated by the generative large model. The step of processing the prompt words corresponding to each evaluation dimension using a pre-trained quality assessment model includes: using the quality assessment model to obtain the evaluation rules for the corresponding evaluation dimension from each prompt word; and using the quality assessment model to process the triples corresponding to the evaluation dimension according to the evaluation rules.

2. The quality assessment method for generative large models according to claim 1, characterized in that, The process of converting the text level into a corresponding score includes: The quality assessment model is used to obtain a preset mapping relationship, which includes the correspondence between each text level and each score; The quality assessment model is used to find the score corresponding to the generated text level in the mapping relationship.

3. The method for quality assessment of generative large models according to claim 1 or 2, characterized in that, The method further includes: Obtain a training dataset, wherein each training sample in the training dataset includes a triplet and n actual scores, wherein the triplet includes a question, the standard answer to the question, and the answer to be evaluated, and the n actual scores correspond to n actual text levels; For each of the n evaluation dimensions, a prompt word is generated for the triple based on the prompt word template corresponding to the evaluation dimension; Create a quality assessment model; For each training sample, the quality assessment model is used to process the n prompt words corresponding to the training sample to obtain n predicted scores; The loss is calculated for the n actual scores and n predicted scores using a preset loss function, and the quality assessment model is fine-tuned and trained based on the calculation results.

4. The quality assessment method for generative large models according to claim 3, characterized in that, The quality assessment model is used to process the n prompt words corresponding to the training samples to obtain n predicted scores, including: The quality assessment model is used to process the n prompt words corresponding to the training samples to obtain n predicted text levels corresponding to n assessment dimensions; The n predicted text levels are converted into corresponding n predicted scores.

5. The quality assessment method for generative large models according to claim 3, characterized in that, The loss function is Log-Softmax.

6. A quality assessment device for generative large models, characterized in that, The device includes: The acquisition module is used to acquire triples, wherein the triples include a question, a standard answer to the question, and an answer to be evaluated generated by a generative large model based on the question; The generation module is used to generate a prompt word for the triplet based on the prompt word template corresponding to the evaluation dimension for each of the n evaluation dimensions, where n≥2; The scoring module is used to process the prompt words corresponding to each evaluation dimension using a pre-trained quality assessment model, generate a text level for each evaluation dimension, and convert the text level into a corresponding score. The text level is one of multiple levels that are divided into regions from good to poor and described in words. The evaluation module is used to perform a comprehensive calculation on the n scores to obtain the quality score of the answer to be evaluated. The quality score is used to reflect the quality of the answer generated by the generative large model. The scoring module is further configured to: obtain the evaluation rules for the corresponding evaluation dimension from each prompt word using the quality assessment model; and process the triples corresponding to the evaluation dimension using the quality assessment model according to the evaluation rules.

7. The quality assessment device for generative large models according to claim 6, characterized in that, The acquisition module is further configured to acquire a training dataset, wherein each training sample in the training dataset includes a triplet and n actual scores, wherein the triplet includes a question, a standard answer to the question, and an answer to be evaluated, and the n actual scores correspond to n actual text levels; The generation module is further configured to generate a prompt word for the triplet based on the prompt word template corresponding to the evaluation dimension for each of the n evaluation dimensions; Create a module for creating quality assessment models; The scoring module is used to process the n prompt words corresponding to each training sample using the quality assessment model to obtain n predicted scores. The training module is used to calculate the loss of the n actual scores and n predicted scores using a preset loss function, and to fine-tune the quality assessment model based on the calculation results.

8. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, which is loaded and executed by a processor to implement the quality assessment method for generative large models as described in any one of claims 1 to 5.

9. A computer device, characterized in that, The computer equipment includes: the quality assessment device for generative large models as described in claim 6 or 7.

Citation Information

Patent Citations

  • Content and form diversity fused Chinese question generation method and system

    CN114970563A

  • Large-model-oriented multi-dimensional question and answer pair generation task evaluation method and system

    CN118093344A