Selection method of vertical subjective question scoring model and scoring method of vertical subjective questions
By constructing multiple scoring prompt templates and weighting the process, the problem of unreasonable selection of scoring models for subjective questions in vertical domains was solved, and multi-angle scoring that is more consistent with expert evaluation was achieved, thereby improving the accuracy and fairness of the scoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD
- Filing Date
- 2024-08-08
- Publication Date
- 2026-05-05
AI Technical Summary
The existing methods for selecting scoring models for subjective questions in vertical domains have low correlation with human evaluation results, leading to unreasonable scoring results.
Multiple scoring prompt templates are constructed, and the prompt angles and scoring dimensions of each template are determined based on expert experience. Single-angle scoring is performed through a large language model, and comprehensive processing is carried out based on the sorting order and weighted weights to obtain multi-angle scores.
This improved the consistency between the vertical domain subjective question scoring model and real expert evaluations, enhancing the accuracy and fairness of the scoring.
Smart Images

Figure CN118797031B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of text processing technology, specifically to a method for selecting a vertical domain subjective question scoring model and a scoring method for vertical domain subjective questions. Background Technology
[0002] To reduce the impact of expert evaluation bias on subjective questions in specific vertical domains and improve the fairness, accuracy and efficiency of scoring, related technical fields have proposed a scheme to use a large language model to score subjective questions in specific vertical domains.
[0003] The premise of using large language models for scoring vertical-domain subjective questions is to reasonably evaluate the merits of large language models based on expert scoring criteria and select those with better scoring quality. However, existing model selection and evaluation methods show low correlation between scoring results and human evaluation results when compared, indicating the irrationality of the reverse selection evaluation method. Summary of the Invention
[0004] To address the problem of unreasonable selection of existing scoring models, this disclosure provides a new method for selecting a scoring model for vertical subjective questions and a scoring method for vertical subjective questions.
[0005] In a first aspect, embodiments of this disclosure provide a method for selecting a scoring model for subjective questions in a vertical domain, including:
[0006] Multiple scoring prompt templates for subjective questions in a vertical domain are constructed, and model inputs are built based on each scoring prompt template, the subjective questions in the vertical domain, and the answers to the subjective questions in the vertical domain. Each scoring prompt template provides a prompt for scoring the answer from a specific prompting perspective. The prompting perspective of each scoring prompt template is determined based on the experience of experts in the vertical domain, and the prompting perspectives of each scoring prompt template are different. Each model input includes a scoring prompt template.
[0007] The input of the model is processed by the large language model to be selected, and the single-angle score of the large language model to be selected for the answer under each scoring prompt template is obtained.
[0008] The single-angle scores obtained from the model inputs of the large language model to be selected, including the same scoring prompt template, are sorted by size to determine the sorting order, and the weighted weight of each single-angle score corresponding to the same scoring prompt template is determined based on the sorting order.
[0009] According to the weighted weights, the single-angle scores output by the candidate large language models are weighted and summed to obtain the multi-angle scores of each candidate large language model.
[0010] Based on the correlation between single-angle scoring and the quality of answers to subjective questions in the vertical domain, a preset number of large language models with the largest or smallest multi-angle scoring are selected as the scoring models for subjective questions in the vertical domain.
[0011] Optionally, the construction of multiple scoring prompt templates for subjective questions in a vertical domain includes:
[0012] Based on expert experience, we determine multiple prompting angles for subjective questions in each vertical domain, scoring dimensions under each prompting angle, and scoring criteria under each scoring dimension.
[0013] Based on the aforementioned prompt angle, the corresponding rating dimension, and the rating criteria under the rating dimension, the multiple rating prompt templates are constructed.
[0014] Optionally, the determination of multiple prompting angles for subjective questions in a specific vertical domain based on expert experience includes:
[0015] The multiple prompting angles determined based on expert experience include at least two of the following angles: holistic angle, accuracy angle, and practicality angle.
[0016] Optionally, it also includes: obtaining reference answers for the subjective questions in the said vertical domain;
[0017] The process of constructing the multiple scoring prompt templates also includes adding the reference answer to each scoring prompt template.
[0018] Optionally, the number of subjective questions in the vertical domain is at least two;
[0019] The model input based on each scoring prompt template, the vertical subjective question, and the answer to the vertical subjective question includes:
[0020] The model input is constructed based on each scoring prompt template, each vertical subjective question, and the answer to each vertical subjective question; or, the model input is constructed based on each scoring prompt template, all vertical subjective questions, and the answer to each vertical subjective question, such that the model input includes all vertical subjective questions and their corresponding answers, as well as the relationship between all vertical subjective questions and their corresponding answers.
[0021] The process of using a large language model to process the model input yields a single-angle score for the answer under each scoring prompt template, including:
[0022] The model output is processed using the large language model to be selected, and the individual scores of the model to be selected for each vertical subjective question and corresponding answer are obtained under each scoring prompt template.
[0023] The average or sum of all individual scores obtained by each candidate large language model under the corresponding scoring prompt template is used as the corresponding single-angle score.
[0024] Optionally, determining the weighted weights of each single-angle score corresponding to the same model input based on the sorting order includes:
[0025] Based on the sorting order, determine the weight amplification factor or weight reduction factor for each single-angle score corresponding to the same model input;
[0026] Based on the weight amplification factor or weight reduction factor, and the baseline weight of each of the scoring prompt templates, the weighted weight of the single-angle score corresponding to the sorting order is determined.
[0027] Optionally, before determining the weighted weights of each single-angle score corresponding to the same model input based on the sorting order, the following steps are included:
[0028] The relative importance of each scoring prompt template was determined based on expert experience, and a comparison matrix was constructed based on the relative importance.
[0029] The comparison matrix is column normalized to obtain a normalized matrix;
[0030] The average value of the matrix elements corresponding to each rating prompt template in the normalized matrix is calculated and used as the weighting weight of the corresponding single-angle rating.
[0031] Secondly, embodiments of this disclosure provide a scoring method for subjective questions within a vertical domain, including:
[0032] The model input is constructed based on multiple pre-built scoring prompt templates, vertical subjective questions, and answers to the vertical subjective questions to be scored; each of the scoring prompt templates provides a prompt for scoring the answer from a specific perspective, and the prompt perspective of each scoring prompt template is determined based on the experience of vertical experts, and the prompt perspectives of each scoring prompt template are different;
[0033] The model inputs are processed by a pre-selected vertical subjective question scoring model to obtain the corresponding single-angle scores;
[0034] A multi-angle score is obtained by weighting and summing the single-angle score and the corresponding weight, and the multi-angle score is used as the score for the answer to be scored.
[0035] Thirdly, embodiments of this disclosure provide a computing device, including a processor and a memory, the memory being used to store a computer program; when the computer program is loaded by the processor, it causes the processor to execute the selection method for the vertical subjective question scoring model as described above and / or the scoring method for the vertical subjective questions as described above.
[0036] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the selection method for the scoring model of vertical subjective questions as described above and / or the scoring method for vertical subjective questions as described above.
[0037] This disclosed embodiment is based on the idea that experts in a specific field subconsciously evaluate answers to subjective questions from different angles, and the resulting score is obtained by comprehensively considering these angles. First, scoring prompt templates with different scoring angles are set for subjective questions in the specific field. These templates are used to prompt the candidate large language models to score the subjective question answers, resulting in single-angle scores. Then, guided by the idea that experts in the specific field will comprehensively process these different angle scores with different weights to obtain the final score, the weighting of each single-angle score is determined according to the ranking order of the scores corresponding to the same scoring prompt template. Finally, the weighted sum of the single-angle scores and the aforementioned weighted weights is used to obtain a multi-angle score. The multi-angle score obtained based on this approach better aligns with the scoring strategies of experts in the specific field for subjective questions. Consequently, the multi-angle scores of the candidate large language models have better comparative value, and the determined scoring model for subjective questions in the specific field is more consistent with the evaluation experience of real experts.
[0038] The subjective question scoring model for the vertical domain selected using the scheme of this embodiment is more in line with real expert experience, and it is more reasonable to use it as a real application model. Attached Figure Description
[0039] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0040] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without any creative effort, wherein...
[0041] Figure 1 This is a flowchart of the method for selecting a scoring model for vertical subjective questions provided in this embodiment of the disclosure;
[0042] Figure 2This is a flowchart of the scoring method for subjective questions in a vertical domain provided in this embodiment of the disclosure;
[0043] Figure 3 This is a schematic diagram of the selection device structure for the vertical subjective question scoring model provided in this public implementation;
[0044] Figure 4 This is a schematic diagram of the structure of the scoring device for subjective questions in the vertical domain provided in this embodiment of the disclosure;
[0045] Figure 5 This is a schematic diagram of the structure of a computing device provided in an embodiment of this disclosure. Detailed Implementation
[0046] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0047] The term "comprising" and its variations as used herein are open-ended inclusion, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. In this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations.
[0048] To address the problem that existing model selection and evaluation methods do not yield consistent results with expert subjective evaluations, leading to unreasonable selection of scoring models for vertical subjective questions, this disclosure provides a new method for selecting scoring models for vertical subjective questions. The method for selecting scoring models for vertical subjective questions provided in this disclosure is executed by a computing device.
[0049] Figure 1 This is a flowchart illustrating the method for selecting a vertical domain subjective question scoring model provided in this embodiment of the disclosure. Figure 1 As shown, the method for selecting the scoring model for vertical subjective questions provided in this embodiment includes steps S110-S150.
[0050] S110: Construct multiple scoring prompt templates for vertical subjective questions, and construct model input based on each scoring prompt template, the vertical subjective question, and the answer to the vertical subjective question.
[0051] Scoring prompt templates are used to instruct or prompt the model on the specific rules to score the answers to subjective questions in a vertical domain. In this embodiment of the disclosure, each scoring prompt template provides a specific prompting angle for scoring the model, and the prompting angles of each scoring prompt template are different. That is to say, the prompting angles corresponding to different scoring prompt templates are not the same.
[0052] The suggestion angles in each of the aforementioned scoring suggestion templates are determined based on the experience of vertical domain experts. Specifically, they are determined in advance based on the vertical domain in which the subjective questions are located and the experience of relevant experts within that domain. The aforementioned determination of scoring suggestion templates based on the experience of vertical domain experts considers the scoring angles that vertical domain experts would use to score answers to subjective questions in that domain, and uses these scoring angles as the suggestion angles adopted in the scoring suggestion templates. These scoring suggestion templates are used to instruct the model to score answers according to the scoring perspective of vertical domain experts.
[0053] For example, in certain vertical fields, experts will score answers from the perspectives of overall quality, accuracy, and practicality. Correspondingly, scoring prompt templates can be constructed from the perspectives of overall quality, accuracy, and practicality.
[0054] In some embodiments, constructing multiple scoring prompt templates for vertical subjective questions includes S111-S112.
[0055] S111: Based on expert experience, determine multiple prompting angles for subjective questions in the vertical domain, the scoring dimensions under each prompting angle, and the scoring criteria under each scoring dimension.
[0056] As analyzed earlier, the ultimate goal of the prompting angle is to score the answers. Accordingly, after determining the prompting angle for subjective questions in a specific vertical domain, corresponding scoring dimensions need to be set for each prompting angle. The scoring dimension refers to the dimension on which the answer is evaluated under the corresponding prompting angle, while the scoring criterion is the standard for determining what score the answer can receive under what circumstances under the corresponding dimension.
[0057] In practice, the scoring dimensions for the overall quality mentioned above include: grammatical level, security, logical reasoning ability, accuracy, comprehensiveness of content, and operability; the scoring dimensions for accuracy include: accuracy of terminology explanation, accuracy of content, and accuracy of factual conformity; and the scoring dimensions for usability include: problem-solving ability, depth of content, and comprehensiveness of content.
[0058] In practice, the corresponding scoring criteria for each scoring dimension of the aforementioned overall quality perspective can be as follows.
[0059] Grammar level: scores of 0, 1, 2, and 3; (1) If the answer is in the wrong order of cause and effect, missing necessary sentence structure, or incorrect word choice leading to misunderstanding, it is recorded as 0; (2) If the answer is incomplete sentence structure, inaccurate word choice, or unreasonable sentence division but does not affect understanding, it is recorded as 1; (3) If the answer is incomplete sentence structure, inaccurate word choice, or unreasonable sentence division but does not affect understanding, it is recorded as 2; (4) If the answer is complete sentence structure, accurate word choice, and fluent and smooth sentence, it is recorded as 3.
[0060] Security: The score can only be 0 or 3 points; (1). If the answer involves any of the following, such as illegal, irregular, pornographic, violent, or malicious links, the score is 0; if the answer does not involve any of these, the score is 3.
[0061] Logical reasoning ability: scored as 0, 1, 2, or 3 points. A score of 0 is given for no logical reasoning ability; 1 for low logical reasoning ability; 2 for basic logical reasoning ability; and 3 for strong logical reasoning ability.
[0062] Accuracy: Score is 0, 1, 2, or 3 points. Accuracy is evaluated based on whether the answer addresses the question, whether proper nouns are correctly explained, and whether the content is accurate and consistent with the facts. A score of 0 is awarded for a very inaccurate answer; 1 for an inaccurate answer; 2 for a mostly accurate answer; and 3 for a very accurate answer.
[0063] Comprehensiveness of Content: Score is 0, 1, 2, or 3 points. Comprehensiveness is evaluated based on whether all key concepts in the question are discussed in detail and whether relevant conclusions are accurately expanded upon. A score of 0 is awarded for a very incomplete answer; 1 for an incomplete answer; 2 for a fairly comprehensive answer; and 3 for a very comprehensive answer.
[0064] Operability: Valued at 0, 1, 2, or 3 points. Operability is evaluated based on the actual operability and usability of the answer. A score of 0 is given if the answer is very impractical; 1 if the answer is impractical; 2 if the answer is basically practical; and 3 if the answer is very practical.
[0065] For each of the aforementioned accuracy-related scoring dimensions, the corresponding scoring criteria can be as follows.
[0066] Accuracy of noun explanation: The score is 0, 1 or 2 points for judging whether the answer explains the relevant noun; (1) If the answer does not explain the relevant proper noun correctly, it is recorded as 0; (2) If the answer explains the relevant proper noun relatively correctly, it is recorded as 1; (3) If the answer explains the relevant proper noun correctly, it is recorded as 2.
[0067] Content accuracy: The score is 0, 1, or 2 points to determine whether the answer is accurate. (1) If the answer is obviously wrong, it is recorded as 0. (2) If the answer is relatively accurate, it is recorded as 1. (3) If the answer is correct, it is recorded as 2.
[0068] Factual accuracy: The score is 0, 1, or 2 points to determine whether the answer is consistent with the facts. (1) If the answer is seriously inconsistent with the facts, it is recorded as 0. (2) If the answer is relatively consistent with the facts, it is recorded as 1. (3) If the answer is completely consistent with the facts, it is recorded as 2.
[0069] From the aforementioned practical perspective, the corresponding standards can be as follows.
[0070] Problem Solving Degree: Determine whether the answer solves the problem raised, with a score of 0, 1, or 2; (1) If the answer does not solve the problem, it is recorded as 0; (2) If the answer partially solves the problem, it is recorded as 1; (3) If the answer completely solves the problem, it is recorded as 2.
[0071] Content depth: The score is 0, 1, or 2 points to determine whether the answer is in-depth. (1) If the answer is only general, it is 0 points. (2) If the answer is relatively in-depth, it is 1 points. (3) If the answer is very in-depth, it is 2 points.
[0072] Comprehensiveness: Determine whether the answer covers all the question keywords and provides an accurate expanded description of all the question keywords. The score is 0, 1, or 2 points. (1) If the answer is not comprehensive, it is recorded as 0. (2) If the answer is relatively comprehensive, it is recorded as 1. (3) If the answer is very comprehensive, it is recorded as 2.
[0073] S112: Based on the prompt angle, the corresponding rating dimension, and the rating criteria under the rating dimension, construct multiple rating prompt templates.
[0074] After obtaining the aforementioned prompts, scoring dimensions, and scoring criteria under the scoring dimensions, the prompts, scoring dimensions, and scoring criteria can be assembled and combined hierarchically to form a scoring prompt template.
[0075] In practice, in addition to the three levels of information mentioned above, the scoring prompt template can also include other restrictive prompts. For example, the following restrictive prompts are included: (1) First, fully understand the meaning of the question and answer, and try not to have redundant output; (2) Output the results directly in JSON format without textual explanation; (3) The evaluation of the answer dimension needs to be faithful to the original text; (4) For each dimension evaluation, only integer scores can be output, which is very important.
[0076] After obtaining the aforementioned scoring prompt templates, the computing device can then input the various scoring prompt templates, the vertical subjective questions, and the answers to the vertical subjective questions into a model. In specific implementations, each model input can only include one scoring prompt template, which may include one vertical subjective question and one answer to that question, or it may include multiple vertical subjective questions and answers to each of those questions. However, if a scoring prompt template includes multiple vertical subjective questions, the vertical subjective questions and their corresponding answers must be in one-to-one correspondence to avoid scoring errors.
[0077] In practice, when there are at least two subjective questions in each vertical domain, the computing device can construct the model input in two ways. One method is to construct the model input separately based on each scoring prompt template, each subjective question in each vertical domain, and the answer to each subjective question, so that each model input includes only one subjective question in the vertical domain and its corresponding answer. The other method is to construct the model input based on each scoring prompt template, all subjective questions in the vertical domain, and the answer to each subjective question, so that the model input includes all subjective questions in the vertical domain and their corresponding answers; however, in this case, the model input should also include the relationship between all subjective questions in the vertical domain and their corresponding answers.
[0078] In some embodiments, in order to enable the model to evaluate the answers more accurately, the computing device can also obtain reference answers for subjective questions in each vertical domain and add the reference answers to each scoring prompt template, using the reference answers as a reference for evaluating the answers.
[0079] S120: The input model is processed using the large language model to be selected, and the single-angle score of the large language model to be selected for the answer under each scoring prompt template is obtained.
[0080] The large language model to be selected is a pre-trained model that can score answers relatively accurately according to the scoring prompt template. In practical applications, the large language model to be selected can be trained by a computing device or other devices. In this embodiment, the number of large models to be selected is at least two.
[0081] The method of using a large language model to process model input involves inputting each model input into the large language model to be selected, so that each large language model to be selected processes its own model input and obtains a single-angle score for the corresponding answer under the corresponding scoring prompt template.
[0082] Following the previous approach, when the computing device processes the model input using the large language model to be selected, it determines the score under the corresponding scoring dimension according to the corresponding scoring criteria, and then adds the aforementioned scores to obtain a single-angle score for a vertical subjective question.
[0083] In practice, to improve the accuracy of the evaluation, the computing device may set up model inputs that include subjective questions and answers from different vertical domains but with the same scoring prompt template, or it may set up multiple subjective questions from different vertical domains and their corresponding difficulty levels in a single model input. In this case, the computing device can determine the single-angle score according to S121-S122 as follows.
[0084] S121: The output of the model is processed using the large language model to be selected, and the individual scores of the model to be selected for each vertical subjective question and the corresponding answer are obtained under each scoring prompt template.
[0085] S122: Calculate the average or sum of all individual scores obtained by each candidate large language model under the corresponding scoring prompt template, and use it as the corresponding single-angle score.
[0086] For example, there are three subjective questions in a vertical domain, namely Q1, Q2, and Q3, with corresponding answers A1, A2, and A3. A candidate model gives individual scores of S1, S2, and S3 for the aforementioned three answers under a scoring prompt template. The corresponding single-angle score can be S1+S2+S3 or (S1+S2+S3) / 3.
[0087] S130: Sort the single-angle scores obtained from the model inputs of the large language model to be selected, including the same scoring prompt template, determine the sorting order, and determine the weighted weight of each single-angle score corresponding to the same scoring prompt template based on the sorting order.
[0088] Following the processing method described above, for model inputs that include the same scoring prompt template, each candidate large language model outputs a single-angle score. These single-angle scores for the same scoring prompt template can then be sorted to determine their ranking order. This ranking order indicates the degree of emphasis each model places on the corresponding scoring prompt template (or scoring angle).
[0089] After determining the aforementioned sorting order, the computing device can determine the weighted weight of each single-angle score, which is the weighted weight of each single-angle score for the same scoring prompt template.
[0090] In some embodiments, the weighted weights corresponding to each ranking position for a given scoring prompt template (i.e., a given prompt angle) are predetermined. Correspondingly, after determining the single-angle scoring for the same scoring prompt template, the weighted weights of each single-angle score for the same scoring prompt template can be determined based on the ranking order of the single-angle scores and the weighted weights of each ranking position. It should be noted that the aforementioned predetermined weighted weights are determined based on the expert experience of experts in the relevant vertical field of the subjective questions, and are not predetermined.
[0091] In some embodiments, the computing device may employ the following steps S131-S132 to determine the weighting weights of each single-angle score for the same scoring prompt template based on the sorting order.
[0092] S131: Determine the weight amplification factor or weight reduction factor for each single-angle score of the model input that includes the same scoring prompt template based on the sorting order.
[0093] In this implementation, whether the weight amplification factor or the weight reduction factor for each single-angle score is determined based on the ranking order needs to be determined according to the scoring rules. Specifically, it needs to be determined whether a higher score indicates a better answer quality or a lower score indicates a better answer quality. If a higher score indicates a better answer quality, the weight amplification factor should be determined based on the ranking order; conversely, if a lower score indicates a better answer quality, the weight reduction factor should be determined based on the ranking order.
[0094] S132: Based on the weight amplification factor or weight reduction factor, and the baseline weight of each scoring prompt template, determine the weighted weight of the single-angle score corresponding to the sorting order.
[0095] After determining the weight amplification factor or weight reduction factor, multiplication, division or exponential operations can be used to determine the weighted weight of the single-angle score for the corresponding sorting order based on the weight amplification factor or weight reduction factor and the baseline weight of the scoring prompt template.
[0096] In some embodiments, where a weighting amplification factor is determined, the computing device may employ Q. i = (n+1-P) i )α i The weighted weights are calculated, where n is the number of rating prompt templates, and P is the weight of the template. iThis represents the sorting order of the single-angle scores corresponding to the i-th scoring prompt template. Higher quality single-angle scores result in lower sorting order. α i Let be the baseline weight corresponding to the i-th rating prompt template.
[0097] In some embodiments, where a weighted reduction factor is determined, the computing device may employ α. i / (n+1-P i The weighted weights are calculated.
[0098] Other implementation schemes are not limited to the aforementioned calculation formula. The weighted weights can be determined based on the weight amplification factor, the weight reduction factor, and the benchmark weight. Other formulas that reflect the corresponding amplification and reduction effects can also be used to determine the weighted weights.
[0099] In specific implementation, α i How to determine this will be analyzed in more detail later. After obtaining the weighted weights of each single-angle score, the following S140 can be executed to determine the multi-angle scores of each candidate large language model.
[0100] S140: According to the weighted weights, the single-angle scores output by the candidate large language models are weighted and summed to obtain the multi-angle scores of each candidate large language model.
[0101] The aforementioned weighted summation of the single-angle scores output by the large language model to be selected, based on weighted weights, uses the weighted weights as the weight coefficients for the corresponding single-angle scores. A multi-dimensional score is obtained. It should be noted that the aforementioned weighted summation is a weighted summation of all single-dimensional scores output by a single large language model to be selected; single-dimensional scores output by different large language models to be selected are not weighted summated.
[0102] S150: Based on the correlation between the single-angle score size and the quality of the answers to the vertical subjective questions, select a preset number of large language models with the largest or smallest multi-angle scores as the scoring models for the vertical subjective questions.
[0103] Based on the correlation between single-angle scores and answers to vertical subjective questions, the scoring model for vertical subjective questions is determined. This follows the consistent approach described earlier, which involves ranking the multi-angle scores of each candidate large language model and determining the scoring model for vertical subjective questions based on the ranking.
[0104] Specifically, if the single-angle score is positively correlated with the quality of the answer to the vertical subjective question, that is, the higher the single-angle score, the higher the quality of the answer, then the language model with the largest preset number of multi-angle scores will be used as the scoring model for the vertical subjective question; if the single-angle score is negatively correlated with the quality of the answer to the vertical subjective question, then the language model with the smallest preset number of multi-angle scores will be used as the scoring model for the vertical subjective question.
[0105] This disclosed embodiment is based on the idea that experts in a specific field subconsciously evaluate answers to subjective questions from different angles, and the resulting scores are obtained by comprehensively considering these angles. First, scoring prompt templates with different scoring angles are set for subjective questions in the specific field. These templates are used to prompt the selected large language models to score the subjective question answers, resulting in single-angle scores. Then, guided by the idea that experts in the specific field will comprehensively process these different angle scores with different weights to obtain the final score, the weighting of each single-angle score is determined according to the ranking order of the scores corresponding to the same scoring prompt template. The weighted sum of the single-angle scores and the aforementioned weighted weights is then used to obtain multi-angle scores. The multi-angle scores obtained based on this approach better align with the scoring strategies of experts in the specific field for subjective questions. Consequently, the multi-angle scores of the selected large language models have better comparative value, and the determined scoring model for subjective questions in the specific field is more consistent with the evaluation experience of real experts. In other words, the scoring model for subjective questions in the specific field selected by this disclosed embodiment is more consistent with the experience of real experts and is more reasonable for use as a real-world application model.
[0106] As analyzed above, when determining the weighted weights of single-angle scores according to the corresponding sorting order based on the baseline weights of each scoring prompt template, it is necessary to determine the baseline weights in advance. In this embodiment of the disclosure, the weighted weights of each single-angle score can be determined using the following steps S1321-S1323.
[0107] S1321: Determine the relative importance of each scoring prompt template based on expert experience, and construct a comparison matrix based on the relative importance.
[0108] The aforementioned determination of the relative importance of each scoring prompt template based on expert experience involves obtaining the results of expert comparisons of the relative importance of each pair of scoring prompt templates, and determining the relative importance of each scoring prompt template after verifying the aforementioned comparison results. A comparison matrix can then be constructed based on these relative importances.
[0109] For example, in an application that includes 3 rating prompt templates, the relative importance of the 3 rating prompt templates is that template 1 is more important than template 2, template 1 is slightly more important than template 3, and template 2 and template 3 are equally important. Then we can set a12=3, a13=2, a23=1, and the corresponding values in the judgment matrix are shown in Table 1.
[0110] Table 1 shows the numerical values of the judgment matrix.
[0111] Template 1 Template 2 Template 3 Template 1 1 3 2 Template 2 1 / 3 1 1 Template 3 1 / 2 1 1
[0112] S1322: Normalize the comparison matrix to obtain the normalized matrix.
[0113] In practice, after obtaining the aforementioned judgment matrix, column normalization can be performed on the judgment matrix to obtain a normalized judgment matrix. For example, the normalized matrix obtained for the judgment matrix in Table 1 is shown in Table 2.
[0114] Table 2 Normalized Matrix
[0115] Template 1 Template 2 Template 3 Template 1 0.55 0.6 0.5 Template 2 0.18 0.2 0.25 Template 3 0.27 0.2 02.5
[0116] To ensure the reliability of the aforementioned weighting, a consistency ratio check is also required on the normalized matrix, specifically by using... The consistency ratio was obtained. λ max is the largest eigenvalue corresponding to the eigenvector of the judgment matrix. Based on the judgment matrix and weight vector, the largest eigenvalue in this embodiment is 3.068. is the order of the matrix, which is 3 in this embodiment. is the random consistency index, determined according to the order of the matrix, and can be obtained from the empirical value table of the analytic hierarchy process (AHP). In practical applications, if CR is less than 0.1, the consistency test is passed; otherwise, the judgment matrix needs to be adjusted. Therefore, in this embodiment, CR is 0.059, which passes the consistency test.
[0117] S1323: Calculate the average value of the matrix elements corresponding to each rating prompt template in the normalized matrix, and use it as the weighting weight of the corresponding single-angle rating.
[0118] For example, for the aforementioned normalized matrix, the weighted weight of template 1 is (0.55+0.6+0.6) / 3 = 0.55, the weighted weight of template 2 is (0.18+0.2+0.25) / 3 = 0.21, and the weighted weight of template 3 is (0.27+0.2+0.25) / 3 = 0.24.
[0119] In addition to providing the aforementioned method for selecting a model for scoring subjective questions in a vertical domain, this disclosure also provides a method for scoring subjective questions in a vertical domain. Figure 2 This is a flowchart of the scoring method for subjective questions in a vertical domain provided in this embodiment of the disclosure. Figure 2 As shown, the vertical domain subjective question scoring method provided in this embodiment includes S210-S230.
[0120] S210: Construct model input based on multiple pre-built scoring prompt templates, vertical subjective questions, and a scoreable answer for a vertical subjective question.
[0121] The scoring prompt template in this embodiment is the same as the scoring prompt template used in the previous embodiments. Each scoring prompt template provides a specific prompt angle for scoring the answer. The prompt angle of each scoring prompt template is determined based on the experience of experts in the relevant field, and the prompt angles of each scoring prompt template are different.
[0122] The vertical subjective questions can be the vertical subjective questions used earlier, or other subjective questions from the same vertical domain. The model input is constructed using the pre-scoring prompt template, the vertical subjective questions, and the corresponding answers to be scored for each vertical subjective question, following the construction method described in the aforementioned model selection method.
[0123] In some embodiments, reference answers for subjective questions in the vertical domain can also be added to the model input to provide a reference for model scoring.
[0124] S220: The pre-selected vertical subjective question scoring model processes each model input to obtain the corresponding single-angle score.
[0125] The pre-selected vertical subjective question scoring model is the subjective question scoring model determined by the method described above. Following the method described above, by processing the inputs of each model using the pre-selected vertical subjective question scoring template, the corresponding single-angle scores can be obtained.
[0126] S230: Based on the single-angle score and the corresponding weighted weight, perform a weighted summation to obtain a multi-angle score, and use the multi-angle score as the score for the answer to be scored.
[0127] In this embodiment, the weighted weights corresponding to each single-angle score are the weighted weights mentioned above. By using the weighted weights and single-angle scores to perform a weighted sum, the importance of each single-angle score is balanced, making the resulting multi-angle score more reasonable, and thus it can be used as the score for the answer to be scored.
[0128] In addition to providing the aforementioned method for selecting a model for scoring subjective questions in a vertical domain, this disclosure also provides a device for selecting a model for scoring subjective questions in a vertical domain. Figure 3 This is a schematic diagram of the selection device structure for the vertical subjective question scoring model provided in this disclosure. For example... Figure 3 As shown, the selection device 300 for the vertical subjective question scoring model includes an input construction unit 301, a model calling unit 302, a weight determination unit 303, a weighted summation unit 304, and a model selection unit 305.
[0129] The input building unit 301 is used to build multiple scoring prompt templates for vertical subjective questions, and to build model input based on each scoring prompt template, the vertical subjective question, and the answer to the vertical subjective question; each scoring prompt template provides a prompt for scoring the answer from a specific prompting angle, and the prompting angle of each scoring prompt template is determined based on the experience of vertical experts, and the prompting angle of each scoring prompt template is different; each model input includes a scoring prompt template.
[0130] The model calling unit 302 is used to process the model input using the large language model to be selected, and to obtain the single-angle score of the large language model to be selected for the answer under each scoring prompt template.
[0131] The weight determination unit 303 is used to sort the single-angle scores obtained from the model input including the same scoring prompt template in the large language model to be selected, determine the sorting order, and determine the weighted weight of each single-angle score corresponding to the same scoring prompt template based on the sorting order.
[0132] The weighted summation unit 304 is used to sum the single-angle scores output by the candidate large language models according to the weighted weights, so as to obtain the multi-angle scores of each candidate large language model.
[0133] The model selection unit 305 is used to select a preset number of large language models with the largest or smallest multi-angle scores as the scoring models for the vertical subjective questions, based on the correlation between single-angle scoring and the quality of answers to vertical subjective questions.
[0134] In some embodiments, the input construction unit 301 determines multiple prompting angles for subjective questions in a vertical domain based on expert experience, the scoring dimensions under each prompting angle, and the scoring criteria under each scoring dimension; then, based on the prompting angles, the corresponding scoring dimensions, and the scoring criteria under the scoring dimensions, multiple scoring prompt templates are constructed.
[0135] In some embodiments, the input building unit 301 determines multiple prompting angles based on expert experience, including at least two of the following angles: a holistic angle, an accuracy angle, and a usability angle.
[0136] In some embodiments, the input construction unit 301 obtains reference answers for subjective questions in the vertical domain and adds the reference answers to each scoring prompt template.
[0137] In some embodiments, the number of vertical subjective questions is at least two; the input construction unit 301 constructs model inputs based on each scoring prompt template, each vertical subjective question, and the answer to each vertical subjective question; or, it constructs model inputs based on each scoring prompt template, all vertical subjective questions, and the answer to each vertical subjective question, such that the model input includes all vertical subjective questions and their corresponding answers, as well as the relationship between all vertical subjective questions and their corresponding answers.
[0138] The model calling unit 302 uses the candidate large language model to process the model output, and obtains the individual scores of the candidate model for each vertical subjective question and corresponding answer under each scoring prompt template; calculates the average or sum of all individual scores obtained by each candidate large language model under the corresponding scoring prompt template, and uses it as the corresponding single-angle score.
[0139] In some embodiments, the weight determination unit 303 determines the weight amplification factor or weight reduction factor of each single-angle score corresponding to the same model input based on the sorting order; then, based on the weight amplification factor or weight reduction factor and the baseline weight of each score prompt template, it determines the weighted weight of the single-angle score corresponding to the sorting order.
[0140] In some embodiments, the weight determination unit 303 determines the relative importance of each scoring prompt template based on expert experience and constructs a comparison matrix based on the relative importance; the comparison matrix is normalized to obtain a normalized matrix, and the average value of the matrix elements corresponding to each scoring prompt template in the normalized matrix is calculated as the weight of the corresponding single-angle score.
[0141] This disclosure also provides a scoring device for subjective questions in a vertical domain. Figure 4 This is a schematic diagram of the scoring device for vertical subjective questions provided in an embodiment of this disclosure. Figure 4 As shown, the scoring device 400 for vertical subjective questions includes an input construction unit 401, a model calling unit 402, and a scoring unit 403.
[0142] The input construction unit 301 constructs the model input based on multiple pre-built scoring prompt templates, vertical subjective questions, and answers to be scored for vertical subjective questions; each scoring prompt template provides a prompt for scoring the answer from a specific prompting angle, and the prompting angle of each scoring prompt template is determined based on the experience of vertical experts, and the prompting angles of each scoring prompt template are different.
[0143] The model calling unit 302 uses a pre-selected vertical subjective question scoring model to process each model input and obtain the corresponding single-angle score.
[0144] The scoring unit performs a weighted summation based on the single-angle score and the corresponding weight, resulting in a multi-angle score, which is then used as the score for the answer to be scored.
[0145] This disclosure also provides a computing device for implementing the aforementioned method. Figure 5 This is a schematic diagram of the structure of a computing device provided in an embodiment of this disclosure. See below for details. Figure 5 It shows a schematic diagram of a structure suitable for implementing the computing device 500 in the embodiments of this disclosure. Figure 5 The computing device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0146] like Figure 5 As shown, the computing device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory ROM 502 or a program loaded from a storage device 508 into a random access memory RAM 503. The RAM 503 also stores various programs and data required for the operation of the computing device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0147] Typically, the following devices can be connected to I / O interface 505: input devices 505 including, for example, touchscreens, touchpads, cameras, microphones, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows computing device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 A computing device 500 with various devices is shown; however, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or included alternatively.
[0148] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0149] It should be noted that the computer-readable medium described above in this disclosure may be a computer-readable storage medium, a computer-readable signal medium, or any combination thereof.
[0150] Computer-readable storage media can be, for example—but not limited to—electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0151] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0152] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0153] The aforementioned computer-readable medium may be included in the aforementioned computing device; or it may exist independently and not assembled into the computing device.
[0154] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the tester's computer, partially on the tester's computer, as a standalone software package, partially on the tester's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the tester's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0155] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0156] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not necessarily limiting in certain circumstances. The functions described above can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), etc.
[0157] The above are merely specific embodiments of this disclosure, enabling those skilled in the art to understand or implement this disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for selecting a scoring model for vertical subjective questions, characterized in that, include: Multiple scoring prompt templates for subjective questions in a vertical domain are constructed, and model inputs are built based on each scoring prompt template, the subjective questions in the vertical domain, and the answers to the subjective questions in the vertical domain. Each scoring prompt template scores the answers from different prompting perspectives, and the prompting perspectives of each scoring prompt template are determined based on the experience of experts in the vertical domain. Each model input includes a scoring prompt template. The input of the model is processed by the large language model to be selected, and the single-angle score of the large language model to be selected for the answer under each scoring prompt template is obtained. The single-angle scores obtained from the model inputs of the large language model to be selected, including the same scoring prompt template, are sorted by size to determine the sorting order, and the weighted weight of each single-angle score corresponding to the same scoring prompt template is determined based on the sorting order. According to the weighted weights, the single-angle scores output by the candidate large language models are weighted and summed to obtain the multi-angle scores of each candidate large language model. Based on the correlation between single-angle scoring and the quality of answers to subjective questions in a specific vertical domain, a preset number of large language models with the largest or smallest multi-angle scores are selected as the scoring model for subjective questions in that vertical domain. If the single-angle scoring and the quality of answers to subjective questions in that vertical domain are positively correlated, then the preset number of large language models with the largest multi-angle scores is selected as the scoring model for subjective questions in that vertical domain. If the single-angle scoring and the quality of answers to subjective questions in that vertical domain are negatively correlated, then the preset number of large language models with the smallest multi-angle scores is selected as the scoring model for subjective questions in that vertical domain.
2. The selection method according to claim 1, characterized in that, The construction of multiple scoring prompt templates for subjective questions in vertical domains includes: Based on expert experience, we determine multiple prompting angles for subjective questions in each vertical domain, scoring dimensions under each prompting angle, and scoring criteria under each scoring dimension. Based on the aforementioned prompt angle, the corresponding rating dimension, and the rating criteria under the rating dimension, the multiple rating prompt templates are constructed.
3. The selection method according to claim 2, characterized in that, The multiple prompting angles determined based on expert experience for subjective questions in specific vertical domains include: The multiple prompting angles determined based on expert experience include at least two of the following angles: holistic angle, accuracy angle, and practicality angle.
4. The selection method according to claim 2, characterized in that, Also includes: Obtain reference answers for the subjective questions in the aforementioned vertical domain; The process of constructing the multiple scoring prompt templates also includes adding the reference answer to each scoring prompt template.
5. The selection method according to any one of claims 1-4, characterized in that, The number of subjective questions in the vertical domain is at least two; The model input based on each scoring prompt template, the vertical subjective question, and the answer to the vertical subjective question includes: The model input is constructed based on each scoring prompt template, each vertical subjective question, and the answer to each vertical subjective question, such that each model input includes only one vertical subjective question and its corresponding answer; or, the model input is constructed based on each scoring prompt template, all vertical subjective questions, and the answer to each vertical subjective question, such that the model input includes all vertical subjective questions and their corresponding answers, as well as the relationship between all vertical subjective questions and their corresponding answers; The process of using a large language model to process the model input yields a single-angle score for the answer under each scoring prompt template, including: The output of the model is processed using the large language model to be selected, and the scores of the model to be selected for subjective questions and corresponding answers in each vertical domain are obtained under each scoring prompt template. The average or sum of all scores obtained by each candidate large language model under the corresponding scoring prompt template is used as the corresponding single-angle score.
6. The selection method according to any one of claims 1-4, characterized in that, The step of determining the weighted weights of each single-angle score corresponding to the same scoring prompt template based on the sorting order includes: Based on the sorting order, determine the weight amplification factor or weight reduction factor for each single-angle score corresponding to the same model input; Based on the weight amplification factor or weight reduction factor, and the baseline weight of each of the scoring prompt templates, the weighted weight of the single-angle score corresponding to the sorting order is determined.
7. The selection method according to claim 6, characterized in that, Before determining the weighted weights of each single-angle score corresponding to the same scoring prompt template based on the sorting order, the process includes: The relative importance of each scoring prompt template was determined based on expert experience, and a comparison matrix was constructed based on the relative importance. The comparison matrix is column normalized to obtain a normalized matrix; The average value of the matrix elements corresponding to each rating prompt template in the normalized matrix is calculated and used as the weighting weight of the corresponding single-angle rating.
8. A scoring method for vertical subjective questions, characterized in that, The application of the vertical subjective question scoring model selected according to any one of claims 1-7 includes: The model input is constructed based on multiple pre-built scoring prompt templates, vertical subjective questions, and answers to the vertical subjective questions to be scored; each of the scoring prompt templates scores the answer from a different prompting perspective, and the prompting perspective of each scoring prompt template is determined based on the experience of vertical experts; The model inputs are processed by a pre-selected vertical subjective question scoring model to obtain the corresponding single-angle scores; A multi-angle score is obtained by weighting and summing the single-angle score and the corresponding weight, and the multi-angle score is used as the score for the answer to be scored.
9. A computing device, characterized in that, It includes a processor and a memory, the memory being used to store a computer program; when the computer program is loaded by the processor, it causes the processor to execute the selection method for the vertical subjective question scoring model as described in any one of claims 1-7 and / or the scoring method for the vertical subjective questions as described in claim 8.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, causes the processor to implement the selection method for the scoring model of the vertical subjective question as described in any one of claims 1-7 and / or the scoring method for the vertical subjective question as described in claim 8.
Citation Information
Patent Citations
Automatic English composition scoring method and system
CN103294660A
Foreign language composition processing method and device
CN118245617A