Method, system, and computer-readable storage medium for aligning question evaluation model, which automatically evaluates questions for evaluating LLM-based system, with human judgment using metric-rubric set

The system aligns question evaluation models with human standards by using an indicator-rubric set to calculate consistency and correlation scores, addressing the lack of clear criteria for LLM-generated questions and improving the efficiency of automated evaluation.

WO2026095206A1PCT designated stage Publication Date: 2026-05-07SELECT STAR INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SELECT STAR INC
Filing Date
2024-12-13
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

The lack of clear evaluation criteria for questions generated by Large Language Models (LLMs) and the inefficiency of human evaluator methods in assessing these questions necessitate an automated system for aligning question evaluation models with human capabilities.

Method used

A method and system for aligning a question evaluation model using an indicator-rubric set, derived from multiple evaluators and a large language model, which includes steps to calculate consistency and correlation scores to refine the evaluation model, enabling it to produce results similar to human evaluations.

Benefits of technology

The system enables automatic evaluation of LLM-generated questions with improved reliability and efficiency, aligning the model's output with human standards by training it on refined indicator-rubric sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024096933_07052026_PF_FP_ABST
    Figure KR2024096933_07052026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for aligning a question evaluation model, which automatically evaluates questions for evaluating a large language model (LLM)-based system, with human judgment using a metric-rubric set. More specifically, the present invention relates to a method and system for aligning a question evaluation model, which automatically evaluates questions for evaluating an LLM-based system, with human judgment using a metric-rubric set by training the question evaluation model on the basis of a metric-rubric set derived from evaluation results obtained from a plurality of evaluators and an LLM-based question evaluation model.
Need to check novelty before this filing date? Find Prior Art

Description

A method, system, and computer-readable storage medium for aligning a question evaluation model that automatically evaluates questions for evaluating a system using LLM with humans using an indicator-rubric set.

[0001] The present invention relates to a method and system for aligning a question evaluation model, which automatically evaluates questions for evaluating a system using LLM, with humans using an indicator-rubric set. More specifically, the invention relates to a method and system for aligning a question evaluation model, which automatically evaluates questions for evaluating a system using LLM, with humans using an indicator-rubric set, wherein the question evaluation model is trained based on an indicator-rubric set derived from the evaluation results of a question evaluation model based on multiple evaluators and a large language model.

[0002]

[0003] With the recent advancement of Large Language Models (LLMs), various types of synthetic data are being generated. Question generation, in particular, is a critical area because it must be configured to align with the purpose and operational intent of the LLM. However, currently, there is a lack of clear evaluation criteria to determine whether questions generated by the LLM meet these objectives, and it is difficult to perform automated evaluations.

[0004] Furthermore, since the existing method of evaluating large-scale data by human evaluators is inefficient in terms of time and cost, there is a need for an automated system capable of systematically evaluating the quality of questions generated by LLMs. To overcome these limitations, a system is required that can automatically evaluate the quality of LLM-generated questions by verifying the LLM-based question evaluation model through direct evaluation based on standardized criteria designed by humans. Under these circumstances, there is a demand for technology that can align a question evaluation model—which automatically evaluates questions for assessing LLM-based systems—with human capabilities.

[0005]

[0006] The present invention aims to provide a method and system for aligning a question evaluation model, which automatically evaluates questions for evaluating a system using LLM, with humans using an indicator-rubric set. More specifically, the invention aims to provide a method and system for aligning a question evaluation model, which automatically evaluates questions for evaluating a system using LLM, with humans using an indicator-rubric set, wherein the question evaluation model is trained based on an indicator-rubric set derived from the evaluation results of a question evaluation model based on multiple evaluators and a big language model.

[0007]

[0008] In order to solve the above problems, in one embodiment of the present invention, a method for improving a question evaluation model that automatically evaluates questions for evaluating a system using LLM using an indicator-rubric set comprises: an indicator-rubric set receiving step of receiving a plurality of indicator-rubric sets for commonly evaluating a plurality of questions from an administrator terminal; a first evaluation result receiving step of receiving a first evaluation result of a plurality of evaluators for the plurality of questions from an evaluator terminal; a preliminary set derivation step of calculating a consistency score for the first evaluation result for each indicator and deriving an indicator-rubric set in which the consistency score is greater than or equal to a preset first criterion as a preliminary indicator-rubric set; and a second evaluation result derivation step of inputting the plurality of questions and the preliminary indicator-rubric set into a question evaluation model and deriving a second evaluation result for each question from the question evaluation model. A method for improving a question evaluation model using an indicator-rubric set is provided, comprising: a final set derivation step of calculating a correlation score for the first evaluation result and the second evaluation result for each indicator, and deriving an indicator-rubric set in which the correlation score is greater than or equal to a pre-set second criterion as a final indicator-rubric set; and a final set transmission step of transmitting the final indicator-rubric set to the question evaluation model.

[0009] In one embodiment of the present invention, the question evaluation model includes a model based on a large language model, and when an indicator-rubric set and a first evaluation result are input to the question evaluation model, a final indicator-rubric set is derived, and the question evaluation model is trained based on the final indicator-rubric set, and when a question is input to the question evaluation model after training, the question can be evaluated based on the final indicator-rubric set.

[0010] In one embodiment of the present invention, the preliminary set derivation step calculates an intraclass correlation coefficient (ICC) for each indicator included in the indicator-rubric set based on the first evaluation result, the intraclass correlation coefficient is included in the consistency score, and an indicator-rubric set in which the consistency score is greater than or equal to a pre-set first criterion is derived as a preliminary indicator-rubric set, and for an indicator-rubric set in which the consistency score is less than the pre-set first criterion, the rubric included in the indicator-rubric set may be modified.

[0011] In one embodiment of the present invention, the step of deriving the final set may be performed by calculating a first average value corresponding to the average of the first evaluation result for each indicator included in the preliminary indicator-rubric set, calculating a correlation coefficient between the first average value and the second evaluation result, and including the correlation coefficient in the correlation score, deriving a preliminary indicator-rubric set in which the correlation score is greater than or equal to a pre-set second criterion as the final indicator-rubric set, and deriving a preliminary indicator-rubric set in which the correlation score is less than the pre-set second criterion as the residual indicator-rubric set for training the question evaluation model.

[0012] In one embodiment of the present invention, a method for improving the question evaluation model using an indicator-rubric set further comprises a question evaluation model learning step for training the question evaluation model based on a learning evaluation result value for a residual indicator-rubric set that does not correspond to the final indicator-rubric set among the preliminary indicator-rubric sets, wherein the learning evaluation result value includes the plurality of questions and the residual indicator-rubric set, and the question evaluation model learning step may train the question evaluation model by inputting the plurality of questions and the residual indicator-rubric set into the question evaluation model.

[0013] In one embodiment of the present invention, the learning evaluation result value further includes the first evaluation result of the plurality of evaluators, and the question evaluation model learning step can train the question evaluation model by inputting the plurality of questions, the residual indicator-rubric set, and the first evaluation result into the question evaluation model.

[0014] In one embodiment of the present invention, the learning evaluation result value further includes extended data generated by inputting the plurality of questions and the residual indicator-rubric set into a language model outside the server system, and the question evaluation model learning step can train the question evaluation model by inputting the plurality of questions, the residual indicator-rubric set, and the extended data into the question evaluation model.

[0015] To solve the above-mentioned problem, in one embodiment of the present invention, a server system that performs a method of improving a question evaluation model, which automatically evaluates questions for evaluating a system using LLM, using an indicator-rubric set, comprises: an indicator-rubric set receiving unit that receives a plurality of indicator-rubric sets for commonly evaluating a plurality of questions from an administrator terminal; a first evaluation result receiving unit that receives a first evaluation result of a plurality of evaluators for the plurality of questions from an evaluator terminal; a preliminary set derivation unit that calculates a consistency score for the first evaluation result for each indicator and derives an indicator-rubric set in which the consistency score is greater than or equal to a preset first standard as a preliminary indicator-rubric set; and a second evaluation result derivation unit that inputs the plurality of questions and the preliminary indicator-rubric set into a question evaluation model and derives a second evaluation result for each question from the question evaluation model. A server system is provided comprising: a final set derivation unit that calculates a correlation score for the first evaluation result and the second evaluation result for each indicator and derives an indicator-rubric set as a final indicator-rubric set in which the correlation score is greater than or equal to a preset second criterion; and a final set transmission unit that transmits the final indicator-rubric set to the question evaluation model.

[0016] In order to solve the above problem, in one embodiment of the present invention, a computer-readable storage medium for implementing a method to improve a question evaluation model that automatically evaluates questions for evaluating a system using LLM performed in a server system using an indicator-rubric set, wherein the computer-readable storage medium comprises computer-executable instructions that cause the server system to perform the following steps, the following steps comprising: an indicator-rubric set receiving step of receiving a plurality of indicator-rubric sets for commonly evaluating a plurality of questions from a manager terminal; a first evaluation result receiving step of receiving a first evaluation result of a plurality of evaluators for the plurality of questions from an evaluator terminal; a preliminary set derivation step of calculating a consistency score for the first evaluation result for each indicator and deriving an indicator-rubric set in which the consistency score is greater than or equal to a preset first criterion as a preliminary indicator-rubric set; and a second evaluation result derivation step of inputting the plurality of questions and the preliminary indicator-rubric set into a question evaluation model and deriving a second evaluation result for each question from the question evaluation model. A computer-readable storage medium is provided, comprising: a final set derivation step of calculating a correlation score for the first evaluation result and the second evaluation result for each indicator, and deriving an indicator-rubric set in which the correlation score is greater than or equal to a preset second criterion as a final indicator-rubric set; and a final set transmission step of transmitting the final indicator-rubric set to the question evaluation model.

[0017]

[0018] According to one embodiment of the present invention, by using an indicator-rubric set, the question evaluation model can be improved to derive an evaluation close to a human evaluation.

[0019] According to one embodiment of the present invention, by using a question evaluation model, it is possible to achieve the effect of automatically evaluating questions for evaluating a system using LLM.

[0020] According to one embodiment of the present invention, it is possible to derive a preliminary indicator-rubric set among indicator-rubric sets based on the evaluation results of a human evaluator regarding a plurality of questions.

[0021] According to one embodiment of the present invention, the effect of deriving a final indicator-rubric set from a preliminary indicator-rubric set based on the evaluation results of a human evaluator and the evaluation results of a question evaluation model for a plurality of questions can be achieved.

[0022] According to one embodiment of the present invention, a question evaluation model is trained through a final indicator-rubric set, thereby enabling the question evaluation model to be improved or aligned with humans.

[0023] According to one embodiment of the present invention, the intraclass correlation coefficient (ICC) can be calculated to verify the consistency of the first evaluation results of multiple evaluators.

[0024] According to one embodiment of the present invention, by calculating a correlation coefficient, it is possible to achieve the effect of verifying the correlation between the first evaluation results of multiple evaluators and the second evaluation results of the question evaluation model.

[0025] According to one embodiment of the present invention, a question evaluation model is trained through a residual indicator-rubric set, thereby enabling the question evaluation model to be improved or aligned with humans.

[0026]

[0027] FIG. 1 schematically illustrates the connection configuration of a server system according to one embodiment of the present invention.

[0028] FIG. 2 schematically illustrates the internal configuration of a server system according to one embodiment of the present invention.

[0029] FIG. 3 schematically illustrates the steps of a method for improving a question evaluation model according to an embodiment of the present invention using an indicator-rubric set.

[0030] FIG. 4 schematically illustrates an indicator-rubric set according to one embodiment of the present invention.

[0031] FIG. 5 schematically illustrates the process of performing the first evaluation result reception step according to one embodiment of the present invention.

[0032] FIG. 6 schematically illustrates the process of performing a preliminary set derivation step according to one embodiment of the present invention.

[0033] FIG. 7 schematically illustrates the process of performing the second evaluation result derivation step according to one embodiment of the present invention.

[0034] FIG. 8 schematically illustrates the process of performing the final set derivation step according to one embodiment of the present invention.

[0035] FIG. 9 schematically illustrates the process of performing the question evaluation model learning step according to one embodiment of the present invention.

[0036] FIG. 10 illustrates the internal configuration of a computing device according to one embodiment of the present invention.

[0037]

[0038] Hereinafter, various embodiments and / or aspects are disclosed with reference to the drawings. For illustrative purposes, numerous specific details are disclosed in the following description to aid in a general understanding of one or more aspects. However, it will also be recognized by those skilled in the art that these aspects may be practiced without such specific details. The following description and the accompanying drawings describe specific exemplary aspects of one or more aspects in detail. However, these aspects are exemplary, and some of the various methods in the principles of the various aspects may be used, and the description is intended to include all such aspects and their equivalents.

[0039] Additionally, terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.

[0040] Furthermore, in the embodiments of the present invention, all terms used herein, including technical or scientific terms, unless otherwise defined, have the same meaning as generally understood by those skilled in the art to which the present invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in the embodiments of the present invention.

[0041] The "administrator terminal" and "evaluator terminal" mentioned below may be implemented as a computer or portable terminal capable of connecting to a server or other terminal via a network. Here, the computer includes, for example, a notebook, desktop, or laptop equipped with a web browser, and the portable terminal may include, for example, any type of handheld-based wireless communication device that ensures portability and mobility, such as a smartphone, PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), Wibro (Wireless Broadband Internet), BLE Beacon (Bluetooth Low Energy Beacon) terminal. In addition, the “network” can be implemented as a wired network such as a Local Area Network (LAN), Wide Area Network (WAN), or Value Added Network (VAN), or as any type of wireless network such as a mobile radio communication network or a satellite communication network.

[0042] FIG. 1 schematically illustrates the connection configuration of a server system (1000) according to one embodiment of the present invention.

[0043] In schematic terms, FIG. 1 illustrates the connection configuration of a server system (1000) that performs a method to improve the question evaluation model (1100) using an indicator-rubric set, a question generation model (2000), an administrator terminal (3000), and an evaluator terminal (4000).

[0044] Specifically, the method of improving the question evaluation model (1100) using an indicator-rubrick set can be performed in a server system (1000) including one or more processors and one or more memories, and as shown in FIG. 1, the server system (1000) communicates with a question generation model (2000), an administrator terminal (3000), and an evaluator terminal (4000) and can perform the method of improving the question evaluation model (1100) using an indicator-rubrick set.

[0045] At this time, the server system (1000) can receive a plurality of questions generated from the question generation model (2000), receive an indicator-rubric set entered by an administrator from the administrator terminal (3000), and receive a plurality of first evaluation results from an evaluator terminal (4000).

[0046] At this time, the question generation model (2000) is a big language model-based model that generates questions based on input information, is located outside or inside the server system (1000), and can derive questions for evaluating a system using LLM. When a document is input into the question generation model (2000), a question can be generated using a big language model based on the document.

[0047] In one embodiment of the present invention, the administrator terminal (3000) may include a terminal used by an administrator who manages or trains the question evaluation model (1100) of the present invention through a method of improving the question evaluation model (1100) using an indicator-rubric set, or a terminal used by an administrator who intends to evaluate a system using LLM by utilizing the question evaluation model (1100), and the evaluator terminal (4000) may include a terminal used by an evaluator who evaluates a plurality of questions transmitted to the evaluator terminal (4000).

[0048] At this time, the evaluator terminal (4000) may include all of the evaluator terminals (4000) of each of the plurality of evaluators, and may correspond to one of the evaluator terminals (4000) of any one of the plurality of evaluators.

[0049] Additionally, the connection configuration between the administrator terminal (3000) and the evaluator terminal (4000) and the server system (1000) may include a form in which the server system (1000) is accessed through the administrator terminal (3000) or the evaluator terminal (4000), which includes a computer or portable terminal used by the administrator or evaluator.

[0050] As illustrated in FIG. 1, the server system (1000) includes a question evaluation model (1100), and in one embodiment of the present invention, the question evaluation model (1100) is a big language model-based model that evaluates a question when a question is input. The server system (1000) can train the question evaluation model (1100) based on an indicator-rubric set to derive an evaluation similar to a human evaluation, and the question evaluation model (1100) can be a model that can automatically evaluate questions for evaluating a system using LLM according to the training.

[0051] In one embodiment of the present invention, the question evaluation model (1100) includes a model based on a large language model, and when an indicator-rubric set and a first evaluation result are input to the question evaluation model (1100), a final indicator-rubric set is derived. The question evaluation model (1100) can be improved by training the question evaluation model (1100) based on the final indicator-rubric set, and when a question is input to the question evaluation model (1100) after training, the question can be evaluated based on the final indicator-rubric set.

[0052] FIG. 2 schematically illustrates the internal configuration of a server system (1000) according to one embodiment of the present invention.

[0053] As illustrated in FIG. 2, the server system (1000) comprises: an indicator rubric set receiving unit (100) that performs an indicator rubric set receiving step of receiving a plurality of indicator rubric sets from an administrator terminal (3000) for commonly evaluating a plurality of questions; a first evaluation result receiving unit (200) that performs a first evaluation result receiving step of receiving a first evaluation result of a plurality of evaluators for the plurality of questions from an evaluator terminal (4000); and a preliminary set derivation unit (300) that performs a preliminary set derivation step of calculating a consistency score for the first evaluation result for each indicator and deriving an indicator rubric set in which the consistency score is greater than or equal to a preset first standard as a preliminary indicator rubric set. A second evaluation result derivation unit (400) that inputs the above multiple questions and the above preliminary indicator-rubric set into a question evaluation model (1100) and performs a second evaluation result derivation step of deriving a second evaluation result for each question from the question evaluation model (1100); a final set derivation unit (500) that calculates a correlation score for the above first evaluation result and the above second evaluation result for each indicator and performs a final set derivation step of deriving an indicator-rubric set in which the correlation score is greater than or equal to a preset second standard as a final indicator-rubric set; and a final set transmission unit (600) that performs a final set transmission step of transmitting the above final indicator-rubric set to the question evaluation model (1100). and includes a question evaluation model learning unit (700) that performs a question evaluation model (1100) learning step for learning the question evaluation model (1100) based on the learning evaluation result value for the remaining indicator-rubrick set that does not correspond to the final indicator-rubrick set among the preliminary indicator-rubrick sets.

[0054] Specifically, each component included in the server system (1000) illustrated in FIG. 2 performs the role of controlling the operation of the server system (1000) that performs a method of improving the question evaluation model (1100) of the present invention using an indicator-rubrick set.

[0055] More specifically, the indicator rubric set receiving unit (100) of the server system (1000) can receive a plurality of indicator-rubric sets from the administrator terminal (3000) for evaluating a plurality of questions in common.

[0056] In one embodiment of the present invention, the indicator-rubric set corresponds to information directly set by a human, as it corresponds to an indicator-rubric set set by an administrator. Therefore, based on the indicator-rubric set set by a human, a human evaluator and an artificial intelligence question evaluation model (1100) can each evaluate multiple questions.

[0057] In addition, the above multiple questions include questions derived from the question generation model (2000) and correspond to questions for evaluating a system using LLM.

[0058] The first evaluation result receiving unit (200) of the above server system (1000) can receive the first evaluation results of a plurality of evaluators for the plurality of questions from the evaluator terminal (4000).

[0059] In one embodiment of the present invention, there are multiple evaluators who evaluate the multiple questions based on the indicator-rubric set, and when a server system (1000) transmits the indicator-rubric set and the multiple questions to multiple evaluator terminals (4000), the multiple evaluators can perform an evaluation of the multiple questions and transmit a first evaluation result back to the server system (1000).

[0060] In this case, the first evaluation result includes the results of the plurality of evaluators evaluating the plurality of questions based on the indicator-rubric set. Accordingly, the first evaluation result may correspond to the evaluation result of each of the plurality of evaluators and the evaluation result of all of the plurality of evaluators.

[0061] The preliminary set derivation unit (300) of the above server system (1000) can calculate a consistency score for the first evaluation result for each indicator and derive an indicator-rubric set in which the consistency score is greater than or equal to a pre-set first standard as a preliminary indicator-rubric set.

[0062] In one embodiment of the present invention, regarding the first evaluation results of the plurality of evaluators received by the server system (1000), a consistency score can be calculated representing the consistency between each evaluator's first evaluation results as a number from 0 to 1, and the higher the consistency score is calculated, the more consistent the data may be.

[0063] Accordingly, if the consistency score for the first evaluation result is greater than or equal to the first established standard, it may be judged to be a reliable indicator-rubric set because it corresponds to highly consistent data, and if the consistency score is less than the first established standard, it may be judged to be a unreliable indicator-rubric set because it corresponds to low consistency data.

[0064] That is, an indicator-rubric set with a consistency score equal to or greater than the first criterion above can be derived as a preliminary indicator-rubric set, and an indicator-rubric set with a consistency score less than the first criterion above can be modified.

[0065] If the indicator-rubric set is modified based on the consistency score, the first evaluation result for the modified indicator-rubric set can be received again to recalculate the consistency score.

[0066] The second evaluation result derivation unit (400) of the above server system (1000) inputs the above plurality of questions and the above preliminary indicator-rubric set into the question evaluation model (1100) and can derive a second evaluation result for each question from the question evaluation model (1100).

[0067] In one embodiment of the present invention, there exists a question evaluation model (1100) that evaluates the plurality of questions based on the plurality of questions and the preliminary indicator-rubric set, and when a server system (1000) inputs the plurality of questions and the preliminary indicator-rubric set to the question evaluation model (1100), the question evaluation model (1100) can perform an evaluation of the plurality of questions and derive a second evaluation result.

[0068] At this time, the preliminary indicator-rubric set corresponds to the preliminary indicator-rubric set derived in the preliminary set derivation step, and the second evaluation result includes the result of the question evaluation model (1100) evaluating the plurality of questions based on the preliminary indicator-rubric set.

[0069] The final set derivation unit (500) of the above server system (1000) can calculate a correlation score for the above first evaluation result and the above second evaluation result for each indicator, and derive an indicator-rubric set in which the correlation score is greater than or equal to a pre-set second standard as a final indicator-rubric set.

[0070] In one embodiment of the present invention, based on the first evaluation result and the second evaluation result, a correlation score can be calculated for each indicator, representing the correlation between the first evaluation result and the second evaluation result as a number from -1 to 1, and the higher the correlation score, the more similar the data may be.

[0071] Accordingly, if the correlation score is greater than or equal to the pre-set second standard, it is determined that the question evaluation model (1100) has produced evaluation results similar to those of the multiple evaluators because the first evaluation result and the second evaluation result correspond to data similar to each other, and thus can be used as data to train the question evaluation model (1100) to produce evaluations similar to human evaluations through the corresponding indicator-rubric set. If the correlation score is less than the pre-set second standard, it is determined that the question evaluation model (1100) has produced evaluation results not similar to those of the multiple evaluators, and thus can be used as data to train the question evaluation model (1100) to produce more improved evaluations through the corresponding indicator-rubric set.

[0072] That is, an indicator-rubric set with a correlation score equal to or greater than the second criterion above can be derived as a final indicator-rubric set, and an indicator-rubric set with a correlation score less than the second criterion above can be derived as a residual indicator-rubric set.

[0073] The final set transmission unit (600) of the above server system (1000) can transmit the final indicator-rubric set to the question evaluation model (1100).

[0074] In one embodiment of the present invention, since the final indicator-rubric set corresponds to reliable data capable of deriving consistent evaluation results among multiple evaluators and similar evaluation results between multiple evaluators and the question evaluation model (1100), the question evaluation model (1100) can be improved so that when the question evaluation model (1100) evaluates a question, it can derive results similar to human evaluation by training the question evaluation model (1100) through the final indicator-rubric set.

[0075] The question evaluation model learning unit (700) of the above server system (1000) can train the question evaluation model (1100) based on the learning evaluation result value for the remaining indicator-rubric set among the preliminary indicator-rubric sets that does not correspond to the final indicator-rubric set.

[0076] In one embodiment of the present invention, since the residual indicator-rubric set is data that produced consistent evaluation results among multiple evaluators but dissimilar evaluation results between multiple evaluators and the question evaluation model (1100), the question evaluation model (1100) can be improved so that when the question evaluation model (1100) evaluates a question, it produces results similar to human evaluation by training the question evaluation model (1100) through the residual indicator-rubric set.

[0077] FIG. 3 schematically illustrates the steps of a method for improving a question evaluation model (1100) according to one embodiment of the present invention using an indicator-rubric set.

[0078] As illustrated in FIG. 3, when an indicator-rubric set for evaluating multiple questions is set (S10), a first evaluation result by multiple evaluators is derived for the indicator-rubric set, and after receiving the first evaluation result from the evaluator terminal (4000), a consistency score for the first evaluation result is calculated to derive a preliminary indicator-rubric set (S20).

[0079] At this time, consistency between the first evaluation results of the plurality of evaluators can be identified through the consistency score, and an indicator-rubric set in which the consistency score is greater than or equal to a pre-set first criterion can be derived as a preliminary indicator-rubric set.

[0080] A second evaluation result (S30) can be derived using a question evaluation model (1100) for the above multiple questions and the above preliminary indicator-rubric set, and a final indicator-rubric set (S40) can be derived by calculating the correlation score between the above first evaluation result and the above second evaluation result.

[0081] Through the correlation score above, the similarity between the first evaluation result and the second evaluation result above can be identified, and an indicator-rubric set in which the correlation score calculated for each indicator among the preliminary indicator-rubric set is greater than or equal to a pre-established second criterion can be derived as a final indicator-rubric set.

[0082] The question evaluation model (1100) can be trained (S50) based on the remaining indicator-rubric set excluding the final indicator-rubric set among the preliminary indicator-rubric sets, and the question evaluation model (1100) can be improved to be aligned with humans (S60) through the final indicator-rubric set.

[0083] Accordingly, the question evaluation model (1100) can be trained so that the second evaluation result derived by the question evaluation model (1100) through the residual indicator-rubric set is similar to the first evaluation result of the plurality of evaluators, thereby causing the residual indicator-rubric set to gradually decrease, and the final indicator-rubric set can be applied to the question evaluation model (1100) so that the question evaluation model (1100) derives a second evaluation result that is highly similar to the first evaluation result through the final indicator-rubric set.

[0084] FIG. 4 schematically illustrates an indicator-rubric set according to one embodiment of the present invention.

[0085] In general, an indicator represents the purpose or criteria for evaluation, and a rubric is a pre-shared standard used to evaluate learning outcomes or the degree of achievement, corresponding to a detailed presentation of criteria for evaluation and standards for performance thereon. Accordingly, based on an indicator-rubric set containing one or more pairs of the indicator and the rubric, the present invention allows a plurality of human evaluators and an artificial intelligence question evaluation model (1100) to evaluate a given question. The indicator-rubric set can be input into a server system (1000) in text form.

[0086] Specifically, through the indicator rubric set reception step, the server system (1000) can receive multiple indicator-rubric sets from the administrator terminal (3000) for commonly evaluating multiple questions. Subsequently, the multiple questions and the indicator-rubric sets can be transmitted to the evaluator terminals (4000) of multiple evaluators so that multiple evaluators can evaluate the multiple questions.

[0087] At this time, the above multiple questions include questions derived from a question generation model (2000) based on a large language model located outside or inside the server system (1000), and may correspond to questions for evaluating a system using LLM.

[0088] As illustrated in FIG. 4, in the indicator-rubric set received from the administrator terminal (3000), the indicator includes answerability and topic relevance, and each indicator has multiple rubrics. For the indicator of answerability, the rubrics include Yes or No, and multiple evaluators can evaluate it as Yes 1 if the answer correctly answers the question, and as No 2 if an incorrect answer is derived or the question is not answered.

[0089] In one embodiment of the present invention, when generating a large amount of questions through a language model including the question generation model (2000), a technique for evaluating the quality of the generated questions may be required to generate accurate and reliable questions. The quality of the generated questions can be evaluated and the language model improved to generate more accurate and reliable questions.

[0090] Therefore, when there is a large number of questions generated by a language model, the manager can set an indicator-rubric set to evaluate the quality of the questions.

[0091] For example, when 100 questions are generated to evaluate a system using LLM by inputting a document into a language model-based question generation model (2000), the quality of the questions can be evaluated, and another system using LLM can be evaluated through the questions with high reliability. At this time, since it takes a lot of time and manpower to manually evaluate all the generated questions, a question evaluation model (1100) that can automatically evaluate multiple questions and evaluate them in a manner similar to humans may be required.

[0092] As shown in FIG. 4, when the server system (1000) receives an indicator-rubric set from the administrator terminal (3000), it can evaluate each of the multiple questions using the corresponding indicator-rubric set. Therefore, the indicator-rubric set can be applied commonly to all questions, and there may be multiple rubrics for a single indicator.

[0093] If multiple evaluators give scores to multiple questions without a precise standard called a rubric, the reliability of the evaluation results among the multiple evaluators may decrease because the standard for which the evaluator assigned the score is unclear. However, if multiple evaluators give scores to multiple questions based on a clear rubric as shown in Fig. 4, the reliability of the evaluation results may increase, and the consistency of the first evaluation results for the multiple evaluators may be higher.

[0094] FIG. 5 schematically illustrates the process of performing the first evaluation result reception step according to one embodiment of the present invention.

[0095] In summary, FIG. 5(a) shows the results of evaluations by multiple evaluators for each question in the form of a table, and FIG. 5(b) shows the first evaluation results of multiple evaluators.

[0096] Specifically, the first evaluation result receiving step can receive the first evaluation results of multiple evaluators regarding the multiple questions from the evaluator terminal (4000). When the multiple evaluators input the first evaluation results, in which they evaluated the multiple questions based on the indicator-rubric set, through the evaluator terminal (4000), the server system (1000) that received the first evaluation results can derive a preliminary indicator-rubric set based on the first evaluation results.

[0097] As illustrated in FIG. 5(a), a plurality of evaluators may derive a first evaluation result for each indicator included in the indicator-rubric set for a given plurality of questions. In one embodiment of the present invention, the first evaluation result may include a score indicated as a specific number according to the indicator-rubric set, the reason for the score being derived, or a subjective evaluation indicated as text.

[0098] In FIG. 5(a), there are a total of N evaluators, from Evaluator #1 to Evaluator #N, and each evaluator can derive a score according to each indicator for multiple questions. In this case, N corresponds to a natural number greater than or equal to 1. Accordingly, the first evaluation result corresponds to the result of a human directly evaluating the multiple questions, and the first evaluation result can be displayed in the form of an indicator-rubric set.

[0099] For example, if the indicator-rubric set includes three indicators from indicator 1 to indicator 3, for question 1, evaluator #1 can derive a first evaluation result with 1 point for indicator 1, 4 points for indicator 2, and 2 points for indicator 3.

[0100] In addition, if there are five rubrics for one indicator and multiple evaluators can derive a first evaluation result with a score from 1 to 5 based on each rubric, each of the multiple evaluators can derive a first evaluation result with a score from 1 to 5 for a given question.

[0101] In one embodiment of the present invention, when 100 questions are generated through the question generation model (2000), the server system (1000) can receive a first evaluation result for 100 questions by transmitting all 100 questions to the evaluator terminal (4000), or the server system (1000) can receive a first evaluation result for 10 questions by sampling 10 questions out of 100 questions and transmitting only the 10 questions to the evaluator terminal (4000).

[0102] As illustrated in FIG. 5(b), the first evaluation result of the plurality of evaluators may include the first evaluation result #1 of evaluator #1 to the first evaluation result #N of evaluator #N, and the first evaluation result may include the first evaluation result #1 to the first evaluation result #N.

[0103] Referring to FIG. 5(a), a first evaluation result is derived for each indicator for a plurality of questions, that is, a first evaluation result according to indicators 1 to 3 is derived for question 1, and a first evaluation result according to indicators 1 to 3 is derived for question 2. The first evaluation result #1 of evaluator #1 can be derived as [1, 4, 2, 2, 2, 4], the first evaluation result #2 of evaluator #2 can be derived as [1, 5, 1, 2, 1, 5], and the first evaluation result #N of evaluator #N can be derived as [1, 4, 2, 1, 2, 5].

[0104] In one embodiment of the present invention, when calculating a consistency score based on a first evaluation result, if multiple evaluators derive various scores from 1 to 5 for one indicator in a first evaluation result such as (b) of FIG. 5, the consistency score may be calculated low and it may be determined that the consistency of the corresponding indicator-rubric set is poor.

[0105] However, if multiple evaluators obtain the same score of 1 for a single indicator, the consistency score is calculated to be high, and the reliability of the indicator-rubric set may be judged to be high.

[0106] FIG. 6 schematically illustrates the process of performing a preliminary set derivation step according to one embodiment of the present invention.

[0107] In summary, the preliminary set derivation step may calculate a consistency score for the first evaluation result for each indicator and derive an indicator-rubric set in which the consistency score is greater than or equal to a pre-established first criterion as a preliminary indicator-rubric set.

[0108] Specifically, the preliminary set derivation step can calculate the intraclass correlation coefficient (ICC) for each indicator included in the indicator-rubric set based on the first evaluation result, and the corresponding intraclass correlation coefficient is included in the consistency score.

[0109] Subsequently, an indicator-rubric set in which the consistency score is greater than or equal to a pre-established first criterion is derived as a preliminary indicator-rubric set, and for an indicator-rubric set in which the consistency score is less than the pre-established first criterion, the rubric included in the corresponding indicator-rubric set can be modified.

[0110] The aforementioned intraclass correlation coefficient is a method for analyzing the degree of agreement among evaluation results from multiple evaluators, and it corresponds to a score that verifies how consistent multiple evaluation results are. For example, the intraclass correlation coefficient is a score that assesses how consistent the evaluation results of multiple evaluators regarding the same subject are, or how consistent the evaluation results are among the evaluators.

[0111] As illustrated in FIG. 6, a consistency score can be calculated based on a first evaluation result including the first evaluation result #1 of evaluator #1 and the first evaluation result #N of evaluator #N, and a preliminary indicator-rubric set among the indicator-rubric sets can be derived based on the consistency score.

[0112] The consistency between the first evaluation results of multiple evaluators can be determined through the above consistency score, and if the consistency score is equal to or greater than a pre-established first criterion, the consistency between the first evaluation results of multiple evaluators is verified to be high, and the corresponding indicator-rubric set can be derived as a preliminary indicator-rubric set as reliable data.

[0113] In one embodiment of the present invention, the consistency score corresponds to a score that can verify whether the trend of the first evaluation result is consistent, or whether the first evaluation results of each of the plurality of evaluators are consistent with one another.

[0114] The consistency score may be calculated higher as the trends between the first evaluation results of each of the multiple evaluators match, or as the first evaluation results match each other, and the higher the consistency score is calculated, the higher the reliability of the corresponding indicator-rubric set may be judged.

[0115] For example, if Evaluator #1 produced a first evaluation result of [1, 1, 1, 1, 2] and Evaluator #2 produced an evaluation result of [2, 2, 2, 2, 3], the first evaluation results of the two evaluators do not match, but the consistency score can be calculated as high because the trends of the evaluations match.

[0116] In addition, if Evaluator #1 produced a first evaluation result of [1, 2, 3, 4, 5] and Evaluator #2 produced a first evaluation result of [1, 2, 3, 4, 5], the above consistency score can be calculated high because the first evaluation results of the two evaluators are completely identical.

[0117] Preferably, an indicator-rubric set in which the consistency score is greater than or equal to a preset first criterion can be derived as a preliminary indicator-rubric set. For example, if the first criterion corresponds to 0.7, an indicator-rubric set in which the consistency score is 0.7 or higher can be derived as a preliminary indicator-rubric set, and an indicator-rubric set in which the consistency score is less than 0.7 can have its rubric modified by the administrator terminal (3000).

[0118] Accordingly, the derived preliminary indicator-rubric set can be input into a question evaluation model (1100) to evaluate multiple questions thereafter, and the remaining indicator-rubric sets excluding the preliminary indicator-rubric set can be excluded from the process, or the rubrics can be modified and used again by multiple evaluators to evaluate the multiple questions.

[0119] FIG. 7 schematically illustrates the process of performing the second evaluation result derivation step according to one embodiment of the present invention.

[0120] In summary, the second evaluation result derivation step can input the plurality of questions and the preliminary indicator-rubric set into the question evaluation model (1100) and derive a second evaluation result for each question from the question evaluation model (1100).

[0121] Specifically, the question evaluation model (1100) of the present invention can derive a second evaluation result for each indicator included in the preliminary indicator-rubric set for a plurality of input questions. In one embodiment of the present invention, the second evaluation result may include a score indicated as a specific number according to the preliminary indicator-rubric set, the reason for the score being derived, or a subjective evaluation indicated as text.

[0122] As illustrated in FIG. 7, the plurality of questions and the preliminary indicator-rubric set are input into a question evaluation model (1100), and the question evaluation model (1100) can derive a second evaluation result for the plurality of questions based on the preliminary indicator-rubric set.

[0123] Accordingly, the question evaluation model (1100) can derive a second evaluation result for each of the multiple questions using an indicator included in the preliminary indicator-rubric set. The second evaluation result can be derived in the same form as the first evaluation result, and since it is derived based on the preliminary indicator-rubric set, the number of data may be smaller compared to the first evaluation result derived based on the indicator-rubric set.

[0124] FIG. 8 schematically illustrates the process of performing the final set derivation step according to one embodiment of the present invention.

[0125] In summary, FIG. 8 (a) shows the results of evaluations by multiple evaluators and a question evaluation model (1100) for each question in the form of a table, and FIG. 8 (b) shows the process of performing the step of calculating correlation scores based on the first evaluation results and the second evaluation results.

[0126] Specifically, the final set derivation step may calculate a correlation score for the first evaluation result and the second evaluation result for each indicator, and derive an indicator-rubric set in which the correlation score is equal to or greater than a pre-established second criterion as the final indicator-rubric set.

[0127] The above final set derivation step may calculate a first average value corresponding to the average of the first evaluation results for each indicator included in the above preliminary indicator-rubric set, calculate a correlation coefficient between the first average value and the above second evaluation result, include the correlation coefficient in the above correlation score, derive a preliminary indicator-rubric set in which the correlation score is greater than or equal to a pre-set second criterion as a final indicator-rubric set, and derive a preliminary indicator-rubric set in which the correlation score is less than the pre-set second criterion as a residual indicator-rubric set for training the above question evaluation model (1100).

[0128] As illustrated in FIG. 8(a), the question evaluation model (1100) can derive a second evaluation result for each indicator included in the preliminary indicator-rubric set for a given plurality of questions. In FIG. 8(a), when there are a total of N evaluators from Evaluator #1 to Evaluator #N and a question evaluation model (1100) based on a large language model exists, each evaluator derives a first evaluation result according to each indicator for a plurality of questions, where N corresponds to a natural number greater than or equal to 1, and the first evaluation result and the second evaluation result can be displayed in the form of an indicator-rubric set.

[0129] For example, if the preliminary indicator-rubric set derived from the indicator-rubric set includes three indicators from indicator 1 to indicator 3, the question evaluation model (1100) for question 1 can derive a second evaluation result of 1 point for indicator 1, 4 points for indicator 2, and 5 points for indicator 3.

[0130] In addition, if there are five rubrics for one indicator and the question evaluation model (1100) can derive a second evaluation result with a score from 1 to 5 based on each rubric, the question evaluation model (1100) can derive a second evaluation result with a score from 1 to 5 for the input question.

[0131] As illustrated in FIG. 8(b), there exists a first evaluation result including the first evaluation result #1 of evaluator #1 and the first evaluation result #N of evaluator #N, and a second evaluation result of the question evaluation model (1100). Referring to FIG. 8(a), a second evaluation result is derived for each indicator for a plurality of questions, that is, a second evaluation result according to indicators 1 to 3 is derived for question 1, and a second evaluation result according to indicators 1 to 3 is derived for question 2. The second evaluation result of the question evaluation model (1100) can be derived as [1, 4, 5, 1, 1, 1].

[0132] At this time, a first average value corresponding to the average of the first evaluation results of multiple evaluators can be calculated, and a correlation score between the first average value and the second evaluation result can be calculated. The correlation score may be calculated to be higher the greater the association between the first average value and the second evaluation result.

[0133] In one embodiment of the present invention, when calculating a correlation score based on a first evaluation result and a second evaluation result, the first average value of the first evaluation result of multiple evaluators for one indicator in the first evaluation result and the second evaluation result are compared, and the greater the difference between the first average value and the second evaluation result, the lower the correlation score is calculated, and it may be determined that the reliability of the corresponding preliminary indicator-rubric set is reduced.

[0134] However, regarding a single indicator, the smaller the difference between the first average value and the second evaluation result, the higher the correlation score is calculated, and the reliability of the corresponding preliminary indicator-rubric set can be judged to be high.

[0135] For example, the smaller the difference between the first average value and the second evaluation result, the more similar the first evaluation result evaluated by multiple evaluators and the second evaluation result evaluated by the question evaluation model (1100) can be judged to be, and thus the question evaluation model (1100) can be judged to have reliability as a model capable of deriving an evaluation similar to a human evaluation including the multiple evaluators.

[0136] More specifically, among the preliminary indicator-rubric sets mentioned above, an indicator-rubric set whose correlation score is equal to or greater than a pre-established second criterion can be derived as a final indicator-rubric set. The correlation between the first evaluation result and the second evaluation result can be determined through the correlation score, and if the correlation score is equal to or greater than the pre-established second criterion, the association between the first evaluation result and the second evaluation result is verified to be high, so the corresponding preliminary indicator-rubric set can be derived as a final indicator-rubric set as reliable data.

[0137] In one embodiment of the present invention, the correlation score is a score capable of identifying the correlation or association between a first evaluation result and a second evaluation result, and corresponds to a score capable of identifying how another evaluation result changes when either of the evaluation results including the first evaluation result and the second evaluation result changes, or identifying the association and similarity between the two evaluation results.

[0138] The greater the association or similarity between the first evaluation result and the second evaluation result, the higher the correlation score may be calculated, and the higher the correlation score, the higher the accuracy and reliability of the corresponding preliminary indicator-rubric set may be judged.

[0139] For example, referring to Figure 8(a), since the first evaluation results for Indicator 1 in Question 1 were all derived as 1, the first average value may correspond to 1. At this time, since the second evaluation result for Indicator 1 in Question 1 corresponds to 1 and the first average value and the second evaluation result are identical, the correlation score may be calculated as high.

[0140] In addition, since the first evaluation result for Indicator 3 in Question 1 was derived as [2, 1, 2], the first average value may be approximately 1.6. At this time, the second evaluation result for Indicator 3 in Question 1 is 5, and since there is a difference of 3.4 between the first average value and the second evaluation result, the correlation score may be calculated to be lower.

[0141] Preferably, a preliminary indicator-rubric set having a correlation score equal to or greater than a pre-established second criterion can be derived as a final indicator-rubric set. For example, if the second criterion corresponds to 0.7, a preliminary indicator-rubric set having a correlation score of 0.7 or higher can be derived as a final indicator-rubric set, and a preliminary indicator-rubric set having a correlation score of less than 0.7 can be derived as a residual indicator-rubric set.

[0142] Accordingly, the derived final indicator-rubric set can be transmitted to the question evaluation model (1100) for application to the question evaluation model (1100), and the remaining indicator-rubric set corresponding to the remainder of the preliminary indicator-rubric set excluding the final indicator-rubric set can be used as a learning evaluation result value to train the question evaluation model (1100).

[0143] If the evaluation result of the artificial intelligence question evaluation model (1100) is less accurate, then the correlation score between the first evaluation result and the second evaluation result is less than the second criterion, it is not that the reliability of the indicator-rubric set is reduced, but that the accuracy of the question evaluation model (1100) is reduced.

[0144] Since the reliability of the above preliminary indicator-rubric set has already been verified by multiple evaluators, the question evaluation model (1100) can be improved so that it can accurately evaluate the said preliminary indicator-rubric set. Accordingly, a second evaluation result can be derived by modifying the prompt input to the question evaluation model (1100), or the question evaluation model (1100) can be trained based on the learning evaluation result value for the said residual indicator-rubric set.

[0145] FIG. 9 schematically illustrates the process of performing the learning step of the question evaluation model (1100) according to one embodiment of the present invention.

[0146] In summary, the question evaluation model (1100) learning step can train the question evaluation model (1100) based on the learning evaluation result value for the remaining indicator-rubric set among the preliminary indicator-rubric sets that does not correspond to the final indicator-rubric set.

[0147] Specifically, a preliminary indicator-rubric set is derived based on the first evaluation results of a plurality of evaluators among the indicator-rubric sets set from the administrator terminal (3000), and a final indicator-rubric set can be derived based on the second evaluation results of the question evaluation model (1100) among the preliminary indicator-rubric sets.

[0148] Therefore, since the above preliminary indicator-rubric set corresponds to data in which consistency among multiple evaluators has been verified, and the above final indicator-rubric set corresponds to data in which the correlation between the multiple evaluators and the question evaluation model (1100) has been verified, the above final indicator-rubric set can be applied to the question evaluation model (1100) to improve the question evaluation model (1100) so that it can derive evaluation results similar to those of humans.

[0149] The above learning evaluation result value includes the plurality of questions and the residual indicator-rubric set, and the question evaluation model (1100) learning step can input the plurality of questions and the residual indicator-rubric set into the question evaluation model (1100) to train the corresponding question evaluation model (1100). Preferably, the above learning evaluation result value may include data input to the question evaluation model (1100) to train the question evaluation model (1100).

[0150] In one embodiment of the present invention, the question evaluation model (1100) learning step involves inputting the plurality of questions and the residual indicator-rubric set into the question evaluation model (1100) to train the question evaluation model (1100), and training the final indicator-rubric set to be close to the preliminary indicator-rubric set so that the question evaluation model (1100) can be improved to be aligned with humans.

[0151] Since the above preliminary indicator-rubric set corresponds to data in which consistency among multiple evaluators has been verified, even if it is not included in the final indicator-rubric set, the data can be used for training the question evaluation model (1100) without being discarded.

[0152] As shown in FIG. 9, (1) corresponds to an indicator-rubrick set received from the administrator terminal (3000), (2) corresponds to a preliminary indicator-rubrick set, and (3) corresponds to a final indicator-rubrick set. In FIG. 9, the shaded portion excluding (2) to (3) corresponds to a residual indicator-rubrick set, and the question evaluation model (1100) can be trained based on the learning evaluation result value for the residual indicator-rubrick set.

[0153] At this time, during the learning stage of the question evaluation model (1100), the plurality of questions and the residual indicator-rubric set can be input into the question evaluation model (1100) to train the question evaluation model (1100). The residual indicator-rubric set corresponds to an indicator-rubric set with low correlation between the first evaluation result and the second evaluation result of the question evaluation model (1100) because the correlation score between the first evaluation result of the plurality of evaluators and the second evaluation result of the question evaluation model (1100) was calculated to be low, and can be utilized as data to improve the question evaluation model (1100) so that the correlation between the first evaluation result and the second evaluation result can be increased.

[0154] More specifically, since the above final indicator-rubric set corresponds to data in which the correlation between the evaluation results derived by multiple evaluators and the question evaluation model (1100) is verified, it can be applied to the question evaluation model (1100) and utilized as data to improve the question evaluation model (1100). At this time, the question evaluation model (1100) can be improved to align with humans, and the question evaluation model (1100) can be improved so that the evaluation results derived by the question evaluation model (1100) are similar to human evaluation results.

[0155] The above residual indicator-rubric set corresponds to data in which reliability among multiple evaluators has been verified, but the correlation between the evaluation results derived by each of the multiple evaluators and the question evaluation model (1100) has not been verified. Therefore, it can be used as data to train the question evaluation model (1100) so that the question evaluation model (1100) can derive evaluation results that are more similar to human evaluation results. At this time, the question evaluation model (1100) can be trained so that (3) in FIG. 9 becomes closer to (2).

[0156] In one embodiment of the present invention, the learning evaluation result value further includes the first evaluation result of the plurality of evaluators, and the question evaluation model (1100) learning step can learn the question evaluation model (1100) by inputting the plurality of questions, the residual indicator-rubric set, and the first evaluation result into the question evaluation model (1100).

[0157] The above residual indicator-rubric set can be used as an incorrect example to train the question evaluation model (1100) to derive a result more similar to the first evaluation result for the above multiple questions.

[0158] In one embodiment of the present invention, the learning evaluation result value further includes extended data generated by inputting the plurality of questions and the residual indicator-rubric set into a language model outside the server system (1000), and the question evaluation model (1100) learning step can train the question evaluation model (1100) by inputting the plurality of questions, the residual indicator-rubric set, and the extended data into the question evaluation model (1100).

[0159] The above extended data includes data that can be generated by inputting the above residual indicator-rubric set into a language model outside the server system (1000) that corresponds to a higher-level model than the above question evaluation model (1100), and when inputting the above extended data into the above question evaluation model (1100), all or part of the above extended data may be input.

[0160] Therefore, the question evaluation model (1100) can be trained to reduce the residual indicator-rubric set among the preliminary indicator-rubric sets.

[0161] Preferably, the question evaluation model (1100) automatically evaluates a first question generated from the question generation model (2000) to evaluate an external LLM system, and automatically evaluates a second question generated from the external LLM system, wherein the external LLM system is located outside the server system (1000) and corresponds to a system using a large language model.

[0162] Accordingly, after the final indicator-rubric set is applied to the question evaluation model (1100) or the question evaluation model (1100) is trained based on the residual indicator-rubric set, the question evaluation model (1100) can automatically evaluate questions for evaluating a system using LLM and can produce evaluation results similar to human evaluation results, thereby becoming an improved question evaluation model (1100) compared to the previous one.

[0163] Preferably, when there is a system using an LLM located outside the server system (1000) and a question generation model (2000) for evaluating the system, the model for evaluating questions generated from the question generation model (2000) corresponds to a question evaluation model (1100) included in the server system (1000), and the system using the LLM can be evaluated by improving the question evaluation model (1100) using an indicator-rubric set set by a human.

[0164] The present invention has the effect of improving the question evaluation model (1100) to be more advanced, and the question evaluation model (1100) improved according to the present invention has the effect of deriving an indicator-rubric set similar to a human, such as an indicator-rubric set input to the administrator terminal (3000). In addition, through the present invention, the question evaluation model (1100), which automatically evaluates questions for evaluating a system using LLM using an indicator-rubric set set by a human, can be aligned with a human.

[0165] Finally, regarding the question evaluation model (1100) enhanced according to the present invention, when an indicator-rubric set and a first evaluation result are input into the question evaluation model (1100), a final indicator-rubric set is derived, and the derived final indicator-rubric set is applied to the question evaluation model (1100) to automatically evaluate multiple questions.

[0166] FIG. 10 illustrates, in an exemplary manner, the internal configuration of a computing device (11000) according to one embodiment of the present invention.

[0167] The server system (1000) mentioned in the description of FIG. 1 may include components of the computing device (11000) illustrated in FIG. 10, which will be described later.

[0168] As illustrated in FIG. 10, the computing device (11000) may include at least one processor (11100), memory (11200), peripheral interface (11300), input / output subsystem (I / O subsystem) (11400), power circuit (11500), and communication circuit (11600).

[0169] Specifically, the memory (11200) may include, for example, high-speed random access memory, a magnetic disk, SRAM, DRAM, ROM, flash memory, or non-volatile memory. The memory (11200) may include software modules, instruction sets, or various other data required for the operation of the computing device (11000).

[0170] At this time, access to the memory (11200) from other components, such as the processor (11100) or the peripheral device interface (11300), can be controlled by the processor (11100). The processor (11100) may be composed of a single or multiple units and may include processors in the form of GPUs and TPUs to improve computational processing speed.

[0171] The above peripheral device interface (11300) can connect input and / or output peripheral devices of the computing device (11000) to the processor (11100) and the memory (11200). The processor (11100) can perform various functions for the computing device (11000) and process data by executing a software module or instruction set stored in the memory (11200).

[0172] The input / output subsystem (11400) may connect various input / output peripheral devices to the peripheral device interface (11300). For example, the input / output subsystem (11400) may include a controller for connecting peripheral devices such as a monitor, keyboard, mouse, printer, or, if necessary, a touchscreen or sensor to the peripheral device interface (11300). According to another aspect, the input / output peripheral devices may be connected to the peripheral device interface (11300) without passing through the input / output subsystem (11400).

[0173] The power circuit (11500) may supply power to all or part of the components of the terminal. For example, the power circuit (11500) may include one or more power sources such as a power management system, a battery or alternating current (AC), a charging system, a power failure detection circuit, a power converter or inverter, a power status indicator, or any other components for power generation, management, and distribution.

[0174] The communication circuit (11600) may enable communication with another computing device using at least one external port. Alternatively, as described above, the communication circuit (11600) may enable communication with another computing device by including an RF circuit and transmitting and receiving an RF signal, also known as an electromagnetic signal, as needed.

[0175] The embodiment of FIG. 10 is merely an example of the computing device (11000), and the computing device (11000) may have some components shown in FIG. 10 omitted, additional components not shown in FIG. 10 added, or a configuration or arrangement that combines two or more components. For example, a computing device for a communication terminal in a mobile environment may include a touchscreen or sensors in addition to the components shown in FIG. 10, and the communication circuit (1160) may include a circuit for RF communication of various communication methods (Wi-Fi, 3G, LTE, 5G, 6G, Bluetooth, NFC, Zigbee, etc.). The components that can be included in the computing device (11000) may be implemented as hardware, software, or a combination of both hardware and software, including one or more integrated circuits specialized for signal processing or applications.

[0176] Methods according to embodiments of the present invention may be implemented in the form of program instructions that can be executed through various computing devices and recorded on a computer-readable medium. In particular, the program according to the present embodiment may be configured as a PC-based program or an application dedicated to a mobile terminal. An application to which the present invention is applied may be installed on a user terminal through a file provided by a file distribution system. For example, the file distribution system may include a file transmission unit (not shown) that transmits the file upon a request from the user terminal.

[0177] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.

[0178] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave in order to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be standardized and stored or executed in a standardized manner on a networked computing device. Software and data may be stored on one or more computer-readable recording media.

[0179] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.

[0180] According to one embodiment of the present invention, by using an indicator-rubric set, the question evaluation model can be improved to derive an evaluation close to a human evaluation.

[0181] According to one embodiment of the present invention, by using a question evaluation model, it is possible to achieve the effect of automatically evaluating questions for evaluating a system using LLM.

[0182] According to one embodiment of the present invention, it is possible to derive a preliminary indicator-rubric set among indicator-rubric sets based on the evaluation results of a human evaluator regarding a plurality of questions.

[0183] According to one embodiment of the present invention, the effect of deriving a final indicator-rubric set from a preliminary indicator-rubric set based on the evaluation results of a human evaluator and the evaluation results of a question evaluation model for a plurality of questions can be achieved.

[0184] According to one embodiment of the present invention, a question evaluation model is trained through a final indicator-rubric set, thereby enabling the question evaluation model to be improved or aligned with humans.

[0185] According to one embodiment of the present invention, the intraclass correlation coefficient (ICC) can be calculated to verify the consistency of the first evaluation results of multiple evaluators.

[0186] According to one embodiment of the present invention, by calculating a correlation coefficient, it is possible to achieve the effect of verifying the correlation between the first evaluation results of multiple evaluators and the second evaluation results of the question evaluation model.

[0187] According to one embodiment of the present invention, a question evaluation model is trained through a residual indicator-rubric set, thereby enabling the question evaluation model to be improved or aligned with humans.

[0188] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or if the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents. Therefore, other implementations, other embodiments, and equivalents to the claims below are also within the scope of the claims.

Claims

1. A method for improving a question evaluation model that automatically evaluates questions for evaluating a system using LLM by using an indicator-rubric set, Indicator rubric set receiving step of receiving multiple indicator-rubric sets from an administrator terminal for commonly evaluating multiple questions; A first evaluation result receiving step of receiving first evaluation results from multiple evaluators regarding the above multiple questions from an evaluator terminal; A preliminary set derivation step of calculating a consistency score for the first evaluation result for each indicator, and deriving an indicator-rubric set in which the consistency score is greater than or equal to a pre-established first criterion as a preliminary indicator-rubric set; A second evaluation result derivation step of inputting the above plurality of questions and the above preliminary indicator-rubric set into a question evaluation model and deriving a second evaluation result for each question from the question evaluation model; A final set derivation step of calculating correlation scores for the first evaluation result and the second evaluation result for each indicator, and deriving an indicator-rubric set in which the correlation score is greater than or equal to a pre-set second criterion as a final indicator-rubric set; and A method for improving a question evaluation model using an indicator-rubric set, comprising: a final set transmission step of transmitting the final indicator-rubric set to the question evaluation model.

2. In Claim 1, The above question evaluation model includes a model based on a large language model, and When an indicator-rubric set and a first evaluation result are input into the above question evaluation model, a final indicator-rubric set is derived, and the above question evaluation model is trained based on the said final indicator-rubric set. A method for improving a question evaluation model using an indicator-rubric set, wherein when a question is input into the question evaluation model after training, the model can evaluate the corresponding question based on the final indicator-rubric set.

3. In Claim 1, The above preliminary set derivation step is, Based on the above first evaluation results, an intraclass correlation coefficient (ICC) for each indicator included in the above indicator-rubric set is calculated, and the corresponding intraclass correlation coefficient is included in the above consistency score, and A method for improving a question evaluation model using an indicator-rubric set, wherein an indicator-rubric set having a consistency score equal to or greater than a pre-set first criterion is derived as a preliminary indicator-rubric set, and for an indicator-rubric set having a consistency score less than the pre-set first criterion, a rubric included in that indicator-rubric set is modified.

4. In Claim 1, The above final set derivation step is, For each indicator included in the above preliminary indicator-rubric set, a first average value corresponding to the average of the above first evaluation results is calculated, and then a correlation coefficient between the said first average value and the said second evaluation result is calculated, and the said correlation coefficient is included in the said correlation score. A method for improving a question evaluation model using an indicator rubric set, wherein a preliminary indicator rubric set having a correlation score equal to or greater than a pre-set second criterion is derived as a final indicator rubric set, and a preliminary indicator rubric set having a correlation score less than the pre-set second criterion is derived as a residual indicator rubric set for training the question evaluation model.

5. In Claim 1, A method to improve the above question evaluation model using an indicator-rubric set is, It further includes a question evaluation model learning step for training the question evaluation model based on the learning evaluation result value for the remaining indicator-rubric set among the preliminary indicator-rubric sets that does not correspond to the final indicator-rubric set, and The above learning evaluation result value includes the above plurality of questions and the above residual indicator-rubric set, and The above question evaluation model learning stage is, A method for improving a question evaluation model using an indicator-rubric set, wherein the above multiple questions and the above residual indicator-rubric set are input into the question evaluation model to train the said question evaluation model.

6. In Claim 5, The above learning evaluation result value is, It further includes the first evaluation results of the aforementioned plurality of evaluators, and The above question evaluation model learning stage is, A method for improving a question evaluation model using an indicator-rubric set, wherein the above plurality of questions, the above residual indicator-rubric set, and the above first evaluation result are input into the question evaluation model to train the said question evaluation model.

7. In Claim 5, The above learning evaluation result value is, It further includes extended data generated by inputting the above multiple questions and the above residual indicator-rubric set into a language model outside the server system, and The above question evaluation model learning stage is, A method for improving a question evaluation model using an indicator-rubric set, wherein the above multiple questions, the above residual indicator-rubric set, and the above extended data are input into the question evaluation model to train the said question evaluation model.

8. A server system that performs a method to improve a question evaluation model, which automatically evaluates questions for evaluating a system using LLM, using an indicator-rubric set, An indicator rubric set receiving unit that receives multiple indicator-rubric sets from an administrator terminal for commonly evaluating multiple questions; A first evaluation result receiving unit that receives the first evaluation results of multiple evaluators regarding the above multiple questions from an evaluator terminal; A preliminary set derivation unit that calculates a consistency score for the first evaluation result for each indicator and derives an indicator-rubric set in which the consistency score is greater than or equal to a pre-set first criterion as a preliminary indicator-rubric set; A second evaluation result derivation unit that inputs the above plurality of questions and the above preliminary indicator-rubric set into a question evaluation model and derives a second evaluation result for each question from the question evaluation model; A final set derivation unit that calculates correlation scores for the first evaluation result and the second evaluation result for each indicator, and derives an indicator-rubric set in which the correlation score is greater than or equal to a pre-set second criterion as a final indicator-rubric set; and A server system comprising: a final set transmission unit that transmits the final indicator-rubric set to the question evaluation model.

9. A computer-readable storage medium for implementing a method to improve a question evaluation model that automatically evaluates questions for evaluating a system using LLM executed on a server system using an indicator-rubric set, The above computer-readable storage medium includes computer-executable instructions that cause the server system to perform the following steps, and The steps below are: Indicator rubric set receiving step of receiving multiple indicator-rubric sets from an administrator terminal for commonly evaluating multiple questions; A first evaluation result receiving step of receiving first evaluation results from multiple evaluators regarding the above multiple questions from an evaluator terminal; A preliminary set derivation step of calculating a consistency score for the first evaluation result for each indicator, and deriving an indicator-rubric set in which the consistency score is greater than or equal to a pre-established first criterion as a preliminary indicator-rubric set; A second evaluation result derivation step of inputting the above plurality of questions and the above preliminary indicator-rubric set into a question evaluation model and deriving a second evaluation result for each question from the question evaluation model; A final set derivation step of calculating correlation scores for the first evaluation result and the second evaluation result for each indicator, and deriving an indicator-rubric set in which the correlation score is greater than or equal to a pre-set second criterion as a final indicator-rubric set; and A computer-readable storage medium comprising: a final set transmission step of transmitting the final indicator-rubric set to the question evaluation model.

Citation Information

Patent Citations

  • Language production ability evaluation system, language production ability evaluation program, and language production ability evaluation method

    JP7521860B1

  • Ensemble-based machine learning characterization of human-machine dialog

    US11861317B1

  • Query evaluation in natural language processing systems

    US11972223B1

  • Automated Evaluation of Free-Form Answers and Generation of Actionable Feedback to Multidimensional Reasoning Questions

    US20240054909A1

  • Conversational Interface for Content Creation and Editing Using Large Language Models

    US20240126576A1