Computer system, method for supporting performance evaluation of large-scale language models, and program
The system addresses the inefficiencies of existing LLM evaluation methods by generating evaluation criteria through prompt-based comparisons and database analysis, enabling accurate and efficient performance assessment of LLMs in business contexts.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- HITACHI LTD
- Filing Date
- 2025-01-08
- Publication Date
- 2026-07-21
AI Technical Summary
Existing methods for evaluating large language models (LLMs) are time-consuming and costly, and there is a lack of direct relevance between the information of LLMs and evaluation perspectives, making it difficult to generate accurate performance evaluations.
A computer system and method that includes a processor, storage, and network interface, which inputs prompts to multiple LLMs, compares outputs, generates evaluation criteria, and calculates relative and absolute performance scores using a database system to evaluate LLMs in business operations.
Enables the generation of evaluation criteria for quantitatively assessing LLM performance, clarifying issues and configurations, and providing accurate, comprehensive evaluation results.
Smart Images

Figure 2026119873000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the performance evaluation of large language models in specific operations.
Background Art
[0002] In the development of large language models (LLMs), relative evaluation using the pairwise comparison method by humans has become the mainstream for evaluating LLMs. Human evaluation is time-consuming and costly. Also, since the evaluation perspectives by humans are not clear, it is difficult to obtain the information necessary for performance improvement.
[0003] As an evaluation technique for large language models in specific operations, for example, the technique described in Patent Document 1 is known. Patent Document 1 describes "an information processing method executed by an information processing system, which determines the perspective of a product or service to be sold based on the information of the product or service to be sold, the perspective being a qualitative perspective, acquires public information of a target person, generates instruction information for requesting an output of evaluation information for evaluating the target person based on the public information from the determined perspective of the product or service, inputs the instruction information into a predetermined large language model, the predetermined large language model outputs the evaluation information, and the evaluation information has a quantitative evaluation value attached to the qualitative perspective."
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] Since there is no direct relevance between the information of large language models and the evaluation perspectives of large language models in operations, it is difficult to generate the evaluation perspectives of the performance of large language models in operations by the method described in Patent Document 1.
[0006] The present invention aims to realize a system and method for generating evaluation criteria for evaluating the performance of large-scale language models in business applications. [Means for solving the problem]
[0007] A representative example of the invention disclosed in this application is as follows: That is, a computer system comprising a processor, a storage device connected to the processor, and a network interface connected to the processor, communicating with a plurality of large-scale language models, holding a plurality of first prompts for causing the large-scale language models to execute tasks, a first process of inputting the plurality of first prompts to a large-scale language model to be evaluated and obtaining a plurality of first outputs which are the results of executing the tasks; a second process of inputting the plurality of first prompts to a large-scale language model to be compared and obtaining a plurality of second outputs which are the results of executing the tasks; a third process of evaluating the superiority or inferiority of the performance of the large-scale language model to be evaluated and the large-scale language model to be compared in the tasks by comparing the first and second outputs, generating a second prompt for outputting the evaluation results as a relative evaluation result along with the reasons for the evaluation, inputting the second prompt to a large-scale language model for evaluation and obtaining the relative evaluation results; and a fourth process of generating a third prompt for generating evaluation perspectives to be used for evaluating the performance of the large-scale language model to be evaluated in the tasks, including a plurality of the relative evaluation results, inputting the third prompt to a large-scale language model for generation and obtaining a plurality of the evaluation perspectives. [Effects of the Invention]
[0008] According to the present invention, evaluation criteria for evaluating the performance of large-scale language models in business operations can be generated. Other issues, configurations, and effects will be clarified by the following description of the embodiments. [Brief explanation of the drawing]
[0009] [Figure 1]This figure shows an example of the system configuration of Example 1. [Figure 2] This figure shows an example of the hardware configuration of the computer that constitutes the large-scale language model evaluation system of Example 1. [Figure 3] This figure shows an example of a large-scale language model database in Example 1. [Figure 4] This figure shows an example of the business database in Example 1. [Figure 5] This figure shows an example of the output database from Example 1. [Figure 6] This figure shows an example of the relative evaluation database in Example 1. [Figure 7] This figure shows an example of the evaluation criteria database for Example 1. [Figure 8A] This figure shows an example of the importance database for Example 1. [Figure 8B] This figure shows an example of the importance database for Example 1. [Figure 9] This figure shows an example of the score database in Example 1. [Figure 10A] This figure shows an example of the screen displayed by the input unit of Example 1. [Figure 10B] This figure shows an example of the screen displayed by the input unit of Example 1. [Figure 10C] This figure shows an example of the screen displayed by the input unit of Example 1. [Figure 11] This flowchart illustrates the overview of the evaluation process for large-scale language models performed by the large-scale language model evaluation system of Example 1. [Figure 12] This flowchart illustrates an example of a business execution process performed by the large-scale language model evaluation system of Example 1. [Figure 13] This flowchart illustrates an example of the relative evaluation process performed by the large-scale language model evaluation system of Example 1. [Figure 14] This flowchart illustrates an example of the evaluation perspective generation process performed by the large-scale language model evaluation system of Example 1. [Figure 15]It is a flowchart for explaining an example of the importance calculation process executed by the large language model evaluation system of Example 1. [Figure 16] It is a flowchart for explaining an example of the absolute evaluation process executed by the large language model evaluation system of Example 1. [Figure 17] It is a flowchart for explaining an example of the comprehensive score calculation process executed by the large language model evaluation system of Example 1. [Figure 18] It is a diagram showing an example of the evaluation result screen presented by the large language model evaluation system of Example 1. [Figure 19] It is a flowchart for explaining an example of the business execution process executed by the large language model evaluation system of Example 2. [Figure 20] It is a flowchart for explaining an example of the relative evaluation process executed by the large language model evaluation system of Example 2. [Figure 21] It is a flowchart for explaining an example of the evaluation perspective generation process executed by the large language model evaluation system of Example 2. [Figure 22] It is a flowchart for explaining an example of the evaluation perspective generation process executed by the large language model evaluation system of Example 3. [Figure 23] It is a flowchart for explaining an example of the relative evaluation process executed by the large language model evaluation system of Example 4. [Figure 24] It is a flowchart for explaining an example of the evaluation perspective generation process executed by the large language model evaluation system of Example 5.
Mode for Carrying Out the Invention
[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the present invention is not to be construed as being limited to the description of the embodiments shown below. It will be readily understood by those skilled in the art that the specific configuration can be changed without departing from the spirit or gist of the present invention.
[0011] In the configuration of the invention described below, identical or similar components or functions are denoted by the same reference numerals, and redundant descriptions are omitted.
[0012] The designations "First," "Second," "Third," etc., used in this specification are for the purpose of identifying constituent elements and do not necessarily limit their number or order. [Examples]
[0013] Figure 1 shows an example of the system configuration of Example 1. Figure 2 shows an example of the hardware configuration of the computer that constitutes the large-scale language model evaluation system of Example 1.
[0014] The system consists of a large-scale language model evaluation system 100 and a large-scale language model system 101. The large-scale language model evaluation system 100 and the large-scale language model system 101 are connected to each other via a network such as a LAN (Local Area Network).
[0015] The large-scale language model system 101 is a system that provides large-scale language models. The large-scale language model system 101 accepts prompt input containing the content of the task to be performed by the large-scale language model, inputs the prompt to the large-scale language model, and responds with the output of the large-scale language model. In Embodiment 1, the large-scale language model system 101 is assumed to hold multiple large-scale language models.
[0016] Furthermore, the system configuration may involve the large-scale language model evaluation system 100 connecting to multiple large-scale language model systems 101, each having a different large-scale language model. Additionally, the large-scale language model evaluation system 100 may hold multiple large-scale language models.
[0017] The large-scale language model evaluation system 100 performs quantitative performance evaluation of large-scale language models in business applications. The large-scale language model evaluation system 100 consists of a computer 200 as shown in Figure 2. The computer 200 has a processor 201, main memory 202, secondary memory 203, and a network interface 204. Each hardware element is connected to the others via a bus 205.
[0018] The processor 201 executes the program stored in the main memory 202. By executing processing according to the program, the processor 201 operates as a functional unit (module) that realizes a specific function. In the following description, when the processing is described with a functional unit as the subject, it indicates that the processor 201 is executing the program that realizes that functional unit.
[0019] The main memory 202 is a memory unit that stores the program executed by the processor 201 and the information used by that program. The main memory 202 is also used as a work area.
[0020] The secondary storage device 203 is a large-capacity storage device such as an HDD (Hard Disk Drive) or SSD (Solid State Drive). The programs and information stored in the main memory 202 may also be stored in the secondary storage device 203. In this case, the processor 201 reads the programs and information from the secondary storage device 203 and loads them into the main memory 202.
[0021] Network interface 204 is an interface for connecting to a network.
[0022] The large-scale language model evaluation system 100 includes an input unit 110, a task execution unit 111, a relative evaluation unit 112, an evaluation perspective generation unit 113, an importance calculation unit 114, an absolute evaluation unit 115, an overall score calculation unit 116, and a display unit 117. The large-scale language model evaluation system 100 also holds a large-scale language model DB 120, a task DB 121, an output DB 122, a relative evaluation DB 123, an evaluation perspective DB 124, an importance DB 125, and a score DB 126.
[0023] Furthermore, regarding the functional units of the large-scale language model evaluation system 100, multiple functional units may be combined into a single functional unit, or a single functional unit may be divided into multiple functional units according to its function.
[0024] Figure 3 shows an example of the large-scale language model DB120 of Example 1.
[0025] The Large-Scale Language Model DB120 is a database for managing the large-scale language models provided by the Large-Scale Language Model System 101. The Large-Scale Language Model DB120 stores, for example, a table 300 as shown in Figure 3. Table 300 stores entries that include a model ID 301 and a model name 302.
[0026] Model ID 301 is a field that stores the ID of the large-scale language model. Model Name 302 is a field that stores the name of the large-scale language model.
[0027] Note that the fields included in an entry are not limited to those mentioned above. An entry may not include any of the aforementioned fields, nor may it include other fields. For example, it may include a field that stores information such as an API for using the large-scale language model system 101 that provides a large-scale language model.
[0028] Figure 4 shows an example of the business database 121 in Example 1.
[0029] Business DB121 is a database for managing the tasks to be executed by the large-scale language model. Business DB121 stores table 400, as shown in Figure 4. Table 400 stores entries that include business ID 401 and business name 402.
[0030] Business ID 401 is a field that stores the ID of the business. Business Name 402 is a field that stores the name of the business.
[0031] Note that the fields included in an entry are not limited to those mentioned above. An entry may not include any of the fields mentioned above, or it may include other fields.
[0032] The business database 121 also stores a set of prompts for executing business processes using a large-scale language model. Furthermore, the business database 121 may also store the correct output for the business process corresponding to each prompt.
[0033] Figure 5 shows an example of the output DB122 from Example 1.
[0034] Output DB122 is a database for managing the output of large-scale language models. Output DB122 stores table 500, as shown in Figure 5. Table 500 is generated when a task is executed by the large-scale language model. Table 500 stores entries including output ID 501, task ID 502, model ID 503, prompt 504, and output 505.
[0035] Output ID 501 is a field that stores the ID of the output. Business ID 502 is a field that stores the ID of the business executed by the large-scale language model. Model ID 503 is a field that stores the ID of the large-scale language model that executed the business. Prompt 504 is a field that stores the prompt entered into the large-scale language model. Output 505 is a field that stores the output, which is the result of the large-scale language model's business execution.
[0036] Note that the fields included in an entry are not limited to those mentioned above. An entry may not include any of the fields mentioned above, or it may include other fields.
[0037] Figure 6 shows an example of the relative evaluation DB123 in Example 1.
[0038] The relative evaluation DB123 is a database for managing the relative performance evaluation of large-scale language models in business applications. The relative evaluation DB123 stores table 600, as shown in Figure 6. Table 600 is generated when an evaluation of the performance of a large-scale language model in business applications is performed. Table 600 stores entries that include evaluation ID 601, model ID 602, output ID 603, relative evaluation 604, and evaluation reason 605.
[0039] Evaluation ID 601 is a field that stores the ID of the relative evaluation. Model ID 602 is a field that stores the ID of the large-scale language model that was evaluated relative. Output ID 603 is a field that stores the ID of the output (business execution result) used in the evaluation. Relative evaluation 604 is a field that stores identification information of the large-scale language model that has superior output (business execution result) as a result of the relative evaluation. Evaluation reason 605 is a field that stores the evaluation reason (basis for evaluation).
[0040] Note that the fields included in an entry are not limited to those mentioned above. An entry may not include any of the fields mentioned above, or it may include other fields.
[0041] Figure 7 shows an example of the evaluation criteria DB124 for Example 1.
[0042] The Evaluation Criteria DB124 is a database for managing evaluation criteria. The Evaluation Criteria DB124 stores Table 700, as shown in Figure 7. Table 700 is generated when the performance of a large-scale language model in a business context is evaluated. Table 700 stores entries including Evaluation Criteria ID 701, Evaluation Criteria 702, Definition 703, Specific Example A 704, and Specific Example B 705.
[0043] Evaluation Criterion ID 701 is a field that stores the ID of the evaluation criterion. Evaluation Criterion 702 is a field that stores the evaluation criterion. Definition 703 is a field that stores the definition of the evaluation criterion. The definition of the evaluation criterion is represented as a string. Specific Example A 704 is a field that stores specific examples of output that satisfy the evaluation criterion. Specific Example B 705 is a field that stores specific examples of output that do not satisfy the evaluation criterion.
[0044] Note that the fields included in an entry are not limited to those mentioned above. An entry may not include any of the fields mentioned above, or it may include other fields.
[0045] Figures 8A and 8B show an example of the importance level DB125 for Example 1.
[0046] Importance DB125 is a database for managing the importance (weights) of evaluation criteria. Importance DB125 stores Table 800, as shown in Figure 8A, and Table 810, as shown in Figure 8B. Table 800 stores the results of pairwise comparisons of evaluation criteria using the Analytic Hierarchy Process, and Table 810 stores the weights representing the demand for each evaluation criterion. Tables 800 and 810 are generated when the performance of a large-scale language model in a business context is evaluated.
[0047] Table 800 stores entries including result ID 801, result 802, model ID 803, target evaluation criterion 804, and comparison evaluation criterion 805.
[0048] Result ID 801 is a field that stores the ID of the result of the pairwise comparison of evaluation criterions. Result 802 is a field that stores the result of the pairwise comparison of evaluation criterions. Model ID 803 is a field that stores the ID of the large-scale language model on which the pairwise comparison of evaluation criterions was performed. Target evaluation criterion 804 is a field that stores the ID of the target evaluation criterion. Comparison evaluation criterion 805 is a field that stores the ID of the evaluation criterion to be compared.
[0049] Table 810 stores entries that include evaluation criterion ID 811 and weight 812.
[0050] Evaluation Criterion ID 811 is a field that stores the ID of the evaluation criterion. Weight 812 is a field that stores the weight representing the importance of the evaluation criterion.
[0051] Note that the fields included in an entry are not limited to those mentioned above. An entry may not include any of the fields mentioned above, or it may include other fields.
[0052] Figure 9 shows an example of the score DB126 in Example 1.
[0053] Score DB126 is a database for managing the evaluation results based on each evaluation criterion. Score DB126 stores table 900, as shown in Figure 9. Table 900 is generated when the performance of a large-scale language model in a business context is evaluated. Table 900 stores entries that include score ID 901, model ID 902, evaluation criterion ID 903, and score 904.
[0054] Score ID 901 is a field that stores the ID of the score related to the evaluation criteria. Model ID 902 is a field that stores the ID of the large-scale language model being evaluated. Evaluation Criterion ID 903 is a field that stores the ID of the evaluation criteria. Score 904 is a field that stores the score of the large-scale language model being evaluated based on the evaluation criteria.
[0055] Figures 10A, 10B, and 10C show examples of screens displayed by the input unit 110 of Embodiment 1.
[0056] The screen 1000 shown in Figure 10A is a screen for registering a large-scale language model. Screen 1000 includes an input field 1001, an add button 1002, and a register button 1003.
[0057] Input field 1001 is for entering the name of the large-scale language model. Add button 1002 is for adding input field 1001. Register button 1003 is for registering the large-scale language model entered in input field 1001 to the large-scale language model DB 120.
[0058] The screen 1010 shown in Figure 10B is a screen for registering tasks. Screen 1010 includes input fields 1011 and 1012, an add button 1013, and a register button 1014.
[0059] Input field 1011 is for entering the name of the task. Input field 1012 is for entering a file containing a set of prompts for executing the task using a large-scale language model. Add button 1013 is for adding input fields 1011 and 1012. Register button 1014 is for registering the task entered in input field 1011 and the set of prompts entered in input field 1012 to the task DB 121.
[0060] The screen 1020 shown in Figure 10C is a screen for setting evaluation settings. Screen 1020 includes input fields 1021, 1022, 1023 and an evaluation button 1024.
[0061] Input field 1021 is for inputting the task to be evaluated (evaluation task). Input field 1022 is for inputting the large-scale language model to be evaluated (evaluation large-scale language model). Input field 1023 is for inputting the large-scale language model to be compared with the evaluation large-scale language model (comparison large-scale language model). Note that input field 1023 may also contain the set of correct outputs for the task. Evaluation button 1024 is a button to instruct the evaluation. When evaluation button 1024 is pressed, the large-scale language model evaluation system 100 starts the evaluation process of the large-scale language model.
[0062] Figure 11 is a flowchart illustrating the overview of the evaluation process of a large-scale language model performed by the large-scale language model evaluation system 100 of Example 1.
[0063] The business execution unit 111 executes business execution processing to cause the evaluation large-scale language model and the comparison large-scale language model to perform evaluation tasks (step S1101).
[0064] The relative evaluation unit 112 performs a relative evaluation process to evaluate the relative performance of the evaluation large-scale language model and the comparison large-scale language model in the evaluation task (step S1102).
[0065] The relative evaluation unit 112 executes an evaluation perspective generation process to generate evaluation perspectives used to quantitatively evaluate the performance of the large-scale evaluation language model in evaluation work (step S1103).
[0066] The importance calculation unit 114 performs an importance calculation process to calculate the importance of the evaluation criteria (step S1104).
[0067] The absolute evaluation unit 115 performs an absolute evaluation process to evaluate the absolute performance of the large-scale language model used in evaluation tasks, based on the evaluation criteria (step S1105).
[0068] The overall score calculation unit 116 performs an overall score calculation process to calculate an overall score, which is a quantitative evaluation of the large-scale language model used in evaluation work (step S1106).
[0069] Figure 12 is a flowchart illustrating an example of a business execution process performed by the large-scale language model evaluation system 100 of Example 1.
[0070] The business execution unit 111 obtains a set of evaluation task prompts from the business database 121 (step S1201).
[0071] The business execution unit 111 inputs each prompt to the evaluation large-scale language model of the large-scale language model system 101 and obtains the execution result (output) of the business (step S1202).
[0072] The business execution unit 111 registers the output in the output DB 122 (step S1203). Specifically, the business execution unit 111 generates table 500. The business execution unit 111 adds one entry to table 500 for each output, sets the ID in the output ID 501 of the entry, sets the ID of the evaluation business in the business ID 502, and sets the ID of the large-scale evaluation language model in the model ID 503. The business execution unit 111 sets the entered prompt in the prompt 504 of each entry, and sets the output for the prompt in output 505.
[0073] The business execution unit 111 inputs each prompt to the comparative large-scale language model of the large-scale language model system 101 and obtains the execution result (output) of the business (step S1204).
[0074] The business execution unit 111 registers the output in the output DB 122 (step S1205). Specifically, the business execution unit 111 generates table 500. The business execution unit 111 adds one entry to table 500 for each output, sets the ID in the output ID 501 of the entry, sets the ID of the evaluation business in the business ID 502, and sets the ID of the comparative large-scale language model in the model ID 503. The business execution unit 111 sets the entered prompt in the prompt 504 of each entry, and sets the output for the prompt in output 505.
[0075] Furthermore, output ID 501 in each table 500 is set to correspond to the output for the same prompt.
[0076] Figure 13 is a flowchart illustrating an example of the relative evaluation process performed by the large-scale language model evaluation system 100 of Example 1.
[0077] The relative evaluation unit 112 obtains the execution results of the evaluation large-scale language model and the comparison large-scale language model from the output DB 122 (step S1301). Specifically, the relative evaluation unit 112 obtains the table 500 stored in the output DB 122.
[0078] The relative evaluation unit 112 generates a relative evaluation execution prompt that includes the execution results of the tasks and allows for a comparison of the execution results of the evaluation large-scale language model and the comparison large-scale language model (step S1302).
[0079] The relative evaluation execution prompt is a prompt that causes the execution of a task that compares the execution results of the evaluation large-scale language model and the comparison large-scale language model for the same evaluation task prompt, and outputs a relative evaluation result that includes identification information of the large-scale language model that performs better in the evaluation task. The relative evaluation result includes identification information of the large-scale language model that performs better in the evaluation task and the reason for the evaluation (basis for the evaluation). When a large-scale language model receives input from the relative evaluation execution prompt, it compares the execution results of the tasks for each evaluation task prompt and outputs the relative evaluation result for each evaluation task prompt.
[0080] The relative evaluation unit 112 inputs a relative evaluation execution prompt to the large-scale language model for evaluation of the large-scale language model system 101 and obtains multiple relative evaluation results (outputs) (step S1303).
[0081] The large-scale language model for evaluation may be pre-configured, specified by the user, or randomly selected by the large-scale language model system 101.
[0082] The relative evaluation unit 112 registers multiple relative evaluation results in the relative evaluation DB 123 (step S1304).
[0083] Specifically, the relative evaluation unit 112 generates a table 600. The relative evaluation unit 112 adds one entry to the table 600 for each relative evaluation result and sets the ID in the evaluation ID 601 of the entry. The relative evaluation unit 112 sets the ID of the large-scale language model used for evaluation in the model ID 602 of each entry, sets the ID of the output (business execution result) of the target large-scale language model in the output ID 603, sets the identification information of the large-scale language model included in the relative evaluation result in the relative evaluation 604, and sets the basis for the evaluation included in the relative evaluation result in the evaluation reason 605.
[0084] Figure 14 is a flowchart illustrating an example of the evaluation perspective generation process performed by the large-scale language model evaluation system 100 of Example 1.
[0085] The evaluation perspective generation unit 113 obtains evaluation reasons from the relative evaluation DB 123 (step S1401). Specifically, the evaluation perspective generation unit 113 obtains evaluation reasons from each entry in the table 600 stored in the relative evaluation DB 123.
[0086] The evaluation perspective generation unit 113 generates an evaluation perspective generation prompt that includes the evaluation reason and allows the generation of evaluation perspectives from the evaluation reason (step S1402).
[0087] The evaluation criteria generation prompt is a prompt that uses evaluation reasons to generate evaluation criteria, definitions of evaluation criteria, specific examples that satisfy the evaluation criteria, and specific examples that do not satisfy the evaluation criteria as evaluation criteria data. When a large-scale language model receives input from the evaluation criteria generation prompt, it extracts evaluation criteria for each evaluation reason, generates evaluation criteria data for each evaluation reason, and outputs multiple evaluation criteria data.
[0088] The evaluation criteria represent the factors used to judge the superiority or inferiority of large-scale language models in business operations, that is, information that expresses the evaluation criteria for large-scale language models in business operations. Therefore, the large-scale language model is tasked with analyzing the evaluation criteria and extracting information related to the evaluation criteria.
[0089] The evaluation perspective generation unit 113 inputs an evaluation perspective generation prompt to the large-scale language model for generating the large-scale language model system 101 and obtains multiple evaluation perspective data (output) (step S1403).
[0090] The large-scale language model for generation may be pre-configured, specified by the user, or randomly selected by the large-scale language model system 101.
[0091] The evaluation perspective generation unit 113 includes multiple evaluation perspective data and generates an integration prompt for integrating similar evaluation perspectives (step S1404).
[0092] The integration prompt is a prompt to generate new evaluation perspective data by integrating evaluation perspective data from similar evaluation perspectives. Evaluation perspective data for which no similar evaluation perspectives exist will be output as is. When the large-scale language model receives input from the integration prompt, it identifies similar evaluation perspectives, integrates the evaluation perspective data from those similar perspectives, and outputs the integrated result.
[0093] The evaluation perspective generation unit 113 inputs an evaluation perspective generation prompt to the large-scale language model for integration of the large-scale language model system 101 and obtains the integration result (output) (step S1405).
[0094] The large-scale language model for integration may be pre-configured, specified by the user, or randomly selected by the large-scale language model system 101.
[0095] The evaluation criteria generation unit 113 registers the evaluation criteria in the evaluation criteria DB 124 based on the integration results (step S1406).
[0096] Specifically, the evaluation perspective generation unit 113 generates table 700. The evaluation perspective generation unit 113 adds one entry to each evaluation perspective data and sets the ID to evaluation perspective ID 701 of the added entry. The evaluation perspective generation unit 113 sets the evaluation perspective in evaluation perspective 702 of each entry, sets the meaning of the evaluation perspective in definition 703, sets an example that satisfies the evaluation perspective in example example A 704, and sets an example that does not satisfy the evaluation perspective in example example B 705.
[0097] Figure 15 is a flowchart illustrating an example of the importance calculation process performed by the large-scale language model evaluation system 100 of Example 1.
[0098] The importance calculation unit 114 obtains evaluation perspective data from the evaluation perspective DB 124 (step S1501). Specifically, the importance calculation unit 114 obtains table 700 stored in the evaluation perspective DB 124.
[0099] The importance calculation unit 114 includes the acquired evaluation perspective data and generates an importance evaluation prompt to perform a pairwise comparison of the evaluation perspectives (step S1502).
[0100] The importance assessment prompt is a prompt to perform a task that involves performing pairwise comparisons on pairs of evaluation criteria and outputting importance data that includes the results of the pairwise comparisons. When a large-scale language model receives input from the importance assessment prompt, it generates pairs of evaluation criteria, performs pairwise comparisons on each pair of evaluation criteria, generates importance data that includes the results of the pairwise comparisons, and outputs multiple pieces of importance data.
[0101] The importance calculation unit 114 inputs an importance evaluation prompt to the large-scale language model for evaluation of the large-scale language model system 101 and obtains multiple importance data (outputs) (step S1503).
[0102] The large-scale language model for evaluation may be pre-configured, specified by the user, or randomly selected by the large-scale language model system 101.
[0103] The importance calculation unit 114 registers importance data in the importance DB 125 (step S1504).
[0104] Specifically, the importance calculation unit 114 generates a table 800. The importance calculation unit 114 adds one entry to each importance data and sets the ID to the result ID 801 of the added entry. The importance calculation unit 114 sets the result of the match / comparison to the result 802 of each entry, sets the ID of the large-scale language model for evaluation to the model ID 803, and sets the evaluation criteria to the target evaluation criteria 804 and the comparison evaluation criteria 805.
[0105] The importance calculation unit 114 calculates the weight of each evaluation criterion using the importance data (step S1505). For example, the importance calculation unit 114 calculates the eigenvector values of importance as weights.
[0106] The importance calculation unit 114 registers the weights of the evaluation criteria in the score DB 126 (step S1506).
[0107] Specifically, the importance calculation unit 114 generates a table 810. The importance calculation unit 114 adds as many entries as there are evaluation criteria to the table 810, sets the evaluation criterion ID 811 in each entry to the ID of the evaluation criterion, and sets the weight of the evaluation criterion in the weight 812.
[0108] By comprehensively extracting evaluation criteria from the reasons for evaluating the superiority or inferiority of the output of large-scale language models in business operations, which is included in relative evaluations, accurate evaluation becomes possible. Furthermore, by integrating similar evaluation criteria, accurate evaluation becomes possible.
[0109] Figure 16 is a flowchart illustrating an example of the absolute evaluation process performed by the large-scale language model evaluation system 100 of Example 1.
[0110] The absolute evaluation unit 115 obtains the execution results of the target large-scale language model's operations from the output DB 122 (step S1601). Specifically, the absolute evaluation unit 115 obtains the table 500 stored in the output DB 122.
[0111] The absolute evaluation unit 115 retrieves evaluation perspective data from the evaluation perspective DB 124 (step S1602). Specifically, the absolute evaluation unit 115 retrieves table 700 stored in the evaluation perspective DB 124.
[0112] The absolute evaluation unit 115 generates an absolute evaluation execution prompt for each combination of the work execution result and evaluation perspective (step S1603).
[0113] The absolute evaluation execution prompt includes the execution result of a task, evaluation criteria, definitions of the evaluation criteria, specific examples that satisfy the evaluation criteria, and specific examples that do not satisfy the evaluation criteria, and is a prompt for evaluating whether the execution result of a task satisfies the evaluation criteria. When the large-scale language model receives input from the absolute evaluation execution prompt, it generates combinations of the execution result of a task and evaluation criteria, and for each combination of the execution result of a task and evaluation criteria, it generates an absolute evaluation result that includes the output, evaluation criteria, and a numerical value indicating whether or not the evaluation criteria are satisfied. In this embodiment, if the evaluation criteria are satisfied, 1 is output, and if the evaluation criteria are not satisfied, 0 is output.
[0114] Since the determination of whether or not the evaluation criteria are met is less dependent on the performance of the large-scale language model, the determination result can be treated as an absolute performance evaluation.
[0115] The absolute evaluation unit 115 inputs an absolute evaluation execution prompt to the large-scale language model for evaluation in the large-scale language model system 101 and obtains multiple absolute evaluation results (outputs) (step S1604).
[0116] The large-scale language model for evaluation may be pre-configured, specified by the user, or randomly selected by the large-scale language model system 101.
[0117] The absolute evaluation unit 115 calculates a score for each evaluation criterion using the absolute evaluation results for the same evaluation criterion (step S1605). For example, the absolute evaluation unit 115 calculates the score as the average of the numerical values.
[0118] The absolute evaluation unit 115 registers the score in the score DB 126 (step S1606).
[0119] Specifically, the absolute evaluation unit 115 generates a table 900. The absolute evaluation unit 115 adds entries to the table 900 equal to the number of evaluation criteria. The absolute evaluation unit 115 sets IDs for the score ID 901 and evaluation criterion ID 903 of each entry, sets the ID of the target large-scale language model for the model ID 902, and sets the score for the score 904.
[0120] Figure 17 is a flowchart illustrating an example of the overall score calculation process performed by the large-scale language model evaluation system 100 of Example 1.
[0121] The overall score calculation unit 116 obtains the weights of the evaluation criteria from the importance database 125 (step S1701). Specifically, the overall score calculation unit 116 obtains the table 810 stored in the importance database 125.
[0122] The overall score calculation unit 116 obtains the scores for the evaluation criteria from the score DB 126 (step S1702). Specifically, the overall score calculation unit 116 obtains the table 800 stored in the importance DB 125.
[0123] The overall score calculation unit 116 calculates the overall score using the weights and scores of each evaluation criterion (step S1703). For example, the overall score calculation unit 116 calculates the overall score as the sum of the product values of the weights and scores. The overall score may be normalized so that the maximum value of the overall score is 100.
[0124] Figure 18 shows an example of the evaluation results screen presented by the large-scale language model evaluation system 100 of Example 1.
[0125] The display unit 117 displays an evaluation result screen 1800, as shown in Figure 18, based on the score DB 126 and the overall score.
[0126] According to Example 1, the large-scale language model evaluation system 100 can generate evaluation criteria for evaluating the performance of a large-scale language model in business operations by using multiple relative evaluations. Furthermore, the large-scale language model evaluation system 100 can perform a quantitative evaluation of the performance of the large-scale language model in business operations based on the evaluation criteria. [Examples]
[0127] Example 2 differs from Example 1 in its method for generating evaluation criteria. The following description focuses on the differences between Example 2 and Example 1.
[0128] The system configuration of Example 2 is the same as that of Example 1. The hardware and software configurations of the large-scale language model evaluation system 100 in Example 2 are the same as those of Example 1.
[0129] The importance calculation process, absolute evaluation process, and overall score calculation process in Example 2 are the same as those in Example 1.
[0130] Figure 19 is a flowchart illustrating an example of a business execution process performed by the large-scale language model evaluation system 100 of Example 2.
[0131] After the processing in step S1203 is executed, the business execution unit 111 selects a comparative large-scale language model (step S1251). In this embodiment, it is assumed that a list of selectable large-scale language models is pre-configured.
[0132] The business execution unit 111 inputs each prompt into the comparative large-scale language model of the large-scale language model system 101 and obtains the execution result (output) of the business (step S1204). The business execution unit 111 registers the output in the output DB 122 (step S1205).
[0133] The business execution unit 111 determines whether processing has been completed for all selectable large-scale language models (step S1252).
[0134] If processing has not been completed for all selectable large-scale language models, the business execution unit 111 returns to step S1251.
[0135] Once processing is complete for all selectable large-scale language models, the business execution unit 111 terminates the business execution process.
[0136] Alternatively, a large-scale language model to be used as a comparative large-scale language model may be selected in advance, and steps S1204 and S1205 may be executed in parallel for each comparative large-scale language model.
[0137] Figure 20 is a flowchart illustrating an example of the relative evaluation process performed by the large-scale language model evaluation system 100 of Example 2.
[0138] The relative evaluation unit 112 selects a comparative large-scale language model (step S1351). In this embodiment, it is assumed that a list of selectable large-scale language models is pre-configured.
[0139] After executing the processes from step S1301 to step S1304, the relative evaluation unit 112 determines whether processing has been completed for all selectable large-scale language models (step S1352).
[0140] If processing is not complete for all selectable large-scale language models, the relative evaluation unit 112 returns to step S1351.
[0141] Once processing is complete for all selectable large-scale language models, the relative evaluation unit 112 terminates the relative evaluation process.
[0142] Alternatively, a large-scale language model to be used as a comparative large-scale language model may be selected in advance, and steps S1301 to S1304 may be executed in parallel for each comparative large-scale language model.
[0143] Figure 21 is a flowchart illustrating an example of the evaluation perspective generation process performed by the large-scale language model evaluation system 100 of Example 2.
[0144] The evaluation perspective generation unit 113 selects a comparative large-scale language model (step S1451). In this embodiment, it is assumed that a list of selectable large-scale language models is pre-configured.
[0145] The evaluation perspective generation unit 113 obtains evaluation reasons from the relative evaluation DB 123 (step S1401). Specifically, the evaluation perspective generation unit 113 obtains evaluation reasons from each entry in the table 600, which is stored in the relative evaluation DB 123 and corresponds to the selected large-scale comparative language model.
[0146] The evaluation perspective generation unit 113 generates an evaluation perspective generation prompt to generate evaluation perspectives from the evaluation reasons (step S1452).
[0147] In Example 2, the evaluation criteria generation prompt includes the evaluation reason and the evaluation criteria generated in the previous loop. Including the evaluation criteria generated in the previous loop in the prompt improves the accuracy of evaluation criterion generation.
[0148] After executing the process in step S1403, the evaluation perspective generation unit 113 determines whether processing has been completed for all selectable large-scale language models (step S1453).
[0149] If processing is not complete for all selectable large-scale language models, the evaluation perspective generation unit 113 returns to step S1451.
[0150] Once processing is complete for all selectable large-scale language models, the evaluation perspective generation unit 113 generates an integration prompt containing multiple evaluation perspective data generated in the final loop processing, for integrating similar evaluation perspectives (step S1454).
[0151] The processes in steps S1405 and S1406 are the same as in Example 1. [Examples]
[0152] Example 3 differs from Example 1 in its method of generating evaluation criteria. The following description focuses on the differences between Example 3 and Example 1.
[0153] The system configuration of Example 3 is the same as that of Example 1. The hardware and software configurations of the large-scale language model evaluation system 100 in Example 3 are the same as those of Example 1.
[0154] The importance calculation process, absolute evaluation process, and overall score calculation process in Example 3 are the same as in Example 1. The business execution process and relative evaluation process in Example 3 are the same as in Example 2.
[0155] Figure 22 is a flowchart illustrating an example of the evaluation perspective generation process performed by the large-scale language model evaluation system 100 of Example 3.
[0156] The evaluation perspective generation unit 113 selects a comparative large-scale language model (step S1451). In this embodiment, it is assumed that a list of selectable large-scale language models is pre-configured.
[0157] After executing the processes from step S1401 to step S1403, the evaluation perspective generation unit 113 determines whether processing has been completed for all selectable large-scale language models (step S1453).
[0158] If processing is not complete for all selectable large-scale language models, the evaluation perspective generation unit 113 returns to step S1451.
[0159] Once processing is complete for all selectable large-scale language models, the evaluation perspective generation unit 113 generates an integration prompt containing multiple evaluation perspective data generated in each loop process, for integrating similar evaluation perspectives (step S1461). Subsequently, the evaluation perspective generation unit 113 executes the processes in steps S1405 and S1406.
[0160] Alternatively, a large-scale language model to be used as a comparative large-scale language model may be selected in advance, and steps S1401 to S1403 may be executed in parallel for each comparative large-scale language model.
[0161] By generating evaluation criteria from the results of relative evaluations with different large-scale language models, the variety of evaluation criteria can be increased. [Examples]
[0162] Example 4 differs from Example 1 in its method for generating evaluation criteria. The following description focuses on the differences between Example 4 and Example 1.
[0163] The system configuration of Example 4 is the same as that of Example 1. The hardware and software configurations of the large-scale language model evaluation system 100 in Example 4 are the same as those of Example 1.
[0164] The evaluation criteria generation process, importance calculation process, absolute evaluation process, and overall score calculation process in Example 4 are the same as in Example 1. The business execution process in Example 4 is the same as in Example 2.
[0165] Figure 23 is a flowchart illustrating an example of the relative evaluation process performed by the large-scale language model evaluation system 100 of Example 4.
[0166] The relative evaluation unit 112 obtains the execution results of the evaluation large-scale language model's tasks and the execution results of each comparison large-scale language model's tasks from the output DB 122 (step S1351).
[0167] The relative evaluation unit 112 generates a relative evaluation execution prompt that includes the execution results of the tasks and allows for the comparison of the execution results of the evaluation large-scale language model and multiple comparison large-scale language models (step S1352).
[0168] The relative evaluation execution prompt in Example 4 is a prompt to execute a task that compares the execution results of the evaluation large-scale language model and multiple comparative large-scale language models for the same evaluation task prompt, and outputs a relative evaluation result that includes identification information of the large-scale language model with superior performance in the evaluation task.
[0169] The relative evaluation unit 112 executes the processes of steps S1303 and S1304.
[0170] By comparing the results of multiple tasks, more detailed evaluation criteria can be obtained. This allows for a greater variety of evaluation perspectives. [Examples]
[0171] Example 5 differs from Example 1 in its method of integrating evaluation criteria in the evaluation criteria generation process. The following description will focus on the differences between Example 5 and Example 1.
[0172] The system configuration of Example 5 is the same as that of Example 1. The hardware and software configurations of the large-scale language model evaluation system 100 in Example 5 are the same as those of Example 1.
[0173] The business execution process, relative evaluation process, importance calculation process, absolute evaluation process, and overall score calculation process in Example 5 are the same as those in Example 1.
[0174] Figure 24 is a flowchart illustrating an example of the evaluation perspective generation process performed by the large-scale language model evaluation system 100 of Example 5.
[0175] After executing the processes from step S1401 to step S1403, the evaluation perspective generation unit 113 calculates a vector which is an embedded representation of the definition of the evaluation perspective (step S1471).
[0176] The evaluation criteria generation unit 113 generates a list of evaluation criteria (step S1472).
[0177] The evaluation perspective generation unit 113 selects one evaluation perspective from the list of evaluation perspectives (step S1473).
[0178] The evaluation perspective generation unit 113 calculates the similarity between the selected evaluation perspective and other evaluation perspectives (step S1474). Specifically, the evaluation perspective generation unit 113 calculates the cosine similarity using a vector representing the definition of the evaluation perspective.
[0179] The evaluation perspective generation unit 113 determines whether or not there are evaluation perspectives similar to the selected evaluation perspective (step S1475). Specifically, the evaluation perspective generation unit 113 determines whether or not there are evaluation perspectives whose cosine similarity is greater than a threshold.
[0180] If no evaluation criteria similar to the selected evaluation criteria exist, the evaluation criteria generation unit 113 proceeds to step S1477.
[0181] If there are evaluation criteria similar to the selected evaluation criteria, the evaluation criteria generation unit 113 deletes the selected evaluation criteria from the list of evaluation criteria (step S1476), and then proceeds to step S1477.
[0182] In step S1477, the evaluation perspective generation unit 113 determines whether processing has been completed for all evaluation perspectives in the list of evaluation perspectives (step S1477).
[0183] If processing has not been completed for all evaluation criteria in the list of evaluation criteria, the evaluation criteria generation unit 113 returns to step S1473.
[0184] Once processing is complete for all evaluation criteria in the list of evaluation criteria, the evaluation criteria generation unit 113 registers the evaluation criteria registered in the list of evaluation criteria in the evaluation criteria DB 124 (step S1478).
[0185] By integrating evaluation criteria using a rule-based approach, the cost of evaluating large-scale language models can be reduced.
[0186] It should be noted that the present invention is not limited to the embodiments described above, and various modifications are included. Furthermore, for example, the embodiments described above are detailed explanations of the configuration in order to clearly illustrate the present invention, and are not necessarily limited to those having all the configurations described. In addition, some of the configurations in each embodiment can be added to, deleted from, or replaced with other configurations.
[0187] Furthermore, each of the above-mentioned configurations, functions, processing units, processing means, etc., may be implemented in hardware, in whole or in part, for example, by designing them as integrated circuits. The present invention can also be implemented by software program code that realizes the functions of the embodiment. In this case, a storage medium on which the program code is recorded is provided to a computer, and the processor of that computer reads the program code stored in the storage medium. In this case, the program code read from the storage medium itself realizes the functions of the embodiment described above, and the program code itself and the storage medium on which it is stored constitute the present invention. Examples of storage media used to supply such program code include flexible disks, CD-ROMs, DVD-ROMs, hard disks, SSDs (Solid State Drives), optical disks, magneto-optical disks, CD-Rs, magnetic tapes, non-volatile memory cards, ROMs, and the like.
[0188] Furthermore, the program code that implements the functions described in this embodiment can be implemented in a wide range of programming or scripting languages, such as assembler, C / C++, Perl, Shell, PHP, Python, and Java (registered trademark).
[0189] Furthermore, the program code for the software that implements the functions of the embodiment may be distributed via a network and stored in a storage means such as a computer's hard disk or memory, or in a storage medium such as a CD-RW or CD-R, and the computer's processor may read and execute the program code stored in the storage means or storage medium.
[0190] In the above-described embodiment, the control lines and information lines shown are those deemed necessary for explanation and do not necessarily represent all control lines and information lines in the actual product. All components may be interconnected. [Explanation of symbols]
[0191] 100 Large-scale language model evaluation systems 101 Large-scale language model systems 110 Input Section 111 Business Execution Department 112 Relative Evaluation Department 113 Evaluation Criteria Generation Unit 114 Importance calculation part 115 Absolute Evaluation Section 116. Overall Score Calculation Section 117 Display section 120 Large-scale language model databases 121 Business DB 122 Output DB 123 Relative Evaluation Database 124 Evaluation Criteria Database 125 Importance DB 126 Score DB 200 calculator 201 Processor 202 Main storage 203 Secondary storage device 204 Network Interfaces Bus 205 1000, 1010, 1020 screens 1800 Evaluation Results Screen
Claims
1. A computer system, The system comprises a processor, a storage device connected to the processor, and a network interface connected to the processor. Connecting to multiple large-scale language models in a communicative manner, The large-scale language model holds multiple first prompts for causing it to perform tasks, A first process involves inputting the multiple first prompts into the large-scale language model to be evaluated and obtaining multiple first outputs, which are the results of executing the task. A second process involves inputting the multiple first prompts into a large-scale language model for comparison and obtaining multiple second outputs, which are the execution results of the aforementioned tasks. A third process involves comparing the first output and the second output to evaluate the performance of the large-scale language model under evaluation and the large-scale language model under comparison in the business, generating a second prompt to output the evaluation results as a relative evaluation result along with the reasons for the evaluation, inputting the second prompt to the large-scale language model for evaluation, and obtaining the relative evaluation result. A fourth process involves generating a third prompt to generate evaluation criteria used to evaluate the performance of the large-scale language model being evaluated in the business, which include multiple relative evaluation results, inputting the third prompt to the large-scale language model for generation, and obtaining multiple evaluation criteria. A computer system characterized by performing the following.
2. A computer system according to claim 1, A fifth process includes generating a fourth prompt that includes multiple evaluation criteria and causes a determination of whether each of the multiple evaluation criteria is satisfied, inputting the fourth prompt to a large-scale language model for determination, and obtaining the determination results for the multiple evaluation criteria. A sixth process that calculates a score for each of the evaluation criteria of the large-scale language model being evaluated based on the results of the evaluation of multiple evaluation criteria, A seventh process, which uses the scores for each evaluation criterion of the large-scale language model to be evaluated to calculate an overall score representing the performance of the large-scale language model to be evaluated in the business, A computer system characterized by performing the following.
3. A computer system according to claim 2, The fourth process includes generating a sixth prompt to integrate similar evaluation perspectives, inputting the sixth prompt into a large-scale language model for integration, and obtaining the result of integrating the multiple evaluation perspectives. The fifth process is characterized by generating the fourth prompt for the multiple evaluation viewpoints after integration.
4. A computer system according to claim 3, An eighth process involves generating a fifth prompt to determine the importance of multiple evaluation criteria, inputting the fifth prompt into a large-scale language model for determination, and obtaining the importance of multiple evaluation criteria. A ninth process which calculates the weights of the multiple evaluation criteria using the importance of the multiple evaluation criteria, Execute, The seventh process is characterized by calculating the overall score using the score for each evaluation criterion of the large-scale language model to be evaluated and the weights of the multiple evaluation criterions.
5. A computer system according to claim 3, The second and third processes are performed on different large-scale language models for comparison. In the fourth process, the process of selecting the large-scale language model to be compared from among the multiple large-scale language models, and the process of inputting the third prompt to the large-scale language model for generation and obtaining the evaluation perspective are repeatedly executed. The computer system is characterized in that the third prompt includes the evaluation viewpoint obtained in the previous iteration.
6. A computer system according to claim 3, A computer system characterized by performing the second, third, and fourth processes on different large-scale language models of comparison.
7. A computer system according to claim 3, The second process is performed on different large-scale language models of the comparison target, A computer system characterized in that the third prompt includes an instruction to compare the first output with the second output obtained by the second processing performed on each of the different large-scale language models being compared.
8. A computer system according to claim 3, The computer system is characterized in that the third prompt includes an instruction to extract a definition of the evaluation criteria, specific examples that satisfy the evaluation criteria, and specific examples that do not satisfy the evaluation criteria.
9. A computer system according to claim 8, The fourth process is, A process of converting the definition of each of the multiple evaluation perspectives into a vector which is an embedded representation, A process for calculating similarity to determine the similarity of multiple evaluation perspectives using the aforementioned vector, This includes a process of integrating similar evaluation criteria based on the aforementioned similarity, The fifth process is characterized by generating the fourth prompt for the multiple evaluation viewpoints after integration.
10. A method for supporting the evaluation of the performance of large-scale language models in business operations, which is executed by a computer system, The aforementioned computer system, The system comprises a processor, a storage device connected to the processor, and a network interface connected to the processor. Connecting to multiple large-scale language models in a communicative manner, The large-scale language model holds multiple first prompts for causing it to perform tasks, The method for supporting the evaluation of the performance of large-scale language models in the aforementioned tasks is: The first step involves the processor inputting the multiple first prompts to the large-scale language model under evaluation and obtaining multiple first outputs, which are the results of executing the task. The processor inputs the multiple first prompts to the large-scale language model to be compared and obtains multiple second outputs which are the results of executing the task. A third step is for the processor to evaluate the performance of the large-scale language model under evaluation and the large-scale language model under comparison in the business by comparing the first output and the second output, generate a second prompt to output the evaluation result as a relative evaluation result along with the reason for the evaluation, input the second prompt to the large-scale language model for evaluation, and obtain the relative evaluation result. A fourth step involves the processor generating a third prompt to generate evaluation criteria used to evaluate the performance of the large-scale language model being evaluated in the business, which include a plurality of the relative evaluation results, inputting the third prompt to the large-scale language model for generation, and obtaining the plurality of evaluation criteria. A method for supporting the evaluation of the performance of large-scale language models in business operations, characterized by including the following:
11. A method for supporting the evaluation of the performance of a large-scale language model in the business described in claim 10, A fifth step in which the processor generates a fourth prompt that includes a plurality of evaluation criteria and causes the processor to determine whether each of the plurality of evaluation criteria is satisfied, inputs the fourth prompt to a large-scale language model for determination, and obtains the result of the determination of the plurality of evaluation criteria. A sixth step in which the processor calculates a score for each of the evaluation criteria of the large-scale language model under evaluation based on the results of the determination of the multiple evaluation criteria, A seventh step in which the processor calculates an overall score representing the performance of the large-scale language model under evaluation in the business, using the scores for each evaluation criterion of the large-scale language model under evaluation; A method for supporting the evaluation of the performance of large-scale language models in business operations, characterized by including the following:
12. A method for supporting the evaluation of the performance of a large-scale language model in the business described in claim 11, The fourth step includes the processor generating a sixth prompt to integrate similar evaluation criteria, inputting the sixth prompt into a large language model for integration, and obtaining the result of integrating the multiple evaluation criteria. The fifth step is a method for supporting the evaluation of the performance of a large-scale language model in business operations, characterized in that the processor generates the fourth prompt for the multiple evaluation viewpoints after integration.
13. A method for supporting the evaluation of the performance of a large-scale language model in the business described in claim 12, The processor generates a fifth prompt for determining the importance of a plurality of evaluation criteria, inputs the fifth prompt to a large-scale language model for determination, and obtains the importance of the plurality of evaluation criteria. The processor performs a ninth step of calculating the weights of the multiple evaluation criteria using the importance of the multiple evaluation criteria, Includes, The seventh step is a method for supporting the evaluation of the performance of a large-scale language model in business, characterized in that the processor calculates the overall score using the scores for each evaluation criterion of the large-scale language model to be evaluated and the weights of the multiple evaluation criterions.
14. A method for supporting the evaluation of the performance of a large-scale language model in the business described in claim 12, The processor performs the second and third steps for different large-scale language models to be compared. In the fourth step, the processor repeatedly performs the process of selecting the large-scale language model to be compared from among the multiple large-scale language models, and the process of inputting the third prompt to the large-scale language model for generation and obtaining the evaluation perspective. A method for supporting the evaluation of the performance of a large-scale language model in business operations, characterized in that the third prompt includes the evaluation viewpoint obtained in the previous iteration.
15. A method for supporting the evaluation of the performance of a large-scale language model in the business described in claim 12, A method for supporting the evaluation of the performance of a large-scale language model in business operations, characterized in that the processor performs the second step, the third step, and the fourth step for different large-scale language models to be compared.
16. A method for supporting the evaluation of the performance of a large-scale language model in the business described in claim 12, The processor performs the second step for different large-scale language models to be compared, A method for supporting the evaluation of the performance of a large-scale language model in business, characterized in that the third prompt includes an instruction to compare the first output with the second output obtained in the second step performed for each of the different large-scale language models being compared.
17. A method for supporting the evaluation of the performance of a large-scale language model in the business described in claim 12, A method for supporting the evaluation of the performance of a large-scale language model in business operations, characterized in that the third prompt includes instructions to extract a definition of the evaluation criteria, specific examples that satisfy the evaluation criteria, and specific examples that do not satisfy the evaluation criteria.
18. A method for supporting the evaluation of the performance of a large-scale language model in the business described in claim 17, The fourth step described above is: The processor performs the steps of converting the definition of each of the multiple evaluation perspectives into a vector which is an embedding representation, The processor performs the steps of calculating a similarity score for determining the similarity of a plurality of evaluation viewpoints using the vector, The processor includes the step of integrating similar evaluation criteria based on the similarity, The fifth step is a method for supporting the evaluation of the performance of a large-scale language model in business operations, characterized in that the processor generates the fourth prompt for the multiple evaluation viewpoints after integration.
19. A program to be executed by a computer, The aforementioned computer is Connecting to multiple large-scale language models in a communicative manner, The large-scale language model holds multiple first prompts for causing it to perform tasks, The aforementioned program A first process involves inputting the multiple first prompts into the large-scale language model to be evaluated and obtaining multiple first outputs, which are the results of executing the task. A second process involves inputting the multiple first prompts into a large-scale language model for comparison and obtaining multiple second outputs, which are the execution results of the aforementioned tasks. A third process involves comparing the first output and the second output to evaluate the performance of the large-scale language model under evaluation and the large-scale language model under comparison in the business, generating a second prompt to output the evaluation results as a relative evaluation result along with the reasons for the evaluation, inputting the second prompt to the large-scale language model for evaluation, and obtaining the relative evaluation result. A fourth process involves generating a third prompt to generate evaluation criteria used to evaluate the performance of the large-scale language model being evaluated in the business, which include multiple relative evaluation results, inputting the third prompt to the large-scale language model for generation, and obtaining multiple evaluation criteria. A program characterized by causing the computer to execute the above.