Computer system and method for evaluating machine learning model

The system evaluates generative AI models by generating multiple prompts with varying expression patterns and comparing answers to expected outputs, addressing the challenge of unknown training data and ensuring consistent performance assessments.

WO2026110397A1PCT designated stage Publication Date: 2026-05-28HITACHI LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/022330
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-21
Filing Date
2025-06-20
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Existing evaluation techniques for generative AI models, particularly Large Language Models (LLMs), struggle to accurately assess their performance in various usage scenarios due to unknown training data, leading to inconsistencies in response accuracy.

Method used

A computer system and method that evaluates generative AI models by generating multiple prompts with varying expression patterns, comparing answers to expected outputs, and calculating similarity scores to determine proficiency across different usage scenarios.

Benefits of technology

Enables accurate evaluation of generative AI responses by considering variations in prompt expression, ensuring consistent and reliable performance assessments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025022330_28052026_PF_FP_ABST
    Figure JP2025022330_28052026_PF_FP_ABST
Patent Text Reader

Abstract

This computer system is communicably connected to a machine learning model, receives an evaluation instruction including information designating a use case of the machine learning model, an input prompt, and an expected answer, generates a plurality of conversion prompts by changing an expression pattern of the input prompt, inputs the input prompt and each of the plurality of conversion prompts to the machine learning model to acquire a plurality of generated answers, compares the expected answer and the generated answer for each of the plurality of generated answers, calculates a similarity score indicating a similarity between the expected answer and the generated answer, and executes statistical processing using the similarity score of the plurality of generated answers, to thereby calculate an index for evaluating the answer accuracy of the machine learning model in the use case.
Need to check novelty before this filing date? Find Prior Art

Description

Computer System and Evaluation Method of Machine Learning Model Incorporation by Reference

[0001] This application claims the priority of Japanese Patent Application No. 2024-203059 filed on November 21, 2024, and incorporates its content by reference into this application.

[0002] The present invention relates to an evaluation technique for generative AI.

[0003] The utilization of generative AI is progressing for various domains and various purposes. The accuracy of the answers (outputs) of generative AI for usage scenarios such as the domains and purposes for introducing generative AI is important.

[0004] As a technique for evaluating the accuracy of a machine learning model, the technique described in Patent Document 1 is known. Patent Document 1 describes "an abnormality degree calculation unit that calculates the abnormality degree of partial time-series data, which is at least a part of time-series data; an abnormality sign determination unit that determines the partial time-series data as abnormality sign data when the abnormality degree of the partial time-series data exceeds a threshold; a normal determination unit that calculates the reliability of the abnormality sign data based on feature data corresponding to the feature amount of the abnormality sign data, and determines the abnormality sign data as normal data when the reliability is unacceptable. The normal determination unit stores the feature data corresponding to the feature amount of the abnormality sign data in a feature data storage unit. The normal determination unit stores the feature data corresponding to the abnormality sign data in a feature data layer, which is a layer corresponding to the range of the value of the feature amount of the abnormality sign data, in the feature data storage unit. For each feature data layer, the normal determination unit stores the cumulative number of abnormality sign data corresponding to the feature data layer as the feature data. The normal determination unit calculates the reliability of the abnormality sign data based on the ratio of the cumulative number of abnormality sign data corresponding to the abnormality sign data to the total number of the cumulative number of abnormality sign data."

[0005] Patent No. 7016450

[0006] The technology described in Patent Document 1 assumes that the training data is known. However, in Large Language Models (LLMs) used to realize generative AI, the training data is often unknown.

[0007] The present invention aims to realize a technology for evaluating the accuracy of responses generated by AI in various usage scenarios.

[0008] A typical example of the invention disclosed in this application is as follows: A computer system comprising a processor, a storage device connected to the processor, and a network interface connected to the processor, which is communicatively connected to a machine learning model that accepts a prompt including the content of a task as input, executes the task defined by the prompt, and generates an answer; the processor accepts an evaluation instruction including information specifying the usage scenario of the machine learning model, an input prompt, and an expected answer; generates a plurality of conversion prompts by changing the expression pattern of the input prompt; inputs each of the input prompt and the plurality of conversion prompts to the machine learning model to obtain a plurality of generated answers; compares each of the plurality of generated answers with the expected answer to calculate a similarity score representing the similarity between the expected answer and the generated answer; calculates an index for evaluating the accuracy of the machine learning model's answer in the usage scenario by performing statistical processing using the similarity scores of the plurality of generated answers; and records a first calculation result including the index associated with the information specifying the usage scenario in the storage device.

[0009] According to the present invention, the accuracy of the generated AI's responses in actual usage scenarios can be evaluated. Other issues, configurations, and effects will be clarified by the following description of the embodiments.

[0010] This figure shows an example of the functional configuration of the evaluation system in Example 1. This figure shows an example of the hardware configuration of the computer constituting the evaluation system in Example 1. This figure shows an example of the screen displayed by the evaluation system in Example 1. This figure shows an example of the information managed by the prompt / answer storage unit of Example 1. This figure shows an example of the information managed by the prompt / answer storage unit of Example 1. This figure shows an example of the information managed by the comparison result storage unit of Example 1. This figure shows an example of the information managed by the proficiency storage unit of Example 1. This is a flowchart illustrating the overview of the processing performed by the evaluation system in Example 1. This is a flowchart illustrating an example of the prompt conversion processing performed by the evaluation system in Example 1. This is a flowchart illustrating an example of the answer comparison processing performed by the evaluation system in Example 1. This is a flowchart illustrating an example of the answer comparison processing performed by the evaluation system in Example 1. This figure shows an example of the similarity calculation method in the answer comparison processing of Example 1. This is a flowchart illustrating an example of the proficiency determination processing performed by the evaluation system in Example 1. This figure shows an example of the screen displayed by the evaluation system in Example 1. This figure shows an example of the functional configuration of the evaluation system in Example 2. This figure shows an example of the screen displayed by the evaluation system in Example 2. This figure shows an example of the information managed by the prompt / answer storage unit of Example 2. This figure shows an example of the information managed by the prompt / answer storage unit of Example 2. This figure shows an example of information managed by the comparison result storage unit of Example 2. This figure shows an example of information managed by the proficiency storage unit of Example 2. This figure shows an example of information managed by the parameter storage unit of Example 2. This figure shows an example of information managed by the fitness storage unit of Example 2. This is a flowchart illustrating the overview of the processing performed by the evaluation system of Example 2. This is a flowchart illustrating an example of the answer comparison processing performed by the evaluation system of Example 2. This is a flowchart illustrating an example of the answer comparison processing performed by the evaluation system of Example 2. This figure shows an example of the method for calculating the similarity score in the answer comparison processing of Example 2. This is a flowchart illustrating an example of the proficiency determination processing performed by the evaluation system of Example 2. This is a flowchart illustrating an example of the proficiency determination processing performed by the evaluation system of Example 2.This is a flowchart illustrating an example of the suitability determination process performed by the evaluation system in Example 2. This is a diagram showing an example of the screen displayed by the evaluation system in Example 2. This is a diagram showing an example of the screen displayed by the evaluation system in Example 2.

[0011] The embodiments of the present invention will be described below with reference to the drawings. However, the present invention is not to be construed as being limited to the embodiments described below. It will be readily apparent to those skilled in the art that the specific configuration can be modified without departing from the spirit or intent of the present invention.

[0012] In the configuration of the invention described below, identical or similar components or functions are denoted by the same reference numerals, and redundant descriptions are omitted.

[0013] The designations "First," "Second," "Third," etc., used in this specification are for the purpose of identifying constituent elements and do not necessarily limit their number or order.

[0014] Figure 1 shows an example of the functional configuration of the evaluation system 100 of Example 1. Figure 2 shows an example of the hardware configuration of the computer that constitutes the evaluation system 100 of Example 1.

[0015] The evaluation system 100 evaluates the accuracy of the generative AI's responses in a given usage scenario. Here, the usage scenario is a concept that includes the domain in which the generative AI is used and the purpose for which the generative AI is used. The evaluation system 100 connects to the natural language processing service 101 via a network (not shown).

[0016] The natural language processing service 101 is a service provided by a system that implements generative AI. This system provides the service using a Language Language Model (LLM). The evaluation system 100 may maintain multiple LLMs. The LLM receives prompts containing task content, such as questions written in natural language, understands the task content, and generates and outputs text that serves as the answer by executing the task. The accuracy of the generative AI's answer is equivalent to the accuracy of the LLM's answer.

[0017] The evaluation system 100 is a system implemented using a computer 200 as shown in Figure 2. The computer 200 has a processor 201, a main memory 202, a secondary memory 203, and a network interface 204. Each hardware element is connected via a bus 205.

[0018] The processor 201 executes a program stored in the main memory 202. By executing processing according to the program, the processor 201 operates as a functional unit (module) that realizes a specific function. In the following description, when the processing is described with a functional unit as the subject, it indicates that the processor 201 is executing a program that realizes that functional unit.

[0019] The main memory 202 is a memory device that stores the program executed by the processor 201 and the information used by the program. The main memory 202 is also used as a work area. The secondary memory 203 is a large-capacity storage device such as an HDD (Hard Disk Drive) or SSD (Solid State Drive). The program and information stored in the main memory 202 may also be stored in the secondary memory 203. In this case, the processor 201 reads the program and information from the secondary memory 203 and loads it into the main memory 202. The network interface 204 is an interface for connecting to a network.

[0020] The evaluation system 100 includes an input unit 110, a prompt conversion unit 111, an answer comparison unit 112, an answer generation unit 113, a proficiency determination unit 114, a display unit 115, a prompt / answer storage unit 120, a comparison result storage unit 121, and a proficiency storage unit 122.

[0021] The prompt / answer storage unit 120 stores combinations of prompts and expected answers. The expected answers are ideal answers output from the LLM. The comparison result storage unit 121 stores the comparison result between the expected answers and the answers output by the LLM (generated answers). The proficiency storage unit 122 stores the proficiency level, which represents the accuracy of the LLM's answers in the usage scenario. If there is a large difference between the expected answers and the generated answers, it indicates that the LLM has not learned enough for the usage scenario and that its proficiency level is low.

[0022] The input unit 110 accepts input of various types of information. The display unit 115 displays various types of information. The display unit 115 also provides various interfaces.

[0023] The prompt conversion unit 111 generates a prompt by changing the expression pattern of the input prompt. Here, the expression pattern is a concept that includes language and grammar (style, word order, etc.). The answer comparison unit 112 compares the expected answer with the generated answer. The answer generation unit 113 sends the prompt to the natural language processing service 101 and obtains the answer. The proficiency determination unit 114 determines the proficiency of the LLM based on the comparison results.

[0024] Furthermore, regarding the functional units of the evaluation system 100, multiple functional units may be combined into a single functional unit, or a single functional unit may be divided into multiple functional units according to its function.

[0025] Figure 3 shows an example of a screen displayed by the evaluation system 100 of Example 1.

[0026] The display unit 115 of the evaluation system 100 displays a screen 300 on a terminal or the like operated by the user. The screen 300 is an interface for setting prompts and expected answers, and includes an input field 301, a selection field 302, an input field 303, and operation buttons 304.

[0027] Input field 301 is for entering the domain in which the generation AI will be introduced, i.e., the usage scenario for the generation AI. Selection field 302 is for selecting the LLM to be evaluated. Input field 303 is for entering prompts and expected answers. Multiple combinations of prompts and expected answers can be entered in input field 303. Operation button 304 is an operation button for outputting evaluation instructions.

[0028] When the user operates the operation button 304, an evaluation instruction including the domain, a list of LLMs, and a combination of prompts and expected answers is output to the evaluation system 100.

[0029] Figures 4A and 4B show an example of the information managed by the prompt / answer storage unit 120 of Embodiment 1.

[0030] The prompt / answer storage unit 120 manages, for example, tables 400 and 410.

[0031] Table 400 is a table for managing metadata of input prompt and expected response combinations. Table 400 stores entries including problem ID 401 and domain 402.

[0032] Problem ID 401 is a field that stores an identifier for identifying the combination of the input prompt and the expected answer. Domain 402 is a field that stores the domain in which the generation AI will be introduced.

[0033] Table 410 is a table for managing combinations of prompts and expected answers. Table 410 stores entries that include a problem ID 411, a representation pattern 412, a prompt 413, and an expected answer 414.

[0034] Problem ID 411 is the same field as Problem ID 401. Expression pattern 412 is a field that stores the expression pattern of the prompt. Prompt 413 is a field that stores the prompt. Expected answer 414 is a field that stores the expected answer.

[0035] Even with the same task content, the generation AI's response changes depending on the expression pattern of the prompt. Therefore, in this embodiment, multiple prompts with different expression patterns are generated from the input prompt, and the accuracy of the generation AI's response is evaluated while taking into account the variation in the response.

[0036] Tables 400 and 410 may be combined into a single table. Furthermore, the prompt / answer storage unit 120 may manage various types of information in a data format other than a table format.

[0037] Figure 5 shows an example of the information managed by the comparison result storage unit 121 in Embodiment 1.

[0038] The comparison result storage unit 121 manages, for example, a table 500. The table 500 stores entries including an ID 501, a question ID 502, a representation pattern 503, a model 504, a generated answer 505, a similarity 506, and a date and time 507.

[0039] The ID 501 is a field for storing an identifier for identifying the comparison result. The question ID 502 is the same field as the question ID 401. The representation pattern 503 is the same field as the representation pattern 412.

[0040] The model 504 is a field for storing the identification information of the LLM. The generated answer 505 is a field for storing the generated answer output by the LLM. The similarity 506 is a field for storing the similarity representing the similarity between the assumed answer and the generated answer. The date and time 507 is a field for storing the date and time when the comparison was made.

[0041] Note that the comparison result storage unit 121 may manage various information in a data format other than the table format.

[0042] FIG. 6 is a diagram showing an example of information managed by the proficiency storage unit 122 of Example 1.

[0043] The proficiency storage unit 122 manages, for example, a table 600. The table 600 stores entries including an ID 601, a model 602, a domain 603, a similarity score 604, a variation score 605, and a date and time 606.

[0044] The ID 601 is a field for storing the identification information of the determination result of the proficiency of the LLM. The model 602 is a field for storing the identification information of the LLM. The domain 603 is a field for storing the domain.

[0045] The similarity score 604 is a field for storing a similarity score representing the similarity between the assumed answer and the generated answer in the usage scenario. The variation score 605 is a field for storing a variation score representing the degree of fluctuation of the generated answer in the usage scenario. The similarity score and the variation score are examples of indices representing proficiency. The date and time 606 is a field for storing the date and time when the proficiency of the LLM was determined.

[0046] Note that the proficiency memory unit 122 may manage various types of information in a data format other than the table format.

[0047] FIG. 7 is a flowchart for explaining the outline of the process executed by the evaluation system 100 of Example 1.

[0048] When the evaluation system 100 receives an evaluation request from a user, it displays the screen 300 and accepts an evaluation instruction (step S101).

[0049] The display unit 115 displays the screen 300 on, for example, a terminal operated by the user. The input unit 110 receives an evaluation instruction via the screen 300. The input unit 110 outputs the domain included in the evaluation instruction, the list of LLMs, and the combination of the prompt and the expected answer to the prompt conversion unit 111.

[0050] The prompt conversion unit 111 adds as many entries to the table 400 as the number of combinations of the prompt and the expected answer, and sets identification information in the question ID 401 of the added entry. Further, the prompt conversion unit 111 sets the domain included in the evaluation instruction in the domain 402 of the added entry.

[0051] The prompt conversion unit 111 adds as many entries to the table 410 as the number of combinations of the prompt and the expected answer, and sets identification information in the question ID 411 of the added entry. Further, the prompt conversion unit 111 sets the prompt and the expected answer in the prompt 413 and the expected answer 414 of the added entry.

[0052] The prompt conversion unit 111 identifies the expression pattern of each prompt, and sets the identified expression pattern in the expression pattern 412 of the corresponding entry. As a method for identifying the expression pattern of the prompt, a method using a preset rule and a method using an LLM can be considered. Note that a column for inputting the expression pattern may be provided on the screen 300.

[0053] The evaluation system 100 executes a prompt conversion process (step S102). The details of the prompt conversion process will be described later.

[0054] The evaluation system 100 performs a response comparison process (step S103). Details of the response comparison process will be described later.

[0055] The evaluation system 100 executes a proficiency assessment process (step S104). Details of the proficiency assessment process will be described later.

[0056] Figure 8 is a flowchart illustrating an example of the prompt conversion process performed by the evaluation system 100 of Example 1.

[0057] The prompt conversion unit 111 refers to the table 400 and selects a problem ID (step S201).

[0058] The prompt conversion unit 111 obtains a combination of prompt and expected answer corresponding to the selected problem ID from the table 410 (step S202).

[0059] The prompt conversion unit 111 generates a list of expression patterns for the problem ID (step S203).

[0060] In this embodiment, it is assumed that a definition list containing predefined expression patterns is set in advance. The prompt conversion unit 111 obtains an expression pattern from the expression pattern 412 of the entry corresponding to the selected problem ID in table 410. The prompt conversion unit 111 generates an expression pattern list by excluding the expression pattern obtained from the definition list.

[0061] The prompt conversion unit 111 selects an expression pattern from the expression pattern list (step S204).

[0062] The prompt conversion unit 111 converts the acquired prompt into a selected expression pattern (step S205). In this embodiment, in the case of grammatical conversion, only the prompt is changed, while in the case of language conversion, both the prompt and the expected answer are converted.

[0063] The prompt conversion unit 111 registers the prompt whose expression pattern has been converted in the table 410 (step S206).

[0064] Specifically, the prompt conversion unit 111 adds an entry to the table 410, sets the selected problem ID in the entry's problem ID 411, and sets the selected expression pattern in the expression pattern 412. The prompt conversion unit 111 sets the prompt 413 of the added entry to the prompt with the converted expression pattern. The prompt conversion unit 111 sets the expected answer 414 of the added entry to the expected answer. In the case of grammar conversion, the acquired expected answer is set as is in the expected answer 414, and in the case of language conversion, the expected answer 414 is set to the expected answer with the language converted.

[0065] The prompt conversion unit 111 determines whether processing has been completed for all expression patterns in the expression pattern list (step S207).

[0066] If processing has not been completed for all expression patterns in the expression pattern list, the prompt conversion unit 111 returns to step S204.

[0067] When processing is complete for all expression patterns in the expression pattern list, the prompt conversion unit 111 determines whether processing has been completed for all problem IDs (step S208).

[0068] If processing is not complete for all problem IDs, the prompt conversion unit 111 returns to step S201.

[0069] Once processing is complete for all problem IDs, the prompt conversion unit 111 terminates the prompt conversion process.

[0070] Figures 9A and 9B are flowcharts illustrating an example of the response comparison process performed by the evaluation system 100 of Example 1. Figure 10 is a diagram showing an example of the method for calculating similarity in the response comparison process of Example 1.

[0071] The response comparison unit 112 selects an LLM from the list of LLMs included in the evaluation instruction (step S301).

[0072] The answer comparison unit 112 refers to the table 400 and selects a problem ID (step S302).

[0073] The answer comparison unit 112 refers to the table 410 and selects one combination from the prompt and expected answer combinations corresponding to the selected problem ID (step S303). At this time, the answer comparison unit 112 sets the evaluation count to 1.

[0074] The response comparison unit 112 obtains a generated response by sending the selected LLM information and the selected combination prompt to the response generation unit 113 (step S304).

[0075] The response generation unit 113 inputs a prompt to the natural language processing service 101, which provides a service using the selected LLM, and obtains a generated response. The response generation unit 113 then transmits the generated response to the response comparison unit 112.

[0076] The response comparison unit 112 identifies the response items of the expected response and the response items of the generated response for the selected combination (step S305).

[0077] In this embodiment, the answer includes multiple answer items. For example, if the expected answer includes answer items in CSV format, the answer comparison unit 112 retrieves the comma-separated answer items. Alternatively, answer items can be retrieved from the generated answer by using a prompt that instructs the generation of the generated answer in CSV format. Furthermore, LLM may be used to retrieve the answer items from both the expected answer and the generated answer.

[0078] The answer comparison unit 112 generates pairs of answer items for the expected answers and answer items for the generated answers (step S306).

[0079] The response comparison unit 112 selects a pair (step S307) and calculates the similarity (individual similarity) between the response items constituting the pair (step S308). Since the response items are strings, the similarity between strings can be calculated as the similarity between response items using known techniques such as Word2Vec. The response comparison unit 112 stores the pairs and their similarities in association.

[0080] The answer comparison unit 112 determines whether processing has been completed for all pairs (step S309).

[0081] If processing is not complete for all pairs, the answer comparison unit 112 returns to step S307.

[0082] Once processing is complete for all pairs, the response comparison unit 112 identifies the maximum individual similarity for each response item of the expected response based on the individual similarity of all pairs (step S310).

[0083] Here, the process of step S310 will be explained using Figure 10. In the matrix 1000 shown in Figure 10, the columns correspond to the answer items of the expected answers, and the rows correspond to the answer items of the generated answers. The cells store the individual similarity score corresponding to the pair of answer items. In step S310, the answer comparison unit 112 identifies the maximum value of the individual similarity score for each column.

[0084] The response comparison unit 112 calculates a similarity score representing the similarity between the expected response and the generated response using the maximum value of the individual similarity score for each response item of the expected response (step S311).

[0085] Specifically, the response comparison unit 112 calculates the similarity score as the average of the maximum values ​​of the individual similarity scores for each response item of the expected responses. Note that the calculation method described above is just one example and is not limited thereto. For example, the similarity between the expected responses and the generated responses may be calculated using LLM (Likely Least Readiness Model).

[0086] The response comparison unit 112 records the calculation result (step S312). At this time, the response comparison unit 112 increments the evaluation count by 1.

[0087] Specifically, the answer comparison unit 112 adds an entry to the table 500 and sets identification information in the ID 501 of the added entry. The answer comparison unit 112 sets the selected problem ID in the problem ID 502 of the added entry, sets the expression pattern of the acquired prompt in the expression pattern 503, and sets the identification information of the selected LLM in the model 504. The answer comparison unit 112 sets the acquired generated answer in the generated answer 505 of the added entry and sets the calculated similarity in the similarity 506. The answer comparison unit 112 sets the start date and time of the answer comparison process in the date and time 507 of the added entry.

[0088] The response comparison unit 112 determines whether the number of evaluations is greater than the threshold (step S313). The threshold is assumed to be set in advance. Note that the threshold can be changed as appropriate.

[0089] In this way, the response comparison unit 112 performs the evaluation of the similarity between the expected response and the generated response multiple times using the same prompt. Since the AI's response may differ even with the same prompt, the accuracy of the evaluation is ensured by performing the evaluation multiple times.

[0090] If the number of evaluations is below the threshold, the response comparison unit 112 returns to step S304.

[0091] If the number of evaluations is greater than the threshold, the response comparison unit 112 determines whether processing has been completed for all prompt and expected response combinations corresponding to the selected problem ID (step S314).

[0092] If processing has not been completed for all prompt and expected answer combinations corresponding to the selected problem ID, the answer comparison unit 112 returns to step S303.

[0093] In this embodiment, the response comparison unit 112 performs the evaluation of the similarity between the expected response and the generated response multiple times using prompts with the same task content but different expression patterns. Even if the task content is the same, if the expression pattern of the prompt is different, the response of the generating AI may be different. Therefore, the accuracy of the evaluation is ensured by performing the evaluation multiple times using prompts with different expression patterns.

[0094] When processing is complete for all prompt and expected answer combinations corresponding to the selected problem ID, the answer comparison unit 112 determines whether processing has been completed for all problem IDs (step S315).

[0095] If processing is not complete for all problem IDs, the answer comparison unit 112 returns to step S302.

[0096] If processing is completed for all problem IDs, the answer comparison unit 112 determines whether processing has been completed for all LLMs in the LLM list included in the evaluation instruction (step S316).

[0097] If processing has not been completed for all LLMs in the LLM list, the answer comparison unit 112 returns to step S301.

[0098] When processing is complete for all LLMs in the LLM list, the answer comparison unit 112 terminates the answer comparison process.

[0099] Figure 11 is a flowchart illustrating an example of the proficiency determination process performed by the evaluation system 100 of Example 1.

[0100] The proficiency determination unit 114 selects an LLM from the list of LLMs included in the evaluation instruction (step S401).

[0101] The proficiency determination unit 114 refers to the table 400 and selects a problem ID (step S402).

[0102] The proficiency assessment unit 114 calculates a similarity score for the problem ID by referring to the table 500 (step S403).

[0103] Specifically, the proficiency determination unit 114 refers to the table 500 and searches for entries in which the identification information of the selected LLM is set in the model 504 and the selected problem ID is set in the problem ID 502. The proficiency determination unit 114 obtains the similarity stored in the similarity 506 of the searched entries and calculates the average of the similarity values ​​as the similarity score of the problem ID.

[0104] The proficiency assessment unit 114 refers to the table 500 and calculates the variation score for the problem ID (step S404).

[0105] The proficiency determination unit 114 calculates the standard deviation and mean value of the similarity obtained in step S403. The proficiency determination unit 114 calculates the variation score of the problem ID by dividing the standard deviation of the similarity by the mean value of the similarity. In this embodiment, the greater the variation in the generated answers, the larger the variation score.

[0106] The proficiency determination unit 114 determines whether processing has been completed for all problem IDs (step S405).

[0107] If processing is not complete for all problem IDs, the proficiency determination unit 114 returns to step S402.

[0108] Once processing is complete for all problem IDs, the proficiency determination unit 114 selects a domain (step S406).

[0109] The proficiency determination unit 114 selects one domain from the list of domains. Note that when step S406 is executed for the first time, the proficiency determination unit 114 generates a list of domains using table 400.

[0110] The proficiency determination unit 114 calculates a similarity score for the usage scenario (step S407). In this embodiment, the domain is the information that specifies the usage scenario.

[0111] Specifically, the proficiency assessment unit 114 refers to the table 400 to identify the problem ID corresponding to the selected domain. The proficiency assessment unit 114 calculates the average of the similarity scores of the identified problem IDs as the similarity score for the usage scenario (domain).

[0112] The proficiency assessment unit 114 calculates the variation score in the usage scenario (step S408).

[0113] Specifically, the proficiency assessment unit 114 refers to the table 400 to identify the problem ID corresponding to the selected domain. The proficiency assessment unit 114 calculates the average of the variation scores of the identified problem IDs as the variation score in the usage scenario (domain).

[0114] The proficiency determination unit 114 records the calculation result (step S409).

[0115] Specifically, the proficiency determination unit 114 adds an entry to the table 600 and sets identification information in the ID 601 of the added entry. The proficiency determination unit 114 sets the identification information of the selected LLM in the model 602 of the added entry and sets the selected domain in the domain 603. The proficiency determination unit 114 sets the calculated similarity score 604 and variation score 605 in the similarity score 604 and variation score 605 of the added entry. The proficiency determination unit 114 sets the start date and time of the response comparison process in the date and time 606 of the added entry.

[0116] The proficiency determination unit 114 determines whether processing has been completed for all domains (step S410).

[0117] If processing is not complete for all domains, the proficiency determination unit 114 returns to step S406.

[0118] If processing is completed for all domains, the proficiency determination unit 114 determines whether processing has been completed for all LLMs in the LLM list included in the evaluation instruction (step S411).

[0119] If processing has not been completed for all LLMs in the LLM list, the proficiency determination unit 114 returns to step S401.

[0120] When processing is completed for all LLMs in the LLM list, the proficiency determination unit 114 terminates the proficiency determination process.

[0121] Figure 12 shows an example of a screen displayed by the evaluation system 100 of Example 1.

[0122] When the display unit 115 of the evaluation system 100 receives a request from a user to display the LLM proficiency level in a specific domain, it displays screen 1200 on the terminal or other device operated by the user. Screen 1200 includes a display field 1201 and a table 1202.

[0123] Display field 1201 displays the scores of each LLM in the specified domain in graph format. Table 1202 is a table that displays the details of each LLM score in the specified domain. The order of LLMs in the graph and table can be changed based on the release date and usage price. Also, even for the same LLM, accuracy improves with version upgrades. Therefore, it is possible to perform evaluations periodically and display the history in chronological order.

[0124] As described above, the evaluation system 100 of Example 1 compares the generated response obtained by inputting prompts with different expression patterns with the expected response, and compares the generated response obtained by inputting the same prompt multiple times with the expected response, thereby enabling the evaluation of the accuracy of the generated AI's response in usage scenarios, taking into account the output characteristics of the generated AI.

[0125] In Example 2, the score calculated as proficiency differs from that in Example 1. Furthermore, Example 2 differs from Example 1 in that it evaluates the accuracy of the generated AI's responses for each purpose (application) and degree of social impact. The following describes Example 2, focusing on the differences from Example 1.

[0126] Figure 13 shows an example of the functional configuration of the evaluation system 100 of Example 2.

[0127] The evaluation system 100 of Example 2 differs from that of Example 1 in that it includes a fitness determination unit 116, a parameter storage unit 123, and a fitness storage unit 124. The hardware configuration of the evaluation system 100 is the same as that of Example 1.

[0128] Figure 14 shows an example of a screen displayed by the evaluation system 100 of Example 2.

[0129] The display unit 115 of the evaluation system 100 displays a screen 1400 on a terminal or the like operated by the user. The screen 1400 is an interface for setting prompts and expected answers, and includes selection fields 1401, 1402, input fields 1403, 1404, 1405, 1406, and operation buttons 1407.

[0130] Selection field 1401 is for selecting the industry in which the generation AI will be introduced. Selection field 1402 is for selecting the field in which the generation AI will be introduced. In Example 2, the combination of industry and field becomes the domain. Input field 1403 is for selecting the purpose of introducing the generation AI. Selection field 1404 is for selecting the magnitude of the social impact that the generation AI's responses will have. Selection field 1405 is for selecting the LLM to be evaluated. Input field 1406 is for inputting prompts and expected answers. Multiple combinations of prompts and expected answers can be entered in input field 1406. Operation button 1407 is an operation button for outputting evaluation instructions.

[0131] When a user operates the operation button 1407, evaluation instructions including the industry, field, purpose of implementation, magnitude of social impact, LLM list, and combinations of prompts and expected answers are output to the evaluation system 100.

[0132] Figures 15A and 15B show an example of the information managed by the prompt / answer storage unit 120 of Embodiment 2.

[0133] The prompt / response storage unit 120 manages, for example, table 1500 and table 1510.

[0134] Table 1500 is a table for managing metadata of input prompt and expected response combinations. Table 1500 stores entries including problem ID 1501, domain 1502, objective 1503, and severity 1504.

[0135] Problem ID 1501 and Domain 1502 are the same fields as Problem ID 401 and Domain 402. In Example 2, Domain 1502 is set to include the industry and field.

[0136] Objective 1503 is a field that stores the purpose of introducing the generation AI. Importance 1504 is a field that stores the magnitude of the social impact.

[0137] Table 1510 is a table for managing combinations of prompts and expected answers. Table 1510 stores entries that include problem ID 1511, expression pattern 1512, prompt 1513, and expected answer 1514.

[0138] Problem ID 1511, expression pattern 1512, prompt 1513, and expected answer 1514 are the same fields as Problem ID 411, expression pattern 412, prompt 413, and expected answer 414.

[0139] Tables 1500 and 1510 may be combined into a single table. Furthermore, the prompt / answer storage unit 120 may manage various types of information in a data format other than a table format.

[0140] Figure 16 shows an example of the information managed by the comparison result storage unit 121 in Embodiment 2.

[0141] The comparison result storage unit 121 manages, for example, a table 1600. The table 1600 stores entries including ID 1601, problem ID 1602, representation pattern 1603, model 1604, generated answer 1605, similarity 1606, error rate 1607, missing answer rate 1608, and date and time 1609.

[0142] ID 1601, Problem ID 1602, Representation Pattern 1603, Model 1604, Generated Response 1605, Similarity 1606, and Date / Time 1609 are the same fields as ID 501, Problem ID 502, Representation Pattern 503, Model 504, Generated Response 505, Similarity 506, and Date / Time 507.

[0143] The error rate 1607 is a field that stores an indicator representing the degree of error in the generated response compared to the expected response. The missing response rate 1608 is a field that stores an indicator representing the degree of missing responses in the generated response compared to the expected response.

[0144] The comparison result storage unit 121 may manage various types of information in a data format other than a table format.

[0145] Figure 17 shows an example of the information managed by the proficiency memory unit 122 of Embodiment 2.

[0146] The proficiency memory unit 122 manages, for example, table 1700. Table 1700 stores entries including ID 1701, model 1702, domain 1703, purpose 1704, importance 1705, similarity score 1706, variation score 1707, incorrect answer score 1708, missing answer score 1709, and date and time 1710.

[0147] ID 1701, Model 1702, Domain 1703, Similarity Score 1706, Variation Score 1707, and Date / Time 1710 are the same fields as ID 601, Model 602, Domain 603, Similarity Score 604, Variation Score 605, and Date / Time 606.

[0148] Objective 1704 is a field that stores the purpose of introducing the generation AI. Importance 1705 is a field that stores the magnitude of the social impact.

[0149] The error response score 1708 is a field that stores the error response score, which represents the degree of error in the generated response in the usage scenario. The missing response score 1709 is a field that stores the missing response score, which represents the degree of missing responses in the generated response in the usage scenario.

[0150] Similarity score, variation score, incorrect response score, and missing response score are examples of indicators that represent proficiency.

[0151] The proficiency level storage unit 122 may manage various types of information in a data format other than a table format.

[0152] Figure 18 shows an example of the information managed by the parameter storage unit 123 of Embodiment 2.

[0153] The parameter storage unit 123 manages, for example, table 1800. Table 1800 stores entries including purpose 1801, importance 1802, parameter (1) 1803, parameter (2) 1804, and parameter (3) 1805.

[0154] Objective 1801 is a field that stores the purpose of introducing the generation AI. Importance 1802 is a field that stores the magnitude of the social impact.

[0155] Parameter (1) 1803, parameter (2) 1804, and parameter (3) 1805 are fields that store parameters used by the fitness determination unit 116.

[0156] The parameter storage unit 123 may manage various types of information in a data format other than a table format.

[0157] Figure 19 shows an example of the information managed by the fitness storage unit 124 of Embodiment 2.

[0158] The fitness storage unit 124 manages, for example, a table 1900. The table 1900 stores entries including a model 1901, a purpose 1902, a severity level 1903, and a fitness level 1904.

[0159] Model 1901 is a field that stores the identification information of the LLM. Purpose 1902 is a field that stores the purpose of introducing the generative AI. Importance 1903 is a field that stores the magnitude of the social impact. Fit 1904 is a field that stores the fit score, which represents the degree to which the generative AI fits the purpose.

[0160] The fitness storage unit 124 may manage various types of information in a data format other than a table format.

[0161] Figure 20 is a flowchart illustrating the overview of the processes performed by the evaluation system 100 in Example 2.

[0162] In Example 2, the evaluation system 100 performs a proficiency determination process, and then performs a suitability determination process (step S151).

[0163] The prompt conversion process in Example 2 is the same as in Example 1, so its explanation is omitted. In Example 2, the answer comparison process and the proficiency determination process differ in some respects from Example 1.

[0164] Figures 21A and 21B are flowcharts illustrating an example of the response comparison process performed by the evaluation system 100 in Example 2. Figure 22 is a diagram showing an example of the method for calculating the similarity score in the response comparison process of Example 2.

[0165] The response comparison unit 112 calculates the similarity and then calculates the response omission rate (step S351). Specifically, the following processes are performed.

[0166] (S351-1) The answer comparison unit 112 determines whether there are pairs with a similarity greater than a threshold for each answer item of the expected answers. That is, for each column of the matrix 2200 shown in Figure 22, it is determined whether there are cells with a similarity greater than a threshold. The threshold is, for example, 80%. Note that a threshold may be set for each domain. In Figure 22, there are no pairs with a similarity greater than the threshold for the answer item "My Number Utilization Administration System".

[0167] In the following explanation, response items for which no pairs with a similarity greater than the threshold exist are referred to as missing response items.

[0168] (S351-2) The response comparison unit 112 calculates the response omission rate by dividing the number of missing response items by the number of expected response items. In this embodiment, a smaller response omission rate indicates fewer missed responses.

[0169] The response comparison unit 112 calculates the rate of incorrect responses (step S352). Specifically, the following process is performed.

[0170] (S352-1) The response comparison unit 112 determines whether there are pairs with a similarity greater than a threshold for each response item in the generated responses. That is, for each row of the matrix 2200 shown in Figure 22, it is determined whether there are cells with a similarity greater than a threshold. The threshold is, for example, 80%. Note that a threshold may be set for each domain. In Figure 22, there are no pairs with a similarity greater than the threshold for the response item "system network".

[0171] In the following explanation, a generated response item for which no pairs with a similarity score greater than the threshold exist will be referred to as an incorrect response item.

[0172] (S352-2) The response comparison unit 112 calculates the error rate by dividing the number of incorrect response items by the number of expected response items in the generated response. In this embodiment, a smaller error rate indicates fewer errors in the answer.

[0173] The response comparison unit 112 records the calculation result (step S353). At this time, the response comparison unit 112 increments the evaluation count by 1.

[0174] Specifically, the answer comparison unit 112 adds an entry to the table 1600 and sets the identification information in the ID 1601 of the added entry. The answer comparison unit 112 sets the selected problem ID in the problem ID 1602 of the added entry, sets the expression pattern of the acquired prompt in the expression pattern 1603, and sets the identification information of the selected LLM in the model 1604. The answer comparison unit 112 sets the acquired generated answer in the generated answer 1605 of the added entry, and sets the integrated similarity 1606, incorrect answer rate 1607, and missing answer rate 1608. The answer comparison unit 112 sets the start date and time of the answer comparison process in the date and time 1609 of the added entry.

[0175] Figures 23A and 23B are flowcharts illustrating an example of the proficiency determination process performed by the evaluation system 100 of Embodiment 2.

[0176] The proficiency determination unit 114 refers to the table 1500 to generate combinations of domain, importance, and objective (step S451). The proficiency determination unit 114 then proceeds to step S401.

[0177] The proficiency assessment unit 114 calculates the variation score for the problem ID, and then calculates the score for missing answers for the problem ID (step S452).

[0178] Specifically, the proficiency determination unit 114 refers to the table 1600 and searches for entries in which the identification information of the selected LLM is set in the model 1604 and the selected problem ID is set in the problem ID 1602. The proficiency determination unit 114 calculates the average of the answer omission rates stored in the answer omission rate 1608 of the searched entries as the answer omission score for the problem ID.

[0179] The proficiency assessment unit 114 calculates the incorrect answer score for the problem ID (step S453).

[0180] Specifically, the proficiency assessment unit 114 refers to the table 1600 and searches for entries in which the selected LLM identification information is set in model 1604 and the selected problem ID is set in problem ID 1602. The proficiency assessment unit 114 calculates the average of the error response rates stored in error response rate 1607 for the searched entries as the error response score for the problem ID.

[0181] Once processing is complete for all problem IDs, the proficiency determination unit 114 selects a combination of domain, importance, and objective (step S454).

[0182] The proficiency determination unit 114 calculates a similarity score for the usage scenario (step S455). In this embodiment, the combination of domain, importance, and purpose is the information that specifies the usage scenario.

[0183] Specifically, the proficiency assessment unit 114 refers to the table 1500 to identify the problem ID corresponding to the selected combination. The proficiency assessment unit 114 calculates the average of the similarity scores of the identified problem IDs as the similarity score in the usage scenario.

[0184] The proficiency assessment unit 114 calculates the variation score in the usage scenario (step S456).

[0185] Specifically, the proficiency assessment unit 114 refers to the table 1500 to identify the problem ID corresponding to the selected combination. The proficiency assessment unit 114 calculates the average of the variation scores of the identified problem IDs as the variation score in the usage scenario.

[0186] The proficiency assessment unit 114 calculates the score for missed answers in the usage scenario (step S457).

[0187] Specifically, the proficiency assessment unit 114 refers to the table 1500 to identify the problem ID corresponding to the selected combination. The proficiency assessment unit 114 calculates the average of the missing answer scores for the identified problem IDs as the missing answer score in the usage scenario.

[0188] The proficiency assessment unit 114 calculates the score for incorrect answers in the usage scenario (step S458).

[0189] Specifically, the proficiency assessment unit 114 refers to the table 1500 to identify the problem ID corresponding to the selected combination. The proficiency assessment unit 114 calculates the average of the incorrect answer scores for the identified problem IDs as the incorrect answer score in the usage scenario.

[0190] The proficiency determination unit 114 records the calculation result (step S459).

[0191] Specifically, the proficiency determination unit 114 adds an entry to the table 1700 and sets identification information in the ID 1701 of the added entry. The proficiency determination unit 114 sets the identification information of the selected LLM in the model 1702 of the added entry and sets the selected combination values ​​in the domain 1703, purpose 1704, and importance 1705. The proficiency determination unit 114 sets the calculated similarity score 1706, variation score 1707, incorrect answer score 1708, and missing answer score 1709 of the added entry. The proficiency determination unit 114 sets the start date and time of the answer comparison process in the date and time 1710 of the added entry.

[0192] The proficiency determination unit 114 determines whether processing has been completed for all combinations (step S460).

[0193] If processing is not complete for all combinations, the proficiency determination unit 114 returns to step S454.

[0194] If processing is completed for all combinations, the proficiency determination unit 114 determines whether processing has been completed for all models (step S411).

[0195] If processing is not complete for all models, the proficiency determination unit 114 returns to step S401.

[0196] Once processing is complete for all models, the proficiency determination unit 114 terminates the proficiency determination process.

[0197] Figure 24 is a flowchart illustrating an example of the fitness determination process performed by the evaluation system 100 of Example 2.

[0198] The suitability determination unit 116 generates a combination of purpose and importance by referring to the table 1500 (step S501).

[0199] The suitability determination unit 116 selects an LLM from the list of LLMs included in the evaluation instruction (step S502).

[0200] The suitability determination unit 116 selects a combination of purpose and importance (step S503).

[0201] The suitability determination unit 116 calculates the suitability of the LLM for the combination of purpose and importance (step S504). Specifically, the following processes are performed.

[0202] (S504-1) The suitability determination unit 116 obtains parameters corresponding to the combination of purpose and importance by referring to the table 1800.

[0203] (S504-2) The suitability determination unit 116 refers to the table 1700 and searches for an entry in which the identification information of the selected LLM is set in Model 1702 and the selected objective and importance are set in Objective 1704 and Importance 1705. The suitability determination unit 116 obtains the variation score, the incorrect response score, and the missing response score from the searched entry.

[0204] (S504-2) The fitness determination unit 116 calculates the fitness using various scores and parameters. For example, the fitness determination unit 116 calculates the fitness using formula (1). Here, S_1 represents the variation score, S_2 represents the incorrect response score, and S_3 represents the missing response score. Also, P_1 represents parameter (1), P_2 represents parameter (2), and P_3 represents parameter (3). Note that formula (1) is just one example of a method for calculating the fitness and is not limited thereto.

[0205]

[0206] The suitability determination unit 116 records the calculation result (step S505).

[0207] Specifically, the fitness determination unit 116 adds an entry to the table 1900, sets the identification information of the selected LLM in the model 1901, and sets the selected purpose and importance in the purpose 1902 and importance 1903. The fitness determination unit 116 sets the calculated fitness 1904 of the added entry in the fitness 1904.

[0208] The suitability determination unit 116 determines whether processing has been completed for all combinations (step S506).

[0209] If processing has not been completed for all combinations, the suitability determination unit 116 returns to step S503.

[0210] If processing is completed for all combinations, the fit determination unit 116 determines whether processing has been completed for all models (step S507).

[0211] If processing is not complete for all models, the fitness determination unit 116 returns to step S502.

[0212] Once processing is complete for all models, the fitness determination unit 116 terminates the fitness determination process.

[0213] Figures 25 and 26 show examples of screens displayed by the evaluation system 100 of Embodiment 2.

[0214] When the display unit 115 of the evaluation system 100 receives a request from the user to display the degree of fit of the model, it displays screen 2500 on a terminal or the like operated by the user. Screen 2500 includes a display field 2501, a table 2502, and operation buttons 2503.

[0215] Display area 2501 is a column that displays the degree of fit of LLM for each purpose in graph format. Table 2502 is a table that displays the degree of fit of LLM for each purpose.

[0216] Operation button 2503 is a button for modifying the parameters used in calculating the degree of fit.

[0217] When the display unit 115 of the evaluation system 100 receives a request from the user to display the LLM proficiency level for a specific combination of domain, purpose, and importance, it displays screen 2600 on the terminal operated by the user. In addition, the response items of the generated response may be highlighted based on the similarity score.

[0218] As described above, the evaluation system 100 of Example 2 can present indicators representing various levels of proficiency. Furthermore, the evaluation system 100 of Example 2 can evaluate the accuracy of the generated AI's responses for each combination of purpose and importance.

[0219] Furthermore, the present invention is not limited to the embodiments described above, and various modifications are included. Also, for example, the embodiments described above are detailed explanations of the configuration in order to clearly illustrate the present invention, and are not necessarily limited to those comprising all the described configurations. In addition, some of the configurations in each embodiment can be added to, deleted from, or replaced with other configurations. Moreover, the generating AI is not limited to LLM (text generation), but includes AI that generates various data formats such as speech.

[0220] Furthermore, each of the above-mentioned configurations, functions, processing units, processing means, etc., may be implemented in hardware, either partially or entirely, by designing them as integrated circuits, for example. The present invention can also be implemented by software program code that realizes the functions of the embodiment. In this case, a storage medium on which the program code is recorded is provided to a computer, and the processor of that computer reads the program code stored in the storage medium. In this case, the program code read from the storage medium itself realizes the functions of the embodiment described above, and the program code itself and the storage medium on which it is stored constitute the present invention. Examples of storage media used to supply such program code include flexible disks, CD-ROMs, DVD-ROMs, hard disks, SSDs (Solid State Drives), optical disks, magneto-optical disks, CD-Rs, magnetic tapes, non-volatile memory cards, ROMs, and the like.

[0221] Furthermore, the program code that implements the functions described in this embodiment can be implemented in a wide range of programming or scripting languages, such as assembler, C / C++, Perl, Shell, PHP, Python, and Java.

[0222] Furthermore, the program code for the software that implements the functions of the embodiment may be distributed via a network and stored in a storage means such as a computer's hard disk or memory, or in a storage medium such as a CD-RW or CD-R, and the computer's processor may read and execute the program code stored in the storage means or storage medium.

[0223] In the above-described embodiment, the control lines and information lines shown are those deemed necessary for explanation and do not necessarily represent all control lines and information lines in the actual product. All components may be interconnected.

Claims

1. A computer system comprising a processor, a storage device connected to the processor, and a network interface connected to the processor, wherein the system is communicatively connected to a machine learning model that accepts prompts containing task content as input, executes the tasks defined by the prompts, and generates answers, the processor accepts evaluation instructions including information specifying the usage scenario of the machine learning model, input prompts, and expected answers, generates a plurality of conversion prompts by changing the expression pattern of the input prompts, inputs each of the input prompts and the plurality of conversion prompts to the machine learning model to obtain a plurality of generated answers, compares each of the plurality of generated answers with the expected answers to calculate a similarity score representing the similarity between the expected answers and the generated answers, calculates an index for evaluating the accuracy of the machine learning model's answers in the usage scenario by performing statistical processing using the similarity scores of the plurality of generated answers, and records a first calculation result including the index associated with the information specifying the usage scenario in the storage device.

2. A computer system according to claim 1, wherein the processor calculates a similarity score as an index representing the similarity between the assumed answer and the generated answer in the usage scenario by performing statistical processing using the similarity of the plurality of generated answers.

3. A computer system according to claim 2, wherein the processor calculates a fluctuation score representing the degree of fluctuation of the generated responses in the usage scenario as an index by performing statistical processing using the similarity of the plurality of generated responses.

4. A computer system according to claim 3, wherein the processor calculates an error score representing the degree of error in the generated responses in the usage scenario and an error score representing the degree of omission in the generated responses in the usage scenario, by performing statistical processing using the similarity of the plurality of generated responses, as the indicators.

5. A computer system according to claim 4, wherein the information specifying the usage scenario includes the purpose of use, the processor calculates a degree of fit representing the degree of fit of the machine learning model to the purpose of use using the fluctuation score, the erroneous response score, and the missing response score included in the first calculation result for which the purpose of use is the same, and records a second calculation result including the degree of fit associated with the purpose of use in the storage device.

6. A computer system according to claim 1, wherein the processor repeatedly performs a process of inputting the same prompt or the same conversion prompt to the machine learning model and obtaining the generated response a predetermined number of times.

7. A computer system according to claim 5, wherein the system maintains parameter information for managing parameters used to calculate the degree of fit for each purpose of use, and the processor selects the purpose of use, and calculates the degree of fit using the parameters corresponding to the selected purpose of use, and the variation score, the error score, and the missing answer score included in the first calculation result for which the purpose of use is the same.

8. A method for evaluating a machine learning model executed by a computer system, wherein the computer system has a processor, a storage device connected to the processor, and a network interface connected to the processor, and is connected in a communicative manner to a machine learning model that accepts prompts including the content of a task as input, executes the task defined by the prompt, and generates an answer, and the method for evaluating the machine learning model is characterized by comprising: a step in which the processor receives an evaluation instruction including information specifying a usage scenario for the machine learning model, an input prompt, and an expected answer; a step in which the processor generates a plurality of conversion prompts by changing the expression pattern of the input prompt; a step in which the processor inputs each of the input prompt and the plurality of conversion prompts to the machine learning model to obtain a plurality of generated answers; a step in which the processor compares each of the plurality of generated answers with the expected answer and calculates a similarity score representing the similarity between the expected answer and the generated answer; a step in which the processor calculates an index for evaluating the accuracy of the machine learning model's answer in the usage scenario by performing statistical processing using the similarity scores of the plurality of generated answers; and a step in which the processor records the calculation results, including the index associated with the information specifying the usage scenario, in the storage device.

Citation Information

Patent Citations

  • Model screening method and device, computer storage medium and electronic equipment

    CN118820454A

  • System and Method for Validating a Black-box Model

    US20240249039A1